Guides
Performance
Every measured number on this site, and an explicit list of what is not measured.
Performance
This page contains every performance claim on this site, and a list of what is not measured. That second list is the point. A docs page that lists six benchmarks and no caveats is a page you should distrust on the other five things it says.
All figures come from the framework’s own benchmarks and docs/STATUS.md,
measured on darwin/arm64 (Apple M1) unless noted. The renderer, buffer and
catalog figures were taken by decisions 1 and 2 at v0.1.0, empirically — a
scratch benchmark module built outside the repository. The colour-quantiser
figures below are v1.0.0 measurements recorded in the framework’s release notes
and buffer benchmarks.
Measured: the renderer
On a 200×60 scene that is 99% static chrome:
| Diff output | 141 bytes |
| Full repaint | 19,979 bytes |
| Reduction | ~141× |
| Cost | ~7,133 ns/op |
| Allocations | 0 allocs/op |
The claim is the ratio, not the absolute number. ADR 0002 originally recorded a byte count of “107 bytes”; that figure was an artefact of a single-byte-glyph benchmark scene, and the 2026-10-04 amendment corrected it to state the ratio.
A frame in which nothing is dirty writes zero bytes and does not call the sink at all — so an idle application costs nothing rather than burning CPU at the frame rate.
Measured: buffer representation
The comparison that overturned a prior leaning toward struct-of-arrays:
| Workload | Packed AoS | SoA |
|---|---|---|
| Row skip, mostly-clean rows | 5,373 ns/op | 5,128 ns/op |
| Row skip, every row dirty | 147.2 ns/op | 636.7 ns/op |
OpenTUI’s SoA advantage is a Zig mem.eql advantage and does not transfer to
Go. AoS ties where SoA was supposed to win and is 4.4× faster when every row is
dirty, because comparing a row is one wide memcmp rather than four.
The dirty-heavy case is the one a TUI actually hits: a dashboard’s data region changes while its chrome does not.
Measured: versus tcell
tcell flush, one-row-dirty workload |
280,814 ns/op |
| TermMosaic two-tier diff, same workload | 7,133 ns/op |
This, plus the fact that tcell’s headless backend cannot expose the cell buffer
that widget tests need, is what settled
ADR 0001.
Measured: the catalog’s flat cost
| Items | List |
Table |
|---|---|---|
| 10,000 | 13,320 ns | 16,801 ns |
| 100,000 | 14,242 ns | 17,885 ns |
Ten times the data for seven percent more time, both at zero allocations.
Measured: colour-quantiser selection (v1.0.0)
Steady-state frame-path cost of mapping a truecolor Colour to the nearest
entry in the 256- and 16-colour rungs, after the Lab/CIEDE2000 replacement:
| Before (redmean) | After (Lab CIEDE2000) | |
|---|---|---|
Nearest256 |
224.8 ns/op | 6.611 ns/op, 0 allocs |
Nearest16 |
16.02 ns/op | 7.126 ns/op, 0 allocs |
The replacement is faster as well as correct, which is not the usual shape
of a fix: selection goes through a per-colour memo, so only the first use of a
colour pays for the exhaustive CIEDE2000 search. The audit’s own selection-error
metric reads 0.000 on both rungs, with 0 of 281,216 colour-rungs regressed.
Source: the framework’s buffer benchmarks and the v1.0.0 release notes.
Measured: wide glyphs
Benchmarked for the first time in v0.1.0, on a 200×60 scene:
| Scene | Full repaint |
|---|---|
| ASCII | 74,078 ns/op |
| Wide glyphs | 74,139 ns/op |
The row-skip tier is indistinguishable, and all paths are 0 allocs/op. The wide paths did not regress.
Specified-and-tested, not measured
These are properties a test pins. They are real, and they are not benchmarks — ADR 0005 and ADR 0008 are explicit about the difference.
| Property | Why it is not a benchmark |
|---|---|
| 0 allocs per frame on the frame path | Asserted rather than benchmarked. The allocator’s behaviour under a real workload is not what the assertion is about; the assertion is that no code path allocates. |
| 0 allocs on the key path | ADR 0005 says input decoding is I/O-bound — a benchmark would measure the operating system. The number that matters is specified and pinned by a test. |
Draw allocates nothing in steady state (TextInput) |
Same reasoning. The visible runs are rebuilt only when the text, caret, styles or rect change — a structural property, not a timing one. |
| The degenerate-size contract | A contract, not a cost: no panic, clip never blank, 0×0 writes zero bytes. |
Derived, not observed
No resize has ever been observed against a real terminal being dragged.
ADR 0007’s drag-resize costs are derived from existing code and from ADR 0002/0003’s measurements. The ADR says so itself, in those terms, and this site does not paper over it.
What does exist for resize is scripted: the degenerate-size sweep in
render/responsive_contract_test.go, and examples/hello’s TestResizeGolden
walking grow → shrink → degenerate → recover with a golden file per step. That
is the scripted resize sweep ADR 0007 risk 6 asks for, and it is not the same
thing as a human dragging a window.
Writing those four tests also found a real defect: Renderer.Render called
Sink.Flush even when it wrote zero bytes, breaking ADR 0007 §4’s contract.
Not measured, and listed
Each of these is a gap, not an omission:
- Per-widget cost for the five viz widgets. There is no benchmark for
Gauge,Meter,Sparkline,BarChartorProgressBar.GaugeandSparklineare the Braille ones, so they are the ones where a benchmark would be most interesting. - Frame pacing under load.
render.Pacerhas a 30–60 fps budget and no benchmark. - Mutation cost. The flat-cost numbers are per-frame. Changing an item list invalidates; doing that every frame is a different cost, unmeasured.
- Resize cost against a real terminal. See above.
- Memory. Flat per-frame cost does not mean the data is free to hold, and nothing here measures a 100,000-row widget’s footprint.
- The capture tool. Nothing measures how long
cmd/capturetakes.
One defect the benchmark found — fixed in v0.2.0
This section said the defect was unfixed. It is fixed. Leaving the text would have been a page asserting a known regression in its own hot path that no longer exists, which is worse than a performance page that quietly omits it.
What the defect was. The diff’s cursor-run suppression assumed one cell per rune, so every wide glyph was preceded by a cursor-position escape — a wide glyph advances the terminal’s cursor by two cells, while the run tracker recorded only the glyph’s own column:
| Scene, 6,000 glyphs | Cursor moves | Bytes |
|---|---|---|
| ASCII | 30 | 6,233 |
| Wide, before the fix | 6,000 | 68,832 |
| Wide, as of v0.2.0 | 60 | 19,443 |
Correct output either way, but about 11× the bytes for identical bytes on screen.
What the fix is. The tracker now advances by the glyph’s cell width rather than assuming one cell per rune — 60 moves and 19,443 bytes, 3.12×, down from 6,000 moves and 68,832 bytes. The ASCII path is unchanged at ~7,200 ns/op with 0 allocs, which is the part worth noting: the defect cost nothing on narrow text, so no ASCII benchmark would ever have surfaced it.
The benchmark that had pinned the defective behaviour now asserts the corrected one. That is the more useful half of this entry — a benchmark written against wrong output is a benchmark that would have failed the fix.
How to reproduce
git clone https://github.com/serkanalgur/termmosaic
cd termmosaic
go test ./... -bench=. -benchmem
The relevant benchmarks are in widgets/data/bench_test.go,
internal/diff/, buffer/wideglyph_test.go and render/wideglyph_bench_test.go.
Raw benchmark output is quoted inline in
ADR 0001 and
ADR 0002, including the workloads that did
not produce a clean result — which is the more useful half of a benchmark
record.
Your numbers on your scene will differ. These are the framework’s, on an M1. The wide-glyph figures are from v0.2.0; the rest were first measured at v0.1.0 and have not changed since.
Reading next
- Renderer and diff — where the numbers come from.
- Virtualization — what “flat cost” claims.
- The ADRs — the full reasoning, verbatim, with the rejected alternatives and what could not be measured.