This document specifies the cosmostrix diff-based terminal rendering engine: its data structures, algorithms, complexity class, worst-case behavior, and design rationale. It is intended as a reference for contributors, downstream TUI authors who want to cite or adapt the design, and reviewers evaluating cosmostrix against alternative rendering strategies.
Scope: interactive rendering path (
src/engine/cosmic_dragon_engine/terminal/draw(),src/engine/cosmic_dragon_engine/frame.rs,src/engine/cosmic_dragon_engine/cloud/rain.rs). Benchmark mode (no terminal writes) is out of scope.
The Cosmic Dragon Diff-Based Rendering Engine lives under
src/engine/cosmic_dragon_engine/ (4 subsystems: cloud/, frame.rs,
terminal/, runtime.rs), re-exported at the crate root via
pub(crate) use cosmic_dragon_engine::{cloud, frame, runtime, terminal};
in main.rs so all crate::cloud::* / crate::frame::* /
crate::runtime::* / crate::terminal::* call sites resolve unchanged.
| File | Role |
|---|---|
src/engine/cosmic_dragon_engine/frame.rs |
Differential frame buffer with double-buffered generation-based dirty tracking |
src/engine/cosmic_dragon_engine/terminal/ |
Raw-mode guard, alternate screen, RLE-batched ANSI diff pipeline, 256 KiB single-syscall flush, /dev/tty fallback |
src/engine/cosmic_dragon_engine/runtime.rs |
Runtime type vocabulary: ColorScheme, ColorMode, BoldMode, ColorPipeline |
Imported by every render-path module across src/. These are not a
subsystem — they are the substrate every rendering path stands on.
Foundations do not get relocated; they get maintained.
A folder wrapper was tried once. The earliest src/engine/cosmic_dragon_engine/
(commit 4e2ebe7 lineage) was created as a pure re-export module. It had
zero callers. It was deleted as dead code in commit 46ba457. The lesson
is now codified in src/cosmic_dragon_incubator/README.md:
An incubator namespace must hold real new code, not re-exports of existing code.
Moving frame / terminal / runtime into an engine/ folder would
be the same mistake at a larger scale: a 54-file churn commit
(crate::frame -> crate::engine::frame) with zero behavior change,
pure rename noise, and a high regression surface — for the cosmetic
gain of "the engine has a folder." The engine does not need a folder.
The engine needs to stay fast, and fast means fewer indirections and
shorter import paths.
This is a hard policy, not a preference.
frame,terminal, andruntimestay at the crate root. Forever. Noengine/folder. Norender/folder. Nocore/folder. They are crate-level primitives, likecrossterm::eventorstd::io.- Patches land in place. New rendering optimizations extend the existing files (under the 800-LOC cap, splitting if needed) — they do not branch into a new namespace.
- Additive growth goes to
src/cosmic_dragon_incubator/. The incubator namespace exists for new v15+ features. The flat engine is not v15+ — it is the foundation. Foundations are not relocated; they are maintained. - A reorganization commit will be rejected at review. If you
find yourself reaching for
git mv src/engine/cosmic_dragon_engine/frame.rs src/engine/frame.rs, stop. Read commit46ba457. Read this section. Open a doc issue instead.
"The Cosmic Dragon Diff-Based Rendering Engine" is the engine's identity. That identity is expressed in:
src/bench/bench_report.rs— the engine name printed in benchmark outputdocs/RENDER_ENGINE.md— this documentdocs/COSMIC_DRAGON_ARCHITECTURE.md— architectural narrativesrc/cosmic_dragon_incubator/README.md— incubator policyCHANGELOG.md— historical milestonesREADME.md— top-level project description
A folder is not one of them. The brand is in the docs and the code's behavior, not in the directory tree. Competitors who want to copy this engine will copy the three files and the docs — not the folder they live in. Keeping the layout flat makes the engine easier to lift, study, and cite, not harder.
Terminal emulators are byte-stream interpreters. Every visual change requires sending ANSI escape sequences over a PTY. For a Matrix rain renderer running at 60 FPS on an 80×24 terminal (1,920 cells), a naive "clear-and-redraw-everything" strategy emits roughly:
- 1,920 × (SGR fg + SGR bg + char) ≈ 1,920 × 25 bytes ≈ 48 KiB/frame
- At 60 FPS: 2.8 MiB/s of ANSI bytes through the PTY
Modern terminal emulators (Alacritty, kitty, wezterm) can sustain this, but older emulators (gnome-terminal, xterm, Termux) throttle at ~200–500 KiB/s, capping effective FPS at 5–15. The CPU cost of parsing SGR sequences in the emulator also dominates battery on laptops.
Goal: minimize bytes-per-frame sent to the terminal while preserving visual fidelity (color, afterglow, depth-of-field, overlay), so cosmostrix can hit 60 FPS on commodity terminals.
cosmostrix tracks the last frame sent to the terminal and, on each
draw() call, emits only the cells that differ from the last frame.
Within the diff, consecutive cells sharing the same (fg, bg, bold)
style tuple are batched into a single SGR + raw-character run.
Frame (src/engine/cosmic_dragon_engine/frame.rs)
├── cells: Vec<Cell> // current frame content (16 B/cell)
├── cell_gen: Vec<u32> // 4 B/cell — content generation stamp
├── gen: u32 // current content generation counter
├── semantic_gen: u32 // bumped on charset/theme/shading change
├── dirty_gen: u32 // dirty generation counter (O(1) clear)
├── dirty_cell_gen: Vec<u32> // 4 B/cell — dirty generation stamp
├── dirty: SmallVec<[usize; 64]> // queue of dirty cell indices (inline 64)
└── dirty_all: bool // fast-path flag for full redraw
Terminal (src/engine/cosmic_dragon_engine/terminal/)
├── last: Option<LastFrame> // snapshot of last sent frame
│ ├── cells: Vec<Cell> // 16 B/cell, mirrors terminal state
│ ├── semantic_gen: u32 // for invalidation detection
│ └── width/height: u16 // for resize detection
├── ansi_buf: Vec<u8> // 256 KiB cap, single write_all per frame
├── dirty_flat: Vec<usize> // reusable sort buffer
├── row_buf: String // reusable RLE accumulator (full redraw)
├── run_buf: String // reusable RLE accumulator (diff redraw)
└── color_cache: Option<ColorCache> // pre-computed SGR bytes per (fg,bg)
Cell (src/types/cell.rs) — 16 bytes
├── ch: char (4 B)
├── fg: Option<Color> (4 B — niche-optimized, same size as Color)
├── bg: Option<Color> (4 B — niche-optimized)
├── bold: bool (1 B)
└── padding (3 B, alignment)
// src/engine/cosmic_dragon_engine/frame.rs
pub fn set(&mut self, x: u16, y: u16, cell: Cell) {
let cur = /* cell from previous frame, or blank */;
if cur == cell {
return; // <- skip: cell unchanged, no dirty mark
}
self.cells[i] = cell;
self.cell_gen[i] = self.gen;
// Double-buffered dirty mark — stamp with current dirty_gen.
// Skip the push if already stamped this frame (no duplicates).
if !self.dirty_all && self.dirty_cell_gen[i] != self.dirty_gen {
self.dirty_cell_gen[i] = self.dirty_gen;
self.dirty.push(i);
}
}Cell derives PartialEq, so the comparison is a 16-byte field-wise
compare. The compiler emits a branch-predictable scalar compare; on
x86-64 this is ~4 cycles per cell. For a 1,920-cell frame, the worst
case is ~7,680 cycles (~2.5 µs at 3 GHz), negligible against the
~16 ms frame budget at 60 FPS.
Why not SIMD? The 16-byte Cell layout fits cleanly into a single
__m128i lane, but the early-exit on the ch field (first 4 bytes)
means most comparisons return on byte 1 anyway. The realistic gain is
<10% on the hot path, not worth the loss of Option<Color> ergonomics.
The dirty system uses double-buffered generations instead of a per-cell bit array. Two counters track two independent notions:
gen+cell_gen: Vec<u32>— content generation. A cell is "live" iffcell_gen[i] == gen. Bumped byclear_with_bgon semantic resets.dirty_gen+dirty_cell_gen: Vec<u32>— dirty generation. A cell is "dirty this frame" iffdirty_cell_gen[i] == dirty_gen. Bumped byclear_dirtyat end of every frame.
The double-buffer trick: instead of memset-clearing a per-cell dirty
flag array (O(N) every frame), clear_dirty does a single u32 bump.
All previous dirty stamps become "stale" (don't match the new counter)
and are instantly "clean" — no memset, no iteration. Cost: one integer
add. At 200×60=12,000 cells, the old memset was ~150 AVX2 stores; the
new bump is 1 add.
Each frame.set() that actually changes a cell:
- Stamps
dirty_cell_gen[i] = dirty_gen(4-byte write, O(1)) - Pushes the cell index onto
dirty: SmallVec<[usize; 64]>(if not already stamped this frame — duplicate-skip via the stamp check)
The renderer does not scan all cells on draw(). It iterates
dirty directly. Worst case dirty.len() == cells.len() triggers a
full-redraw fast path (dirty_all flag).
// src/engine/cosmic_dragon_engine/terminal/draw.rs draw()
let dirty_count = frame.dirty_indices().len();
let dirty_is_large = dirty_count >= total_cells / DIRTY_THRESHOLD_RATIO;
let do_full_redraw = !can_reuse_last || frame.is_dirty_all() || dirty_is_large;DIRTY_THRESHOLD_RATIO = 8 (in constants.rs) is set so that once
12.5% of cells are dirty, the per-cell overhead of the diff path exceeds the amortized cost of a full redraw. The threshold was tuned empirically.
For each dirty cell (sorted row-major):
1. Compare against last.cells[idx] — skip if equal (already sent)
2. Start a run: record (fg0, bg0, bold0)
3. Scan forward while:
- next dirty index == prev + 1 (contiguous)
- same row (no wrap)
- same (fg, bg, bold)
4. Emit:
- MoveTo(x0, y0) [only if cursor not already there]
- SGR(fg0, bg0) [only if (fg0,bg0) != current SGR state]
- bold toggle [only if bold0 != current bold state]
- raw chars [run_buf as UTF-8 bytes]
5. Update cur_fg, cur_bg, cur_bold, cur_pos
Critical optimization: SGR state (cur_fg, cur_bg, cur_bold)
is tracked across runs within a single draw() call. If run N+1 has
the same colors as run N, no SGR escape is emitted — only the raw
characters. This is where the bandwidth win comes from: a column of
50 green-on-black characters that are all dirty emits:
- Without state tracking: 50 ×
\x1b[38;2;0;255;0m\x1b[48;2;0;0;0m<ch>≈ 1,250 bytes - With state tracking: 1 × SGR + 50 ×
<ch>≈ 60 bytes (20× reduction)
ColorCache (built once after palette initialization) maps
(Option<Color>, Option<Color>) -> &[u8] SGR byte sequences. The
diff path's emit_sgr() checks the cache first; only cache misses
fall through to write_sgr_colors_buf() (which formats the SGR
on-the-fly via push_u8 — no heap allocation).
For the built-in palettes, the cache hit rate varies significantly
by scene and visual effects enabled. Measured on AMD Ryzen 7 5800HS
with --perf-stats (historical — see docs/BENCHMARKING.md §5 for
the current v50 4-scene reference matrix):
| Configuration | Hit Rate | Notes |
|---|---|---|
| Monolith scene, Cosmos palette, Subtle glitch (3%) | 18.1% | DoF + phosphor generate many intermediate shades |
| Matrix scene, Cosmos palette, Default glitch (10%) | 38.2% | More palette-color cells, but glitch still misses |
The Monolith scene's lower hit rate is expected, not a bug: depth-of-field
blends layer-0 colors 35% toward black, phosphor afterglow generates
intermediate shades between palette entries, and glitch generates random
colors. All of these produce colors that aren't palette entries, so they
miss the cache and fall through to write_sgr_colors_buf() (which is
still allocation-free — just slower than a cache hit).
The Matrix scene hits higher (38.2%) because it uses more palette-color cells directly (classic glyph rain), but still misses on glitch (10%) and bold/shading variations.
The cache remains valuable because the palette-color cells (rain
heads, bright trail cells) still hit, and those are the most frequently
re-rendered cells. The miss path is optimized: write_sgr_colors_buf
uses push_u8 (no heap alloc) and writes directly into ansi_buf.
Some changes invalidate the meaning of a cell, not just its
content. For example, switching charset from binary to katakana
means a cell containing '1' may now need to display a different
character even if '1' was the "same" logical position. The
semantic_gen counter handles this:
// src/engine/cosmic_dragon_engine/frame.rs
pub fn invalidate_semantic(&mut self, bg: Option<Color>) {
self.semantic_gen = self.semantic_gen.wrapping_add(1);
self.clear_with_bg(bg); // bump gen, set dirty_all
}On draw(), if last.semantic_gen != frame.semantic_gen, a full
redraw is forced. This prevents stale glyphs from lingering after a
charset or theme switch.
Some operations require unconditional full redraw even when no
semantic change occurred: window resize, focus regain, HUD toggle-off.
The cloud.force_draw_everything() method sets a flag that, on the
next rain_at() call, triggers frame.clear_with_bg() — bumping the
generation counter and setting dirty_all = true.
This is the escape hatch that prevents "stale cell residue" bugs (e.g., HUD text remaining visible after toggle-off in regions the rain didn't write this frame).
| Operation | Time | Space |
|---|---|---|
frame.set(x, y, cell) — no change |
O(1) | O(1) |
frame.set(x, y, cell) — changed |
O(1) | O(1) amortized |
frame.clear_with_bg() |
O(W×H) | O(1) |
frame.invalidate_semantic() |
O(W×H) | O(1) |
terminal.draw() — diff path |
O(D log D) + O(D) | O(D) |
terminal.draw() — full redraw |
O(W×H) | O(1) |
Where:
- W, H = terminal dimensions
- D = number of dirty cells (D ≤ W×H)
The O(D log D) on the diff path comes from sorting dirty_flat. In
practice, dirty cells are usually already near-sorted (rain advances
top-to-bottom, left-to-right), so the sort is close to O(D).
Worst case: theme switch triggers invalidate_semantic() ->
clear_with_bg() -> dirty_all = true -> full redraw O(W×H). At
200×60 = 12,000 cells, this is ~12,000 cell copies + ~12,000 SGR
emissions (with RLE, fewer). Measured at ~8 ms on a 3 GHz CPU, well
within the 16 ms frame budget.
| Scenario | Bytes per cell |
|---|---|
| Unchanged (skipped) | 0 |
| Changed, same style as previous run | 1–4 (UTF-8 char only) |
| Changed, new style, cache hit | ~20 (SGR + char) |
| Changed, new style, cache miss | ~25 (SGR + char) |
| Full redraw, RLE batched (N cells same style) | ~20 + N (amortized ~1/cell) |
cosmostrix emits combined fg+bg SGR in a single escape:
\x1b[38;2;R;G;B;48;2;R;G;Bm
This is 19 bytes for true-color fg+bg. The ColorCache pre-formats
these for all palette color pairs, so the hot path is a hashmap lookup
extend_from_slice.
For ANSI-256 colors (when true-color isn't supported), the format is:
\x1b[38;5;V;48;5;Vm
For reset/default colors: \x1b[39;49m (6 bytes).
MoveTo(x, y) is emitted as \x1b[Y+1;X+1H (variable length, 6–10
bytes). The diff path tracks cur_pos and skips the MoveTo if the
cursor is already at the target position (e.g., consecutive runs on
the same row). This saves ~8 bytes per skipped MoveTo.
On terminals that support ESC[?2026h (kitty, wezterm, foot), each
frame is wrapped in sync markers:
\x1b[?2026h<frame bytes>\x1b[?2026l
This tells the terminal to buffer the frame and present it atomically, eliminating visible tearing. The overhead is 12 bytes/frame, negligible.
Clear screen + emit every cell every frame.
- Pro: dead simple, no state, no bug class from stale tracking
- Con: 5–20× more bandwidth; caps FPS at ~15 on gnome-terminal
- When it wins: very short runs (<5 seconds) where the steady-state bandwidth savings don't amortize the state-tracking overhead
cosmostrix chose diff-based because the target use case is long-running (24/7 screensaver, multi-hour hacking session). The bandwidth savings compound over time.
Track each droplet's head + tail position; emit only 2–3 cursor moves
-
character writes per droplet per frame.
-
Pro: minimal bandwidth (~3 cells/droplet, no equality check)
-
Con: cannot represent background haze, phosphor afterglow, or depth-of-field — these require writing cells that aren't droplet heads/tails
-
When it wins: classic matrix rain with no visual effects
cosmostrix's signature Monolith Rain + depth-of-field + phosphor afterglow all require writing non-droplet cells, which per-droplet targeting cannot express.
Set per-column scroll regions; let the terminal emulator handle motion.
- Pro: zero CPU cost for motion (offloaded to terminal)
- Con: only uniform vertical motion; no glitch, no color cycle, no horizontal motion, no overlay
- When it wins: 1990s BBS-era matrix rain on serial terminals
Modern terminals still support scroll regions, but the visual limitations make it unsuitable for cosmostrix's feature set.
Render rain as a bitmap; send as an image.
- Pro: pixel-perfect, true Gaussian blur for depth-of-field
- Con: terminal support is spotty (kitty/wezterm yes, gnome-terminal/xterm/Termux no); feels like an image, not terminal text; loses the "character grid" aesthetic
- When it wins: art installations, not CLI tools
cosmostrix targets universal terminal support, which rules out graphics protocols.
Spawn N threads, each renders one column, mux output.
- Pro: parallelism
- Con: ANSI escape interleaving corrupts; no shared palette state; synchronization nightmare; the bottleneck is PTY bandwidth (single writer), not CPU, so parallelism doesn't help
- When it wins: never, for terminal rendering
Internal benchmark (cosmostrix --benchmark --json) and interactive
--perf-stats on AMD Ryzen 7 5800HS, Linux 6.18, 120×40 terminal:
| Metric | Value |
|---|---|
| Avg FPS (no terminal I/O) | 28,029 |
| Peak FPS | 39,971 |
| Total frames in 10s | 280,293 |
| p99 frame time (ms) | 0.043 |
| Peak RSS | 4.4 MiB |
| Metric | Monolith | Matrix | Notes |
|---|---|---|---|
| Avg dirty cells/frame | 332.8 | 1,103.5 | Matrix is 3.3× denser |
| Avg ANSI bytes/frame | 7,134.8 | 31,537.9 | Matrix is 4.4× more bandwidth |
| Naive full-redraw equivalent | ~48 KB | ~48 KB | Same terminal size |
| RLE compression vs naive | 6.7× | 1.5× | Monolith benefits more from RLE |
| Bandwidth to terminal | 418.1 KiB/s | 1,847.4 KiB/s | Matrix nears gnome-terminal limit |
| SGR cache hit rate | 18.1% | 38.2% | Matrix uses more palette colors |
| Avg frame time | 0.109 ms | 0.450 ms | Both well under 16.67ms budget |
| Max frame time | 0.303 ms | 2.474 ms | No visible jank |
| Endurance health | 89.8/100 | 57.8/100 | Matrix triggers "investigate" |
Key takeaways:
- Engine ceiling 28K FPS = 467× headroom over 60 FPS target. Visual effects consume <0.25% of frame budget.
- Monolith is the efficiency champion: 6.7× RLE compression, 418 KiB/s bandwidth. Sparse structured rain means most cells are stable per frame, so diff-based + RLE shines.
- Matrix is bandwidth-heavy: 1.8 MiB/s (4.4× Monolith). Dense classic rain means more dirty cells per frame. Still under Alacritty/kitty capacity (~10 MiB/s), but approaches gnome-terminal's limit (~2 MiB/s).
- Both scenes stay under 16ms frame budget — no visible jank even at peak (max 2.47ms << 16.67ms).
- SGR cache hit rate is scene-dependent: 18–38% measured. Not the ~95% originally estimated, but the miss path is allocation-free so the cost is acceptable. See §2.5 for analysis.
For competitor comparison data (cosmostrix vs cmatrix vs unimatrix),
see benchmark/benchmark.sh and the results table in
benchmark/HIST_BENCH.md.
Symptom: overlay text or old glyphs remain visible after the overlay is removed or the scene changes.
Cause: diff-based rendering only refreshes cells the rain actively writes. Cells in "dead zones" (no active droplet this frame) keep their previous content.
Defense: force_draw_everything() escape hatch. Called on HUD
toggle-off, focus regain, resize, paste events. Triggers
frame.clear_with_bg() which sets dirty_all = true, forcing every
cell to be re-sent.
Symptom: background fills with random characters after a full redraw.
Cause: phosphor_base_ch[] (the afterglow character buffer) is
separate from frame.cells[]. A full redraw that clears cells but
not phosphor_base_ch exposes stale afterglow characters as visible
background.
Defense: force_draw_everything() in cloud/rain.rs explicitly
clears phosphor_base_ch (or calls reset_phosphor_state() for
Monolith scene) before the redraw pass.
Symptom: after a charset switch, some cells still show the old charset's characters.
Cause: diff-based rendering compares cell content, not cell
"meaning". A cell containing '1' in the binary charset and '1'
in the hex charset would be considered "unchanged" even though the
intended display differs.
Defense: semantic_gen counter. Charset/theme/shading changes
bump the counter, which forces a full redraw via invalidate_semantic().
Symptom: garbage characters at terminal edges after a resize.
Cause: crossterm reports the new size, but the terminal emulator may not have finished resizing its internal buffer. Raw values can be degenerate (0×0 or 65535×65535) during the transition.
Defense: resize values are clamped to MIN_TERMINAL_COLS.. MAX_TERMINAL_COLS and MIN_TERMINAL_LINES..MAX_TERMINAL_LINES before
use. A debounce window (RESIZE_DEBOUNCE_MS) coalesces rapid resize
events (e.g. window drag) into a single cloud.reset().
Identified but not yet implemented:
-
Per-row hash fast path: 64-bit hash per row, updated incrementally. Skip entire row on hash match. Estimated 10–20% CPU reduction in idle mode, low impact in active mode.
-
Cell compaction to 16 bytes: pack
Colortou32with high-bit "is_set" flag. Enables SIMD compare (__m128i). Estimated 2–3× fasterframe.set()hot path, but early-exit onchfield limits real-world gain to <10%. -
SGR cache hit-rate instrumentation — DONE in v13.3.0. The
ColorCachenow tracks atomic hit/miss counters, exposed viaTerminal::encoding_stats()and the--perf-statsexit report's ENCODING section. TheTerminalalso tracks total ANSI bytes flushed and frame count, so the report shows actualavg_bytes_per_frameandbandwidth (KiB/s)instead of the previous estimate. -
Adaptive dirty threshold: dynamically tune
DIRTY_THRESHOLD_RATIObased on observed full-redraw frequency and dirty cell distribution. Currently static.
These are listed for transparency; no commitment to implement.
If you adapt or reference the cosmostrix render engine design in academic work or another project, please cite:
@software{cosmostrix,
author = {rezky\_nightky (oxyzenQ)},
title = {cosmostrix: Professional-grade cinematic Matrix rain renderer},
year = {2026},
url = {https://github.com/oxyzenQ/cosmostrix},
note = {Diff-based terminal rendering engine, v50.x},
}For the specific design rationale in this document, link to
docs/RENDER_ENGINE.md at the specific commit you reference.