Performance architecture
HyperMarkdown parses the change, not the conversation.
Sub-block caching
Document-level memoization cannot help when one giant code block or table remains active for thousands of chunks. HyperMarkdown settles smaller units:
- complete code lines
- complete table rows
- complete list items
- the changing trailing frontier
Those units retain their parsed output and React identity. New input only pays for work that can still change.
Traditional streaming renderer
new token
↓
growing active block
↓
parse the active block again
↓
render again
HyperMarkdown:
new token
↓
active block
│
├── settled code lines → cached
├── settled table rows → cached
├── settled list items → cached
└── changing frontier → parse
A 1,000-line code block does not become a 1,000-line parsing problem every time another token arrives.
Integration guidance
To preserve the architecture's advantage:
- Write deltas rather than accumulated snapshots.
- Keep one renderer mounted for the duration of an answer.
- Finalize exactly once.
- Build plugin objects and component maps once.
- Avoid mirroring the active Markdown buffer through React state.
- Leave word animation off unless you need it; it grows the DOM and pauses streamed highlighting.
- Coalesce tiny tokens (rAF or 16–32 ms) so parse+commit is per frame, not per character.
See Benchmark methodology for how the numbers on the homepage were measured.