Streaming
Rendering a reply while the model is still generating it. Streaming in gpt_markdown is data, not a Stream — rebuild the widget with a longer string as tokens arrive, and say whether more is coming via isStreaming.
Accumulating text
Buffer incoming tokens in a StringBuffer, call setState on each chunk, and flip isStreaming to false on completion:
1// Streaming is data, not a Dart Stream.
2// Rebuild the widget with a longer string as tokens arrive.
3class ReplyView extends StatefulWidget {
4 const ReplyView({super.key, required this.stream});
5 final Stream<String> stream; // your transport layer
6
7
8 State<ReplyView> createState() => _ReplyViewState();
9}
10
11class _ReplyViewState extends State<ReplyView> {
12 final _buffer = StringBuffer();
13 bool _generating = true;
14
15
16 void initState() {
17 super.initState();
18 widget.stream.listen(
19 (chunk) => setState(() => _buffer.write(chunk)),
20 onDone: () => setState(() => _generating = false),
21 onError: (_) => setState(() => _generating = false), // always flip on error
22 );
23 }
24
25
26 Widget build(BuildContext context) => SingleChildScrollView(
27 padding: const EdgeInsets.all(16),
28 child: GptMarkdown(
29 _buffer.toString(),
30 animation: GptMarkdownAnimation.fade,
31 isStreaming: _generating,
32 ),
33 );
34}Animation modes
Character reveal supports none, typewriter, fade, blurIn, and wave. The default none mode renders immediately with no ticker or reveal wrapper:
1// Character reveal modes:
2// none — immediate output; no ticker or reveal wrapper
3// typewriter — paced characters with no per-character visual transition
4// fade — arriving text fades from transparent
5// blurIn — arriving text resolves from blur while fading in
6// wave — an accent-colour crest travels across the newest characters
7GptMarkdown(
8 message.text,
9 animation: GptMarkdownAnimation.blurIn,
10 isStreaming: message.isGenerating,
11 charactersPerSecond: 300,
12 revealFadeSeconds: 0.25,
13)isStreaming lifecycle
isStreaming the widget cannot tell whether the text got longer (new token) or replaced (regenerate). It also cannot know when to fast-forward the remaining reveal.1// isStreaming tells the widget two things:
2// true → more text is coming; keep the reveal animation running
3// false → the reply is complete; fast-forward any remaining reveal
4//
5// Without it the widget cannot distinguish:
6// • text got longer → new token, continue revealing
7// • text got replaced → regenerate/branch, restart reveal
8//
9// Static history messages render with zero animation cost:
10GptMarkdown(
11 message.text,
12 animation: GptMarkdownAnimation.fade,
13 isStreaming: message.isGenerating, // false for everything already sent
14)Block entrances and lazy rendering
In 1.3, character reveal and block entrance animation are separate. blockAnimation applies to atomic blocks such as headings, lists, tables, and code fences. GptMarkdownremains eager; use SliverGptMarkdown inside a CustomScrollView when offscreen blocks should wait for the viewport.
1// Character reveal and block entrance are independent.
2GptMarkdown(
3 reply,
4 animation: GptMarkdownAnimation.fade,
5 blockAnimation: GptMarkdownBlockAnimation.growIn,
6 blockAnimationDuration: const Duration(milliseconds: 200),
7 blockAnimationCurve: Curves.easeOut,
8 isStreaming: generating,
9)
10
11// Options: none, fadeIn, growIn, slideUp, and scaleIn.
12// For a viewport-lazy message list, use the sliver explicitly:
13CustomScrollView(
14 slivers: [
15 SliverGptMarkdown(reply),
16 ],
17)SliverGptMarkdown for long documents
Use SliverGptMarkdown in CustomScrollView.slivers when a long document has its own scrollable viewport. Segmentation and custom syntax boundary matching happen eagerly, but built-in AST, span, and widget construction waits for the viewport and its cache extent. The widget has no scroll controller and composes with other slivers.
1SelectionArea(
2 child: CustomScrollView(
3 slivers: [
4 const SliverAppBar(title: Text('Answer')),
5 SliverGptMarkdown(
6 document,
7 config: const GptMarkdownConfig(
8 style: TextStyle(fontSize: 16),
9 ),
10 ),
11 ],
12 ),
13)
14
15// Source updates appear immediately. Character reveal is intentionally
16// available on GptMarkdown, not SliverGptMarkdown.components or inlineComponents force one SliverToBoxAdapter, preserving legacy whole-document behavior but disabling viewport laziness.Low-level StreamingMarkdown widget show
Most apps should configure streaming directly on GptMarkdown. The exported StreamingMarkdown widget is the lower-level split-and-cache layer for an integration that supplies its own Markdown renderer. Its builder renders both the settled prefix and live tail; seamGap supplies the block spacing between them.
1StreamingMarkdown(
2 text: reply,
3 isStreaming: generating,
4 charactersPerSecond: 300,
5 seamGap: 16,
6 builder: (context, slice) => MyMarkdownRenderer(slice),
7)Advanced streaming behavior show performance and edge cases
Performance
The modern parser normalizes the source and re-splits only the previous tail plus newly appended text. Pure-Dart document structure is cached separately from styled spans, and reveal work is limited to segments intersecting the active tail. This keeps frequent updates proportional to the changed region in the common case:
| Mode | Cost per token | Notes |
|---|---|---|
| animation: none | 14.6 ms / token | Rebuilds whole document. Gets worse as reply grows. |
| animation: fade | 11.0 ms / token | Settled prefix cached. Stayed flat in the measured reply. |
Results depend on source shape and device. The branch documentation reports rendering-work measurements rather than FPS or response latency; mobile, web, selection, RTL, and full scroll-through behavior need separate profiling.
charactersPerSecond pacing
charactersPerSecond (default 300) is a baseline, not a cap. Pick the speed that reads well when the model is slower than the reveal — lag adaptation handles the opposite case automatically. revealFadeSeconds defaults to 0.25 and controls how long an arriving character takes to settle; it does not affect typewriter. Block entrances default to 200 milliseconds with Curves.easeOut.
1// charactersPerSecond (default: 300) is the baseline reveal speed.
2// Three behaviours layer on top of it automatically:
3//
4// 1. Lag adaptation — if the backlog would take > 0.4 s, the reveal
5// speeds up to clear it in that window. The animation never falls
6// further behind on a long reply.
7//
8// 2. Fast-forward — when isStreaming turns false, the remainder
9// lands within 0.15 s. A finished reply never trickles.
10//
11// 3. Restart — text that replaces rather than extends restarts the
12// reveal (regenerate or branch switch).
13GptMarkdown(
14 reply,
15 animation: GptMarkdownAnimation.fade,
16 isStreaming: generating,
17 charactersPerSecond: 220, // calmer default; adaptation still applies
18)Reduced motion
When MediaQuery.disableAnimationsOf(context) is true, the reveal is skipped and the finished document renders immediately. No ticker runs. Nothing to configure — but test it, because it is the path users with motion sensitivity get:
1// The widget checks MediaQuery.disableAnimationsOf(context).
2// When true, the reveal is skipped and the document renders immediately.
3// No ticker runs — the fast path is taken regardless of 'animation'.
4//
5// Test this path explicitly, because it is what users with motion
6// sensitivity get in production:
7MediaQuery(
8 data: const MediaQueryData(disableAnimations: true),
9 child: GptMarkdown(
10 reply,
11 animation: GptMarkdownAnimation.fade,
12 isStreaming: generating,
13 ),
14)Incomplete code fences
The split never cuts inside a fenced code block. An unterminated fence in the settled prefix would render as literal text until the closing fence arrived — a visible flicker. Use the closed flag in codeBuilder to show a lighter UI mid-stream:
1// While closed = false on a code block, the closing fence hasn't arrived.
2// gpt_markdown never cuts the settled prefix inside a code fence —
3// a split inside an unterminated fence would render as literal text
4// and cause a visible flicker.
5//
6// Use the closed flag in codeBuilder to show a lighter UI mid-stream:
7GptMarkdown(
8 reply,
9 animation: GptMarkdownAnimation.fade,
10 isStreaming: generating,
11 codeBuilder: (context, name, code, closed) {
12 if (!closed) {
13 return Container(
14 width: double.infinity,
15 padding: const EdgeInsets.all(12),
16 color: Colors.grey.shade900,
17 child: Text(
18 code,
19 style: const TextStyle(fontFamily: 'monospace', color: Colors.white70),
20 ),
21 );
22 }
23 return MyHighlightedCodeBlock(language: name, code: code);
24 },
25)Fast-forward without visible animation
To get the split-document caching without a visible reveal, use a very high charactersPerSecond. The reveal finishes within a single frame while the prefix stays cached:
1// Set a very high charactersPerSecond to get the split-document caching
2// without a visible reveal animation. The prefix stays cached while the
3// reveal completes within a single frame.
4//
5// Useful when you want the performance benefit of the split but a
6// 'no animation' appearance.
7GptMarkdown(
8 reply,
9 animation: GptMarkdownAnimation.fade,
10 isStreaming: generating,
11 charactersPerSecond: 100000, // finishes within one frame
12)Selection after reveal
The live tail is rebuilt every frame while revealing, so it is not a stable selection target. Selection returns the moment isStreaming becomes false and the fast-forward completes. History messages with isStreaming: false are fully selectable.
1// Selection is unavailable on the live tail while the reveal runs.
2// The tail is rebuilt every frame, so it is not a stable selection target.
3// Selection returns the moment isStreaming becomes false and the
4// fast-forward completes.
5//
6// Static history messages are fully selectable — wrap the whole chat in
7// a single SelectionArea:
8SelectionArea(
9 child: ListView.builder(
10 itemBuilder: (ctx, i) => GptMarkdown(
11 messages[i].text,
12 animation: GptMarkdownAnimation.fade,
13 isStreaming: messages[i].isGenerating,
14 ),
15 ),
16)Streaming semantics
While a live reply is rebuilding, accessibility semantics collapse to one stable node so assistive technology is not presented with a changing tree on every token. Structured semantics reopen after the reply becomes quiet and reveal completion settles. Assert the final structure only after setting isStreaming: false and pumping through the quiet period.
Chat list and scroll pitfalls
ListView above rebuilds every bubble. Make the generating message itself listenable so only that bubble rebuilds.isStreaming: true after the reply finishes. The ticker keeps running and the tail keeps rebuilding for nothing. Always flip it in onDone — including error paths.1// The split keeps *this widget* cheap. It cannot help if the parent
2// ListView rebuilds every bubble on each token.
3//
4// Make only the generating message listenable:
5class _ChatListState extends State<ChatList> {
6 // Only _messages[_generatingIndex] changes during streaming —
7 // use a ValueNotifier or Riverpod so the rest of the list stays stable.
8
9 Widget build(BuildContext context) {
10 return ListView.builder(
11 itemCount: _messages.length,
12 itemBuilder: (ctx, i) {
13 final msg = _messages[i];
14 return GptMarkdown(
15 msg.text,
16 animation: GptMarkdownAnimation.fade,
17 isStreaming: msg.isGenerating,
18 );
19 },
20 );
21 }
22}
23
24// Auto-scrolling pitfall:
25// The reply grows continuously — jumping to the bottom on every token
26// fights the reveal animation. Instead, animate to the extent, or
27// only pin-scroll while the user is already at the bottom.Run the streaming demo
The package example includes a simulated reply with independent model-speed and reveal-speed sliders. Set the model faster than the reveal to watch lag adaptation catch up, press Stop mid-reply to see fast-forward, and switch to none to compare the non-animation path.
1cd example
2flutter run -d macos -t lib/streaming_demo.dart