You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Render very long completed messages in slices so the app does not freeze
#17666
Opening a thread that holds a very long completed assistant message parses and renders all of it in one pass on the main thread. For a 1 MiB Markdown message that blocks input, scrolling and painting for about 1.5 s in a production build, and for about 5 s under vp dev. The freeze scales with size: about 0.18 s at 128 KiB, and a 16 KiB message renders in about 40 ms.
Proposal
Messages of 128 KiB or more that mount already complete are parsed in a Web Worker with the same unified pipeline the app uses now. The result is mounted 20 top-level blocks at a time on a scheduler that keeps each turn near 12 ms, so input and paint keep running. Shorter messages, streaming messages and messages that streamed in and then finished keep the current synchronous path. If the worker fails, the message renders in one pass as it does today.
Recording (T3 Code in dark mode with the Iris theme): main on the left, this branch on the right. Same 1 MiB message in a copy of the app run by vp dev, one unmeasured take each, so only the shape is comparable. On main the page is frozen for about 4.4 s and the message appears at once at about 4.8 s. On the branch the longest freeze is about 0.24 s and content starts after about 1.8 s, but the message finishes after about 12 s and scrolling is choppy while slices mount. The hostname in the details panel is blurred.
compare-dark.mp4
Production-build measurements (the table below has the same numbers):
The tradeoff you would be approving
The longest freeze gets much shorter and content appears sooner, but the whole message takes longer to finish and uses more CPU. A very long message now appears from the top in slices instead of all at once after a freeze.
Production build, headless Chrome 154, Linux 7.2.9 (WSL2, x86_64), browser pinned to 4 CPUs. Medians of 10 samples per side from two accepted batches; the two batches agreed within max(10%, 16.7 ms) on every compared metric.
Metric
main
Branch
Benefit
1 MiB: longest main-thread stall (ms)
1,507
70
21.5× shorter (−95%)
1 MiB: first content in the page (ms)
1,068
598
1.8× sooner (−44%)
1 MiB: whole message in the page (ms)
1,506
2,361
1.6× longer (+57%), worse
1 MiB: renderer CPU (ms)
2,370
3,310
+40%, worse
128 KiB: longest main-thread stall (ms)
179
50
3.6× shorter (−72%)
128 KiB: whole message in the page (ms)
178
237
1.3× longer (+33%), worse
16 KiB (unchanged path): whole message in the page (ms)
37
33
no clear change
What I would like to know
Is this direction acceptable at all, given that long messages finish later?
If so, is a 128 KiB cutoff reasonable? I chose it, I did not tune it. I measured 16 KiB, 128 KiB and 1 MiB and nothing in between.
Is a message that streamed in and then finished staying on the synchronous path acceptable? At 1 MiB that completion freezes for about 0.8 s. I prototyped swapping to the sliced path at completion; it cut the longest stall to about 0.4 s but took about 5× as long to finish, so I left it out.
Not covered: several long messages in one thread, scrolling while slices mount (slices are 50–175 ms tasks in the dev server, so scrolling is choppy meanwhile), and a single top-level block larger than a slice, which is still mounted whole. I will open a PR only after a maintainer approves a direction here.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
Opening a thread that holds a very long completed assistant message parses and renders all of it in one pass on the main thread. For a 1 MiB Markdown message that blocks input, scrolling and painting for about 1.5 s in a production build, and for about 5 s under
vp dev. The freeze scales with size: about 0.18 s at 128 KiB, and a 16 KiB message renders in about 40 ms.Proposal
Messages of 128 KiB or more that mount already complete are parsed in a Web Worker with the same unified pipeline the app uses now. The result is mounted 20 top-level blocks at a time on a scheduler that keeps each turn near 12 ms, so input and paint keep running. Shorter messages, streaming messages and messages that streamed in and then finished keep the current synchronous path. If the worker fails, the message renders in one pass as it does today.
I have a working branch (one commit, 10 files, about +840/−110, no new setting or wire field): main...bompus:t3code:perf/progressive-chat-markdown
Before and after
Recording (T3 Code in dark mode with the Iris theme):
mainon the left, this branch on the right. Same 1 MiB message in a copy of the app run byvp dev, one unmeasured take each, so only the shape is comparable. Onmainthe page is frozen for about 4.4 s and the message appears at once at about 4.8 s. On the branch the longest freeze is about 0.24 s and content starts after about 1.8 s, but the message finishes after about 12 s and scrolling is choppy while slices mount. The hostname in the details panel is blurred.compare-dark.mp4
Production-build measurements (the table below has the same numbers):
The tradeoff you would be approving
The longest freeze gets much shorter and content appears sooner, but the whole message takes longer to finish and uses more CPU. A very long message now appears from the top in slices instead of all at once after a freeze.
Production build, headless Chrome 154, Linux 7.2.9 (WSL2, x86_64), browser pinned to 4 CPUs. Medians of 10 samples per side from two accepted batches; the two batches agreed within max(10%, 16.7 ms) on every compared metric.
mainWhat I would like to know
Not covered: several long messages in one thread, scrolling while slices mount (slices are 50–175 ms tasks in the dev server, so scrolling is choppy meanwhile), and a single top-level block larger than a slice, which is still mounted whole. I will open a PR only after a maintainer approves a direction here.
All reactions