You said:
ChatGPT said:
I’ll inspect the mintty patch history and rendering hot paths first, then compare against wezterm/ghostty rendering/update strategies and pull out concrete candidate experiments rather than code changes.
You said:
ChatGPT said:
Noted. With Ligatures=0, the ASCII Uniscribe bypass and ASCII run merging are directly relevant. With DisplaySpeedup=0, I’ll ignore frame-dropping/throttling as a primary “optimization” and focus on reducing work per full refresh / update.
I found several plausible candidates. The biggest one is not compiler flags; it is avoiding repainting/copying work during scroll/full-refresh cases, then shaving CPU work inside the ASCII path you now hit with Ligatures=0.
I only did static inspection / grepping, not code changes or benchmarks.
What your patch already attacks
The current patch series already hits the obvious hot paths:
| Area | Current patch |
|---|---|
| Tearing / coherent updates | wintext: double-buffer display updates, wintext: blit only the painted row span from the back buffer |
| ASCII text | Ligatures=0 path renders printable ASCII runs through ExtTextOutW instead of Uniscribe |
| Blank text | all-blank runs skip Uniscribe |
| Run count | term: merge output chunks of unshaped ASCII text |
| GDI fills | stock DC_BRUSH for per-run background fills |
| Width calls | win_char_width memoization |
| Build | -O3, LTO, but you already found -mtune / -march are not the lever |
With your settings, Ligatures=0 means unshaped_ascii is true in src/term.c, so the ASCII run merge and ExtTextOutW bypass are active. DisplaySpeedup=0 means the frame-skip path in do_update() is not a performance strategy here.
Best candidate: resurrect display scrolling, but into the back buffer
There is an inactive scroll-optimization scaffold in src/term.c:
-
use_display_scrollingis disabled. -
disp_scroll()/disp_do_scroll()exist. -
disp_do_scroll()wants to callwin_do_scroll(topscroll, botscroll, lines). -
I did not find a
win_do_scrollimplementation.
That is the highest-upside target for “large terminals doing full refreshes” and high-DPI screens.
Right now, scroll-heavy output appears to do roughly this:
-
terminal model scrolls,
-
lines_scrolled += lines, -
paint eventually walks/redraws visible rows,
-
back buffer is blitted to the window.
For normal full-width scroll regions, a better fast path is:
-
scroll/copy the affected rectangle inside
paint_buf_dc, -
rotate/update
term.displines, -
invalidate only newly exposed rows,
-
still do one coherent final
BitBltto the real window.
That preserves your tearing fix because the window still only sees the final buffer copy, but it avoids redrawing all shifted rows. This should matter a lot for huge windows and HiDPI because the shifted area is mostly identical pixels.
I would start with a deliberately narrow fast path:
-
full-width scroll region only,
-
no left/right margin mode,
-
normal
LATTR_NORMlines only, -
no background image / sixel / emoji overlay interaction at first,
-
no search-bar exclusion complication,
-
no double-height / double-width line attributes,
-
copy one contiguous scroll rectangle in the back buffer,
-
mark exposed rows dirty.
This is probably the most mintty-shaped version of what GPU terminals get “for free” by retaining/copying scene data.
Next candidate: dirty row bitset instead of one dirty top/bottom span
Your buffered path tracks:
-
paint_dirty_top -
paint_dirty_bot
and win_paint_buffer_end() blits one contiguous span.
That is simple and good, but it over-blits whenever dirty rows are sparse. Cursor blink, selection updates, search highlights, IME/composition, or small changes near top and bottom can force a tall HiDPI blit.
Candidate: track a dirty row bitset, then in win_paint_buffer_end() coalesce contiguous dirty row runs:
rows dirty: 3,4,5,30,31,80 blits: 3-5, 30-31, 80
Use a heuristic:
-
if dirty row runs are few, issue multiple
BitBlts; -
if many rows are dirty, collapse to the existing single span.
For full-screen refreshes this will not help much, but for mixed workloads on huge terminals it prevents “almost full-window” copies caused by two distant dirty rows.
ASCII hot path: avoid repeated bidi_class() for boring ASCII
Even after the ASCII merge, the drawing loop in src/term.c still does per-cell classification work. Around the render loop, it calls bidi_class(xtchar) before deciding whether to break runs. bidi_class() in src/minibidi.c is a table search over Unicode ranges.
With Ligatures=0, most printable ASCII cells do not need full bidi classification for the “continue same ASCII run” case. The current code already has logic like “if both previous and current are ASCII, don’t break on bidi class,” but it still computes the class first.
Candidate:
-
add an ASCII fast path before
bidi_class(); -
for
unshaped_ascii && current_ascii && previous_ascii, skip full bidi classification; -
only compute the class at ASCII/non-ASCII boundaries or when non-ASCII is involved.
This is a CPU-side improvement, not GDI-side. It should show up most on full repaint of ASCII-heavy terminals.
Line-level render cache, inspired by wezterm/ghostty
WezTerm and Ghostty both avoid rebuilding more render state than necessary. Ghostty’s render state has retained row data and dirty tracking, with a two-phase update path designed to keep terminal locking and update work bounded; it also has helpers that scan grouped cells for dirty/special handling. GitHub
GitHub
WezTerm goes further with CPU-side caches:
-
line shape hash,
-
shaped line element cache,
-
line quad cache,
-
glyph cache.
Mintty cannot copy the GPU architecture, but the useful idea is still valid: cache per-line render decomposition.
Mintty currently appears to rebuild runs from terminal cells during paint. A line-level generation/hash could let you cache:
-
run boundaries,
-
font/style attributes,
-
“boring ASCII” classification,
-
whether line contains RTL / wide / combining / emoji / overhang-sensitive content,
-
maybe precomputed
dxsshape for simple runs.
Then invalidate on:
-
cell edit,
-
scroll exposure,
-
font / DPI / config generation,
-
selection/cursor/search overlay changes,
-
emoji/image/combining interactions.
This is more invasive than the scroll-copy path, but it addresses the “large full refresh” CPU cost directly.
A smaller first version: add only a per-termline sequence number and counters to learn how many “repainted” rows are logically unchanged.
GDI state-call memoization inside win_text()
win_text() still does many per-run GDI state calls:
-
SelectObject(dc, font) -
SetTextColor(dc, fg) -
SetBkColor(dc, bg) -
SetBkMode(...) -
SetDCBrushColor(dc, bg)
On large ASCII refreshes, the number of runs is reduced but still non-trivial. Tracking the last selected font / fg / bg / brush color / background mode per paint DC could avoid redundant calls.
This is low-risk if scoped carefully:
-
reset cache when the HDC changes,
-
reset after helper functions that use temporary DCs,
-
reset around any path that may mutate DC state unexpectedly,
-
keep correctness over cleverness.
Expected gain is modest, but this is the kind of GDI overhead that accumulates.
Guarded lpDx == NULL experiment for “boring monospaced ASCII”
win_text() builds an int dxs[len] array and passes it through text output. For truly boring primary-font ASCII, that may be unnecessary.
Microsoft documents that ExtTextOut can use default spacing when lpDx is NULL; the lpDx array is for explicit character-spacing/origin control. Microsoft Learn
Candidate fast path:
-
primary font,
-
printable ASCII only,
-
no combining,
-
no overstrike,
-
no wide chars,
-
no char spacing option,
-
no scaling/stretch case,
-
normal cell width equals font advance,
-
Ligatures=0, -
not using Uniscribe.
Then try ExtTextOutW(..., lpDx=NULL) or TextOutW.
This needs pixel-exact tests. Terminal alignment bugs are easy here. But if it works, it avoids filling dxs[] and gives GDI the simplest possible text path.
Back buffer: test CreateDIBSection vs compatible bitmap
The double buffer currently uses CreateCompatibleDC() plus CreateCompatibleBitmap(). That usually means a device-dependent bitmap unless the source DC already has a DIB section selected. Microsoft documents CreateDIBSection as creating a bitmap with app-accessible bits. Microsoft Learn
Candidate experiment:
-
use a top-down 32-bit DIB section for the persistent paint buffer;
-
keep rendering text with GDI into the memory DC;
-
measure:
-
text output into DIB section,
-
background fills,
-
final
BitBlt, -
resize behavior,
-
HiDPI bandwidth.
-
This could be faster or slower. DIBs are attractive for predictable memory layout and direct clearing, but GDI text into a DIB section may underperform a device-compatible bitmap on some systems. I would benchmark it, not assume it wins.
Non-ASCII shaping: cache or bypass more, but lower priority for your config
Your patch already avoids Uniscribe for blank and printable ASCII. That matters because ScriptStringAnalyse does substantial work: it allocates internal buffers and performs itemization, shaping, placing, and breaking. Microsoft Learn
Remaining candidates:
-
cache “this run definitely does not need Uniscribe” classification;
-
bypass Uniscribe for box drawing / block elements / simple CJK runs if correctness tests pass;
-
cache shaped results for repeated non-ASCII runs.
I would be careful caching SCRIPT_STRING_ANALYSIS itself. The object is opaque, allocated by Uniscribe, tied to analysis parameters, and must be freed. A mintty-side cache of simpler metadata is probably safer than retaining Uniscribe objects.
Background fill batching
Even with stock DC_BRUSH, there is still a fill per background run. Large full refreshes with lots of color changes will pay for that.
Possible architecture:
-
background pass: merge adjacent same-background rectangles per row;
-
text pass: draw foreground only;
-
special decorations pass.
This is invasive because mintty currently folds background, overhang, clipping, underline, selection, and text output together. It might help color-heavy workloads but is lower priority than scroll-copy and ASCII CPU reductions.
Low-confidence: PolyTextOutW
PolyTextOut can draw multiple strings with the currently selected font and colors. Microsoft Learn
It might batch foreground text runs after a separate background pass. But because mintty has per-run clipping, backgrounds, dx arrays, Uniscribe fallback, overhang handling, wide chars, combining chars, and font fallback, this looks like a larger refactor with uncertain payoff.
I would not chase it first.
Instrument before changing much
The useful counters would be:
-
term_paintrows scanned, -
dirty rows found,
-
rows actually drawn,
-
win_text()calls, -
chars passed to
win_text(), -
ASCII vs non-ASCII chars/runs,
-
bidi_class()calls, -
win_char_width()cache hits/misses, -
ScriptStringAnalyse()calls and time, -
ExtTextOutW()calls and time, -
FillRect()calls and pixels, -
BitBlt()calls and pixels, -
number of dirty row spans,
-
scroll lines eligible for scroll-copy,
-
scroll lines rejected and reason.
For your workload, I would bet the ranking is:
-
buffered scroll-copy / revive display scrolling safely;
-
ASCII
bidi_class()fast path; -
dirty row bitset / multi-span blit;
-
GDI state memoization;
-
guarded
lpDx=NULLboring ASCII path; -
line-level render-run cache;
-
DIB-section back buffer experiment;
-
non-ASCII shaping cache/bypass;
-
background batching /
PolyTextOutW.
The scroll-copy path is the one I would investigate first because the disabled scaffold already exists and the payoff matches your exact problem: large terminals, full refreshes, and HiDPI pixel bandwidth.
You said:
ChatGPT said:
I’ll make a focused diagnostic-only patch against the uploaded tree: low-overhead counters/timers, no rendering behavior changes, and one per-process log under C:\Temp.
I found the buffered paint/render-loop boundaries and will aggregate one compact record per update, with hooks around term_paint(), win_text(), GDI primitives, Uniscribe, and width caching.
I’m adding a MINTTY_PERF=0 kill switch and key=value log records, after resetting prior scratch edits so the diff stays instrumentation-only.
Created the instrumentation-only patch:
It adds src/perf.c / src/perf.h and instruments term.c + wintext.c. No render behavior changes intended.
It logs to:
C:\Temp\mintty.perf.{PID}.log
It creates C:\Temp if needed, appends to the per-PID log, flushes periodically, and closes/flushed at exit. It is enabled by default; disable with:
BashMINTTY_PERF=0 mintty
The log format is one key=value line per update, plus skip= lines for skipped updates. It includes counters/timings for:
update / WM_PAINT timing term_paint timing win_text calls/chars/phases ExtTextOutW vs ScriptStringOut / ScriptStringAnalyse FillRect calls/pixels/time BitBlt calls/pixels/time dirty rows/cells back-buffer use/reject/recreate/stale-full-repaint scroll calls/visible scroll opportunity lines bidi_class calls and ASCII/non-ASCII split run-break reasons win_char_width cache hits/misses/uncached time invalidate calls/cells/full-screen invalidations
I also checked:
git apply --check /mnt/data/mintty-perf-instrumentation.patch git diff --check
and syntax-checked perf.c with stub Win32 headers. I did not do a real mintty build here because this container does not have the Cygwin/Win32 build headers.
You said:
ChatGPT said:
I parsed the full log now. The headline already changed: the slow tail is dominated by non-ASCII/box-line-art rendering fragmentation, not Uniscribe shaping and not BitBlt.
The report-generation wrapper tripped over an invalid Python syntax draft; I’m rerunning with a plain Python script and the already-parsed metrics unchanged.
I parsed the full log with a script, not just the truncated preview: 4,043 records, 4,018 source=update records, 25 source=paint records, total measured update time 107.7s.
Generated files:
Main finding
The slow path is Unicode / line-art fragmentation in win_text(), not BitBlt, not Uniscribe shaping, and not normal ASCII full refresh.
The fancy tig line-art suspicion looks real on the terminal side. I cannot say how much CPU tig itself used, but the mintty log shows very large terminal paint/update costs during non-ASCII updates:
| cohort | updates | total time | p50 | p95 | max |
|---|---|---|---|---|---|
| ASCII-only | 2,505 | 18.6s | 5.6ms | 21.3ms | 57.5ms |
| non-ASCII present | 1,513 | 89.2s | 16.4ms | 362.2ms | 663.7ms |
| non-ASCII >100ms | 214 | 65.7s | 306.5ms | 528.1ms | 663.7ms |
Those 214 slow non-ASCII updates are only 5.3% of updates but account for 61% of total measured update time.
Where the time went
Across all update records:
| metric | total |
|---|---|
term_paint_ticks_us | 100.2s |
win_text_ticks_us | 78.8s |
ext_text_out_ticks_us | 8.4s |
fillrect_ticks_us | 6.8s |
bitblt_ticks_us | 3.7s |
script_string_out_ticks_us | 0.28s |
uniscribe_analyse_ticks_us | 0.13s |
For the bad non-ASCII >100ms cohort:
| metric | total |
|---|---|
| total update time | 65.7s |
term_paint | 65.2s |
win_text | 63.6s |
FillRect | 5.8s |
ExtTextOutW | 0.47s |
BitBlt | 0.38s |
ScriptStringOut | 0.16s |
ScriptStringAnalyse | 0.06s |
So: Uniscribe is basically irrelevant in this capture. The problem is that line-art-ish non-ASCII causes massive call fragmentation:
| cohort | win_text calls | text chars | chars / win_text call | fill calls |
|---|---|---|---|---|
| ASCII-only | 87,487 | 14.6M | 167.3 | 1,631 |
| non-ASCII present | 4.23M | 24.0M | 6.1 | 1.37M |
| non-ASCII >100ms | 3.71M | 5.08M | 1.8 | 1.27M |
That is the smoking gun.
Updated target ranking
1. Prime target: line-art / box-drawing “boring non-ASCII” fast path
This moves to the top.
The generic “non-ASCII shaping cache/bypass” idea should become something narrower:
Detect box drawing / block / common TUI line-art characters as monospaced boring glyphs and keep them in longer runs.
Specifically, investigate reducing breaks at ASCII/non-ASCII boundaries and attr boundaries where safe. In the slow cohort, the code is effectively doing per-cell or near-per-cell work. Even if each GDI call is cheap, millions of tiny win_text/fill calls are not.
Candidate shape:
same attr + same font + same cell width + no combining + no emoji + no bidi hazard => draw as one direct GDI run, even if it contains ASCII + U+2500 box chars
Do not chase Uniscribe first. The log says Uniscribe is not the bottleneck.
2. Prime target: background/fill batching for fragmented TUI rows
FillRect is not the largest standalone timer, but the slow cohort has 1.27M fill calls and 5.8s in FillRect. More importantly, those fill calls probably correlate with the unmeasured overhead inside win_text.
For line-art screens, a background pass that merges same-background rectangles before foreground text could help. This is now more interesting than it looked before.
3. Add deeper win_text() sub-instrumentation before changing too much
There is a large gap inside win_text:
slow non-ASCII cohort: win_text 63.6s ExtTextOutW 0.47s ScriptOut 0.16s Uniscribe 0.06s FillRect 5.8s
That leaves a lot of time inside win_text not explained by the GDI calls I instrumented. Before optimizing blindly, I would add sub-counters around:
font fallback / get_font text buffer conversion dx array setup rect / clipping setup overhang handling phase loop overhead background path decisions wide/non-ASCII branch decisions
But the direction is clear: reduce call count and per-cell path overhead.
4. Still worthwhile: buffered scroll-copy / revive display scrolling
This remains a good optimization, but it is not the main explanation for the painful spikes in this log.
There were 416,799 visible scroll lines across 731 scroll-visible updates, so scroll-copy still has real opportunity. But those updates totaled only 12.9s, and the worst non-ASCII spikes had scroll_visible_lines=0.
So I would still do it, but after the line-art fragmentation work.
5. ASCII bidi_class() fast path: useful but not top priority
The log shows huge scan volume:
paint_cells_seen 115.4M bidi_class_calls 115.4M bidi_ascii_cells 108.9M
Even sparse one-cell updates scan/classify a full screen. For example, the sparse <=10 dirty cells cohort has 1,168 updates and 5.3s total time.
So yes, add the ASCII short-circuit. But this will improve the baseline, not the painful Unicode spikes.
6. Dirty-row bitset / multi-span blit: lower priority now
Sparse updates are frequent, but cheap enough:
<=10 dirty cells: 1168 updates 5.3s total p50 3.0ms p95 12.1ms
A better dirty row/span model would help polish idle/cursor/small updates, but it is not where the big time is.
7. lpDx == NULL ASCII experiment: lower priority
ASCII-only is already reasonably healthy:
ASCII-only p50: 5.6ms ASCII-only p95: 21.3ms ASCII-only max: 57.5ms
It may still help full ASCII refreshes, but the log does not justify doing it before the non-ASCII fragmentation work.
8. GDI state memoization: not compelling from this run
Totals are tiny:
select_font_ticks_us 151ms set_text_color_ticks_us 91ms set_bk_color_ticks_us 91ms
There may be hidden overhead not captured here, but the measured state-change calls are not a primary target.
9. DIB-section back buffer: not compelling from this run
BitBlt moved 13.6 billion pixels but took only 3.7s total. It matters, but it is not the current wall.
Practical next move
I would pursue this sequence:
-
Instrument
win_text()internals more deeply, because the slow path is inside it but not in the obvious GDI calls. -
Prototype a box-drawing / block-character fast path that treats common TUI Unicode as direct monospaced text and merges runs aggressively.
-
Batch/merge background fills for same-background cells/runs.
-
Then do scroll-copy into the back buffer.
-
Then do ASCII
bidi_class()skip and small-update scan reductions.
The current data says the fastest win is not “make GDI faster”; it is stop feeding GDI and win_text() millions of tiny fragments for Unicode TUI screens.
You said:
ChatGPT said:
I’ll apply the prior instrumentation patch to a fresh mintty tree, inspect win_text() structure, add deeper counters without changing render behavior, then emit a second patch that applies cleanly on top of the previous one.
The likely missing bucket is now clearer from the code: self-drawn box/block graphics allocate/copy/free a temporary text buffer, create pens/brushes/clip regions, and run GDI line/fill primitives, none of which the first patch timed deeply.
Done:
This patch applies on top of mintty-perf-instrumentation.patch.
I checked:
Bashgit apply /mnt/data/mintty-perf-instrumentation.patch git apply --check /mnt/data/mintty-perf-wintext-deep-instrumentation.patch git apply /mnt/data/mintty-perf-wintext-deep-instrumentation.patch git diff --check -- src/perf.h src/perf.c src/wintext.c
I also syntax-checked src/perf.c with a fake Win32 header. I still could not do a real Cygwin/Win32 mintty build in this container.
The new fields focus on the missing win_text() buckets:
font/attribute setup origtext alloc/free for self-drawn chars dxs and ulen setup Uniscribe decision overhead background clearing / background image path coordinate transforms SetBkMode / SetDCBrushColor plain vs combining textout path self-drawn boxpower / boxcoded / vt52 timings clip region creation/selection/deletion self-drawn FillRect pixels/ticks line/rect/polygon/chord/AngleArc/SetPixel operation counts GDI pen/brush/region/select/delete object counts
This should make the next log identify whether the Unicode TUI spikes are mostly:
per-cell origtext allocation/free per-run pen/brush/clip region churn LineTo / manual box drawing volume block/rect fill volume ordinary win_text setup overhead from fragmentation
You said:
tig on linux/ is towards the end. please just summarize it and think about what might be good to fix; don't start patching mintty yet.ChatGPT said:
I’ll parse the full log into cohorts, then separately compare the tail section against the earlier/non-tail activity so the tig linux/ period does not get washed out by normal usage.
I parsed the full mintty.perf.43752.log with a key/value log parser. The log has 1,234 records, including 1,230 source=update records. I treated update >= 853 as the tig linux/ tail because that is where the sustained >100ms non-ASCII/self-drawn cluster begins.
Summary
The deeper data confirms the previous suspicion more sharply:
The slow tig period is dominated by mintty’s self-drawn box/line-art path, specifically win_text_selfdraw_boxpower, plus per-cell GDI object churn.
It is not primarily:
BitBlt Uniscribe ScriptStringOut ExtTextOutW scrolling ASCII text
Overall profile
| cohort | updates | total | p50 | p95 | max |
|---|---|---|---|---|---|
| all updates | 1,230 | 46.1s | 10.4ms | 283.6ms | 619.0ms |
| ASCII-only | 730 | 5.9s | 5.1ms | 19.5ms | 42.5ms |
| non-ASCII present | 500 | 40.3s | 20.8ms | 382.9ms | 619.0ms |
update >= 853 tail | 382 | 34.8s | 6.0ms | 402.3ms | 619.0ms |
| slow >100ms | 105 | 32.4s | 308.8ms | 505.1ms | 619.0ms |
The 105 slow updates are only 8.5% of updates, but account for 70.3% of total update time.
The tig tail is 31.1% of updates, but 75.5% of total update time.
The culprit bucket
Across all updates:
| bucket | total |
|---|---|
term_paint | 43.8s |
win_text | 37.0s |
win_text_selfdraw | 29.7s |
win_text_selfdraw_boxpower | 29.6s |
win_text_textout_path | 3.0s |
ExtTextOutW | 2.7s |
BitBlt | 1.3s |
ScriptStringOut | 0.14s |
ScriptStringAnalyse | 0.06s |
So almost all of the pain is:
term_paint -> win_text -> selfdraw -> boxpower
The tail makes this even clearer:
| tail bucket | total |
|---|---|
| total update time | 34.8s |
term_paint | 34.2s |
win_text | 32.1s |
win_text_selfdraw | 27.4s |
win_text_selfdraw_boxpower | 27.4s |
BitBlt | 0.31s |
ExtTextOutW | 0.52s |
Uniscribe analyse | 0.036s |
Why it is slow
The bad path is doing self-drawn box-power rendering at near per-cell granularity.
For the 105 slow updates:
| metric | total |
|---|---|
win_text_calls | 1,774,898 |
win_text_boxpower_calls | 1,200,206 |
| line ops | 1,470,530 |
GDI CreatePen | 1,800,309 |
GDI CreateBrush | 600,103 |
GDI CreateRectRgn | 600,103 |
GDI DeleteObject | 3,000,515 |
GDI SelectObject | 1,200,206 |
GDI SelectClipRgn | 1,200,206 |
That averages per slow update to roughly:
~11,431 boxpower calls ~14,005 line ops ~17,146 CreatePen calls ~5,715 CreateBrush calls ~5,715 CreateRectRgn calls ~28,576 DeleteObject calls
This is not a text-shaping problem. It is a “draw thousands of tiny vector-ish box glyph fragments with fresh GDI resources/clips” problem.
The deeper win_text() breakdown
For the slow >100ms cohort:
| sub-bucket | total |
|---|---|
win_text_selfdraw | 27.14s |
| selfdraw resource setup | 7.07s |
| selfdraw teardown | 3.13s |
| clip handling | 0.07s |
| ordinary textout path | 0.47s |
| BitBlt | 0.19s |
There is still about 16.9s inside selfdraw not accounted for by the named resource/teardown/fill/clip timers. That is likely the actual drawing loop and GDI primitive calls themselves: line drawing, arcs, and internal per-glyph math/branching.
The important part: resource churn alone is already ~10.2s in the slow cohort:
Create*/setup: 7.07s teardown: 3.13s
That is a very good optimization target.
What seems worth fixing
1. Stop creating/deleting GDI objects per box glyph
This is now the top target.
The current path is creating pens/brushes/regions at huge scale. Cache or reuse these within a paint/update, probably keyed by:
fg color bg color line width / style cell geometry clip rectangle shape
Even a simple per-update cache for the common pen/brush cases could remove millions of Create* / DeleteObject calls.
This should be the first real optimization.
2. Avoid per-cell clip region setup for boxpower glyphs
The slow path has:
1.2M SelectClipRgn calls 600k CreateRectRgn calls
If boxpower drawing can rely on coordinate clipping from the row/cell rectangle, or if a clip region can be reused per row/run rather than per glyph, this should help materially.
The clip timer itself is only 72ms, but that likely excludes some region object churn counted under resource/teardown.
3. Batch adjacent box-drawing cells into row/run drawing
The deeper data says the problem is not just that Unicode causes short runs. It is that self-drawn boxpower emits many primitives per cell.
A better architecture would recognize adjacent line-art cells and emit longer horizontal/vertical strokes per row or per connected segment.
Example target:
───────
should be one line/rect draw, not N glyph renders with N pens/regions/clips.
This is probably more invasive than caching GDI resources, but likely has the highest ceiling.
4. Consider using font glyphs for common box drawing instead of selfdrawing
If acceptable visually, common Unicode box drawing could go through the normal text path rather than selfdraw. The current ExtTextOutW path is comparatively cheap.
This might be configurable or limited to common box-drawing characters used by TUIs:
U+2500..U+257F box drawing U+2580..U+259F block elements
The risk is alignment/appearance regressions. But the reward is large: ordinary ASCII text is healthy, and ordinary textout is not the bottleneck.
5. Keep scroll-copy lower in priority
Scroll-copy was a good idea from the first log, but this run says it will not fix the tig pain:
pre-tail scroll_visible_lines: 111,100 tail scroll_visible_lines: 8 slow >100ms scroll lines: 0
So scrolling optimization is still worthwhile for output-heavy workloads, but it will not address the line-art stall.
6. ASCII bidi_class() skip is still a baseline cleanup, not the main fix
There are still a lot of full-screen scans:
bidi_class_calls: 35.2M wcw_calls: 26.9M
But the expensive tail is not dominated by those counters. This is worth doing later for general polish.
Updated priority list
-
Cache/reuse GDI pens, brushes, and regions in
win_textselfdraw. -
Avoid per-cell clip region churn for boxpower glyphs.
-
Batch adjacent box-drawing cells into longer strokes/runs.
-
Experiment with rendering common box drawing via font glyphs instead of selfdraw.
-
Then revisit scroll-copy into the back buffer.
-
Then do ASCII scan/bidi fast paths and small-update cleanup.
The fix direction is now very specific: attack win_text_selfdraw_boxpower, especially its per-cell GDI resource lifecycle and primitive emission.
You said:
ChatGPT said:
I’ll inspect the self-drawn box path in the uploaded tree and prepare a small patch series focused first on removing per-glyph GDI resource churn.
You said:
ChatGPT said:
I’ll stack the changes on the current patched base, verify the exact patch application point, and keep any batching or font-path experiments conservative and separate.
I’ll start with localized resource reuse in linedraw, lines, trio, rectdraw, and clip setup, preserving the deep counters for before/after validation.
I’ve confirmed the current base already includes the earlier performance commits and both instrumentation layers, so the first optimization patch can focus purely on cached GDI resources.
Built a small optimization series on top of the existing instrumentation + deep instrumentation patches.
Files:
The series has two patches:
0001-wintext-cache-self-drawn-GDI-resources.patch 0002-wintext-avoid-dummy-text-allocation-for-self-drawn-g.patch
What it targets
This is aimed directly at the slow tig/Unicode line-art path from the deeper log: term_paint -> win_text -> win_text_selfdraw -> win_text_selfdraw_boxpower, where the log showed huge per-cell GDI churn: pens, brushes, regions, clip selection, and deletes.
Patch 1: cache self-drawn GDI resources
Adds a tiny selfdraw resource cache in wintext.c:
512 cached pens 256 cached brushes 1 reusable clip region
It replaces the hot selfdraw-path patterns like:
CCreatePen(...) CreateSolidBrush(...) CreateRectRgn(...) DeleteObject(...)
with cached helpers for:
boxpower common pens boxpower heavy/round pens line helper pens triangle/chord brushes rect fill brushes boxcoded/vt52 pens selfdraw clip region
The clip region is created once and then updated with SetRectRgn, so repeated CreateRectRgn / DeleteObject churn should mostly disappear.
Expected counter movement in the next log:
win_text_gdi_create_pen_calls sharply down win_text_gdi_create_brush_calls sharply down win_text_gdi_create_rgn_calls near 1-ish after warmup win_text_gdi_delete_object_calls sharply down win_text_selfdraw_resource_ticks_us down win_text_selfdraw_teardown_ticks_us down
SelectClipRgn calls will still exist; this patch removes region allocation churn, not clipping itself.
Patch 2: avoid dummy text allocation for self-drawn glyphs
The selfdraw path was allocating a dummy text = newn(wchar, len) buffer, filling it with X, then skipping normal text output anyway. This removes that allocation/free and just keeps origtext = text.
Expected counter movement:
win_text_origtext_alloc_calls 0 win_text_origtext_alloc_chars 0 win_text_origtext_alloc_ticks_us 0 win_text_origtext_free_ticks_us 0
Checked
I verified the patch sequence on a fresh unzip with:
Bashgit apply /mnt/data/mintty-perf-instrumentation.patch git apply /mnt/data/mintty-perf-wintext-deep-instrumentation.patch git apply --check /mnt/data/mintty-render-fix-series.patch git apply /mnt/data/mintty-render-fix-series.patch git diff --check
I did not do a real Cygwin/Win32 mintty build in this container; it only has Linux gcc, not the Windows/Cygwin headers/toolchain needed for a real compile.
What I deliberately did not include yet
I did not attempt the higher-risk architectural changes yet:
batch adjacent box-drawing cells into long strokes render box drawing via font glyphs instead of selfdraw skip clip selection for common boxpower cases rewrite dirty-row or scroll-copy behavior
Those could produce larger wins, but the two patches here attack the measured hot churn with a much smaller correctness surface. The next perf log should tell whether remaining time is mostly primitive drawing itself, clipping, or run fragmentation.
You said:
tig stuff is at the end again. not sure, but it might have felt a little faster after scrolling over the same regions. I guess it still needs more patchesChatGPT said:
I’ll compare this log to the previous deep run, verify the cache effects, then inspect the patched hot path to choose the next narrow fix.
The resource-cache patches landed on target: create/delete churn is no longer the bottleneck. The remaining hotspot is per-cell boxpower drawing, so the next fix must reduce primitives or bypass selfdraw.
Worst-frame time dropped from about 619ms to 391ms, and per-run background fills also remain visible. I’m focusing next on vector primitive emission, not further caching.
The two previous patches worked: they removed the GDI object churn almost completely. In the new slow cohort, CreatePen/CreateBrush/CreateRectRgn/DeleteObject are essentially gone, but the remaining cost is still win_text_selfdraw_boxpower: many tiny self-drawn box/line-art runs and millions of primitive operations. The new log has 2,596 update records; slow >100ms updates are 170 records and account for 39.3s, with 30.0s in selfdraw boxpower.
Files:
# mintty perf 42120 summary
Records: 2,599 total, 2,596 update records, 3 paint records.
## Cohorts
| cohort | updates | total ms | mean ms | p50 ms | p95 ms | max ms | win_text ms | selfdraw ms | boxpower ms | boxpower calls | line ops | pen creates | brush creates | region creates | deletes |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| all | 2596 | 70920.063 | 27.319 | 12.525 | 161.815 | 391.436 | 50150.495 | 35300.705 | 35265.713 | 1927932 | 2725074 | 889 | 19 | 1 | 844 |
| ASCII-only | 1550 | 14957.696 | 9.650 | 6.559 | 22.477 | 68.526 | 2860.254 | 0.000 | 0.000 | 0 | 0 | 815 | 0 | 0 | 815 |
| non-ASCII | 1046 | 55962.367 | 53.501 | 19.119 | 284.188 | 391.436 | 47290.241 | 35300.705 | 35265.713 | 1927932 | 2725074 | 74 | 19 | 1 | 29 |
| slow >100ms | 170 | 39332.538 | 231.368 | 228.434 | 350.201 | 391.436 | 37729.109 | 30020.676 | 29988.889 | 1754744 | 2246784 | 9 | 3 | 0 | 0 |
| tail update>=1900 | 700 | 44659.815 | 63.800 | 6.467 | 296.951 | 391.436 | 41430.266 | 31555.895 | 31525.042 | 1805876 | 2568748 | 20 | 3 | 0 | 8 |
## Conclusions
The previous cache patches worked: GDI object creation/deletion is no longer the hot issue. The remaining hot path is `win_text_selfdraw_boxpower`, with millions of small boxpower runs and millions of line/shape operations. The next patch series targets primitive count, clip/pen state churn, phase splitting, and ASCII baseline scanning.
What changed after the last patches
The good news:
slow >100ms cohort: CreatePen 1,800,309 -> 9 CreateBrush 600,103 -> 3 CreateRectRgn 600,103 -> 0 DeleteObject 3,000,515 -> 0 selfdraw resource setup: ~7.07s -> ~0.097s selfdraw teardown: ~3.13s -> ~0.023s
The remaining bad news:
slow >100ms cohort: total update time: 39.3s win_text: 37.7s win_text_selfdraw: 30.0s win_text_selfdraw_boxpower: 30.0s boxpower calls: 1,754,744 selfdraw line ops: 2,246,784
So the next series targets the remaining primitive/state overhead rather than object allocation.
New patch series
It applies on top of the existing instrumentation + deep instrumentation + the last two selfdraw resource patches.
0001-wintext-batch-repeated-horizontal-box-lines.patch 0002-wintext-draw-axis-aligned-box-strokes-as-fills.patch 0003-term-skip-bidi-lookup-for-merged-ASCII-pairs.patch 0004-wintext-select-box-drawing-pens-lazily.patch 0005-wintext-avoid-clip-regions-for-simple-self-drawn-box.patch 0006-term-avoid-overlay-splitting-for-self-drawn-graphics.patch
Expected counter movement
The main counters to watch in the next log:
win_text_selfdraw_line_ops win_text_selfdraw_boxpower_ticks_us win_text_gdi_select_object_calls win_text_gdi_select_clip_calls win_text_clip_set_calls win_text_clip_clear_calls win_text_phase1_calls win_text_phase2_calls bidi_class_calls
The current log’s object counters are already fixed, so I would not expect further movement there.
Patch notes
Patch 1: batch repeated horizontal box lines
This catches runs of repeated:
─ U+2500 ━ U+2501 ═ U+2550
and draws them as one or two filled rectangles for the whole run instead of one glyph at a time. This should help rows with long separators, table borders, and similar TUI output.
Patch 2: draw axis-aligned box strokes as fills
After the previous GDI-object cache, the remaining boxpower cost is mostly the drawing itself. For ordinary horizontal/vertical box strokes, LineTo plus geometric pen machinery is overkill. This patch draws axis-aligned strokes as FillRect rectangles while keeping diagonals/arcs on the existing path.
Patch 3: skip ASCII bidi lookup for merged ASCII pairs
This is the earlier baseline cleanup. It will not fix tig, but the new log still has 75M bidi_class calls. This avoids the Unicode range lookup for the common merged printable-ASCII case.
Patch 4: select box drawing pens lazily
After patch 2, many common box glyphs no longer need a selected pen at all. This avoids selecting/restoring a pen for fill-only selfdraw runs.
Patch 5: avoid clip regions for simple self-drawn boxes
The new log still has about 1.9M clip selections in the boxpower path. This skips clip-region setup for simple box/block characters that draw within the cell using fills. It keeps clipping for rounded corners, diagonals, and uncommon/non-safe paths.
Patch 6: avoid overlay splitting for self-drawn graphics
Self-drawn graphics are not font glyphs with italic/overhang behavior. This avoids the extra background/foreground overlay split for those runs, while keeping cursor-mark splitting intact.
Checked
I verified the series like this:
Bashgit apply /mnt/data/mintty-perf-instrumentation.patch git apply /mnt/data/mintty-perf-wintext-deep-instrumentation.patch git apply /mnt/data/mintty-render-fix-series.patch git apply --check /mnt/data/mintty-next-boxperf-series.patch git am /mnt/data/mintty-next-boxperf-patches/*.patch git diff --check
Still not a real Cygwin/Win32 build here; this container lacks the Windows/Cygwin headers/toolchain.