|
| 1 | +--- |
| 2 | +layout: single |
| 3 | +title: "Condense Tools Live Validation on tcad-scraper" |
| 4 | +date: 2026-02-25 |
| 5 | +author_profile: true |
| 6 | +categories: [ast-grep-mcp, code-condensation] |
| 7 | +tags: [mcp, condense, typescript, ast-grep, token-reduction, live-testing, bugfix] |
| 8 | +excerpt: "Live validation of all 6 condense MCP tools against the tcad-scraper codebase (219 files, 1.1MB) — uncovered and fixed a per-language stats serialization gap." |
| 9 | +header: |
| 10 | + image: /assets/images/cover-reports.png |
| 11 | + teaser: /assets/images/cover-reports.png |
| 12 | +--- |
| 13 | + |
| 14 | +**Session Date**: 2026-02-25<br> |
| 15 | +**Project**: ast-grep-mcp<br> |
| 16 | +**Focus**: End-to-end live validation of condense feature tools<br> |
| 17 | +**Session Type**: Validation | Bugfix |
| 18 | + |
| 19 | +## Executive Summary |
| 20 | + |
| 21 | +Ran all 6 condense MCP tools against the production tcad-scraper codebase (219 files, 1.1MB TypeScript/JS/Python). Every tool executed successfully with zero errors. During validation, discovered and fixed a serialization gap where `per_language_stats` in `condense_pack` output was missing byte-level metrics — only line counts were emitted. After the fix, all 117 condense tests continue to pass. |
| 22 | + |
| 23 | +The `ai_chat` strategy achieved 15.7% actual reduction (43,827 tokens saved) against a theoretical 85% estimate, confirming the JS/TS surface extractor's brace-matching heuristic retains more code than the theoretical model assumes — documented as a future improvement area. |
| 24 | + |
| 25 | +## Key Metrics |
| 26 | + |
| 27 | +| Metric | Value | |
| 28 | +|--------|-------| |
| 29 | +| Tools validated | 6 / 6 | |
| 30 | +| Target codebase | tcad-scraper (219 files, 1.1 MB) | |
| 31 | +| Errors encountered | 0 | |
| 32 | +| Bug found and fixed | 1 (per-language byte stats) | |
| 33 | +| Tests passing | 117 / 117 | |
| 34 | +| Best actual reduction | 15.7% (ai_chat) | |
| 35 | +| Tokens saved (ai_chat) | 43,827 | |
| 36 | + |
| 37 | +## Tool-by-Tool Results |
| 38 | + |
| 39 | +### 1. condense_normalize (batch, 212 TS files) |
| 40 | + |
| 41 | +Processed in 5 batches of 50. Completed in 0.4s. |
| 42 | + |
| 43 | +| Metric | Value | |
| 44 | +|--------|-------| |
| 45 | +| Files processed | 212 | |
| 46 | +| Normalizations applied | 11,494 | |
| 47 | +| Files with changes | 210 / 212 | |
| 48 | +| Byte delta | +1,774 (0.16% expansion) | |
| 49 | + |
| 50 | +Quote canonicalization was the dominant transform. Net byte expansion is expected — normalization targets compression consistency, not direct size reduction. Top file: `continuous-batch-scraper.ts` with 1,159 normalizations. |
| 51 | + |
| 52 | +### 2. condense_strip (batch, 219 files) |
| 53 | + |
| 54 | +| Metric | Value | |
| 55 | +|--------|-------| |
| 56 | +| Files processed | 219 | |
| 57 | +| Lines removed | 164 | |
| 58 | +| Line reduction | 0.40% | |
| 59 | +| Files with removals | 20 / 219 | |
| 60 | +| Elapsed | 0.6s | |
| 61 | + |
| 62 | +Removed `console.log`, `debugger`, `print()`, and `pdb.set_trace` statements. Top file: `setup-test-db.ts` (37 lines removed). Codebase is relatively clean — only 0.4% dead code. |
| 63 | + |
| 64 | +### 3. condense_extract_surface (212 TS files) |
| 65 | + |
| 66 | +| Metric | Value | |
| 67 | +|--------|-------| |
| 68 | +| Files processed | 212 | |
| 69 | +| Condensed lines | 33,938 | |
| 70 | +| Reduction | 15.0% | |
| 71 | +| Output size | 946,780 chars | |
| 72 | +| Elapsed | 22.5s | |
| 73 | + |
| 74 | +Kept only `export` declarations with brace-matched blocks. Test files with `describe`/`it` (no `export` prefix) fall back to keeping all lines, limiting reduction. |
| 75 | + |
| 76 | +### 4. condense_pack (all 4 strategies) |
| 77 | + |
| 78 | +| Strategy | Condensed | Reduction | Tokens (est) | Time | |
| 79 | +|----------|-----------|-----------|-------------|------| |
| 80 | +| ai_chat | 938,923 B | 15.7% | 234,730 | 10.6s | |
| 81 | +| ai_analysis | 1,102,420 B | 1.1% | 275,605 | 19.2s | |
| 82 | +| archival | 1,102,420 B | 1.1% | 275,605 | 14.9s | |
| 83 | +| polyglot | 938,923 B | 15.7% | 234,730 | 15.5s | |
| 84 | + |
| 85 | +`ai_chat` and `polyglot` produce identical output (all files are code, no config/text routing divergence). `ai_analysis` and `archival` are identical (both lossless, normalize+strip only). |
| 86 | + |
| 87 | +Per-language breakdown (ai_chat): |
| 88 | + |
| 89 | +| Language | Files | Reduction | |
| 90 | +|----------|-------|-----------| |
| 91 | +| TypeScript | 212 | 15.6% | |
| 92 | +| JavaScript | 4 | 26.4% | |
| 93 | +| Python | 3 | 23.5% | |
| 94 | + |
| 95 | +### 5. condense_estimate |
| 96 | + |
| 97 | +| Strategy | Est. Bytes | Est. Tokens | Theoretical Reduction | |
| 98 | +|----------|-----------|-------------|----------------------| |
| 99 | +| ai_chat | 167,134 | 41,783 | ~85% | |
| 100 | +| ai_analysis | 668,538 | 167,134 | ~40% | |
| 101 | +| archival | 779,961 | 194,990 | ~30% | |
| 102 | +| polyglot | 389,980 | 97,495 | ~65% | |
| 103 | + |
| 104 | +Top reduction candidates: `continuous-batch-scraper.ts` (1,769 lines, 4.3% of codebase). |
| 105 | + |
| 106 | +### 6. condense_normalize on ~/reports/ (32 files) |
| 107 | + |
| 108 | +Also validated against the reports site (JS, Python, CSS files): |
| 109 | + |
| 110 | +| Metric | Value | |
| 111 | +|--------|-------| |
| 112 | +| Files processed | 32 | |
| 113 | +| Normalizations applied | 59 | |
| 114 | +| Files with changes | 16 / 32 | |
| 115 | +| Byte reduction | 75 (0.02%) | |
| 116 | + |
| 117 | +## Bug Found and Fixed |
| 118 | + |
| 119 | +**Problem**: `per_language_stats` in `condense_pack` output only serialized `files_processed`, `original_lines`, `condensed_lines` — missing byte-level metrics entirely. This caused per-language stats to appear as all zeros when accessing `original_bytes`/`condensed_bytes` keys. |
| 120 | + |
| 121 | +**Root cause**: `LanguageCondenseStats` dataclass had no byte fields, and `condense_pack_impl` only aggregated line counts per language. |
| 122 | + |
| 123 | +**Fix** (2 files): |
| 124 | + |
| 125 | +`src/ast_grep_mcp/models/condense.py:11-12` — Added fields: |
| 126 | +```python |
| 127 | +original_bytes: int = 0 |
| 128 | +condensed_bytes: int = 0 |
| 129 | +``` |
| 130 | + |
| 131 | +`src/ast_grep_mcp/features/condense/service.py:399-400` — Aggregate bytes: |
| 132 | +```python |
| 133 | +stats.original_bytes += file_result["original_bytes"] |
| 134 | +stats.condensed_bytes += file_result["condensed_bytes"] |
| 135 | +``` |
| 136 | + |
| 137 | +`src/ast_grep_mcp/features/condense/service.py:431-436` — Serialize with computed reduction: |
| 138 | +```python |
| 139 | +"original_bytes": s.original_bytes, |
| 140 | +"condensed_bytes": s.condensed_bytes, |
| 141 | +"reduction_pct": round((1.0 - s.condensed_bytes / s.original_bytes) * 100, 1) |
| 142 | +``` |
| 143 | + |
| 144 | +All 117 condense tests pass after the fix. |
| 145 | + |
| 146 | +## Estimate vs Actual Gap |
| 147 | + |
| 148 | +| Strategy | Estimated Reduction | Actual Reduction | Gap | |
| 149 | +|----------|-------------------|-----------------|-----| |
| 150 | +| ai_chat | ~85% | 15.7% | 69.3pp | |
| 151 | +| ai_analysis | ~40% | 1.1% | 38.9pp | |
| 152 | + |
| 153 | +The estimator uses theoretical `STRATEGY_REDUCTION_RATIOS` constants. The actual JS/TS surface extractor keeps entire brace-matched export blocks (including function bodies), and test files with `describe`/`it` fall back to keeping everything. This is the primary improvement target for the next phase. |
| 154 | + |
| 155 | +## Files Modified |
| 156 | + |
| 157 | +| File | Change | |
| 158 | +|------|--------| |
| 159 | +| `src/ast_grep_mcp/models/condense.py:11-12` | Added `original_bytes`, `condensed_bytes` fields | |
| 160 | +| `src/ast_grep_mcp/features/condense/service.py:399-400` | Aggregate byte counts per language | |
| 161 | +| `src/ast_grep_mcp/features/condense/service.py:429-436` | Serialize byte metrics + reduction_pct | |
| 162 | + |
| 163 | +## Git Context |
| 164 | + |
| 165 | +``` |
| 166 | +d97d782 refactor(condense): remove unused CondenseDefaults constants and standardize field naming |
| 167 | +1ffc15b feat(condense): implement P9 — condense_train_dictionary tool (zstd) |
| 168 | +9a09893 fix(condense): address critical/high code review findings |
| 169 | +``` |
| 170 | + |
| 171 | +### 6. condense_train_dictionary (TypeScript) |
| 172 | + |
| 173 | +| Metric | Value | |
| 174 | +|--------|-------| |
| 175 | +| Dictionary path | `.condense/dictionaries/dict_typescript.zdict` | |
| 176 | +| Dictionary size | 112,640 B (110 KB) | |
| 177 | +| Samples used | 200 | |
| 178 | +| Total sample bytes | 999,080 B (~1 MB) | |
| 179 | +| Est. compression improvement | 15.0% | |
| 180 | +| Elapsed | 13.2s | |
| 181 | + |
| 182 | +Trained a zstd dictionary on 200 TypeScript files from tcad-scraper, written to `tcad-scraper/.condense/dictionaries/dict_typescript.zdict`. The dictionary captures repeated cross-file patterns (import paths, type annotations, test boilerplate) that standard zstd cannot exploit. Usage: |
| 183 | + |
| 184 | +```bash |
| 185 | +zstd -D .condense/dictionaries/dict_typescript.zdict <file> |
| 186 | +``` |
| 187 | + |
| 188 | +The 15% estimated improvement applies on top of standard zstd compression ratios — most effective for small-to-medium files (<100KB) with consistent coding patterns across the codebase. |
| 189 | + |
| 190 | +## References |
| 191 | + |
| 192 | +- [CLAUDE.md](/Users/alyshialedlie/code/ast-grep-mcp/CLAUDE.md) — Project quick start |
| 193 | +- [docs/BACKLOG.md](/Users/alyshialedlie/code/ast-grep-mcp/docs/BACKLOG.md) — Remaining work items |
| 194 | +- [docs/CODE-CONDENSE-PHASE-2.md](/Users/alyshialedlie/code/ast-grep-mcp/docs/CODE-CONDENSE-PHASE-2.md) — Phase 2 plan |
0 commit comments