Skip to content

Commit d12f29e

Browse files
bai-uipathclaude
andcommitted
docs(evalboard): trim the comments added by this branch
Comment-only. Cuts the narration of what the code used to do and the restatements of what it now does, keeping the reasons that aren't recoverable from reading it: why the per-run cache is keyed per run rather than per source, why the ad-hoc load key and the sort key differ, why hover moved to a marker class, and why the model vote normalizes. Net 57 lines. No behavior change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 3aa3f0c commit d12f29e

12 files changed

Lines changed: 91 additions & 148 deletions

File tree

‎evalboard/app/_components/harness-selector.tsx‎

Lines changed: 7 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -42,14 +42,13 @@ export function HarnessSelector({
4242
const qs = p.toString();
4343
router.replace(qs ? `${pathname}?${qs}` : pathname, { scroll: false });
4444
};
45-
// Every known harness always gets a segment, whether or not it turned up in
46-
// the discovery window. A weekly harness (delegate) drops out of that window
47-
// between firings, and a control that quietly loses an option reads as "this
48-
// harness was removed" rather than "it hasn't run lately". Scoping to one
49-
// with no runs in the window shows an empty result, which is honest and one
50-
// click from recoverable; a missing segment is neither. `current` is unioned
51-
// in too, so a deep-linked `?h=` outside the known set still reads as
52-
// selected. orderHarnesses fixes the order, so segments never reshuffle.
45+
// Every known harness gets a segment whether or not it turned up in the
46+
// discovery window. A weekly harness (delegate) drops out of that window
47+
// between firings, and a control that quietly loses an option reads as
48+
// "removed" rather than "hasn't run lately"; an empty result is honest and
49+
// one click from recoverable. `current` is unioned in so a deep-linked `?h=`
50+
// outside the known set still reads as selected, and orderHarnesses fixes
51+
// the order so segments never reshuffle.
5352
const opts = orderHarnesses([
5453
...KNOWN_HARNESSES,
5554
...harnesses,

‎evalboard/app/api/download/route.ts‎

Lines changed: 3 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -14,11 +14,9 @@ export const dynamic = "force-dynamic";
1414
// blobs first, so this mirrors what the page would load.
1515
//
1616
// `task` is REQUIRED. Omitting it used to mean "zip the whole run", which for a
17-
// nightly meant ~10k blobs / ~400 MB: an uncapped fan-out of blob downloads, a
18-
// stat-per-file walk over Azure Files, and the entire archive buffered in
19-
// memory before any of it was sent. That path is gone, along with the run
20-
// page's button for it — a task folder is a handful of files and returns in
21-
// well under a second. To inspect a whole run, use the blob container directly.
17+
// nightly meant ~10k blobs / ~400 MB fetched uncapped and buffered in memory
18+
// before any of it was sent. That path and its button are gone; to inspect a
19+
// whole run, use the blob container.
2220
export async function GET(req: Request) {
2321
const url = new URL(req.url);
2422
const runId = url.searchParams.get("run");

‎evalboard/app/api/refresh/__tests__/route.test.ts‎

Lines changed: 2 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -128,10 +128,8 @@ describe("POST /api/refresh", () => {
128128
await expect(fs.access(runDir)).rejects.toThrow();
129129
});
130130

131-
// Evicting only the on-disk copy is not enough: the front page reads a
132-
// memoized projection of run.json that a settled run holds for a day
133-
// (lib/overview.ts::perRunRevalidate), so without this the refresh
134-
// button is a no-op for anything older than 24h.
131+
// Without this the button is a no-op for a settled run, whose
132+
// projection is held for a day (lib/overview.ts::perRunRevalidate).
135133
test("valid run -> the memoized projection is invalidated too", async () => {
136134
const POST = await loadPost();
137135
await POST(post("2026-06-01_04-04-22"));

‎evalboard/app/api/refresh/route.ts‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -40,9 +40,8 @@ export async function POST(req: Request) {
4040
return NextResponse.json({ error: "invalid run" }, { status: 400 });
4141
}
4242
// Evicting the on-disk copy is only half the job: the front page reads a
43-
// memoized PROJECTION of it (lib/overview.ts::cachedLoadPerRunFor), and a
44-
// settled run's projection is held for a day. Drop that entry too, or the
45-
// refresh silently does nothing for every run older than 24h.
43+
// memoized projection of it (lib/overview.ts::cachedLoadPerRunFor), held for
44+
// a day on a settled run. Without this the refresh would no-op for those.
4645
revalidateTag(runCacheTag(source.id, runId));
4746
// Re-download is lazy on next render. The run page is force-dynamic and
4847
// reads the sidecars straight from disk, so it is fresh immediately.

‎evalboard/app/globals.css‎

Lines changed: 10 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -42,11 +42,9 @@ details[open] > summary .group-chevron {
4242
}
4343

4444
/* Filter/tag chips (app/runs/[id]/chips.tsx). The composed utility string was
45-
~130 chars and rendered ~11k times on a nightly run page (and ~6k on
46-
/trends): 1.4 MB of identical class text in one response, from 120 distinct
47-
values across 56k attributes. The utilities live here now so the emitted
48-
attribute is four short tokens. The component still decides WHICH tokens to
49-
emit; the variant/state matrix is carried by compound selectors below. */
45+
~130 chars and rendered ~11k times on a nightly run page: 1.4 MB of identical
46+
class text in one response. Emitting four short tokens instead moves it here,
47+
where the browser parses it once and caches it across navigations. */
5048
@layer components {
5149
.chip {
5250
@apply rounded border transition-colors;
@@ -57,8 +55,8 @@ details[open] > summary .group-chevron {
5755
.chip-md {
5856
@apply text-xs px-2 py-0.5;
5957
}
60-
/* Idle. Exactly one of idle/selected is ever emitted, so these never
61-
collide and neither needs to out-specify the other. */
58+
/* Idle. Exactly one of idle/selected is ever emitted, so neither needs to
59+
out-specify the other. */
6260
.chip-skill {
6361
@apply bg-indigo-50 text-indigo-700 border-indigo-200 font-medium;
6462
}
@@ -69,8 +67,7 @@ details[open] > summary .group-chevron {
6967
@apply bg-gray-50 text-gray-500 border-gray-200;
7068
}
7169
/* Hover rides on .chip-act, which only the interactive (button) branch
72-
emits — a non-clickable span must not show an affordance it can't
73-
honor. Replaces the old runtime hover-utility strip. */
70+
emits, so a non-clickable span shows no affordance it can't honor. */
7471
.chip-act.chip-skill:hover {
7572
@apply bg-indigo-100;
7673
}
@@ -96,17 +93,16 @@ details[open] > summary .group-chevron {
9693
@apply bg-studio-blue text-white border-studio-blue;
9794
}
9895

99-
/* Per-row stat key in the grid's stacked card layout — 5.2k renders on a
100-
nightly run page at 49 chars each. */
96+
/* Per-row stat key in the grid's stacked card layout; 5.2k renders. */
10197
.stat-k {
10298
@apply text-[10px] uppercase tracking-wide text-gray-400;
10399
}
104-
/* Numeric table cell on /trends — 5.4k renders at 47 chars each. */
100+
/* Numeric table cell on /trends; 5.4k renders. */
105101
.num-cell {
106102
@apply py-2 px-3 tabular-nums text-right text-gray-700;
107103
}
108-
/* One bar of a /trends sparkline — 9.9k renders. The status color stays a
109-
utility on the element; only the geometry is shared. */
104+
/* One bar of a /trends sparkline; 9.9k renders. The status color stays a
105+
utility on the element, so only the geometry is shared. */
110106
.spark-bar {
111107
@apply w-[6px] h-full rounded-sm;
112108
}

‎evalboard/app/runs/[id]/chips.tsx‎

Lines changed: 5 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -9,11 +9,9 @@ export type ChipVariant = "skill" | "review" | "tag";
99
export type ChipSize = "sm" | "md";
1010

1111
// The colour/size utilities behind these tokens live in app/globals.css under
12-
// `@layer components`. A chip renders ~11k times on a nightly run page, so the
13-
// composed utility string (~130 chars each) was the single largest contributor
14-
// to that page's HTML — 1.4 MB of identical class text. Emitting short tokens
15-
// instead moves that text into the stylesheet, where the browser parses it once
16-
// and caches it across navigations.
12+
// `@layer components`. A chip renders ~11k times on a nightly run page, and the
13+
// composed utility string (~130 chars each) was 1.4 MB of identical class text
14+
// in one response; short tokens move it into the cached stylesheet.
1715
const STYLES: Record<
1816
ChipVariant,
1917
{
@@ -76,9 +74,8 @@ export function ChipButton({
7674
const s = STYLES[variant];
7775
const activeCls =
7876
variant === "tag" && size === "md" ? TAG_MD_ACTIVE : s.active;
79-
// `chip-act` arms the hover rules, and only the interactive branch gets it —
80-
// otherwise the non-interactive <span> shows a hover affordance it can't
81-
// honor. (This replaces stripping `hover:` utilities out of the string.)
77+
// `chip-act` arms the hover rules, and only the interactive branch gets it:
78+
// a non-clickable <span> must not show an affordance it can't honor.
8279
const idleCls = onClick ? `${s.idle} chip-act` : s.idle;
8380
const stateCls = active ? activeCls : idleCls;
8481
const baseCls = `chip ${SIZE_CLS[size]} ${stateCls}`;

‎evalboard/app/runs/[id]/run-view.tsx‎

Lines changed: 4 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -296,14 +296,10 @@ export function RunView({
296296

297297
// Commit a filter change to the URL WITHOUT a server round-trip.
298298
//
299-
// Every reader of `tags` / `rtags` / `q` on this page is client-side — the
300-
// grid is filtered by `filtered` below, out of rows the server already
301-
// sent. The page's server component reads only `src`. So `router.replace`
302-
// bought nothing and cost a full re-render of a force-dynamic route: two
303-
// separate parses of the same multi-MB run.json (readRunSummary and
304-
// readRunTasks each read it), the activation sub-run, the review index, and
305-
// the mature-source scan's serial walk back through earlier runs — all to
306-
// return markup this component recomputes locally anyway.
299+
// Every reader of `tags` / `rtags` / `q` on this page is client-side (see
300+
// `filtered` below); the server component reads only `src`. So
301+
// `router.replace` bought nothing and cost a full re-render of a
302+
// force-dynamic route to return markup this component recomputes anyway.
307303
//
308304
// The native History API keeps the URL shareable and the back button
309305
// working; Next syncs `useSearchParams` off pushState/replaceState, so this

‎evalboard/lib/blob.ts‎

Lines changed: 3 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -297,11 +297,9 @@ export async function ensureRunReviewIndex(
297297
});
298298
}
299299

300-
// NOTE: there is deliberately no whole-run fetch. `ensureRunDir` used to exist
301-
// for the run page's download-as-zip button and pulled EVERY blob under the run
302-
// prefix with no concurrency cap — ~10k blobs / ~400 MB for a nightly, issued as
303-
// one unbounded Promise.all. Both it and the button are gone; per-task download
304-
// (below) is the supported shape.
300+
// There is deliberately no whole-run fetch. `ensureRunDir` used to pull EVERY
301+
// blob under the run prefix with no concurrency cap (~10k blobs / ~400 MB for a
302+
// nightly). Per-task download, below, is the supported shape.
305303

306304
// Narrow fetch: run.json + just one task subdir. Used by the per-task
307305
// detail page so opening a deep link to a 50-task run doesn't pull every

‎evalboard/lib/overview.ts‎

Lines changed: 29 additions & 52 deletions
Original file line numberDiff line numberDiff line change
@@ -325,14 +325,13 @@ async function loadPerRunForId(
325325
}
326326

327327
// First `YYYY-MM-DD` (optionally `_HH-MM-SS`) anywhere in a run id. Ad-hoc ids
328-
// are not date-SHAPED — `parseRunIdDate` anchors, so it rejects them — but
329-
// almost all of them still CARRY their date: `adhoc-2026-09-02_21-59-13`,
330-
// `2026-05-28_skills-full-codex-gpt54`, `haiku45-skills-suite-2026-05-28`.
328+
// aren't date-SHAPED (parseRunIdDate anchors, so it rejects them) but almost all
329+
// still carry their date: `adhoc-2026-09-02_21-59-13`, `skills-2026-05-28`.
331330
const ADHOC_DATE_RE = /(\d{4})-(\d{2})-(\d{2})(?:_(\d{2})-(\d{2})-(\d{2}))?/;
332331

333-
// Cheap, id-only start date for an ad-hoc run; null when the id carries no date.
334-
// A date-only id resolves to the start of its day, which is right for "which
335-
// runs are newest" and can only lose a same-day tie-break. Exported for testing.
332+
// Id-only start date for an ad-hoc run; null when the id carries no date. A
333+
// date-only id resolves to the start of its day, which can only lose a same-day
334+
// tie-break.
336335
export function adhocRunDate(id: string): Date | null {
337336
const m = ADHOC_DATE_RE.exec(id);
338337
if (!m) return null;
@@ -362,43 +361,33 @@ export function adhocRunDate(id: string): Date | null {
362361
// `YYYY-MM-DD_HH-MM-SS`, so a source-blind key would let a Scribe run and a
363362
// skills run with the same id serve each other's projection.
364363
//
365-
// Per-RUN rather than one loader per source, because `unstable_cache` fixes its
366-
// revalidate and tags at construction: a shared loader can only hold one TTL
367-
// for every run, and cannot carry a per-run tag for the refresh button to
368-
// invalidate. Both of those matter — see `perRunRevalidate` below.
364+
// Per-run rather than per-source because `unstable_cache` fixes revalidate and
365+
// tags at construction: one shared loader can hold neither a per-run TTL nor a
366+
// per-run tag for the refresh button to evict.
369367
type PerRunLoader = (id: string) => Promise<PerRun>;
370368

371-
// Keyed `<source>:<run>`. Bounded by the run count of every container the
372-
// process has served, so it grows with history rather than with traffic.
369+
// Keyed `<source>:<run>`; grows with history, not with traffic.
373370
const perRunLoaders = new Map<string, () => Promise<PerRun>>();
374371

375-
// A run's run.json is written once, at the end of the run (the upload is the
376-
// last thing the runner does), so a run that finished yesterday is immutable.
377-
// Its SIDECARS are not — meta.json (title/description), reviews and analysis can
378-
// be edited in blob afterwards — which is what `runCacheTag` is for.
372+
// A run.json is written once, at the end of the run, so a finished run is
373+
// immutable. Its sidecars (meta.json, reviews, analysis) are not, which is what
374+
// `runCacheTag` is for.
379375
const SETTLED_AFTER_MS = 24 * 60 * 60 * 1000;
380-
// A run whose id says it is still recent. Re-read often, because it may be
381-
// today's nightly landing while the page is open.
376+
// Recent enough that today's nightly may still be landing.
382377
const FRESH_REVALIDATE_SECONDS = 300;
383-
// Everything older. The read this avoids is a multi-MB run.json off an Azure
384-
// Files mount, and its content cannot change on its own.
378+
// Older than that: the read this avoids is a multi-MB run.json off Azure Files.
385379
const SETTLED_REVALIDATE_SECONDS = 24 * 60 * 60;
386380

387-
// Cache tag for one run's projection, so POST /api/refresh can evict it. Without
388-
// this, a settled run edited in blob would keep serving a stale projection to
389-
// the front page for a day.
381+
// Cache tag for one run's projection, so POST /api/refresh can evict it;
382+
// without it the button would no-op for anything past SETTLED_AFTER_MS.
390383
export function runCacheTag(sourceId: string, runId: string): string {
391384
return `evalboard-run:${sourceId}:${runId}`;
392385
}
393386

394-
// Settled runs are cached for a day, everything else for 5 minutes. An id that
395-
// carries no date at all is treated as fresh: those are hand-uploaded ad-hoc
396-
// runs, re-uploaded far more often than pipeline runs, and there are a handful.
397-
//
398-
// Evaluated once per run per process, when the loader is first built, so a run
399-
// that settles while the process is up keeps the 5-minute TTL until the next
400-
// restart. That is the safe direction (it is today's behavior) and not worth a
401-
// timer to correct.
387+
// Settled runs cache for a day, everything else for 5 minutes. An id carrying no
388+
// date at all is treated as fresh: those are hand-uploaded and re-uploaded far
389+
// more often than pipeline runs. Evaluated once per run per process, so a run
390+
// that settles while the process is up keeps the short TTL until restart.
402391
function perRunRevalidate(id: string): number {
403392
const started = parseRunIdDate(id) ?? adhocRunDate(id);
404393
if (started == null) return FRESH_REVALIDATE_SECONDS;
@@ -1234,28 +1223,19 @@ export function buildAdhocRows(
12341223
};
12351224
}
12361225

1237-
// Extra candidates loaded beyond `limit`, to cover ids that turn out to have no
1238-
// readable overview (aborted uploads, and the `deploys/` prefix, which is not a
1239-
// run at all) and so get dropped from the rows.
1226+
// Extra candidates loaded beyond `limit`, covering ids with no readable overview
1227+
// (aborted uploads, the `deploys/` prefix) that then drop out of the rows.
12401228
const ADHOC_LOAD_SLACK = 10;
12411229

12421230
// The Ad-hoc runs section (front page, below the daily listing). "Ad-hoc" here
12431231
// means "not a daily-pipeline run" — i.e. the id isn't date-shaped, which is
12441232
// exactly the set listRunIdsInWindow excludes from the chart and main table.
12451233
//
12461234
// Only the newest `limit + ADHOC_LOAD_SLACK` candidates are loaded, ordered by
1247-
// the date in the id. This section used to load EVERY ad-hoc candidate before
1248-
// sorting, on the premise that the set was "small by construction (manual
1249-
// uploads only)". That premise expired: nothing expires the ad-hoc prefixes in
1250-
// the `runs` container (unlike `runs-gha`'s 14-day rule), so the set had grown
1251-
// to 165 runs / ~294 MB of run.json read on every cold front-page render, to
1252-
// show ten rows.
1253-
//
1254-
// The authoritative sort key stays run.json's `start_time` — the id key only
1255-
// decides what to LOAD, and the two can only disagree for a run whose id date
1256-
// contradicts its own start_time. Ids carrying no date at all are always loaded
1257-
// (there are a handful, e.g. `sdk-live-r2-final`), so they can never be ordered
1258-
// out of the section by a key they don't have.
1235+
// the date in the id; loading all 165 to show ten rows cost ~294 MB per cold
1236+
// render. run.json's `start_time` stays the authoritative sort, so the id only
1237+
// decides what to READ. Ids with no date are always loaded, so they can never be
1238+
// ordered out by a key they don't have.
12591239
export async function getAdhocRunListing(
12601240
limit: number | null,
12611241
source: Source = DEFAULT_SOURCE,
@@ -1271,8 +1251,7 @@ export async function getAdhocRunListing(
12711251
else dated.push({ id, at: at.getTime() });
12721252
}
12731253
dated.sort((a, b) => b.at - a.at);
1274-
// null limit = "load everything", which the page never asks for but the
1275-
// signature allows; a finite limit takes the newest slice plus slack.
1254+
// null limit = "load everything": allowed by the signature, never used.
12761255
const budget = limit == null ? dated.length : limit + ADHOC_LOAD_SLACK;
12771256
const loadedAll = budget >= dated.length;
12781257
const toLoad = [...undated, ...dated.slice(0, budget).map((d) => d.id)];
@@ -1282,9 +1261,7 @@ export async function getAdhocRunListing(
12821261
cachedLoadPerRunFor(source),
12831262
);
12841263
const listing = buildAdhocRows(perRun, limit);
1285-
// `total` drives the "Show more" affordance, so while the load is truncated
1286-
// it has to report the candidate count rather than what we happened to read
1287-
// — otherwise the section caps itself at the first page. Once everything is
1288-
// loaded it reverts to the exact readable-row count.
1264+
// `total` drives "Show more", so a truncated load must report the candidate
1265+
// count or the section caps itself at the first page.
12891266
return loadedAll ? listing : { ...listing, total: ids.length };
12901267
}

‎evalboard/lib/pricing.ts‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -128,9 +128,8 @@ const _ROUTING_PREFIXES = [
128128
"openrouter/",
129129
];
130130
const _REGION_PREFIXES = ["eu.", "us.", "apac.", "global."];
131-
// Exported so the run header's model tally can group on the same key pricing
132-
// looks up on. Without it a single row that recorded the qualified id counts
133-
// as a second model and the header claims a mixed-model run.
131+
// Exported so the run header's model tally groups on the same key pricing looks
132+
// up on (lib/runs.ts::tallyModels).
134133
export function normalizeModel(model: string): string {
135134
let m = model.trim();
136135
for (const pre of _ROUTING_PREFIXES) {

0 commit comments

Comments
 (0)