Skip to content

R5 options: six candidate headline metrics, measured -- and 28 of 1,194 events validate against event_v1 #64

Description

@PipFoweraker

R5 was bounced back with instructions: do not retire the headline number, bring alternatives. "I believe it is a headline metric in pdoom. So we replace number of records with a number of record sets, or lists, or find better ways... just propose some options for me."

Six options below, no recommendation — that was the instruction. Every number is measured today, so this is a choice between real figures rather than shapes.

The number being replaced

1,194 events. It appears in the README, in manifests, in a badge, and in cross-repo briefs.

One measurement, taken while costing these options, is worth having before you choose:

28 of 1,194 timeline events validate against event_v1 — the schema they are published under. The other 1,166 fail it.

So the current headline is not merely flattering. By this repo's own published contract, it is wrong by a factor of forty-two, and nothing was checking. That is a separate finding and it is filed as its own issue; it is here because it changes what "replacing the metric" means. Any option below is more accurate than the incumbent, including the ones that look unflattering.


The six

A — Record sets · 4

candidates, frontier_labs, reviewed, timeline_events.

Your own suggestion. Counts breadth of coverage rather than volume.

Over-states: a collection of 46 records counts the same as one of 3,434. Four is also a number that grows by adding collections, which is a new way to game it.
Under-states: everything that happens inside a collection. Two years of deepening one dataset would never move it.

B — Attributed human judgements · 518 ruled on, 337 accepted

Records a named person has ruled on, across 1 reviewer and 3 passes.

Ties the headline directly to the constraint you named — attention, not size.

Over-states: a rough scan verdict counts identically to a considered one, and today's stability measurement was 62%. It also counts unsure and reject as "ruled on", which is honest but reads as bigger than the useful subset.
Under-states: perfectly good unreviewed data reads as zero. A large, well-sourced, untouched corpus looks like no work at all.

C — Maturity distribution · 1 gold, 3 wood

The ladder you ruled this morning, reported as a distribution rather than a count.

Grows only when something genuinely improves, which is the incentive you said you wanted something to aim at.

Over-states: nothing much — it is the strictest option here.
Under-states: volume entirely, and effort that does not cross a rung boundary is invisible. Moving candidates from wood to bronze is one schema file; moving it to gold is months. The metric cannot tell those apart.

D — Reaches a consumer · 20 of 1,194 · 1.7%

Events that actually open a decision window in pdoom1. The rest are demoted to a feed line.

The most externally honest number available: it counts what a player could encounter.

Over-states: nothing.
Under-states: it is hostage to a consumer's configuration, which this repo does not control. pdoom1 changing one demotion rule would move it by a thousand without anything here changing. Publishing it makes another team's setting into our headline.

E — Evidence-backed claims · 43 of 46 labs · 3,428 of 3,434 candidates carry a source

Claims carrying a verbatim source quote or a source URL.

Measures the thing this repo is actually differentiated on, and it is the number a researcher would ask for.

Over-states: a source URL is a much weaker guarantee than a verbatim quote, and this conflates them. The 3,428 figure is nearly the whole corpus and so barely discriminates.
Under-states: curation and selection effort, which is most of the work.

F — Schema-valid records · 28 of 1,194

Records that validate against the schema they are published under.

The honest version of the incumbent — same shape, same place in a sentence, correct.

Over-states: nothing.
Under-states: it will read as catastrophic to anyone who saw the old badge, and the fall from 1,194 to 28 needs explaining every time it is quoted. It also goes to zero for any collection without a schema, which is currently three of four.


Two combinations, since a headline can be a pair

C + B"1 collection at gold, 3 at wood · 518 records ruled on by a named human." Quality and attention, no volume. Neither number can be inflated by ingesting more.

F + D"28 schema-valid events, 20 of which reach a player." Brutal, externally checkable, and the two numbers being close is itself the argument that the corpus is not the bottleneck.

What is cheap and what is not

A, B, D and F are computable today from files already in the repo. C is computable today via check_maturity.py, which was built this morning. E needs a decision about whether a URL and a verbatim quote count as the same thing — they should not, and splitting them is half a day.

Whichever you pick, it should be emitted by a script rather than typed into a README, or it becomes the next stale copy. That is the same failure that put a false printer fact into the registry this week.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions