Skip to content

quotes: 30 verified difficulty candidates, 8 servable today - #101

Merged
PipFoweraker merged 6 commits into
mainfrom
feat/quote-mining-difficulty
Aug 30, 2026
Merged

PipFoweraker merged 6 commits into
mainfrom
feat/quote-mining-difficulty

Conversation

@PipFoweraker

Copy link
Copy Markdown
Owner

The corpus, filled. 30 quotes: 8 servable on licence alone with nobody to ask, 22 awaiting a permission request, 0 refused. All 34 gating checks pass.

Every quote verified twice, by different routes

The research agents read the pages. This seat then re-fetched all 30 independently and matched the exact string, using a different path wherever the first was blocked: greaterwrong for LessWrong, ea.greaterwrong for the EA Forum, PDF text extraction for the 2008 Yudkowsky chapter. That is pdoom1#1075 applied to quotations, and it is why this file can be trusted in a way the reaction fields never could be.

Both quotes an agent had flagged as medium-confidence held up under re-check.

Servable today, nobody to ask (8)

Licence read from the source, not assumed from the platform:

source licence confirmed by
Zvi Mowshowitz x4 bespoke attribution licence his about page: quotation with credit and a link back
Distill x3 CC-BY-4.0 each article footer, read individually
EA Forum, Aschenbrenner CC-BY-4.0 published 2023-03-29, after the 2022-12-01 mandatory date

The Distill check earned itself. The sourcing brief warned the version varies per article, 2.0 on a 2017 piece and 4.0 later, so all three footers were read rather than assumed. All three state CC-BY 4.0, and the wording is "Diagrams and text", so prose is covered.

The best line found anywhere is among these, and it is the brief's own thesis said by someone else years before the game existed:

"A response of 'yes this is an impossible problem but I'll solve it anyway' seems great."
Zvi Mowshowitz, 2022

Held for an ask despite being licence-clear on date

David Seiler's EA Forum comment is dated after the cutoff, but the Forum's announcement speaks of content posted to the Forum, and whether that covers comments identically is not established. Both the terms page and its mirror refused automated fetch, so it could not be checked. An unresolved licence is not a basis. Cheap to fix by asking him.

Two licence-clear quotes deliberately left out

Gwern's (CC0) and one of Zvi's have no Wayback snapshot, and a save request did not complete. The gate requires an archive before anything is servable and is right to: a personal site can change or vanish, and we could not later show we quoted correctly. Both are recorded in EDITORIAL_BRIEF.md with their verified text; retry the capture and they go straight in.

One of them is the best "no retries" line in the set, which is conspicuous to be missing from a death screen that exists to offer one.

The ask list (22)

16 authors, grouped so one email covers several quotes. Recorded per record: Shlegeris's EA Forum post predates the licence cutoff; the Garrabrant pair is co-authored with Abram Demski so both must be asked; Russell's and Olah's sit behind 80,000 Hours' explicitly restrictive terms, making those two permissions each; and Katja Grace's concession appears inside a piece arguing against the risk case, which is what makes it worth having.

Best first three: Karnofsky, Cotra, Christiano. Each has a working contact route and more than one usable quote, and Christiano is the useful bellwether for the MIRI-adjacent group.

One flagged rather than dropped

Bensinger's "we don't seem to have a much better idea now than we did 10 years ago" is accurate and leans stagnation rather than defiance, which the brief warns against as doom without agency. Its editorial_note says so. Wants a human read before he is asked.

Four further candidates were dropped on editorial rather than accuracy grounds: an ambiguous podcast line, a section header rather than prose, a 2026 quote past the predates-2024 preference, and one too jargon-heavy to survive being read alone.

Pip Foweraker and others added 6 commits August 30, 2026 09:31
…cted

Pip's tone rule in his own framing, because it is what a candidate quote is
judged against and it was the thing the schema could not supply:

  This is hard. You're going to lose. We told you you were going to lose before
  Pip started making the game. And the real-world version of this is STILL
  harder than the game.

Call of Duty's death screen is the reference, and it is not mocking the player:
it says losing is not a personal failure because the thing itself is hard and
always has been. The effect wanted is the defiance that makes someone hit
respawn. It is written FOR the player who is not already safety-pilled, who
lost a management game and needs to hear from someone other than the author
that trying and losing at this is the heroic thing.

One constraint keeps it honest: A QUOTE MUST WORK EVEN IF THE GAME IS BAD. If a
line only lands because the player was moved by the preceding twenty minutes,
it is doing the game's job rather than its own. It should survive a screenshot,
read cold, in a version of the game that did not turn out well. That is also
the honest position, since the game is not finished and a quote that quietly
asserts otherwise is overselling in the way this estate has already retracted
once.

First pass is DIFFICULTY ONLY, ruled today. Accountability comes later in
development with much more guarding; the schema supports that tier and the gate
enforces it, and nothing is collected into it on purpose.

MEASURED AND RECORDED: the easiest source is the worst one. The 1,129 arXiv
abstracts are CC0, verified, already fetched and need no permission at all.
Scanned for difficulty language: 232 candidate sentences across 198 papers, and
essentially none survive the brief. The best of them are "Neural networks are
still hard to design" and "CNN representations are notoriously difficult for
humans to interpret" -- true, hedged by academic convention, and about narrow
subproblems rather than about the thing being hard.

The licence was never the constraint. The prose is. Written down so the next
person does not re-run the scan expecting a different answer.

The quotes that work will come from essays, talks and forum posts, where
someone was writing to be understood rather than to be precise, and most of
those need permission. That is the real cost of this feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
…g an ask

The ask list. Every one is basis not_yet_asked, which is the state a found
quote is supposed to be in, and the gate correctly reports 0 servable.

16 authors. The strongest for the brief are Christiano's "some years of
trying-and-failing to find a plausible failure story", Ngo's "even if we knew
exactly what we wanted a superintelligent agent to do, we don't currently know
(even in theory) how to make an agent which actually tries to do that", and
Russell's "there are difficult research problems that we have to solve, but
we've solved quite a lot of problems already" -- which is the only one that
does difficulty AND defiance in one breath, which is the whole brief.

EVERY QUOTE WAS VERIFIED TWICE, BY DIFFERENT ROUTES. The research agent read
the pages; this seat then re-fetched all 21 independently and matched the exact
string, using a different path where the first was blocked -- greaterwrong for
LessWrong, ea.greaterwrong for the EA Forum, and extracting the PDF text for
the 2008 Yudkowsky chapter. That is the pdoom1#1075 rule applied to quotations:
the check does not derive what to look for from the thing being checked. Both
quotes the agent had flagged as medium-confidence held up.

Four candidates from the research were DROPPED on editorial grounds rather than
accuracy: a Christiano podcast line the agent itself noted reads ambiguously
out of context, a Yudkowsky section header rather than prose, a 2026 Rohin Shah
line well past the predates-2024 preference, and a Garrabrant decision-theory
line too jargon-heavy to survive being read alone.

One is flagged in its own editorial_note rather than dropped: Bensinger's
"we don't seem to have a much better idea now than we did 10 years ago" is
accurate and leans stagnation rather than defiance, which the brief warns
against as doom without agency. It wants a second read before anyone asks him.

Also recorded per record: that Shlegeris's is EA Forum but PRE 2022-12-01, so
the site licence does not apply; that the Garrabrant pair is co-authored with
Abram Demski and both must be asked; that Russell's and Olah's sit behind
80,000 Hours' explicitly restrictive terms, so that is two permissions each;
and that Grace's concession appears inside a piece arguing AGAINST the risk
case, which is what makes it worth having.

archive_url is null on all 21. The gate requires it before anything is
servable, and capturing them is part of the ask rather than before it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
… twice

The corpus, filled. 30 quotes: 8 usable today with nobody to ask, 22 on the
ask list, 0 refused.

EVERY QUOTE VERIFIED TWICE BY DIFFERENT ROUTES. The research agents read the
pages; this seat then re-fetched all 30 independently and matched the exact
string, using a different path wherever the first was blocked -- greaterwrong
for LessWrong, ea.greaterwrong for the EA Forum, PDF text extraction for the
2008 Yudkowsky chapter. That is pdoom1#1075 applied to quotations, and it is
why this file can be trusted in a way the reaction fields never could be. Both
quotes an agent flagged medium-confidence held up.

SERVABLE NOW, licence read from the source rather than assumed: four from Zvi
Mowshowitz under his bespoke attribution licence, three from Distill, one EA
Forum post from 2023-03-29.

The Distill check earned itself. The sourcing brief warned the CC BY version
varies per article -- 2.0 on a 2017 piece, 4.0 later -- so all three footers
were read individually. All three state CC-BY 4.0 and the wording is "Diagrams
and text", so prose is covered.

The best line found anywhere is among the servable ones, and it is the brief's
own thesis said by someone else years before the game existed:

  "A response of 'yes this is an impossible problem but I'll solve it anyway'
  seems great."  -- Zvi Mowshowitz, 2022

HELD FOR AN ASK DESPITE BEING LICENCE-CLEAR ON DATE: David Seiler's EA Forum
COMMENT is dated after the 2022-12-01 cutoff, but the Forum's announcement
speaks of content posted to the Forum and whether that covers comments
identically is not established. Both the terms page and its mirror refused
automated fetch, so it could not be checked. An unresolved licence is not a
basis. Cheap to fix by asking him.

TWO LICENCE-CLEAR QUOTES DELIBERATELY LEFT OUT, recorded in the brief with
their verified text: Gwern's under CC0 and one of Zvi's have no Wayback
snapshot and a save request did not complete. The gate requires an archive
before anything is servable and is right to, since a personal site can change
or vanish. One of the two is the best "no retries" line in the whole set, which
is a conspicuous thing to be missing from a death screen that exists to offer
one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
Checked while following up a warm introduction to Katja Grace.
wiki.aiimpacts.org states CC0 1.0. aiimpacts.org, where the quote actually
lives, carries no licence statement at all. Same organisation, two properties,
different terms, and a reader who saw the wiki's CC0 and generalised would have
published on a licence that does not cover this.

Recorded on the record rather than in a session log, because the next person to
look at an AI Impacts URL will have the same thought.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…d quote

Answering the right question: the main surface is the game, but this repository
is PUBLIC, so it is a second surface and the 22 unapproved quotes are already
on GitHub. Storing is not publishing, but a public repo is not a notebook
either. So the two surfaces get separated properly rather than trusted.

project_quotes.py emits data/serveable/api/quotes/approved.jsonl and copies
across only records with a named permission basis. There is no flag to make it
emit anything else. The curated file stays a research index; the served file is
the publication; and the boundary is a build step rather than a rule somebody
remembers. Copying a line out by hand is still possible, since nothing stops a
determined person, but it cannot happen by ACCIDENT.

THE SERVED ZONE CARRIES NO WITHHELD TEXT. Blocked records appear in the lineage
as an id and a status and nothing else. A file that announced "these 22 are not
approved" and then printed all 22 verbatim would republish exactly what it was
refusing to publish. The reasons are useful to a consumer; the words are not
theirs yet. Same argument as project_watch_accepted.py, which reports a blocked
atom with its reason rather than dropping it silently.

Currently: 30 curated, 8 served, 22 withheld, all not_yet_asked.

NOTICE.md is the other half, and it is the part that makes this respectful
rather than merely defensible. Anyone who finds their own words in the
directory gets told plainly that found is not asked and not used, that the game
reads a different file, and how to have the record withdrawn without giving a
reason. Withdrawal is permanent and keeps its note, so nobody who never saw the
message can quietly reinstate it. Narrower requests are supported too:
correction, different attribution, a note that their views have changed, or
removal from one placement but not another.

Keeping the candidates visible is deliberate. The alternative is a private list
of other people's words, which is worse, and it would also contradict the
lesson of #87: judgement held on one machine is judgement that does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
…fies

CC BY does not say "credit the authors", it says credit them in the manner the
author specifies. Distill states its own: "For attribution in academic
contexts, please cite this work as Olah, et al." So the three Distill records
carry that first-author-et-al form in display_preference, and the projection
emits it as a `credit` line. That is both correct licence compliance and the
only form that fits on a screen; the full seven-author list would have been
generous and wrong.

FINAL CONFIDENCE PASS, all 8 verified against TWO independent copies: the live
page and the Wayback snapshot, matching the exact string in both. The single
live-fetch failure is EA Forum's Cloudflare block, and that record's archive
copy confirms it. So every quote in the served zone has been checked against a
source that cannot have drifted since capture.

Served: four from Zvi Mowshowitz under his bespoke attribution licence, three
from Distill under CC-BY-4.0 read from each article's own footer, and one EA
Forum post from 2023-03-29, after the date its CC BY licence became mandatory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
@PipFoweraker
PipFoweraker merged commit 1e436f0 into main Aug 30, 2026
4 checks passed
@PipFoweraker
PipFoweraker deleted the feat/quote-mining-difficulty branch August 30, 2026 07:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant