quotes: 30 verified difficulty candidates, 8 servable today - #101
Merged
Merged
Conversation
…cted Pip's tone rule in his own framing, because it is what a candidate quote is judged against and it was the thing the schema could not supply: This is hard. You're going to lose. We told you you were going to lose before Pip started making the game. And the real-world version of this is STILL harder than the game. Call of Duty's death screen is the reference, and it is not mocking the player: it says losing is not a personal failure because the thing itself is hard and always has been. The effect wanted is the defiance that makes someone hit respawn. It is written FOR the player who is not already safety-pilled, who lost a management game and needs to hear from someone other than the author that trying and losing at this is the heroic thing. One constraint keeps it honest: A QUOTE MUST WORK EVEN IF THE GAME IS BAD. If a line only lands because the player was moved by the preceding twenty minutes, it is doing the game's job rather than its own. It should survive a screenshot, read cold, in a version of the game that did not turn out well. That is also the honest position, since the game is not finished and a quote that quietly asserts otherwise is overselling in the way this estate has already retracted once. First pass is DIFFICULTY ONLY, ruled today. Accountability comes later in development with much more guarding; the schema supports that tier and the gate enforces it, and nothing is collected into it on purpose. MEASURED AND RECORDED: the easiest source is the worst one. The 1,129 arXiv abstracts are CC0, verified, already fetched and need no permission at all. Scanned for difficulty language: 232 candidate sentences across 198 papers, and essentially none survive the brief. The best of them are "Neural networks are still hard to design" and "CNN representations are notoriously difficult for humans to interpret" -- true, hedged by academic convention, and about narrow subproblems rather than about the thing being hard. The licence was never the constraint. The prose is. Written down so the next person does not re-run the scan expecting a different answer. The quotes that work will come from essays, talks and forum posts, where someone was writing to be understood rather than to be precise, and most of those need permission. That is the real cost of this feature. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
…g an ask The ask list. Every one is basis not_yet_asked, which is the state a found quote is supposed to be in, and the gate correctly reports 0 servable. 16 authors. The strongest for the brief are Christiano's "some years of trying-and-failing to find a plausible failure story", Ngo's "even if we knew exactly what we wanted a superintelligent agent to do, we don't currently know (even in theory) how to make an agent which actually tries to do that", and Russell's "there are difficult research problems that we have to solve, but we've solved quite a lot of problems already" -- which is the only one that does difficulty AND defiance in one breath, which is the whole brief. EVERY QUOTE WAS VERIFIED TWICE, BY DIFFERENT ROUTES. The research agent read the pages; this seat then re-fetched all 21 independently and matched the exact string, using a different path where the first was blocked -- greaterwrong for LessWrong, ea.greaterwrong for the EA Forum, and extracting the PDF text for the 2008 Yudkowsky chapter. That is the pdoom1#1075 rule applied to quotations: the check does not derive what to look for from the thing being checked. Both quotes the agent had flagged as medium-confidence held up. Four candidates from the research were DROPPED on editorial grounds rather than accuracy: a Christiano podcast line the agent itself noted reads ambiguously out of context, a Yudkowsky section header rather than prose, a 2026 Rohin Shah line well past the predates-2024 preference, and a Garrabrant decision-theory line too jargon-heavy to survive being read alone. One is flagged in its own editorial_note rather than dropped: Bensinger's "we don't seem to have a much better idea now than we did 10 years ago" is accurate and leans stagnation rather than defiance, which the brief warns against as doom without agency. It wants a second read before anyone asks him. Also recorded per record: that Shlegeris's is EA Forum but PRE 2022-12-01, so the site licence does not apply; that the Garrabrant pair is co-authored with Abram Demski and both must be asked; that Russell's and Olah's sit behind 80,000 Hours' explicitly restrictive terms, so that is two permissions each; and that Grace's concession appears inside a piece arguing AGAINST the risk case, which is what makes it worth having. archive_url is null on all 21. The gate requires it before anything is servable, and capturing them is part of the ask rather than before it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
… twice The corpus, filled. 30 quotes: 8 usable today with nobody to ask, 22 on the ask list, 0 refused. EVERY QUOTE VERIFIED TWICE BY DIFFERENT ROUTES. The research agents read the pages; this seat then re-fetched all 30 independently and matched the exact string, using a different path wherever the first was blocked -- greaterwrong for LessWrong, ea.greaterwrong for the EA Forum, PDF text extraction for the 2008 Yudkowsky chapter. That is pdoom1#1075 applied to quotations, and it is why this file can be trusted in a way the reaction fields never could be. Both quotes an agent flagged medium-confidence held up. SERVABLE NOW, licence read from the source rather than assumed: four from Zvi Mowshowitz under his bespoke attribution licence, three from Distill, one EA Forum post from 2023-03-29. The Distill check earned itself. The sourcing brief warned the CC BY version varies per article -- 2.0 on a 2017 piece, 4.0 later -- so all three footers were read individually. All three state CC-BY 4.0 and the wording is "Diagrams and text", so prose is covered. The best line found anywhere is among the servable ones, and it is the brief's own thesis said by someone else years before the game existed: "A response of 'yes this is an impossible problem but I'll solve it anyway' seems great." -- Zvi Mowshowitz, 2022 HELD FOR AN ASK DESPITE BEING LICENCE-CLEAR ON DATE: David Seiler's EA Forum COMMENT is dated after the 2022-12-01 cutoff, but the Forum's announcement speaks of content posted to the Forum and whether that covers comments identically is not established. Both the terms page and its mirror refused automated fetch, so it could not be checked. An unresolved licence is not a basis. Cheap to fix by asking him. TWO LICENCE-CLEAR QUOTES DELIBERATELY LEFT OUT, recorded in the brief with their verified text: Gwern's under CC0 and one of Zvi's have no Wayback snapshot and a save request did not complete. The gate requires an archive before anything is servable and is right to, since a personal site can change or vanish. One of the two is the best "no retries" line in the whole set, which is a conspicuous thing to be missing from a death screen that exists to offer one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
Checked while following up a warm introduction to Katja Grace. wiki.aiimpacts.org states CC0 1.0. aiimpacts.org, where the quote actually lives, carries no licence statement at all. Same organisation, two properties, different terms, and a reader who saw the wiki's CC0 and generalised would have published on a licence that does not cover this. Recorded on the record rather than in a session log, because the next person to look at an AI Impacts URL will have the same thought. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…d quote Answering the right question: the main surface is the game, but this repository is PUBLIC, so it is a second surface and the 22 unapproved quotes are already on GitHub. Storing is not publishing, but a public repo is not a notebook either. So the two surfaces get separated properly rather than trusted. project_quotes.py emits data/serveable/api/quotes/approved.jsonl and copies across only records with a named permission basis. There is no flag to make it emit anything else. The curated file stays a research index; the served file is the publication; and the boundary is a build step rather than a rule somebody remembers. Copying a line out by hand is still possible, since nothing stops a determined person, but it cannot happen by ACCIDENT. THE SERVED ZONE CARRIES NO WITHHELD TEXT. Blocked records appear in the lineage as an id and a status and nothing else. A file that announced "these 22 are not approved" and then printed all 22 verbatim would republish exactly what it was refusing to publish. The reasons are useful to a consumer; the words are not theirs yet. Same argument as project_watch_accepted.py, which reports a blocked atom with its reason rather than dropping it silently. Currently: 30 curated, 8 served, 22 withheld, all not_yet_asked. NOTICE.md is the other half, and it is the part that makes this respectful rather than merely defensible. Anyone who finds their own words in the directory gets told plainly that found is not asked and not used, that the game reads a different file, and how to have the record withdrawn without giving a reason. Withdrawal is permanent and keeps its note, so nobody who never saw the message can quietly reinstate it. Narrower requests are supported too: correction, different attribution, a note that their views have changed, or removal from one placement but not another. Keeping the candidates visible is deliberate. The alternative is a private list of other people's words, which is worse, and it would also contradict the lesson of #87: judgement held on one machine is judgement that does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
…fies CC BY does not say "credit the authors", it says credit them in the manner the author specifies. Distill states its own: "For attribution in academic contexts, please cite this work as Olah, et al." So the three Distill records carry that first-author-et-al form in display_preference, and the projection emits it as a `credit` line. That is both correct licence compliance and the only form that fits on a screen; the full seven-author list would have been generous and wrong. FINAL CONFIDENCE PASS, all 8 verified against TWO independent copies: the live page and the Wayback snapshot, matching the exact string in both. The single live-fetch failure is EA Forum's Cloudflare block, and that record's archive copy confirms it. So every quote in the served zone has been checked against a source that cannot have drifted since capture. Served: four from Zvi Mowshowitz under his bespoke attribution licence, three from Distill under CC-BY-4.0 read from each article's own footer, and one EA Forum post from 2023-03-29, after the date its CC BY licence became mandatory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCAwQ2dTCtpR8rTrEfv5ee
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The corpus, filled. 30 quotes: 8 servable on licence alone with nobody to ask, 22 awaiting a permission request, 0 refused. All 34 gating checks pass.
Every quote verified twice, by different routes
The research agents read the pages. This seat then re-fetched all 30 independently and matched the exact string, using a different path wherever the first was blocked:
greaterwrongfor LessWrong,ea.greaterwrongfor the EA Forum, PDF text extraction for the 2008 Yudkowsky chapter. That ispdoom1#1075applied to quotations, and it is why this file can be trusted in a way the reaction fields never could be.Both quotes an agent had flagged as medium-confidence held up under re-check.
Servable today, nobody to ask (8)
Licence read from the source, not assumed from the platform:
The Distill check earned itself. The sourcing brief warned the version varies per article, 2.0 on a 2017 piece and 4.0 later, so all three footers were read rather than assumed. All three state CC-BY 4.0, and the wording is "Diagrams and text", so prose is covered.
The best line found anywhere is among these, and it is the brief's own thesis said by someone else years before the game existed:
Held for an ask despite being licence-clear on date
David Seiler's EA Forum comment is dated after the cutoff, but the Forum's announcement speaks of content posted to the Forum, and whether that covers comments identically is not established. Both the terms page and its mirror refused automated fetch, so it could not be checked. An unresolved licence is not a basis. Cheap to fix by asking him.
Two licence-clear quotes deliberately left out
Gwern's (CC0) and one of Zvi's have no Wayback snapshot, and a save request did not complete. The gate requires an archive before anything is servable and is right to: a personal site can change or vanish, and we could not later show we quoted correctly. Both are recorded in
EDITORIAL_BRIEF.mdwith their verified text; retry the capture and they go straight in.One of them is the best "no retries" line in the set, which is conspicuous to be missing from a death screen that exists to offer one.
The ask list (22)
16 authors, grouped so one email covers several quotes. Recorded per record: Shlegeris's EA Forum post predates the licence cutoff; the Garrabrant pair is co-authored with Abram Demski so both must be asked; Russell's and Olah's sit behind 80,000 Hours' explicitly restrictive terms, making those two permissions each; and Katja Grace's concession appears inside a piece arguing against the risk case, which is what makes it worth having.
Best first three: Karnofsky, Cotra, Christiano. Each has a working contact route and more than one usable quote, and Christiano is the useful bellwether for the MIRI-adjacent group.
One flagged rather than dropped
Bensinger's "we don't seem to have a much better idea now than we did 10 years ago" is accurate and leans stagnation rather than defiance, which the brief warns against as doom without agency. Its
editorial_notesays so. Wants a human read before he is asked.Four further candidates were dropped on editorial rather than accuracy grounds: an ambiguous podcast line, a section header rather than prose, a 2026 quote past the predates-2024 preference, and one too jargon-heavy to survive being read alone.