Teek builds models of real, named, living people and uses them to predict behavior. That is a serious thing to do. This document is the constraint, not the disclaimer. It is read before anyone is added, not after something goes wrong.
Legal specifics (right of publicity, state AI likeness statutes, EU AI Act transparency duties) are pending the landscape research pass and will be merged into the Legal Boundary section below. Treat that section as incomplete until this note is removed.
Would you be comfortable showing this profile to the person it describes?
If the answer is no, one of three things is true: a claim is unsourced, a claim is uncharitable beyond what the evidence supports, or the person should not be in the repo. All three are fixable. None of them are acceptable to ship.
This test is not about flattery. Andrew Chen should be able to read his own profile and say "that is a fair reading of what I have published, including the parts I would not have volunteered." That is the bar. Fair, sourced, and unflattering is fine. Unsourced is not.
1. Public figures, public statements. The subject must be a public figure by any reasonable standard, and every source must be material they published, said on the record, or that a credible outlet published about them. No private individuals. No private material. No leaked documents. No material behind a login that was obtained by pretending to be a user.
The public-figure bar is not "has a public account." It is closer to: this person has voluntarily entered public discourse on the subject being modeled, at volume, and would expect their published views to be analyzed. A founder with 1,600 tweets and a small following does not clear it. A managing partner who has published 650 essays about how he invests does.
2. Every claim cited, or it does not exist.
An uncited claim about a real person is a fabrication with their name attached. The evidence ledger exists so that any claim can be traced to a dated document and a verbatim excerpt. Claims marked inferred must be labeled as inference in any output that uses them, and must never be simulated as fact.
3. No clinical language, ever. Teek measures decision style, risk posture, values ordering, and communication pattern. It does not measure mental health. No diagnoses. No disorder names. No "narcissistic," "paranoid," "manic," or any other term with clinical meaning, even colloquially. The professional norm here (psychiatry's Goldwater Rule) exists for good reason: diagnosis at a distance is not valid, and it is defamatory when wrong.
This rule survives even when a source uses such language. If a journalist called someone a narcissist, the claim in the ledger is "a 2019 profile in {outlet} characterized him as narcissistic," attributed and dated. It is never "he is narcissistic."
4. Adversarial sections require external sources. The Cognitive Biases and Narrative Gap sections describe the gap between how someone sees themselves and how others see them. These are the most valuable sections and the most dangerous. They may only be populated from published external analysis, cited. They may never be generated by asking a model what it thinks this person's blind spots are. That produces confident, plausible, uncheckable slander.
5. Simulate to rehearse, never to publish. Teek output is for private research, decision rehearsal, and evaluation. It is not for producing content in a real person's voice for publication, impersonation, or any use where a reader could mistake it for the actual person. Nothing Teek generates gets posted anywhere as if the person said it.
6. No account-linked collection. Corpus collection uses public, anonymous, sanctioned routes: publisher sites, RSS, sitemaps, public archives. Never a logged-in session, never a personal account's cookies, never a route that impersonates a real user of a platform. This is both an ethics rule and an operational one; account-linked scraping puts a real person's account at risk and taints the provenance of everything it touches.
7. Right of reply.
If a living subject asks to see their profile, they see it. If they dispute a claim, the dispute is recorded in the ledger as a contradicted status with their statement cited, regardless of whether we agree. If they ask to be removed, they are removed.
Not every layer carries the same risk. The deeper the profile, the higher the bar.
| Layer | What it contains | Consent posture |
|---|---|---|
| L0 routing card | Public role, stated criteria | Public record. No consent needed. |
| L1 corpus | Their published words, stored | Public record. Respect robots.txt and terms. |
| L2 evidence ledger | Cited claims with status | Public record, but disputes honored on request. |
| L3 coding | Scored decision-style traits | Internal only for living subjects unless they consent. |
| L4 profile | Full cognitive synthesis | Internal only for living subjects unless they consent. |
| L5 evaluation | Predictive accuracy scores | Internal only. |
The line sits between L2 and L3. Everything up to a cited claim ledger is recognizably journalism or research. A scored psychological trait profile of a named living person is a different object, and it does not get published about a living subject without that subject's agreement.
This is why the historical cohort matters beyond its scientific value. Deceased subjects with public-domain corpora let the method be demonstrated in full, openly, with no consent problem at all.
- They ask.
- The public-figure bar turns out not to be met.
- The corpus was collected by a banned route and cannot be rebuilt cleanly.
- The profile cannot be maintained: sources have rotted, claims have expired, and nobody is refreshing it. A stale profile of a real person is a misrepresentation that gets worse over time.
Incomplete. Pending research pass.
The working assumption, to be verified: analysis and commentary about public figures using their published statements sits in well-protected territory. Commercial simulation of a named person's likeness or voice does not, and several jurisdictions have recently tightened this considerably. The distance between "a cited analysis of how Bill Gurley evaluates companies" and "a Bill Gurley chatbot you can subscribe to" is the entire legal question, and Teek should assume the second requires consent and a license.
Teek's method is grounded in psychological assessment research that the author has no training in. The research was done by an AI, the method was implemented by an AI, and the profiles are generated by an AI. That is disclosed openly rather than obscured, because the alternative is implying an expertise that is not there.
This is not a small caveat. It is a live limitation with a specific failure mode: a model asked to code a famous person's text can retrieve its impression of that person from training rather than reading the document in front of it. Every automated coding pass must be built to detect that, and docs/method.md treats it as the central engineering problem, not a footnote.