A log of non-obvious design choices for .table/ and the reasoning
behind them. Every entry is a settled call; pivots get a new entry, not
edits to an old one.
Every row has an id (nanoid) that is system-level — minted by the
writer, never edited by the user. primaryKey (if present) is a
separate domain-level uniqueness constraint over user-facing fields.
Why: cross-table relations need a stable target that doesn't break
when users rename their domain key. id provides that; primaryKey
is purely about user-facing uniqueness validation.
rows.ndjson ends with \n (POSIX-correct).
Why: makes wc -l accurate, ensures the last line is well-formed
on append, plays nicely with line-oriented Unix tools.
"relation": { "table": "projects", "field": "id" }Why: path-based references break when the workspace is reshuffled (folder rename, restructure). Name-based resolution is the app's problem (scan workspace, walk parent dir) but the format itself stays self-contained and rename-resilient.
Personal/per-user view state (last opened, scroll position, focused
row) lives in app-local storage outside the .table/ directory.
Why: .table/ is meant to round-trip cleanly through git and
between users. Personal state would create noisy diffs and would be
meaningless to other users. A future views/ directory convention
could be added non-breakingly if needed.
When a field has an enum constraint, its values sort and group by the enum's declared order, not alphabetical order.
Why: enums encode semantic order (todo → doing → done).
Alphabetical defeats the purpose. Reordering enum values is therefore
a behaviour change and SHOULD bump schema-version.
index.sqlite is a performance layer only. Readers fall back to
scanning rows.ndjson when the cache is absent or stale.
Apps decide their own caching strategy.
Why: keeps the format self-contained as plain text; lets apps opt into the performance tier they need. Removes the spec burden of prescribing cache freshness rules beyond the fallback contract.
Field declared attachment: true; row holds the bare filename;
reader resolves under attachments/. Recommended write convention to
avoid collisions: {nanoid}-{original-name}.ext.
Why: the filename in the row is human-readable and grep-friendly.
Path resolution is centralised in one place (attachments/) so
moves/renames of the table don't cascade through every row.
Final.
Why: .table is unclaimed in practice (SQL keywords and HTML
elements aren't file extensions), reads correctly in conversation
("send me your suppliers.table"), and follows the same naming
pattern as Airtable / Google Tables / Notion — category-format named
after the most recognisable presentation, even when the format
supports many other layouts (board, gallery, list, calendar).
The format is not a "Frictionless Table Schema superset." It owns
its own schema vocabulary. Field type names (string, integer,
enum, etc.) follow conventions used across text-first data formats
generally — they're the right names for these things, not a
compatibility claim with any specific format.
Why: the original spec claimed Frictionless alignment, but
.table/'s extensions (per-field relations, views, attachments,
schema versioning, deprecation) carry the format's actual value;
Frictionless contributes only field-type names. Calling it a "superset"
was dressing on a different format. Frictionless tooling lives in
data-publishing communities; the target user (information worker) has
not heard of it. The interop benefit didn't accrue to the actual user.
Frictionless / CSVW / Obsidian Bases conversion now belongs in optional downstream packages, not the core spec.
foreignKeys (the Frictionless-style table-level constraint) is not
carried. Cross-table links are field-typed via the relation
annotation.
Why: Airtable-style relations are field-typed, not table-level. The original spec carried both "for interop", but the duplication was sunk cost — nobody was cashing in on the Frictionless side, and maintaining two relation primitives for the same concept added mental load.
Every write stamps format: "table" and formatVersion: 1 onto
meta.json. Readers MAY use them to validate that a directory is a
.table/.
Why: a directory called projects.table/ should self-identify.
Without a manifest field, a reader has only the extension as a signal,
which is filesystem-level and unreliable for tooling that operates on
pipes/streams.
Optional. Each body is a standalone markdown file named by the row's
stable system id. The Notion every-row-is-also-a-page pattern,
made line-diffable.
For shorter inline markdown content (a description, a one-paragraph
summary), use a string field with format: "markdown" instead.
Why: NDJSON's one-line-per-row promise breaks down for long-form
content — a 500-line markdown body becomes a giant single line in
rows.ndjson with \n-escaped newlines, and git diff shows the
whole row as changed when a single paragraph edits. Splitting bodies
into per-row files preserves line-diffable history for prose while
keeping rows.ndjson tight and scannable. Each body is also
independently readable as a standalone .md file (cat bodies/p1.md
returns clean markdown).
Default branch is develop. Work lands via PR from feature
branches. Open spec questions get filed as GitHub issues so they're
discoverable.
Why: PR-driven workflow gives the spike a clean review surface
and makes the rationale for each change auditable in the commit log.
Topic-branch naming follows the standard feat/<slug> / fix/<slug>
/ docs/<slug> / chore/<slug> convention.
The format does not specify a data-versioning model — no per-row history, no edit log, no concurrency/conflict semantics, no audit trail format. The spec covers two distinct version axes only:
formatVersioninmeta.json— versions the spec itself (currently1).schema-versioninschema.json— versions the user's data schema (bumps on structural change: field added/deprecated, enum reordered, constraints tightened).
Everything else — undo/redo, "what did this cell look like yesterday," real-time collaboration, audit logs — is left to the consuming app.
Why: every NDJSON-row design choice in .table/ (line-diffable
rows, per-row bodies in separate files, attachments by filename,
append-friendly ordering) exists so git is the version-control
substrate. For any consumer that uses git, the format gets full
history, branching, merging, diffing, and authorship for free.
Specifying a parallel in-format versioning system would reinvent what
git already does well and would bias the format toward one app's UX
needs over others.
Where each app-level need belongs:
| Need | Owner | How |
|---|---|---|
| Undo/redo within an editing session | App | In-memory stack of inverse edits |
| Async collaboration (two users, different times) | Git | Line-diffable rows merge cleanly |
| Real-time collaboration (two users, same moment) | App | CRDTs / OT — separate research area |
| "What did this row look like yesterday?" | App or git | git log / git blame for git-backed; app-side log otherwise |
| Audit / compliance trail | App | Append-only event log; see reserved extension below |
| Schema migration tracking | Format | schema-version bumps on structural change |
Not yet specified, not yet implemented. Reserved for if/when a
consuming app needs a portable, in-format edit log. Likely shape when
it lands: an optional history.ndjson at the directory root, one
event per line:
{"id":"e_xyz","at":"2026-04-28T15:00:00Z","by":"leslie","op":"set","row":"p1","field":"status","from":"todo","to":"doing"}Properties (intended):
rows.ndjsonremains canonical; the log is purely additive.- Append-only, line-diffable, plays nicely with git on top.
- Apps that care implement it; apps that don't, ignore it.
- Readers MUST tolerate absence; MAY ignore unknown
opvalues.
This shape is not normative until specced. Don't write tooling against it yet.
Why park rather than spec now: specifying a versioning model before a real consumer has built against it bakes assumptions we don't have signal for. The format stays lean; the first concrete history UX in a consuming app gets to drive what the extension needs to look like.
The index.sqlite contract (SPEC section 8): staleness is detected by
content hash (schema_hash + rows_hash stored in a _meta table
inside the index), rebuilds are whole-file inside one transaction,
and writers build into a temp file then atomically rename. FTS5
indexes string fields and body contents.
Why hashes, not mtimes: git rewrites mtimes on checkout/pull even when content is unchanged — an mtime check rebuilds after every git operation. Hashing a multi-MB NDJSON costs ~50ms, paid once per open.
Why full rebuild: at this format's size class (Airtable caps at 50k rows/base) a transactional rebuild is sub-second. Incremental indexing is deferred until a real consumer outgrows that — the format is already incremental-friendly (append-only edits detectable by prefix hash + byte watermark) if that day comes.
Why FTS over bodies: searchRows already searches bodies; an
index that omitted them would silently return fewer hits than the
fallback it replaces.
The original locked interface was queryIndex(dirPath, sql). Changed
(while still a stub, before any consumer existed) to a structured
IndexQuery — the filter / sort shapes from views.json plus a
search string — compiled to SQL internally.
Why: raw SQL freezes the cache's internal layout (column naming, encodings, FTS config) into a public contract — the very thing "never the source of truth" exists to prevent. The AST keeps the indexed path and the in-memory fallback speaking one query language so results can't diverge, and user search text never reaches an SQL string.
For multi-writer sync (Hypercore/Autobase per workspace-p2p-spike),
the intended model is an op log per writer, Autobase linearisation,
per-field last-writer-wins, and the .table/ directory materialised
from the log as a deterministic checkout. Within a synced workspace
the log is canonical and files are derived — which D14 already
permits, since the authority model belongs to the consuming app.
Full reasoning in docs/STORAGE-AND-SYNC.md.
Three knock-on rules recorded here because they constrain the format, not just consumers:
- Canonical write order (SPEC section 3): identical state must produce byte-identical files, or convergence can't be hash-checked and git sees phantom diffs.
- Append-only schema evolution is load-bearing for sync, not just for compatibility: field adds commute between concurrent writers; renames/removes would not. Don't relax it.
modified_atis stamped on user-initiated writes only (SPEC section 5): a sync engine touching it on every apply makes every replica differ by timestamp alone.
The reserved history.ndjson (D14) and the sync op log are one
design: if the extension lands, it is the at-rest serialisation of
the same event vocabulary, not a parallel format.
An enum entry is either a bare string or an object
{ value, color?, label? }. Only value participates in validation
and enum-ordered sort/group; color (symbolic 8-color palette) and
label are display-only. Readers coerce strings to { value } via
enumOptions(); both forms may mix in one array.
Why: the flat string array couldn't capture author intent
("active = green"), so every consumer picked chip colors
algorithmically and two consumers disagreed. Putting the intent in
the schema makes it round-trip. Object-or-string (rather than a
parallel enumConfig map keyed by value) keeps the common minimal
case a plain string and the declaration in one place. Colors are
symbolic, not hex, so each consumer themes them for light/dark.
Field format tokens come from a fixed per-type set (number:
decimal:N / percent / currency:<ISO> / …; date: short /
long / relative / …; string: markdown / url / email / …).
Stored values stay raw; formatValue() renders locale-aware via
Intl. Unknown tokens fall back to plain.
Why: arbitrary format strings (Excel's #,##0.00;[Red]) are a
mini-language every reader must reimplement identically or diverge —
and they bake locale assumptions into the data file. A closed enum of
semantics ("this is USD", "show this date short") lets each
consumer render correctly for its own locale. Reuses the existing
format key (already "markdown" on strings) rather than adding a
second display-hint field.
Relations stay a single primitive (see D10). Multi-target is a
cardinality: "one" | "many" flag on the existing relation
declaration; "many" means the row value is an array of ids.
Defaults to "one" — backwards compatible.
Why: "one project" and "many tags" are the same relation concept
at different arity, not two different field types. A flag keeps the
one relation primitive intact and the address grammar unchanged (each
id still resolves to <path>#row=<id>), so navigation works for both
shapes with no new machinery.
A computed: { expr, dialect } field declaration is reserved in the
spec and the Field type, but no evaluator ships. Results are
defined as never persisted (recomputed on read); readers tolerate
the field's presence and render it empty until an evaluator exists.
Why: the storage decision (computed = derived, never written)
is safe to lock now and prevents a "stored value disagrees with
recomputation" staleness class. The evaluation decisions — which
expression dialect (CEL / JSONata / bespoke), the standard library,
whether cross-row aggregation is ever in scope — need a real consumer
to drive them; committing early bakes a language we can't yet
justify. Same park-with-reserved-shape posture as D14's
history.ndjson. Explicitly out of scope even when built:
spreadsheet-style range references (A1:A10) — position-based
addressing is wrong for an id-keyed row model.
The schema-version counter is meaningful only without concurrent
writers (one device, or git with human-resolved merges). Under
multi-writer sync it is advisory: the sync layer's linearisation
order is authoritative for schema supersession, and consumers must
not treat counter equality as schema equality across replicas —
compare the field set itself. (Resolves issue #45; follows from the
D17 op-log model.)
Why: a monotonic counter cannot converge — two offline peers can both bump 1 → 2 with different schemas, and on sync there are two distinct "v2"s. The honest options were: derive the version from content (loses the human-readable number), hand it to the sync layer (leans on a layer that isn't built), or scope the counter to the substrate where it works. The append-only schema rule (D18) already makes concurrent schema edits commute; only the number was lying. Scoping it is zero-code, honest, and correct for the 1.0 substrate (git / single device). When the sync layer lands, it owns supersession by log position — the counter stays what it is today: a human-facing "the schema changed" signal.
Writer-minted ids are 25 chars of lowercase base36 (0-9a-z,
~129 bits) via customAlphabet — no uppercase anywhere in the
alphabet. Validity stays loose: any non-empty string unique within
the table is a legal id, so hand-authored ids (p1) remain fine.
(Resolves issue #42.)
Why: ids are used verbatim as filenames (bodies/{id}.md), and
the default nanoid alphabet mixes case — on case-insensitive
filesystems (default APFS, NTFS, most sync targets) two ids differing
only in letter case resolve to the same path and silently overwrite
each other's bodies. Silent data loss is the worst failure mode, so
the fix removes the failure class (case leaves the alphabet)
rather than adding a filename-encoding layer both writer and parser
would carry forever. Length went 21 → 25 to keep entropy at parity
with the old base64url alphabet. Breaking for previously-minted
mixed-case ids, which is acceptable pre-1.0 — no production data
exists, and fixtures use hand-authored ids that were never affected.
The loose-validity rule is deliberate (maintainer call): a .table/
outside a managed workspace should be writable by a human or an
agent without opaque-id ceremony.
writeTable stages every file to a <name>.tmp sibling, commits
each with an atomic per-file rename(), and performs deletions
(stale bodies, emptied bodies/) only after every rename has
landed. A failure during staging aborts with the previous table
byte-for-byte intact. (Resolves issue #43.)
Why: the old writer wrote each file in place, sequentially — a
crash, a concurrent reader, or a file-sync client could observe a
new schema.json alongside the old rows.ndjson (a torn table),
and a crash mid-bodies could leave bodies deleted but not
rewritten. For a format whose pitch is local-first files under git
and sync clients, the writer must never make a reader's view
inconsistent. Full-directory swap (staging a sibling directory and
renaming it into place) was considered and rejected for 1.0: it
breaks open file handles and watchers on every save and costs a
directory copy for unchanged attachments; per-file rename shrinks
the torn window from "the whole serialisation" to a handful of
renames, which is proportionate. Durability (fsync) is explicitly
out of scope — the contract is about what readers can observe, not
about power loss; the same posture as the index.sqlite rule
(D15).
parseTable skips a malformed NDJSON line, a row without a system
id, or a malformed optional file, and reports each as a diagnostic
(ParsedTable.diagnostics, ValidationError shape, rowIndex =
zero-based line number, or −1 for file-level). Valid rows always
load. The single fatal case is a missing or malformed schema.json.
(Resolves issue #44.)
Why: the format's pitch is hand-editable, line-diffable text
under git. Fail-fast meant one typo or a merge-conflict marker made
the entire table unreadable — an exception, with no way to see the
surviving data. That punishes exactly the users the plain-text
design courts. Skip-and-collect matches NDJSON's own design (each
line independently parseable, SPEC section 3) and the format's
existing tolerance posture (optional files may be absent; unknown
root files are ignored). Strictness remains available one level up:
a consumer can refuse to proceed when diagnostics is non-empty.
Schema stays fatal because every downstream interpretation depends
on it — degrading there would fabricate meaning.
The on-disk format is frozen at formatVersion: 1, tagged
format-v1. Changes from here are additive only — new optional
fields, annotations, or files that existing readers safely ignore
(the tolerance rules in SPEC section 1 make this cheap). Breaking
changes require a major formatVersion bump. The reference
library's TypeScript API is versioned separately and may still move.
Why now: every open vocabulary question is implemented or reserved-with-a-shape (D18–D21), the storage/sync contracts are written (D15–D17, D22), the reader/writer behaviour contracts are implemented and tested (D23–D25), the architectural review (docs/REVIEW.md) found no remaining freeze-blockers after the access-model reconciliation landed, and the first real consumer (the Workspace app) is waiting on a stable target. A frozen format with an evolving library is the correct boundary: consumers bet on bytes, not on npm semver.
<name>.table.zip contains exactly one root directory,
<name>.table/ (SPEC section 13). The reference reader parses the
zip entirely in memory — never extracting to disk — and the writer
emits byte-deterministic archives. Zip handling is ~200 lines over
node:zlib; no archive dependency. (Resolves issue #56, requested
by the Workspace integration.)
Why nested, not files-at-root: Finder and CLI unzip both then
produce the .table/ directory directly, nothing scatters loose
files into the extraction directory, and the table keeps its name
when the archive file is renamed. Why in-memory: it makes the
zip-slip vulnerability class structurally impossible in the
reference path (entry names are still validated for extracting
consumers), avoids temp-directory lifetime questions, and a
.table/ that fits an email fits memory. Why zero-dep: the
format's pitch is implementable-in-an-afternoon; stored + deflate
over a fixed layout doesn't justify an archive stack. Constraints
accepted: no zip64, no encryption — readers MAY reject >4 GiB
archives.
readTableArchive / writeTableArchive run on Node, browsers, and
React Native (Hermes): no Node globals, bytes-only API (the platform
supplies bytes from a path, fetch, or document picker), string
encoding via fflate's helpers rather than assuming global
TextEncoder/TextDecoder. The raw-deflate codec is fflate (pure JS,
zero transitive deps); the container logic — layout, security
posture, determinism — remains bespoke. Supersedes D27's
zero-dependency claim.
Why: a shared archive arrives on every platform — mail on a
phone, upload in a browser, Finder on a Mac — so a Node-only reader
served exactly one of the format's three UI targets. Hermes has
neither node:zlib nor DecompressionStream, which leaves pure-JS
deflate as the only implementation that runs everywhere; writing and
maintaining our own inflate/deflate is not where this format's value
lives. fflate is small, audited, and dependency-free, and only the
codec crosses the boundary — everything the spec normatively
constrains stays in-repo.