Status: Draft / pre-implementation. This document specifies the source-text
format that rawast grammars are written in. The engine bootstraps via JSON
grammar files; .rawast is the canonical hand-authoring format. The two
serialisations describe the same in-memory grammar data; swapping between
them is a loader-level concern that does not affect the engine, the .jast
container, or any downstream consumer of parsed data.
A .rawast file is a sequence of rule definitions, each defining one
named non-terminal of a grammar. The definitions are loaded together into a
Grammar object — the same in-memory data model produced by loading the
equivalent JSON grammar file. .rawast and .json grammar files are two
serialisations of the same data; either can produce any grammar the engine
supports.
A small example:
start: <VALUE>
VALUE: choice {
<STRUCT>,
<LIST>,
string,
int,
float,
"null":null,
"true":true,
"false":false
}
LIST: sequence array {
"[",
repeat <VALUE> separator ",",
"]"
}
PAIR: sequence {
string:@=,
":",
<VALUE>
}
STRUCT: sequence dict {
"{",
repeat <PAIR> separator ",",
"}"
}
That defines a complete JSON grammar in 22 lines. The JSON-form equivalent is roughly 50 lines and considerably less readable.
UTF-8. The engine treats input as a byte stream; non-ASCII bytes inside quoted literals and identifiers are preserved verbatim.
Whitespace (space, tab, carriage return, newline) is allowed between any two tokens and is consumed by the engine's ignore-list mechanism. There is no significant whitespace; indentation has no meaning.
Two comment styles are supported, both placed in the ignore list:
// line comment to end of line
/* block comment, possibly
spanning multiple lines */
identifier ::= [a-zA-Z_][a-zA-Z0-9_]*
Identifiers are case-sensitive. The convention used throughout this document and the bundled grammars:
- UPPER_CASE identifiers name grammar rules (e.g.
VALUE,STRUCT). - lower_case identifiers name terminal parsers registered on the
grammar (e.g.
int,float,string,identifier).
The convention is not enforced by the engine — rule and parser names live in separate registries — but mixing them within a file is discouraged.
Two quoting forms, both with backslash-pass-through for embedded
quotes. The forms have distinct semantics in KEY_EXPR (see §4.2 and
§4.2a) but are otherwise syntactically interchangeable wherever the
grammar reads a string token:
"hello" // double-quote — byte-prefix Key
"with \"embedded\" quotes"
"//\\/* etc */"
'hello' // single-quote — strict (word-bounded) Key
'with \'embedded\' quotes'
The backslash itself is preserved verbatim (this is the engine's existing
pass-through escape behaviour from DoubleQuoteStringParser /
SingleQuoteStringParser; higher-level escape interpretation is a future
concern).
The following identifiers have grammatical meaning and may not be used as parser names or as bare-identifier expressions:
sequence choice repeat separator array dict null true false
indent tab space newline tail use
The first row are the structural keywords (§4.6–4.8 and §4.3 constants). The second row are the save-side pretty-print postfix attributes (§4.6).
A future revision may add optional and other keywords.
@ is a placeholder meaning "the value at this position." It is used
in binding suffixes to flow the result of an expression into the
surrounding catcher.
@= is the dict-key binding marker: emit the value as the next
dict key (equivalent to is_name=true on the producing node).
Both are explained fully in §4.
A .rawast file may declare which terminal-parser groups it needs
by listing them at the top of the file:
use: gdsii
use: standard
start: <LIBRARY>
LIBRARY: sequence dict { ... }
Multiple groups can be combined comma-separated:
use: gdsii, standard
When the loader encounters a use: directive it looks up each named
group in a process-wide registry of parser-group factories and applies
each to the target Grammar before processing any rule definitions. If
a referenced group has not been registered, the load fails with a clear
diagnostic ("use: parser group 'X' not registered") — much friendlier
than the cryptic parse-time failure that would otherwise occur when an
unknown parser is referenced.
Built-in groups (compile-time registered in librawast):
| Group name | Provides |
|---|---|
std |
Common terminal parsers — identifier, qualified_identifier, int, uint, float, string, whitespace, line_comment, block_comment |
gdsii |
All 47 GDSII binary record parsers (header, bgnlib, …) — referenced bare or group-qualified (gdsii.header) |
lefdef |
LEF/DEF-specific identifier (accepts hyphens, slashes per real-world naming) and line_comment (# to EOL) |
tcl |
Tcl terminals modelled on Dodekalogue rules — hspace, newline, comment, brace_group, quoted_string, bracket_sub, bare_word, expand_marker, var_name, until_paren, escape, literal_run |
Additional groups can be registered from host C++ code via
rawast::register_parser_group("name", register_fn); see
include/rawast/parsers_registry.hpp. Runtime plugins loading
.so/.dll files at startup is a potential M5+ extension; the static
registry is sufficient for M1–M4 scope.
A .rawast file is a sequence of rule definitions:
rule_definition ::= identifier ':' expression
Rule definitions are separated by whitespace only — no explicit terminator. Each definition registers its identifier as a named rule.
The special pseudo-rule start designates the top-level entry point of
the grammar:
start: <VALUE>
start must be a reference (<NAME>) to another rule.
A rule definition may carry an ignore-list attribute between the rule name and the colon. The named parsers become the active ignore set whenever the parse driver is inside that rule (or any rule it calls that doesn't carry its own override):
// Grammar-level "default ignore" — attach it to the start rule.
start: <SCRIPT>
SCRIPT ignore tcl.hspace: sequence dict { ... }
// Override on a different rule — its sub-tree uses a different
// ignore set. Useful when one part of the grammar needs to treat
// newlines as significant (here: command separators) and another
// part needs them ignored (here: expression context).
EXPR ignore tcl.hspace tcl.newline: sequence dict { ... }
// Explicit empty list — override to "ignore nothing", useful for
// token-internal contexts where whitespace is part of the data.
WORD_SEGMENTS ignore: sequence dict { ... }
The list is space-separated, not comma-separated, and is
terminated by the :. Empty list (RULE ignore: …) explicitly
overrides to "ignore nothing"; rules without the attribute inherit
their caller's active ignore.
Inheritance semantics: the engine maintains an ignore-stack. On
entering a rule with an explicit override, the override is pushed;
on exit it is popped, restoring whatever the caller had active. A
rule with no ignore attribute simply uses the top of the stack.
This composes naturally — a SCRIPT-context rule that calls a sub-
rule sees the sub-rule run under SCRIPT's ignore, unless the sub-
rule overrides.
The standalone top-level ignore: parser, parser, … directive
(from earlier proposal drafts) is gone — attaching the ignore
to the start rule is the single canonical way to express grammar-
level default ignore. Existing grammars (json, gdsii, lef, def)
were migrated to the new form in lockstep with this change.
Combined with the :subparse="<RULE>" binding (§4.5), the
ignore-stack lets a single grammar express multiple sub-languages
each with their own whitespace policy. The Tcl grammar uses this
pattern: a script context (tcl.hspace ignored, newlines
structural), an expression context (tcl.hspace + tcl.newline
both ignored), and a word-internals context (nothing ignored)
all in one file.
expression ::= reference
| literal
| literal_with_constant
| parser
| binding
| sequence_expr
| choice_expr
| repeat_expr
| optional_expr
reference ::= '<' identifier '>'
A reference to another rule defined in the same file. References are resolved lazily at grammar-load time, so forward references are allowed:
A: <B>
B: "ok"
literal ::= string
A literal token matched byte-for-byte from the input. Produces a Key node. By default, literals are structural — they consume the input text but emit no value to the surrounding catcher.
"[", "{", ":"
strict_literal ::= single_quote_string
A literal that additionally requires a word boundary after the
matched bytes. The match succeeds only when the byte immediately
following the literal is not a word character (ASCII alphanumeric
or _) — or when the literal lands at end-of-input. Produces a Key
node with Node.strict = true.
'not' // matches "not", "not foo", "not\n", "not;"
// does NOT match the prefix of "notch" or "notty"
'END' // matches "END;" or "END\n"
// does NOT match the prefix of "ENDSECTION"
The boundary check fires only when the last character of the literal
is itself a word character. For punctuation literals like '+', '(',
';', the strict form has no effect (and there is no reason to write
them as strict) — the engine still emits the Key node but skips the
boundary check at parse time.
When to use which form.
Reach for "text" (byte-prefix, current default) when:
- The literal is punctuation or an operator that may legitimately
abut following content (
"+","<",";"). - Adjacent tokens are packed without whitespace and the literal is a
short opener that the next item will consume past (e.g.
"["in a bracketed list). - You're capturing a fixed prefix that introduces opaque content
(
"LEF58_":@for vendor-extension property names).
Reach for 'text' (strict, word-bounded) when:
- The literal is a closed-keyword that must not be confused with an
identifier — language reserved words (
'if','not','while'), spec keywords ('END','MACRO','PIN'). - A grammar Choice has alternatives whose literals share a prefix
(
'SPACING'vs'SPACINGTABLE'); strict mode lets the engine pick the right alternative without depending on hand-ordering. - The same identifier-suffix can legitimately appear elsewhere in the
grammar (
'do'in a language that also accepts identifiers likedoneordoc).
Round-trip.
The save direction preserves the surface form. A grammar source that
used 'token' will save back to 'token'; one that used "token"
saves back to "token". Internally the discriminator is the
Node.strict flag carried through the parse-Node-save chain.
JSON-form equivalent.
In JSON form, the strict variant is expressed by either of these equivalent shapes — the loader accepts both:
{"type": "strict_key", "key": "not"}
{"type": "key", "key": "not", "strict": true}The DSL 'not' always lowers to the first form; the second form
remains available as an escape hatch for machine-generated JSON
that prefers a single type value.
literal_with_constant ::= string ':' constant
constant ::= 'null' | 'true' | 'false' | '@' | other_literal
A literal whose successful match emits a constant into the surrounding catcher. This is the discriminator pattern: the literal text in the input selects a branch, and the constant identifies which branch was taken.
"null":null // matches "null", emits the null singleton
"true":true // matches "true", emits the true singleton
"false":false // matches "false", emits the false singleton
"sequence":@ // matches "sequence", emits the string "sequence" itself
The @ form is sugar for "emit the matched literal as a string value." It
saves repeating the literal:
"sequence":@ // equivalent to "sequence":"sequence"
A bare identifier (not in reserved words, not enclosed in <...>) names
a terminal parser registered on the grammar. The dotted form
group.identifier is the same parser referenced via its group-qualified
alias — equivalent in every respect to the bare form:
int // SignedIntParser (from `use: std`)
uint // UIntParser
float // FloatParser
string // DoubleQuoteStringParser
identifier // IdentifierParser
qualified_identifier// QualifiedIdentifierParser (ident('.' ident)*)
gdsii.header // GdsiiHeaderParser — same as `header` once `use: gdsii`
The set of available parsers is determined by what's been registered
when the grammar is loaded — primarily by use: directives at the top
of the file. Each group member is registered under both its bare
local_name and the group.local_name alias, so either form resolves
to the same parser.
binding ::= expression ':' bind_target
bind_target ::= '@' -- pass value through (default)
| '@=' -- emit value as dict key
| identifier ['[]'] '=' bind_rhs -- emit (identifier, rhs) pair into dict
bind_rhs ::= '@' -- parsed expression's value
| literal -- constant value
literal ::= string | int | float | 'true' | 'false' | 'null'
The optional [] after identifier turns the binding into a
list-append — see §4.5a.
The binding suffix wraps an expression and controls how its produced value is routed into the surrounding catcher.
X:@ is the no-op binding — X produces a value, the value flows
into the catcher unchanged. Rarely needed in practice; the :@ can be
omitted entirely.
X:@= flags the value as a dict-key name (equivalent to setting
is_name = true on X). The next non-name value emitted will be paired
with this name in the surrounding dict catcher.
PAIR: sequence {
string:@=, // the parsed string becomes the dict key
":",
<VALUE> // the next value becomes the corresponding dict value
}
X:name=@ emits the pair (constant string name, value-of-X) into
the surrounding dict catcher. Useful for fixed field names:
NREPEAT: sequence dict {
choice {"sequence":@, "choice":@}:type=@, // type field
?choice {"array":@, "dict":@}:container=@, // optional container field
<ITEMS>:items=@ // items field
}
X:name=<literal> emits the pair (constant string name,
literal-value) into the surrounding dict catcher, discarding X's
parsed value. The literal can be a string, int, float, true, false,
or null. The expression X is still parsed (its side effect — typically
matching a discriminator token — is what makes the surrounding rule
fire); only its produced value is discarded.
This is the discriminator-with-constant-emit idiom, the natural pattern for binary formats whose record types are zero-payload markers that label the surrounding element. Example from GDSII:
BOUNDARY: sequence dict {
gdsii.boundary:type="boundary", // match the discriminator; emit ("type","boundary")
?gdsii.elflags:flags=@,
?gdsii.plex:plex=@,
gdsii.layer:layer=@,
gdsii.datatype:datatype=@,
gdsii.xy:xy=@
}
Parser references support two surface forms — bare (boundary) and
group-qualified (gdsii.boundary). Both resolve to the same parser:
the registry registers each group member under its bare local_name
and under group.local_name. The dotted form is parsed as a single
token by std.qualified_identifier, so the AST shape stays flat
({type: "gdsii.boundary"}).
In the JSON-grammar-format equivalent, the binding desugars to a
wrapper dict carrying "type": "binding" (for name=@) or
"type": "binding_const" (for name=<literal>); the loader expands
either into the appropriate Value-kind children. See §8.
X:subparse="<RULE>" is a special binding key recognised by the
engine. After X (a Parse-terminal returning a StringValue) succeeds,
the engine re-invokes the parse loop on the captured string with the
named rule as the new start. The resulting sub-tree replaces X's
string value. Same grammar, same engine, different entry point —
composes languages-within-languages without splitting them into
separate grammar files:
// Inside a hypothetical Tcl-script context:
IF_CMD: sequence dict {
"if":type="if",
tcl.brace_group:cond=@:subparse="EXPR", // brace content re-parsed
tcl.brace_group:body=@:subparse="SCRIPT" // through EXPR / SCRIPT
}
The engine resolves the subparse target's rule name to a NodeId at grammar-load time; a missing rule fails the load with a clear error. Subparse triggers create a fresh ignore-stack so per-rule ignore overrides in the inner context don't leak into the outer parse.
The plain name=value binding writes (or overwrites) a single key
on the surrounding dict. The name[]=value form appends
each match to a list under name instead. Subsequent matches of
the same name accumulate into the same list rather than overwriting
the previous one.
PIN_PROPERTY: choice {
<PIN_DIRECTION>,
<PIN_USE>,
...
<PIN_ANTENNA_PROP>:antennas[]=@ // each ANTENNA match appends one entry
}
For the input
PIN A
ANTENNAGATEAREA 0.01 ;
ANTENNADIFFAREA 0.02 LAYER met1 ;
ANTENNAPARTIALMETALAREA 0.05 LAYER met1 ;
the engine produces
PIN A: {
...
antennas: [
{kind: "GateArea", value: 0.01},
{kind: "DiffArea", value: 0.02, layer: "met1"},
{kind: "PartialMetalArea", value: 0.05, layer: "met1"},
],
}
The [] is the grammar-author signal — it never appears in the
output dict; the key is antennas, not antennas[]. Mechanism: at
parse time, the loader folds the [] suffix back into the binding
name string; the dict-assembly stage strips it and lazily
instantiates an ArrayValue under the base name on first hit.
Semantics worth knowing:
- Lazy creation. If no match fires, the key is absent from
the dict. A
name[]binding never produces an empty list as a default — the field is simply not there. Downstream code should model the field asOptional[list[X]](or its equivalent). - Catcher-context only.
name[]is meaningful only when the binding ends up in a dict-container's catcher; the engine has no other use for it. - No collision with
name. A grammar that mixesname=@andname[]=@on the same key in the same dict is malformed — the scalar write will overwrite the list (or vice versa) depending on match order. The linter does not yet catch this; treat the two forms as mutually exclusive per name.
The same primitive is what enables multi-OBS per MACRO, multi-
PROPERTY-per-PIN, multi-FOREIGN-per-MACRO, and any other
spec-allowed multi-instance pattern in lefdef.rawast.
Inside a sequence body, the bare * token consumes raw input
bytes from the current cursor until the next sibling literal
matches at the cursor — without consuming it. The captured prefix
is emitted as a StringValue; the next sibling (which must be a
literal "…" key) is then matched normally in its own iteration.
BEGINEXT_BLOCK: sequence dict {
"BEGINEXT":type="BeginExt" space,
string:name=@ newline,
*:body=@, // captures every byte up to (but not including) "ENDEXT"
"ENDEXT" newline
}
For the input
BEGINEXT "spec_vendor"
vendor_directive option_a = 42 ;
another_thing "string" ;
ENDEXT
the parsed dict has body set to the literal text between the
BEGINEXT opener and the ENDEXT closer, embedded whitespace and
newlines included.
Constraints. * is only meaningful inside a sequence body and
must be followed by a Key literal in that same sequence — the
literal tells the engine where to stop scanning. The loader rejects
the grammar at load time if either condition is violated; the
linter flags the same issue with a friendlier message during
rawast lint <grammar>. Common error messages:
- "raw consume (
*) must be followed by a literal key in the same sequence; nothing follows here" —*was the last item in the sequence with no sibling after it. - "raw consume (
*) must be followed by a literal key in the same sequence; next sibling is not a Key node" — the item after*is a reference, a parser, another*, etc. — anything that isn't a"…"literal.
Ignore-set bypass. * does not run the surrounding rule's
ignore list before scanning. Whitespace, newlines, and # line
comments that would normally be skipped are part of the captured
payload. This is what makes the round-trip byte-exact: whatever was
between the two literals comes back out unchanged on save.
Save side. The save direction pulls the StringValue bound to
the * and writes it verbatim; the following Key emits its
literal. No special bookkeeping — the round-trip is mechanical.
Each item in an items list may carry zero or more pretty-print
postfix attributes after its expression (and after any binding). They
are pure save-side metadata: the parse direction ignores them entirely.
postfix_attr ::= 'indent' -- depth+1 for this Node's scope
| 'tab' -- emit depth × indent_step before content
| 'space' -- emit ' ' after content
| 'newline' -- emit '\n' after content
| 'tail' '=' string -- emit string after content (escape-interpreted)
Attributes are space-separated and may appear in any order. The save direction applies them in a fixed order, regardless of source order:
[ tab → depth × indent_step ] [ content ] [ tail ] [ space ] [ newline ]
indent bumps the save-time depth counter for the entire scope of
the Node it sits on. Nested indent flags accumulate. The bump happens
before tab fires, so an item with both indent and tab emits the
indent at the new (bumped) depth — convenient for the common
"indented-line" pattern.
tab emits the current depth's indentation (depth × indent_step,
where indent_step is a Grammar-level setting, default two spaces). It
fires at the beginning of the Node's content. Place tab only where
indentation should actually appear in the output — typically at the
start of lines.
space and newline emit " " and "\n" respectively after
the content. They are the common-case sugars; for anything else (;\n,
\\\n, ; , custom separators) use tail.
tail="..." emits the string after content, with C-style escape
sequences interpreted at load: \n, \t, \r, \\, \", \0.
Convention: use the bare newline flag for plain newlines; reserve
tail for non-newline strings or for combinations like backslash-
newline (tail="\\" newline).
Examples:
"{" newline, // "{" then "\n"
":" space, // ":" then " "
"PIN" space, // "PIN" then " "
",;" tail="; ", // (illegal — combining attrs needs space/newline keywords)
"$\\" newline, // emit `$\` then "\n" — but use tail="\\" newline
identifier tail=";" newline, // identifier then "; then "\n"
<PAIR> tab indent, // emit indent before PAIR; depth+1 inside
repeat <PAIR> tab indent separator ",", // each iteration indented at depth+1
The classic JSON-pretty pattern uses indent on the iterated item and
puts a trailing empty-key + newline before the closing brace:
STRUCT: sequence dict {
"{" newline,
repeat <PAIR> tab indent separator "," newline,
"" newline, // trailing newline before "}"
"}" tab // close brace at outer depth (auto-indented)
}
In the JSON-grammar-format equivalent, each pretty-print attribute is a
field on the item dict: {"type": "key", "key": "{", "newline": true},
{"type": "key", "key": "}", "tab": true}, etc. The engine's Node
data type stores all five attributes uniformly; the two surface
serialisations are interchangeable.
indent_step is a Grammar-level setting (default " "). Switch to
tabs via Grammar::set_indent_step("\t") in C++ before parsing.
Runtime compact/pretty toggle. Grammar::save(out, value, pretty)
takes a final pretty parameter (default true). When set to false,
the save direction skips tab, indent (no depth bump), and newline
emissions but still emits space and tail (since the grammar author
may rely on them for round-trip parseability — e.g. a space between two
adjacent identifiers). One grammar covers both pretty and compact
output; no separate "compact grammar" is required.
sequence_expr ::= 'sequence' container? '{' items '}'
container ::= 'array' | 'dict'
items ::= expression (',' expression)*
A sequence of sub-expressions matched in order. The optional container
keyword annotates the surrounding catcher behaviour:
- no container: child values flow through unchanged to the parent's catcher (transparent grouping).
array: at end of frame, accumulated values are materialised into anArrayValue.dict: at end of frame, alternating (name, value) pairs are materialised into aDictValue.
Items inside the braces are separated by commas. Trailing commas are permitted.
choice_expr ::= 'choice' '{' items '}'
Ordered alternation: alternatives are tried in source order. Each
alternative attempt is wrapped in input-cursor mark() / reject() —
if an alternative partially matches and then fails, the input position
is restored and the next alternative is tried from the same position.
This is standard PEG ordered-choice semantics and applies to every
Choice node by default; there is no opt-in attribute and no opt-out.
Alternatives with overlapping first-token signatures parse correctly
via the alt-failure recovery above. The linter emits informational
warnings on such Choices when the LL(k) lookahead can't prove
disjointness within bounded depth — see Grammar::lint() and
docs/AGENTS.md. The warnings are design feedback (the alt-failure
cost is real, even if small), not bugs. There is no flag to suppress
them; either restructure the alternatives to diverge earlier, or
accept the warning as a permanent design note.
repeat_expr ::= 'repeat' ('+' int?)? expression ('separator' expression)?
Iteration of the given expression, optionally separated between iterations by the separator expression. Produces no container of its own; the surrounding sequence's container catches the iteration results.
The quantifier suffix sets a minimum required iteration count:
repeat <X> // min=0 (zero-or-more)
repeat+ <X> // min=1 (one-or-more)
repeat+2 <X> // min=2 (at-least-two)
repeat+5 <X> // min=5 (at-least-five), etc.
repeat+ is shorthand for repeat+1. The save direction canonicalises
back to the same surface form on round-trip (min=1 emits repeat+,
min=N for N≥2 emits repeat+N).
Examples:
repeat <VALUE> separator "," // 0+ values, for arrays
repeat <PAIR> separator "," // 0+ pairs, for dicts
repeat+ <ITEM> // 1+ items (classic PEG `+`)
repeat+2 <ARG>:args[]=@ separator "," // at least 2 args (e.g. a binary op)
optional_expr ::= '?' expression
Zero-or-one match of the given expression. On miss, the expression contributes nothing to the parent's catcher.
?<MODIFIER> // optional modifier
?choice {"array":@, "dict":@}:container=@ // optional field
At parse time, each frame on the engine's parse stack maintains a
_values list of emitted values streaming up from its children. Frames
whose grammar node has container=Array or container=Dict materialise
their accumulated values into an ArrayValue or DictValue at end-of-frame.
For dict containers, accumulated values must alternate name (with
is_name=true) and value entries. The order in which children are
written in the .rawast source must produce names and values in the
correct alternating order.
For dict containers used to build grammar-format trees (like NREPEAT
above), the name=@ and @= binding forms make it straightforward to
emit explicit (name, value) pairs.
start: <VALUE>
VALUE: choice {
<STRUCT>,
<LIST>,
string,
float,
int,
"null":null,
"true":true,
"false":false
}
LIST: sequence array {
"[", repeat <VALUE> separator ",", "]"
}
PAIR: sequence {
string:@=,
":",
<VALUE>
}
STRUCT: sequence dict {
"{", repeat <PAIR> separator ",", "}"
}
The convention used by the bundled grammars (and recommended for new
ones) is to name the start rule after what the parsed value tree
actually is — VALUE for JSON, LIBRARY for GDSII / LEF, DESIGN
for DEF, SCRIPT for Tcl, GRAMMAR for the rawast meta-grammar. A
CSV file is a table of rows, so TABLE:
start: <TABLE>
TABLE: sequence array {
repeat <ROW> separator "\n"
}
ROW: sequence array {
repeat <FIELD> separator ","
}
FIELD: choice {
<QUOTED_FIELD>,
unquoted_field
}
QUOTED_FIELD: string
// unquoted_field is a custom terminal parser, registered separately.
A self-hosting fragment from the prototype's json.ast:
NREPEAT: sequence dict {
choice {"sequence":@, "choice":@}:type=@,
?choice {"array":@, "dict":@}:container=@,
<ITEMS>:items=@
}
REPEAT: sequence dict {
"repeat":type=@,
?choice {"array":@, "dict":@}:container=@,
<ITEM>:item=@,
?sequence {"separator", <ITEM>:separator=@}
}
CMD: sequence dict {
?"?":optional=true,
choice {<NREPEAT>, <REPEAT>}
}
Every .rawast construct maps to a Node in the in-memory grammar tree.
The mapping is direct:
.rawast construct |
Node |
|---|---|
NAME: EXPR |
Registers EXPR's node under NAME in the grammar |
<NAME> |
Node{kind=Ref, value="NAME"} |
"text" |
Node{kind=Key, value="text"} |
"text":CONSTANT |
Key with a Value-kind child holding CONSTANT |
identifier |
Node{kind=Parse, value="identifier"} |
X:@= |
X with is_name=true |
X:name=@ |
A sub-Sequence emitting (name-marker name, X) |
X:name=<literal> |
(X parser, name-marker name, Value-kind literal) — X still parses, its value discarded; literal emitted instead |
sequence { items } |
Node{kind=Sequence, container=None} + children |
sequence array { items } |
Node{kind=Sequence, container=Array} + children |
sequence dict { items } |
Node{kind=Sequence, container=Dict} + children |
choice { items } |
Node{kind=Choice} + children |
repeat X |
Node{kind=Repeat} + [X] |
repeat X separator Y |
Node{kind=Repeat, has_separator=true} + [Y, X] |
?X |
X with is_optional=true |
X indent |
X with depth_in=true (depth+1 for its scope) |
X tab |
X with indent_emit=true (emit indent at content) |
X space |
X with space_after=true (emit " " after) |
X newline |
X with newline_after=true (emit "\n" after) |
X tail="..." |
X with tail string (escape-interpreted) |
After loading, the in-memory grammar tree is indistinguishable from one
loaded from an equivalent JSON file. The .jast container format does
not record which serialisation produced its bundled grammar, because
the data model is the only thing the container needs.
This format spec is finalised at the design level. Implementing the
.rawast loader requires:
- A
.rawastgrammar definition (in JSON form initially, since the engine boostraps from JSON). The grammar produces anValuetree shaped identically to a JSON grammar file's parsed output. - The JSON-grammar loader — walks an
Valuetree and constructs aGrammarvia the builder API. - An
IdentifierParserterminal ([a-zA-Z_][a-zA-Z0-9_]*) with a "not followed by an identifier character" check, used to match reserved words (sequence,choice,repeat, etc.) without consuming partial identifiers likesequencing. - Value-kind nodes with
is_name=true— the engine currently absorbs Value-kind children into the catcher with hardcodedis_name=false. To support thename=@binding form, the absorber needs to honour the Node'sis_nameflag.
Once these four pieces land, .rawast files can be loaded and the
engine self-hosts: future grammars are written in .rawast syntax,
parsed via the JSON-loaded .rawast grammar, and contribute to the
community grammar repository.
The following were open during drafting and are now settled:
- Trailing commas inside
{ items }lists are allowed. Improves editor ergonomics; doesn't introduce ambiguity. start:may appear anywhere in the file. Not required to be the first rule definition. Convention: place at the top for readability, but the loader doesn't enforce it.- String literals are single-line. Embedded literal newlines are
rejected. Newlines inside strings can be expressed via escape
passthrough —
"\n"is the two characters backslash-n, the intended payload for higher-level escape interpretation. - Unterminated block comments are a parse error. Reaching EOF
inside a
/* …comment fails the parse with a max-progress error pointing at the comment's opening position. - Pretty-print postfix attributes (
indent,tab,space,newline,tail="...") attach to the preceding expression with no separator. They are part of the same item; commas separate sibling items (§4.5b). The save direction emits them in the fixed order tab → content → tail → space → newline. indentbumps depth beforetabfires on the same Node, so<X> tab indentemits the indent at the bumped depth — the natural reading of "this item starts a deeper line."tailstrings are escape-interpreted at load time.\n,\t,\r,\\,\",\0are recognised; unknown escapes pass through verbatim. Use the barenewlineflag for plain newlines and reservetailfor non-newline content (or for combinations liketail="\\" newlinewhich emits backslash-newline for line-continuation grammars).
Rule reference for the .rawast meta-grammar itself, generated by
walking grammars/rawast.rawast and emitting EBNF for each rule.
Regenerate after grammar changes:
rawast docs grammars/rawast.rawast --title Rules --heading-level 3 \
>> docs/rawast-format.mdConventions:
"literal"— keyword/punctuation matched by thekeynode kind.Name— reference to another rule in this file.*name*— terminal parser supplied by theuse:group (e.g.*string*,*identifier*come fromstd).A?optional,A*zero-or-more,A+one-or-more.A ( sep A )*— a separator-form repeat.A | B | C— choice.
A companion view describing the value-tree shape the grammar produces
(dict fields, array element shapes, choice alternatives, optional /
constant bindings) is available via rawast schema:
rawast schema grammars/rawast.rawast --title "rawast — value-tree shape" \
> docs/rawast-shape.mdThe two views share a single source (the grammar dict): rawast docs
reads it as input syntax for consumers; rawast schema reads it as
output shape for producers.
Uses: std
Start: GRAMMAR
BINDING := ":" BIND_EXPRBINDINGS := BINDING+BIND_EXPR := NAME_BIND | VALUE_BINDBIND_VAL := "@" | CONSTANTCHOICE_EXPR := "choice" "{" ITEMS "}"CONSTANT := *string* | *float* | *int* | "null" | "true" | "false"CONTAINER_KIND := "array" | "dict"EXPR := SEQUENCE_EXPR | CHOICE_EXPR | REPEAT_EXPR | KEY_EXPR | REF | PARSE_EXPRGRAMMAR := GRAMMAR_ENTRY*GRAMMAR_ENTRY := USE_DECL | START_DECL | RULE_DEFIDENT_LIST := *qualified_identifier* ( "," *qualified_identifier* )*IGNORE_LIST := *qualified_identifier**ITEM := "?"? EXPR BINDINGS? POSTFIX_ATTR*ITEMS := ITEM ( "," ITEM )*KEY_EXPR := *string* KEY_VALUE?KEY_VALUE := ":" ( "@" | CONSTANT )NAME_BIND := "=" BIND_VALPARSE_EXPR := *qualified_identifier*POSTFIX_ATTR := "indent" | "tab" | "space" | "newline" | TAIL_ATTRREF := "<" *identifier* ">"REPEAT_EXPR := "repeat" "+"? ITEM SEPARATOR?RULE_BODY := RULE_IGNORE_ATTR? ":" "?"? EXPR BINDINGS? POSTFIX_ATTR*RULE_DEF := *identifier* RULE_BODYRULE_IGNORE_ATTR := "ignore" IGNORE_LISTSEPARATOR := "separator" ITEMSEQUENCE_EXPR := "sequence" CONTAINER_KIND? "{" ITEMS "}"START_DECL := "start" ":" START_VALSTART_VAL := REFTAIL_ATTR := "tail" "=" *string*USE_DECL := "use" ":" IDENT_LISTVALUE_BIND := *identifier* "=" BIND_VAL