Conversation
P-24.5: indicated hydrogen belonging to an individual ring component is not
carried into the completed spiro system unless the final structure actually
needs it, and the spiro atom never does -- it is the junction, so it bears no
hydrogen. openclatura carried it in regardless, so a fluorene joined at C-9
was cited as 9H-fluorene for a ring assembly whose indicated-hydrogen quota is
zero.
spiro[9H-fluorene-9,1'-cyclopropane]-2-amine
-> spiro[fluorene-9,1'-cyclopropane]-2-amine
spiro[9H-fluorene-9,9'-(9H-benzo[a]indene)]
-> spiro[fluorene-9,9'-benzo[a]indene]
spiro[2H-1-benzopyrane-2,2'-(2H-benzo[b]pyran)]
-> spiro[1-benzopyrane-2,2'-benzo[b]pyran]
Only the citation at the junction goes: a component keeps the rest, so
spiro[...-7,3'-(1H,3H-benzo[b]pyrrole)] joined at 3' keeps its 1H for the NH.
The citation is load-bearing in the name/graph binding metadata, which is why
it cannot simply be deleted. Three things state the same hydrogen and have to
let go of it together: the citation, its hydro operation, and the component's
own name. The spiro atom stays bound because the parent binding covers every
parent atom -- but the parent's term has to be rebuilt from the respelled name,
and the retired operation's own binding removed, or the final audit looks for
text the name no longer carries. The rebuild is deliberately
refresh_parent_binding rather than a full refresh: by this point the spiro
sides have been split out of parts.substituents, so refreshing everything would
drop their atoms and report them unnamed. The side component is respelled at
its source in split_spiro_substituents for the same reason -- its binding term
is its own name.
Measured over the 5000-molecule OPSIN-verified corpus: 10 names change, all 10
still round-trip through OPSIN, none is lost or newly named, and corpus-wide
over-citation drops from 38 names to 36.
One expectation changes. spiro[3H-indole-3,4'-piperidine] becomes
spiro[indole-3,4'-piperidine], which is unambiguous: the alternative completion
with N-1 H and C-2 saturated is a different compound and is named
spiro[indoline-3,4'-piperidine]. Both still verify against OPSIN, and the new
spelling is what the test's own name -- indicated hydrogen follows the graph
tautomer but not a hydrogen-free spiro carbon -- describes.
All 27 QM9 rows that regressed against the stored baseline are restored. None came from the recent indicated-hydrogen work -- all three defects reproduce at f0b5726, before it. They are independent bugs that happen to share the same symptom, a name OPSIN cannot read back. Primed replacement prefix on a dispiro middle component (24 rows). A spiro assembly primes the locants of its non-first components, but a component cited inside the brackets keeps its own name, built from its own numbering. Priming the prefix there detached it from the name it belongs to: dispiro[cyclopropane-1,2'-(1'-azabicyclo[2.1.0]pentane)-3',1''-cyclopropane] is unreadable, where 1-azabicyclo[2.1.0]pentane parses. Only the spiro junction locants around the component carry primes. The enclosing marks are harmless either way, so they are left alone. Pi capacity at a trivalent ring nitrogen (1 row). A neutral ring nitrogen already holding three skeletal bonds has no valence left for a parent pi bond, but the guard saying so ran only when a retained template's indicated hydrogen had been relocated. On the default path the model offered a double bond at pyrrolizine's bridgehead N-4, so an oxo beside it looked as though it had consumed that bond and the added hydrogen was cited on an atom that cannot hold one: 1H-pyrrolizin-5(4H)-one, now 1H-pyrrolizin-5(7aH)-one. Two bridged retained parents pick up lower hydro locants as a consequence -- 2,3-dihydro rather than 3,4-dihydro for 3,5-methano-1H-pyrrolizine, and 1,6,7,8,9,9a rather than 5,6,7,8,9,9a for the ethano/propano quinolizine -- both still OPSIN-exact. Stranded endpoint of an externally consumed pi bond (1 row). Composition that would cite the far endpoint of a bond consumed by an oxo was gated on every carbon in the ring being pi-eligible, so one intrinsic-H carbon elsewhere suppressed it and the endpoint's hydrogen went unnamed -- 6aH-cyclopenta[d][1,3]oxazol-4-one denotes a different tautomer. A ring-fusion endpoint left holding a single uncited hydrogen now forces the composition, giving cyclopenta[d][1,3]oxazol-4(3aH,6aH)-one. The condition is deliberately narrow: a non-fusion CH2 endpoint has a hydro partner to pair with, and widening it past that disturbs the cyclopenta[b]pyrazinone family the original guard protects. Over the 5000-molecule OPSIN-verified corpus only four names change, all four still round-trip through OPSIN, and none is lost or newly named.
dcf644a forced external-pi composition when a ring-fusion endpoint of the consumed bond was left holding an uncited hydrogen. That fixed cyclopenta[d][1,3]oxazol-4-one, but forcing the composition remaximises the parent and can drop an intrinsic carbon-H citation elsewhere, which is what the original all-carbons guard was protecting. On [1,3]dioxolo[4,5-b]pyrrol-6(6aH)-one it dropped the 5H of a ring CH2, leaving a name OPSIN cannot read -- trading one wrong name for another. Both molecules put a single-hydrogen fusion endpoint beside the consumed bond, so that is not the discriminator. Telling them apart needs the composed state to be validated against the hydrogen it actually accounts for, and the indicated-hydrogen citations are not decided until the planner runs, so the check does not belong here. Revert to the state that at least names [1,3]dioxolo[4,5-b]pyrrole correctly and leave the oxazolone defect standing, recorded rather than half-fixed. The dispiro replacement-prefix and trivalent-nitrogen fixes in dcf644a are unaffected: 26 of the 27 original QM9 regressions stay fixed.
The engine proves hydrogen accounting per molecule, so a change to one of its guards can be right on the molecule that motivated it and wrong across the family that guard governs. The rest of the suite does not notice, because it pins individual names rather than families: two recent commits passed all 7366 tests while being wrong on molecules the suite does not contain, and only an external evaluation caught them. Enumerate one such family from first principles -- ortho-fused bicycles with an oxo on a carbon beside a ring-fusion atom, swept over ring size, heteroatom placement and every choice of parent pi bonds, restricted to ordinary-valence neutral species. 13507 members, none drawn from a dataset, with OPSIN as the oracle, so this stays an independent check rather than a snapshot of current behaviour. 252 members do not round-trip today and are pinned in tests/data/fused_oxo_family_known_failures.json. A member outside that list failing is a regression and fails the test; shrinking the list is progress. The gate is real: re-applying a change that had looked correct on its motivating molecule and passed the whole suite makes this test fail with 720 broken members, naming them. The default run samples 2500 members and takes about a minute. Set OPENCLATURA_RUN_FULL_FUSED_OXO_SWEEP=1 to sweep all 13507. The pinned failures fall into three classes, all pre-existing: 79 where the name understates hydrogen at a ring-fusion carbon and so denotes a different tautomer, 17 that render a malformed singleton hydro prefix such as "5-hydro", and 156 others not yet classified.
A fused mancude ketone saturates two ring atoms: the suffix carbon and the far endpoint of the parent pi bond it consumed. When that endpoint is a ring-fusion atom the name never mentioned it, so 6aH-cyclopenta[d][1,3]oxazol-4-one dropped the hydrogen on C-3a and read back through OPSIN as a different tautomer. Two conditions in the hydrogen-site pass assumed a non-fusion atom: A site counted as an indicated-hydrogen position only when its ring bonds summed to two, which is the saturated state of an atom with two ring bonds. A ring-fusion atom has three, so a declared indicated hydrogen sitting on one never entered the candidate pool, the surplus came out one short, and the P-14.7 rule that cites the remaining saturated site never fired. A fusion site now counts too, but only when it actually holds hydrogen -- a bridgehead nitrogen has three single ring bonds and none, and counting it dropped the hydro prefix from 5,6,7,8-tetrahydroimidazo[1,2-a]pyridine. The test for whether a fusion carbon still wants a hydro prefix counted one the name already declares. With fusion sites in the pool that split the pool as though a declared hydrogen were surplus, spelling an impossible "2,3a,6a-trihydro". Declared sites are spelled already and are excluded. Measured on the enumerated family in tests/roundtrip/test_fused_oxo_family.py: 133 of its 13507 members gain a correct fusion name and none regresses, so the pinned failures drop from 252 to 119. These are real fusion names, not a fallback: cyclopenta[d][1,3]oxazol-4(3aH)-one, 3aH-inden-7(4H)-one, 4aH-cyclopenta[b]pyridin-7(7aH)-one, 1H-pyrrolizin-5(7aH)-one. Over the 5000-molecule corpus two names change, both still OPSIN-exact, none lost. Both earlier reported QM9 regression sets stay fixed.
P-3 treats a retained partially saturated parent such as indoline as stating a definite hydrogen population: Table 3.1 defines it as 1H-indoline, the same compound as 2,3-dihydro-1H-indole, and a different hydrogenation state has to go back to the mancude system and use hydro/dehydro terms. Every parent hydrogen therefore stays unless a named operation replaces or consumes it -- substitution and suffixes do, dehydrogenation does not. The template matcher never checked that, so indoline's N-1 accepted a ring C=N and O=C1C=CC=C2CCN=C12 came out as indolin-7-one. Those are not the same compound: indolin-7-one is C8H9NO and the molecule is C8H7NO, two hydrogens poorer, and the name expresses no dehydrogenation. It is 3H-indol-7(2H)-one. Only a template that fixes saturation of its own is constrained. A mancude parent is exempt because P-58.2.3.1 lets its indicated hydrogen be reassigned to accommodate a suffix, including onto the carbon the suffix operates at -- the Blue Book's own 7H-1-benzopyran-7-one. So 1H-carbazol-1-one keeps its imine N-9, where the single indicated hydrogen of 9H-carbazole has moved to C-1. The check is deliberately about dehydrogenation, not about hydrogen being present: 2-methylindoline and N-substituted indolines legitimately replace a parent hydrogen, and a ring double bond is the one thing no operation in the name accounts for. Narrow by design. One member of the 13507-molecule family sweep changes, none regresses, the 5000-molecule corpus is untouched, and the pinned failures drop from 119 to 118. An earlier attempt to reject any donor arriving as an imine cleared 107 of them, but that was collateral -- it discarded sound retained parents and stumbled onto better routes -- and it cost carbazole, benzotriazole, purine and 1,5-benzodiazepine their retained names.
…n bear P-58.2.3.1 lets a mancude parent's indicated hydrogen be cited at another position of the ring system to accommodate the structure. A retained name that spells it where the molecule carries a ring double bond asserts a hydrogen that is not there, and nothing moved it: 4H-1,4-benzothiazine over an imine N-4 rendered 3,4-dihydro-1,4-benzothiazin-5(2H)-one for a compound whose only saturated ring positions are C-2 and C-3. It is 2H-1,4-benzothiazin-5(3H)-one. Move the citation to a position that is actually saturated, lowest locant first. A divalent chalcogen is skipped: benzothiazine's S-1 is a donor with no room for hydrogen, and relocating onto it merely moved the false assertion. The relocated locants then count as declared, or the hydro pass cites one of them a second time and spells "2,4a,5,8a-tetrahydro-2H-...". This is the same principle as the preceding commit seen from the other side. A retained partially saturated parent may not quietly change its hydrogen population; a mancude parent may move its indicated hydrogen, but only onto a position the structure can hold it. Measured on the 13507-molecule family sweep: 106 members gain a correct name and none regresses, so the pinned failures drop from 118 to 12. The 5000-molecule corpus is untouched and the suite is unchanged.
A component name records which tautomer of the *isolated* component was
cited: "pyran" is 2H-pyran, and its spec marks that carbon saturated. A
saturated atom has no pi capacity, and a shared fusion atom takes the
minimum over the components that meet there -- so when the cited carbon
lands on a ring fusion, the fused parent inherits a junction that can
never bear a double bond.
That is what furo[3,4-b]pyran does. Pyran fuses at its 2,3-bond, so
pyran's C-2 is C-7a of the fused system. Frozen, it strands both bonds
incident on it; C-7 then has no carbon neighbour left, and the entire
furan half of the parent drops out of the pi system. Every derivative
loses C-7a's hydrogen without saying so:
O=C1C=COC2COCC12 -> furo[3,4-b]pyran-4(4aH,5H)-one
for a compound whose furan ring is fully saturated. OPSIN reads that
name back as a different structure.
P-25.7.1.3 assigns indicated hydrogen to the *completed* ring system,
after maximising non-cumulative double bonds over the whole skeleton.
component_parent_atoms already released such a carbon safely -- only
when the release leaves the component's own mancude bond count intact,
so the component keeps its identity. The defect was the outer gate,
which refused to release whenever releasing raised the completed
system's pi budget. Raising it is precisely the recomputation the rule
asks for, so the gate is gone.
This empties the fused-oxo family's known-failure list: 13507 members,
252 -> 0. The 5000-molecule corpus is unchanged, name for name.
The gate's budget mismatch was also the only trigger the late
junction-release path ever had, so released_fusion_carbon_sites now
finds nothing to reconcile -- candidates and graph agree by
construction. The machinery is left in place (its scope=False branch is
untouched) and its test now builds the pre-relaxation state explicitly.
The existing family sweep always hangs an oxo beside a ring-fusion atom,
so it only ever exercises suffix-driven hydrogen accounting. This one
drops the oxo and varies the ring alone -- element placement and which
parent pi bonds are present, from fully mancude down to fully saturated.
That makes it the direct gate on how a fused parent hydride is chosen
and how its indicated hydrogen and hydro prefixes are cited. 9426
members, enumerated from first principles rather than any dataset, with
OPSIN as the oracle.
It found 16 failures the oxo sweep is structurally unable to see. Two
fixes clear 12 of them:
Hold the locant the stem already spells. The additive-hydrogen branch
took the lowest locant as the indicated hydrogen while the stem spelt a
different one, leaving the spelt position in the hydro prefix as well,
so the name cited one position twice:
3,4,6,7,8,8a-hexahydro-4H-1,4-benzothiazine
hydrogenates the N-4 that the 4H already saturates. It now reuses
_declared_indicated_hydrogen_split, whose docstring describes this same
failure for the sibling branch below it.
Do not emit a hydro prefix that cites nothing. The parent's own
indicated hydrogen can account for every saturated site, leaving no
surplus; the empty prefix reached the multiplier table as "hydro" of
zero and raised, so 2H,8aH-1,3-benzodioxine got no name at all. That was
a regression from c76d7bf, which the oxo sweep could not have caught.
The four still pinned are all fully mancude systems with no indicated
hydrogen, where several maximum double-bond arrangements of one skeleton
exist. IUPAC names the ring system rather than an arrangement, so the
name cannot distinguish them and OPSIN returns a different Kekule form
of the same compound. Verify before removing them: the generated names
look correct already.
Both sweeps compared canonical SMILES for equality, but the engine's own round-trip check does not: verify_with_opsin goes through resonance_compare.equivalent_smiles, which already accepts an alternate Kekule drawing. The sweeps were therefore stricter than the product they test, and reported four names as failures that the engine itself calls verified. Those four are the aza-pentalenes -- cyclopenta[b]pyrrole, cyclopenta[c]pyrazole, pyrrolo[3,2-b]pyrrole, pyrrolo[3,4-b]pyrrole. Eight ring atoms pair into four non-cumulative double bonds with none left over, so no position is sp3 and the name carries no indicated hydrogen with which to pick one arrangement over the other. Two maximum arrangements exist; OPSIN returns the one we did not write. No name could have made these compare equal, and none needed to. Reusing equivalent_smiles rather than adding a second comparison keeps one policy for the product and its tests, and inherits the stricter guard: it pins isotope, formal charge, radical count and hydrogen count per atom and rejects charge-separated contributors, so a wrong tautomer is still a failure. Confirmed against this branch's own defects -- all six pre-fix furo/thieno-pyranone tautomers are still rejected, and their targets have a unique Kekule form, so the tolerance changes nothing for them. It is not separately unit-tested here because it already is, in tests/unit/test_kekule_resonance_comparison.py. Both families now sweep clean with no molecule excused: 13507 + 9426 members, both known-failure lists empty.
P-54.4.3.2: indoline is not a preferred IUPAC name; the PIN is the mancude
parent plus hydro prefixes, 2,3-dihydro-1H-indole. Both templates already
carried a complete policy saying exactly that -- preferred name, base
parent, hydro locants and the indicated-hydrogen locant -- but scoped to
"unsubstituted_parent". That is why the bare ring system was already right
while every derivative reverted to the retained stem, and why isoindoline,
which has no retained template to fall back on, was right throughout.
Widening the two contexts to "all" is the whole of the data change.
The fixed 2,3-dihydro-1H- split is not right once a suffix stands on a
hydrogenated position. P-58.2.3.1 gives that position the indicated
hydrogen, so oxindole is 1,3-dihydro-2H-indol-2-one. The hydrogenated set
is invariant -- only its division between citation and prefix moves -- so
RetainedHydrogenationPolicy.relocated rebuilds the split rather than
recounting hydrogen, and additive.py applies it once numbering exists,
which is the first point the suffix locant is known. Only a suffix that
takes the position's hydrogen counts: a ring ketone does, while
indoline-2-carboxylic acid hangs off C-2 and stays 2,3-dihydro-1H-.
Removing the two templates instead was measured and rejected. It changed
71 corpus names against this change's 28, left 30 on a non-preferred
3H-indole parent, regressed bare indoline, and dropped four molecules off
retained naming altogether into 2,3-dihydro-1H-cyclopentabenzene and
benzo[b]pyrrole -- systematic fusion names standing in for indene and
indole, which are themselves PINs. Keeping the templates and demoting only
their output avoids all of it.
The reconstruction audit had to learn the new spellings or it would abstain
on names it used to confirm ("retained parent '2,3-dihydro-1H-indole' not
modelled"): they denote the template's own structure, so they resolve to it.
Corpus: 28 of 5000 renamed, no errors, none stops round-tripping, no
retained indolin/indan stem left. Spiro components now cite the preferred
component name, each OPSIN-verified before its expectation was updated.
P-24.5.1 orders the two components of a ring-assembly spiro name alphanumerically, so spiro[9H-fluorene-9,1'-cyclopropane]-2-amine is spiro[cyclopropane-1,9'-fluorene]-2'-amine. The order decides which component is primed, so every locant either one contributes moves with it: substituents, suffixes, stereodescriptors and hydro prefixes. Two constraints fix where the decision belongs. It needs the parent's own component name, which parent_stem_and_terminal can render early, and it has to precede _hoist_side_substituent_prefixes, which moves the side's operations into the parent's record already primed. Deciding later left 5'-methylidene where 5-methylidene belongs. Both hold in split_spiro_substituents, before the record is touched. An assembly being named as a component of an enclosing spiro keeps parent-first: the outer system numbers against the name returned for it, and re-orienting it there silently drops a component from the outer name. spiro_subgraph marks that projection so the reordering declines. The order is decided from the component parent hydride names. Indicated hydrogen and hydro prefixes describe the completed system and are cited outside the brackets once it is built -- P-24.5.1's own 1'H-spiro[imidazolidine-4,2'-quinoxaline], P-31.2.3.3.3's hexahydro-1'H-spiro[...] -- so neither decides which component is cited first, and 1H-pyrrolo[3,2-b]pyridine sorts under "p". P-14.5 compares Roman letters first and admits locants only as a tie-break, so the key is a pair. The further P-24.5.3 discriminators are not modelled. Five places assumed the side component is always the primed one and had to learn otherwise: the replacement-prefix and inline-prefix helpers, the italic element locant of an N-substituted suffix (N', which OPSIN reports as no such locant), the binding refresh, and enclosing marks, which the first cited component must not carry. The sixth was the token-span resolver, and there the assumption is gone rather than taught: a spiro side's replacement prefixes reached the name as bare strings, so "5-aza" had no binding and the resolver guessed from punctuation -- a locant followed by a prime addresses the other component -- which could only ever return that whole component. They now carry the provenance the side's substituents always had, so the locant resolves exactly, to the nitrogen rather than to the ring, and in either order. Corpus: 12 of 5000 renamed, no errors, none stops round-tripping.
__version__ now comes from ._version, which owns both the distribution lookup and its PackageNotFoundError fallback, so the importlib.metadata import left behind in the package __init__ names nothing.
Naming the 5000-molecule corpus takes 53.9s against 28.2s before the
fusion engine, and the cost is concentrated: 37% of molecules reach the
fusion planner and account for 81% of the time, while the slowest 1.2%
account for 42%. One molecule alone took 2.25s.
That molecule builds 4612 citation candidates and rendered every one,
because the rendered name was the last element of the candidate's score.
The planner consumes exactly one: it returns on the first candidate whose
numbering audit passes, and across a 200-molecule sample 80 of 87
enumerations yield a single AST. Ordering by the structural score and
then by the rendered name gives the same sequence the combined score did,
so rendering now happens only inside a tie -- and a generator the caller
abandons never renders past the group it reached.
The layout search validates every partial extension, so its per-call work
is multiplied by the search tree. It rescanned model.faces to resolve a
face id, and tested face adjacency with any({a, b} == {c, d} ...), both
inside loops over faces; those are a dict and a frozenset now. Its edge
crossing test ran four cross products on every pair of drawn edges, 7.2
million of them per thousand molecules; pairs whose bounding boxes are
disjoint cannot cross and are rejected before the exact predicate.
Re-stamping prime depth allocated through dataclasses.replace, which
re-derives the field list on each call; ComponentLocant and FusionSide
carry three fields and are now named directly, same class and same
__post_init__. A component whose requested template is the one its spec
already holds returns that spec instead of rescanning its templates --
verified to return an identical spec for all 60 registered components.
Corpus 53.9s -> 47.9s with no name changed: the 5000 generated names are
identical, 7369 tests pass, and both family sweeps still sweep clean.
Two pure functions of a component spec were recomputed on every citation candidate. component_sides rebuilds the directed peripheral walk from peripheral_order, and interface classification reads it once per call -- 98549 calls across a thousand molecules, over fifteen distinct specs. component_variant_identity re-derives a structural identity by sorting the template's bonds and rings, and keys several groupings per candidate. A spec is frozen and its template is immutable, so both results are properties of the object rather than of the call. They are retained the way the spec already retains its seniority key: computed at most once per instance, reset by dataclass replacement, and available to the first molecule of a cold process rather than accumulating over a run. The variant identity alone drops the engine's sorted() calls from 3.8M to 2.2M per thousand molecules. Interface classification also built its attached-atom inverse map on entry, before four checks that reject most interfaces outright; it is built where it is first read. Corpus 47.9s -> 45.2s, and 53.9s -> 45.2s across the series. All 5000 generated names are unchanged, 7369 tests pass, both family sweeps sweep clean.
Selecting a citation needs each candidate's score; it does not need the name tree behind it. The planner returns on the first candidate whose numbering audit confirms, so of 5382 candidates built across a thousand molecules it reads 61. The score's join-preference key is (join.order, side_rank, attached locant text), and stamping prime depth changes none of those -- it rewrites locants' prime_depth and revalidates each rewritten interface. So the score can be taken from the unprimed joins, and the stamping, the descriptors and the citation plan move behind a thunk the ordering runs only for a candidate it yields. A tie still materialises its group, because breaking it needs the rendered names. This is worth little on ordinary molecules, which build a handful of candidates: the corpus is 45.2s either way. It is worth a third of the worst case, where one system builds 4612 candidates across five pools and drops from 1.47s to 1.01s. Keeping it for that: the engine should not spend a second of a caller's time constructing name trees for citations it has already decided against. All 5000 generated names are unchanged, 7369 tests pass, both family sweeps sweep clean.
d3cb0bb added component_bond_locant_sets, which memoises a component's undirected local bonds on the spec, but staged only descriptor.py. The field it writes through object.__setattr__ was left in the working tree, so at d3cb0bb every read of spec._bond_locant_sets raises AttributeError inside the planner, which catches it as a failed citation and abstains. The effect is silent: name_smiles returns an empty string for every fused ring system rather than failing loudly. furo[2,3-b]furan and benzofuro[2,3-b]pyridine both name correctly again with the field back. That commit is already on origin/fusion_nomenclature_base, so this follows it rather than amending it.
Each of these derives from something the surrounding loop holds fixed, so the work is repeated without the answer ever changing. system_locant_sort_key parses a locant label into its typed form on every comparison a sort makes. Fusion numbering sorts constantly and one ring system spells a handful of labels, so 824k calls draw on a few hundred distinct strings. It is memoised the way retained_locant_sort_key beside it already is, which wraps this very function for this very reason. _match_all_retained_fused_template asked whether the template fixes its own saturation from inside the candidate comprehension, which visits every locant-atom pair: a template property evaluated O(n^2) times per match, 2.3M times over the corpus. It is now settled once per match. _partial_layout_is_valid rebuilt the face index and the face-adjacency set on each of its 114k calls, though both derive from the FaceModel alone -- the comment above them already claimed they were built once per model. FaceModel now retains them like FusionComponentSpec retains its sides. _best_tree_mapping_candidate walks the Cartesian product of the children's local maps, and proved each child's join inside that product. A join is proved from the child's own map and the host's, never from a sibling's, so the same proof was re-derived once per combination of the siblings ahead of it. Each map is now proved on the first combination that reaches it and reused for the rest of that visit, which is scoped to one selected host map. The state counter still counts every combination, so the search budget trips on exactly the state it did before. Interface classifications fall 273k -> 162k and any() calls 9.9M -> 5.3M. The 5000-molecule corpus goes 42.7s -> 38.7s with all 5000 names unchanged; 7369 tests pass.
find_ring_systems routed every fused block through plan_fusion_parent so the block could be numbered from the fusion proof instead of a von Baeyer descriptor search. Routing is what keeps that search off the large fused systems that cannot afford it, and test_polycycle_routing pins exactly that: a nine-triangle block must not reach _polyspiro_or_von_baeyer_ candidate. Only a block with a third ring ever reaches that search. A bicyclic block is proved by _proven_monospiro_or_bicyclo_system, which is cheap and complete, so planning one here bought no routing at all -- it bought a full fusion proof for every fused bicycle in every molecule, whether or not that block was ever going to be the parent. Naphthalene, quinoline, indole and benzofuran each paid for a fusion proof and then took a retained name. The proof is not lost where it is wanted: parent_pipeline plans the selected parent itself, and the planner memoises on the molecule, so the molecules that are named from a fusion parent still are. All 611 names this branch improves over 53c1076 are unchanged, and so are the other 4389. _confirmed_fusion_bicycle_descriptor only ever served the bicyclic route and goes with it. _partial_layout_is_valid keys an undirected pair as the ordered tuple of its ends rather than a frozenset, compares a repeated edge against both orientations instead of building two sets, walks the face cycle by index instead of concatenating a rotated copy, and iterates the pairwise loops by index instead of copying a list slice per step. Same predicate, same answers. Corpus 38.7s -> 30.4s, base 24.1s, so the gate's paired median falls from +62.16% to +25.97%. 7369 tests pass; all 5000 names unchanged.
number_parent chose the preferred numbering by comparing paths: its comparator took two paths and scored both, so the running best was re-scored against every remaining candidate. Scoring is the expensive half -- heteroatom priorities, indicated hydrogen, hydro sites, bond locants and a stereochemistry walk -- and 48% of the 255k scorings the corpus performed were an exact repeat of one already done for the same molecule. The comparator now takes two scores and the loop keeps the best one alongside the best path, so each candidate is scored exactly once. The comparison rules are untouched, including the empty-criterion handling that keeps this from being a plain tuple order. Scorings fall 255k -> 155k, 2.81s -> 1.86s. _match_all_retained_fused_template asked whether an atom is chemically admissible for a locant before asking whether their degrees even agree. Both tests are pure, the degree test is two dict lookups, and the chemistry test walks the atom's bonds and element policy, so the cheap one now goes first: 2.3M chemistry tests become 1.6M for the same candidates. _best_tree_mapping_candidate sorted each leaf's children on four criteria, two of which ask only about a child's own spec and so cannot vary anywhere in that search. They are settled once per visit. Gate paired median +25.97% -> +21.43%. 7369 tests pass, all 5000 names unchanged, and all 611 the branch improves over 53c1076 still improve.
… once The layout search extends an arrangement one face at a time and re-proved the whole thing on every extension: every drawn edge against every other, every face against every other. The faces already placed keep their proof, because the only thing that happens to them between depths is a rescale by one positive factor, and that moves no crossing and no containment. So an extension now proves only the pairs its new face introduces, with the settled faces laid down first so those pairs are a suffix of each list. Callers that have nothing to stand on -- the final audit -- omit the argument and prove everything, as before. The incremental answer was checked against the full one on all 64242 calls the corpus makes. parse_locant is memoised on the rendered text rather than on its argument. Memoising the argument would be wrong: DisplayLocant is an int subclass, so DisplayLocant(4, "4a") hashes and compares equal to 4 while having to parse differently. The text is the real argument, and the pool of texts is tiny next to the 932k calls. lexical_tokens is memoised outright. Assembly re-tokenises the same stems, prefixes and rendered fragments while composing a name, 1.1M times over the corpus, and it is a pure function from text to an immutable tuple. Gate paired median +21.43% -> +14.44%, which is the first reading under the 15% threshold. 7369 tests pass, all 5000 names unchanged, all 611 improvements over 53c1076 intact.
Choosing a parent numbering compares candidates on eight ordered criteria and settles on the first that differs. Over the corpus 96% settle within three of them: 39% on heteroatom priorities, 56% on principal locants. The key was materialised whole regardless, so most candidates had their hydro sites split, their bond locants collected, their substituent citations ordered and their stereochemistry walked so that a comparison already decided could ignore all of it. number_parent now scores through a preference that holds what a criterion is derived from instead of the criteria, and a key that derives each entry the first time it is compared and keeps it for the rest of that candidate's comparisons. The criteria, their order and the comparison rules are untouched. NumberingPreference stays as it was. It is the documented shape of a score and is built from values elsewhere; the deferred form is for the search, which builds one per candidate and reads three entries. The deferred key was checked entry for entry against the eager one on all 155,580 keys the corpus builds. Gate paired median +14.44% -> +10.22% (26.57s against a 24.17s base). 7369 tests pass, all 5000 names unchanged, all 611 improvements over 53c1076 intact.
match_retained_graph_templates memoises its result per molecule under a key that carries the family and the admission flags, because the matches depend on them. The topology key does not: it reads the atom and bond graph alone. Every family and flag combination asked for the same block therefore rebuilt it, and 63.5% of the 86k the corpus builds were a repeat of one built for the same atoms moments earlier. It is now kept on the molecule beside the match cache, keyed on the atoms and cleared with it whenever the graph changes. Corpus 25.97s -> 25.35s; gate paired median +10.22% -> +9.74%. 7369 tests pass, all 5000 names unchanged.
…ing walk _independent_face_sets treated a full edge cover as proof that the face set was complete, and returned. That holds for a catafused system, where every face owns at least one perimeter edge. A pericondensed one can enclose a face entirely: every edge it has is shared with a neighbour, so it contributes none of its own and the cover completes before it is chosen. The search then abandoned the branch one face short of the cycle rank and found nothing at all. Coronene is the smallest case, and today it has no face model whatsoever. It still gets named only because "coronene" is a retained template, so nothing asks the fusion engine; aza-coronene, or a B/O-doped nanographene, has no path at all. With the enclosed faces reachable, coronene resolves to 7 faces over 6 interior atoms and 12 audited layouts. When there is no uncovered edge to pivot on, the faces not yet chosen are the candidates. The independence and twice-covered tests already decide which of them can belong to the model, so nothing else changes: all 5000 corpus names are identical, and the search costs 0.06s over the corpus. The other half is the failure that exposed this. Once fusion nomenclature declines a large fused block, find_ring_systems falls through to the von Baeyer descriptor search, whose main-ring walk enumerates every simple cycle from every atom -- exponential in the cycle rank, and the one search in the engine with no budget. A 21-ring block did not finish. The widest block in the corpus spends 332k states, so the walk is bounded at 2M, a sixfold margin over anything von Baeyer nomenclature is really asked to describe. Exhausting it abstains exactly the way finding no cycle does, which the caller already handles. The nanographene that ran without end now abstains in about two minutes. 7369 tests pass; all 5000 names unchanged; timing unchanged (interleaved A/B: 30.82s with, 31.03s without, 30.22s with).
…annot get_von_baeyer_descriptor_and_path finds the main ring by enumerating every simple cycle from every atom and keeping the longest. That is exponential in the number of cycles: measured on fused hexagon clusters it costs 1.99x per added ring, so it reaches about rank 13 inside its two-million-state budget and a 65-atom, 21-ring block needs on the order of 1e9 states. Such a block ran for 658 s. The cycle space is only as wide as the cycle rank. Every cycle of a graph is a GF(2) sum of the fundamental cycles of a spanning forest, so exclusive-or over the 2**rank subsets of that basis reaches every cycle the walk reaches -- two million subsets for that same block rather than a billion states, one exclusive-or and one population count each, and a subset is only inspected when it beats the longest cycle already proved. The 21-ring block now yields its main ring in 1.04 s: the same 60-atom ring the walk took 658 s to find. This is the answer, not an approximation. The basis spans the whole cycle space, and a subset is kept only when its edges really do form one simple cycle. Nothing here assumes planarity, so it covers cages as well as fused ring systems -- the test agrees with the walk on adamantane, cubane and norbornane alongside the benzenoids. A perimeter shortcut would not have done: this block's perimeter is 40 atoms and its main ring is 60, which weaves through the interior. It runs only where the walk gives up. The walk picks its main ring from the cycles in enumeration order and keeps the first of equal rank, an order a cycle basis cannot reproduce, so replacing it outright would quietly change which ring wins a tie. Entering on _CycleWalkExhausted instead makes "no change to anything that already works" structural rather than something to test for; all 5000 corpus names are unchanged and the gate is unmoved at +9.60%. Bounded at rank 28 -- about 2.7e8 subsets, tens of seconds -- so it abstains honestly rather than waiting, on the same terms as the walk's own budget. The new test compares the two searches on every system where both are affordable: hexagon clusters from 2 to 12 rings, peri-fused arenes, a bowl whose pentagon is enclosed, and three cages. All twenty-one agree. One more confirms a seventeen-hexagon block, past the walk's reach, now produces a descriptor. 7386 tests pass.
build_desc cites secondary bridges by decreasing length and then decreasing
bridgehead locant, and the reconstruction reads each bridge's interior atoms
out of the numbering in that order. add_extra_nodes fed the numbering from
the bridge list instead, which is sorted by length alone and breaks ties by
the order components were discovered. The two agreed only while no two
secondary bridges were the same length. Once two were, they entered the
numbering in the opposite sequence from the one the descriptor named, and
the reconstruction joined each one's interior to the other's bridgeheads.
Seventeen fused hexagons is the smallest cluster that leaves two secondary
bridges of equal length, and its descriptor rebuilt six bonds between the
wrong atoms; sixteen leaves one, and passed. The bridgeheads of a secondary
bridge are main-ring atoms, so their locants are settled before any interior
is appended, and ordering the interiors by the descriptor's own key needs
nothing the base path does not already fix.
Two things this recovers. A dense 25-atom cage of cycle rank 12 was pinned
by a test as naming to the empty string -- it was the cycle walk exhausting
its budget, and with the cycle-space main ring behind it the cage now names
as dodecacyclo[15.3.1.1^{3,17}...]pentacosane, which OPSIN parses back to
exactly that structure. The test now asserts the name.
And the molecule this started from -- a 140-atom B/O-doped nanographene whose
65-atom, 21-ring core is a corannulene-type bowl that fusion nomenclature
cannot orient -- now names, in 120s:
9,23,33,51,59-pentakis(2,4,6-tris(propan-2-yl)phenyl)-2,12,26,36,54,56,62,
63,64,65-decaoxa-3,13,27,37,55-pentaborahenicosacyclo[55.3.1....]
pentahexaconta-1(60),4,6,...,58-pentacosaene
Its bracket carries 22 numbers for 21 rings, summing to 63 for a 65-atom
skeleton, over a main ring of 55+3+2 = 60 -- the longest cycle in the block.
7393 tests pass; all 5000 corpus names unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.