Make bracketed span conditional on valid attributes - #143
Conversation
Per the docs, text in square brackets that is not a link or image is
treated as a generic span only when followed immediately by an
attribute. The parser committed the span matches as soon as it saw
'[...]{', so when the attribute parse subsequently failed (or hit end
of input), the braces were reparsed as literal text but the span
remained: '[x]{#a<b}' rendered as '<span>x</span>{#a<b}'.
Track the speculative span and revert its open/close matches to
literal brackets when the attribute parse fails, so the whole
construct stays literal: '[x]{#a<b}'.
|
On reflexion, there may have been a reason I didn't originally do this, connected to performance. |
|
I think there are two separate questions here: what the rule costs as a spec rule, and what it costs in djot.js. They come out differently. On the spec level you are right: "span only when followed by a valid attribute" means a For djot.js specifically, though, this PR does not add any lookahead. The full attribute parse plus backtracking already happens today for every element followed by The pathological cases are also already guarded against, independent of this PR:
Growth is linear on both branches and the difference is within noise. So whatever performance reason there originally was, I do not think it applies to the current code base: the expensive part (attempt + reparse) has been there all along. That leaves the question as a purely semantic one, and there I see the trade like this:
Worth noting the relaxation would also reverse the guidance from #137 and jgm/djot#399 (where the fully-literal rendering was called the intended one), and djot-php and the cdjot-derived tests would then need to change to match djot.js instead of the other way around. I am fine with either outcome - just giving it some more context. Personally, I think the stricter approach makes more sense and is more clear to the writer. |
|
OK, good, thanks for the additional analysis. |
Follow-up to the discussion in #137: per the docs, "Text in square brackets that is not a link or image and is followed immediately by an attribute is treated as a generic span." The parser committed the span matches as soon as it saw
[...]{, so when the attribute parse subsequently failed (or ran into end of input), the braces were reparsed as literal text but the span itself remained:rendered as
With this change the span is tracked as speculative and its open/close matches are reverted to literal brackets when the attribute parse fails, so the whole construct stays literal:
Inline content between the brackets is still parsed normally (
[*x*]{#a<b}gives[<strong>x</strong>]{#a<b}), and the end-of-input case ([x]{#a) is covered via the same reparse path. Valid attributes are unaffected.Tests: three new cases in
test/spans.test(invalid attribute, invalid attribute with inline formatting, unterminated attribute at end of input). Full suite passes.Also resolves the discussed bug in jgm/djot#399 here.