Repository navigation
Conversation
…ntary-plane character CIDSubset.getChars builds a char[] with StringBuilder.appendCodePoint, so a character outside the BMP occupies two slots. PDFToUnicodeCMap derived each character selector from the array position, so every selector after such a character was written one too high, and the text after an emoji or a mathematical letter extracted as its neighbour. Measured on this branch, DejaVu Math TeX Gyre, the text "A" U+1D400 "BZ". The content stream uses selectors 3 4 5 6; the CMap said before <0003> <0041> <0004> <d835dc00> <0006> <0042> <0007> <005a> after <0003> <0041> <0004> <d835dc00> <0005> <0042> <0006> <005a> Before, selector 5 (the B) had no entry and selector 6 (the Z) was published as B. The CMap is now built from one destination per selector, a String, so a surrogate pair is one entry of length two rather than two array positions. The char[] constructor converts (toDestinations) and callers are unchanged. The range logic needs no surrogate special cases any more: an entry may join a bfrange when it is one code point, and two entries are consecutive when their code points are and their selectors share a 256 block. A destination of several characters is written as a bfchar with a string, which the format allows and FOP-3345 uses. PDFToUnicodeCMapTestCase pinned the drift (surrogatePairTest expected the entry after the pair at 0x63 to be 0x65); its expectations change accordingly, in surrogatePairTest, surrogatePairRangeTest, surrogatePairsRangeTest and rangeSizeSurrogateTest, the last of which also ran its low surrogates past U+DFFF and now starts them at U+DC00. ToUnicodeCharacterisationTestCase records the writer's output for the common shapes, so a change to the range packing has to be deliberate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…itten to ToUnicode as the radical
With complex script features enabled, CJK text is painted correctly but extracts,
searches and reads out as the Kangxi radicals which share the ideographs' glyphs.
One FO, one font (Source Han Sans CN), stock FOP, only the complex-script flag
changed:
features on 0003=U+2F63 0004=U+2F45 0005=U+2F08 0006=U+FF0C 0007=U+724B
features off 0003=U+751F 0004=U+65B9 0005=U+4EBA 0006=U+FF0C 0007=U+724B
The glyphs drawn are identical either way - glyph 3,4,5,6,7 at x=0,12,24,36,48,
advance 1 each. U+724B, which no radical shares, is right either way: the control.
MultiByteFont.performSubstitution runs the font's layout tables over characters -
chars to glyphs, GSUB, then mapGlyphsToChars - and mapGlyphsToChars takes each
glyph's character from findCharacterFromGlyphIndex, which returns the first code
point in the cmap mapped to that glyph ("if more than one correspondence exists,
then the first one is returned"). A CJK font maps a radical and the ideograph it
is the radical of to one glyph: in Source Han Sans CN, U+2F63 and U+751F are both
glyph 18742, U+2F45 and U+65B9 are both 14819, U+2F08 and U+4EBA are both 8966.
The radical is the lower code point, so the reverse lookup hands back the radical,
and that is what reaches the PDF's ToUnicode.
mapGlyphsToChars now prefers the character the glyph came from, which the
GlyphSequence already carries in its CharAssociation, wherever the substitution
left the glyph alone - the association covers exactly one character and the
character map maps that character to this same glyph. Anything the substitution
did produce, a ligature or a glyph with no character of its own, still goes
through findCharacterFromGlyphIndex as before, and a supplementary plane
character still comes back as its surrogate pair.
MultiByteFontTestCase covers all four cases against a hand-built character map
holding the three radical/ideograph pairs above, so it needs no font installed.
Without the fix its first case fails with the radicals for the ideographs, which
is the defect above.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ae8dWi7jSKijUffjpHWdr4
…stands for
MultiByteFont.performSubstitution maps characters to glyphs, runs GSUB, and maps the
glyphs back to characters, because layout works on characters. A glyph that substitution
produced has no character of its own unless the font's cmap happens to map one to it, so
mapGlyphsToChars minted a private-use code point for it (createPrivateUseMapping, from
U+E000). The painter hands that code point to the subset, and the ToUnicode CMap is built
from the subset's characters. So wherever a ligature, a contextual form or a decomposition
was applied, the PDF's text layer said U+E000 upward, or a presentation form, rather than
the letters: search, copy and paste and screen readers lose the text, and nothing in the
rendering shows it.
Measured on this branch, text extracted with pdftotext:
Carlito, "fifty ti office"
before U+FB01 U+E000 "y " U+E001 " o" U+FB03 "ce"
after "fifty ti office"
Noto Sans Arabic, U+0639 U+0644 U+064A U+0643 U+0645
before U+FEDC U+FEE2 U+E001 U+E000 U+FECB U+FEE0 (presentation forms, private use)
after U+0643 U+0645 U+E001 U+0639 U+0644 U+064A
Every word box is identical before and after: the code point used for layout is unchanged,
so no glyph, advance or position moves.
The association each glyph carries out of substitution already names the characters it
came from. MultiByteFont now records them per glyph (getGlyphMeaning) as it maps glyphs
back, CIDSet.getUnicodeSequences returns one string per selector (the recorded characters,
or the selector's own code point), and PDFFactory builds the CMap from those, which the
writer of FOP-3346 accepts: a ligature glyph publishes its letters as a bfchar with a
string destination, an Arabic contextual form its letter.
What it deliberately does not do. The second and later glyphs that a decomposition
produces from one character keep their private-use code point (the U+E001 above), because
a CMap cannot say that several glyphs share one character, and an empty destination reads
as U+FFFD, a space or a raw control character depending on the reader. A glyph seen with
two different meanings in one document publishes neither. And the stand-in glyph drawn for
a missing character records no meaning.
The private-use mapping itself is untouched, so MultiByteFont.hasPrivateUseSubstitutions()
(FOP-3337, AFP) answers as before.
This branch carries two other fixes it is built on: FOP-3346 (the CMap writer takes one
destination per selector) and FOP-3340 (an unsubstituted glyph maps back to the character
it came from), each with its own pull request.
MultiByteFontTestCase: a ligature records its characters, a contextual form its
character, an unsubstituted glyph nothing, the second glyph of a one-character cluster
nothing, a glyph with two meanings neither, the missing-character stand-in nothing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…precomposed letter The first glyph of a multiple substitution recorded the source character for ToUnicode. A font whose ccmp decomposes precomposed letters (Cambria Regular: ά into alpha and a tonos mark, ü into u and uni0308) then had its base glyph recorded as the precomposed letter, and since a plain letter reaches the same glyph through the cmap and records nothing, every plain α, o, e, u, c and A drawn with that glyph was published accented. Where one character is split into exactly as many glyphs as its canonical decomposition (NFD) has characters, each glyph now records its piece in order: the base its plain letter and the mark its combining character, even a mark no character maps to. Any other split keeps the first-glyph rule, so an Arabic letter drawn as a dotless base and its dots (no canonical decomposition) is unchanged. Tests: MultiByteFontTestCase, three cases; the two decomposition cases fail without the change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Pushed 6968306, a correction to this change found on a corpus of real documents. The first glyph of a multiple substitution recorded the source character. Cambria Regular's Where one character is split into exactly as many glyphs as its canonical decomposition (NFD) has characters, The line
The drawing is byte-identical; only the ToUnicode CMap changes. |
|
@plutext Jason, Great work using Claude to ferret out long-standing bugs in FOP, especially as relate to complex scripts, fonts, and Unicode! |
Fixes FOP-3345.
MultiByteFont.performSubstitution maps characters to glyphs, runs GSUB, and maps the
glyphs back to characters, because layout works on characters. A glyph that substitution
produced has no character of its own unless the font's cmap happens to map one to it, so
mapGlyphsToChars minted a private-use code point for it (createPrivateUseMapping, from
U+E000). The painter hands that code point to the subset, and the ToUnicode CMap is built
from the subset's characters. So wherever a ligature, a contextual form or a decomposition
was applied, the PDF's text layer said U+E000 upward, or a presentation form, rather than
the letters: search, copy and paste and screen readers lose the text, and nothing in the
rendering shows it.
Measured on this branch, text extracted with pdftotext:
Every word box is identical before and after: the code point used for layout is unchanged,
so no glyph, advance or position moves.
The association each glyph carries out of substitution already names the characters it
came from. MultiByteFont now records them per glyph (getGlyphMeaning) as it maps glyphs
back, CIDSet.getUnicodeSequences returns one string per selector (the recorded characters,
or the selector's own code point), and PDFFactory builds the CMap from those, which the
writer of FOP-3346 accepts: a ligature glyph publishes its letters as a bfchar with a
string destination, an Arabic contextual form its letter.
What it deliberately does not do. The second and later glyphs that a decomposition
produces from one character keep their private-use code point (the U+E001 above), because
a CMap cannot say that several glyphs share one character, and an empty destination reads
as U+FFFD, a space or a raw control character depending on the reader. A glyph seen with
two different meanings in one document publishes neither. And the stand-in glyph drawn for
a missing character records no meaning.
The private-use mapping itself is untouched, so MultiByteFont.hasPrivateUseSubstitutions()
(FOP-3337, AFP) answers as before.
This branch carries two other fixes it is built on: FOP-3346 (the CMap writer takes one
destination per selector) and FOP-3340 (an unsubstituted glyph maps back to the character
it came from), each with its own pull request.
MultiByteFontTestCase: a ligature records its characters, a contextual form its
character, an unsubstituted glyph nothing, the second glyph of a one-character cluster
nothing, a glyph with two meanings neither, the missing-character stand-in nothing.
Commits
Three commits: the first two are FOP-3346 (#114) and FOP-3340 (#109), which this change is built on and which have pull requests of their own; the third is this fix. Once those two are merged this branch will be rebased to the one commit.
🤖 Generated with Claude Code