Repository navigation
Conversation
…ntary-plane character CIDSubset.getChars builds a char[] with StringBuilder.appendCodePoint, so a character outside the BMP occupies two slots. PDFToUnicodeCMap derived each character selector from the array position, so every selector after such a character was written one too high, and the text after an emoji or a mathematical letter extracted as its neighbour. Measured on this branch, DejaVu Math TeX Gyre, the text "A" U+1D400 "BZ". The content stream uses selectors 3 4 5 6; the CMap said before <0003> <0041> <0004> <d835dc00> <0006> <0042> <0007> <005a> after <0003> <0041> <0004> <d835dc00> <0005> <0042> <0006> <005a> Before, selector 5 (the B) had no entry and selector 6 (the Z) was published as B. The CMap is now built from one destination per selector, a String, so a surrogate pair is one entry of length two rather than two array positions. The char[] constructor converts (toDestinations) and callers are unchanged. The range logic needs no surrogate special cases any more: an entry may join a bfrange when it is one code point, and two entries are consecutive when their code points are and their selectors share a 256 block. A destination of several characters is written as a bfchar with a string, which the format allows and FOP-3345 uses. PDFToUnicodeCMapTestCase pinned the drift (surrogatePairTest expected the entry after the pair at 0x63 to be 0x65); its expectations change accordingly, in surrogatePairTest, surrogatePairRangeTest, surrogatePairsRangeTest and rangeSizeSurrogateTest, the last of which also ran its low surrogates past U+DFFF and now starts them at U+DC00. ToUnicodeCharacterisationTestCase records the writer's output for the common shapes, so a change to the range packing has to be deliberate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes FOP-3346.
CIDSubset.getChars builds a char[] with StringBuilder.appendCodePoint, so a character
outside the BMP occupies two slots. PDFToUnicodeCMap derived each character selector from
the array position, so every selector after such a character was written one too high, and
the text after an emoji or a mathematical letter extracted as its neighbour.
Measured on this branch, DejaVu Math TeX Gyre, the text "A" U+1D400 "BZ". The content
stream uses selectors 3 4 5 6; the CMap said
Before, selector 5 (the B) had no entry and selector 6 (the Z) was published as B.
The CMap is now built from one destination per selector, a String, so a surrogate pair is
one entry of length two rather than two array positions. The char[] constructor converts
(toDestinations) and callers are unchanged. The range logic needs no surrogate special
cases any more: an entry may join a bfrange when it is one code point, and two entries are
consecutive when their code points are and their selectors share a 256 block. A destination
of several characters is written as a bfchar with a string, which the format allows and
FOP-3345 uses.
PDFToUnicodeCMapTestCase pinned the drift (surrogatePairTest expected the entry after the
pair at 0x63 to be 0x65); its expectations change accordingly, in surrogatePairTest,
surrogatePairRangeTest, surrogatePairsRangeTest and rangeSizeSurrogateTest, the last of
which also ran its low surrogates past U+DFFF and now starts them at U+DC00.
ToUnicodeCharacterisationTestCase records the writer's output for the common shapes, so a
change to the range packing has to be deliberate.
🤖 Generated with Claude Code