You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
readString decoded each \uXXXX escape as its own rune, so a UTF-16 surrogate pair ("\ud83d\ude00", the way JSON encoders write non-BMP characters) came out as two U+FFFD. Now, when a \u escape is a high surrogate and is immediately followed by a \u escape for a matching low surrogate, the two are combined with utf16.DecodeRune, which matches encoding/json. Lone or mismatched surrogates still decode to U+FFFD like they do today, and a bad hex digit after \u still returns the existing "Bad \u char" error. The hex-digit parsing moved into a small hexValue helper so it can be reused for the lookahead.
Tests: added TestUnicodeSurrogatePair, which covers a pair (lower and upper case), a pair inside an object value, and lone/reversed/unpaired surrogates.
go test -run TestUnicodeSurrogatePair . # before the fix: FAIL (3 pair cases give "\ufffd\ufffd")
go test ./... # with the fix: ok
Thanks for the PR! I moved test to the asset files, they are easier to exchange between different implementations of Hjson. I removed tests that resulted in \ufffd output, that's not really a desired output, could be considered a bug.
Thanks @trobro — the shared asset files are clearly the better home for these.
On the \ufffd cases: those are lone/unpaired surrogates, and I mapped them to U+FFFD on purpose since that's what Go's own unicode/utf8 and encoding/json do with invalid surrogates. Fine to leave out of the shared suite until Hjson settles the canonical behavior — happy to switch it to error on a lone surrogate instead if you'd prefer.
Makes sense. Keeping v1's U+FFFD behaviour avoids breaking existing inputs. Thanks for merging!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #83
readStringdecoded each\uXXXXescape as its own rune, so a UTF-16 surrogate pair ("\ud83d\ude00", the way JSON encoders write non-BMP characters) came out as two U+FFFD. Now, when a\uescape is a high surrogate and is immediately followed by a\uescape for a matching low surrogate, the two are combined withutf16.DecodeRune, which matchesencoding/json. Lone or mismatched surrogates still decode to U+FFFD like they do today, and a bad hex digit after\ustill returns the existing "Bad \u char" error. The hex-digit parsing moved into a smallhexValuehelper so it can be reused for the lookahead.Tests: added
TestUnicodeSurrogatePair, which covers a pair (lower and upper case), a pair inside an object value, and lone/reversed/unpaired surrogates.