Error column is wrong when non-ASCII characters come earlier on the line: gcc reports bytes (on Windows), Sublime counts characters
Version: SublimeLinter-gcc from Package Control, gcc 16.1.0 (MinGW-w64 UCRT), SublimeLinter 4, Sublime Text 4215 on Windows. Reproduced in a clean install.
gcc_probe.c (UTF-8, tab-indented):
int main(void) {
int a = 1; int b = undefined_one;
char *s = "é😀"; int c = undefined_two;
return 0
}
The linter runs gcc -c -Wall -O0 -x c -o <file>.o - with the source on stdin. On line 3 the errors are
<stdin>:3:30: error: 'undefined_two' undeclared (first use in this function)
<stdin>:3:26: warning: unused variable 'c' [-Wunused-variable]
undefined_two starts at character 26 (1-based) as Sublime counts, but gcc says 30, because the é is two bytes and the 😀 four in UTF-8: gcc's column is a byte offset. In Sublime the highlight of the error starts four characters too far to the right (start 29, expected 25, 0-based), on fined_two instead of undefined_two, and the warning for c is highlighted on undefined_two instead of c. Line 2, without non-ASCII characters, is correct (tabs count as one).
Adding -fdiagnostics-column-unit=byte gives the same numbers here, so on this setup byte is what the default gives. The documentation says the default is display (a column in terminal cells, where a wide character such as an emoji is two and a tab expands), which I did not test on Linux or macOS with a UTF-8 locale, so the unit may differ per platform.
Suggested fix: pass -fdiagnostics-column-unit=byte (gcc 11 and later, older gcc rejects the option, so it would have to be conditional) so the unit is always known, and convert the byte offset to a character offset in reposition_match(self, line, col, m, vv) with the source line from vv.select_line(line): encode the line to UTF-8, cut it at the byte offset, decode, and take the length. The same fix was merged for SublimeLinter-xmllint and is proposed for SublimeLinter-cppcheck (byte columns as well). I can prepare a PR if you would like one.
Error column is wrong when non-ASCII characters come earlier on the line: gcc reports bytes (on Windows), Sublime counts characters
Version: SublimeLinter-gcc from Package Control, gcc 16.1.0 (MinGW-w64 UCRT), SublimeLinter 4, Sublime Text 4215 on Windows. Reproduced in a clean install.
gcc_probe.c(UTF-8, tab-indented):The linter runs
gcc -c -Wall -O0 -x c -o <file>.o -with the source on stdin. On line 3 the errors areundefined_twostarts at character 26 (1-based) as Sublime counts, but gcc says 30, because theéis two bytes and the😀four in UTF-8: gcc's column is a byte offset. In Sublime the highlight of the error starts four characters too far to the right (start 29, expected 25, 0-based), onfined_twoinstead ofundefined_two, and the warning forcis highlighted onundefined_twoinstead ofc. Line 2, without non-ASCII characters, is correct (tabs count as one).Adding
-fdiagnostics-column-unit=bytegives the same numbers here, so on this setup byte is what the default gives. The documentation says the default isdisplay(a column in terminal cells, where a wide character such as an emoji is two and a tab expands), which I did not test on Linux or macOS with a UTF-8 locale, so the unit may differ per platform.Suggested fix: pass
-fdiagnostics-column-unit=byte(gcc 11 and later, older gcc rejects the option, so it would have to be conditional) so the unit is always known, and convert the byte offset to a character offset inreposition_match(self, line, col, m, vv)with the source line fromvv.select_line(line): encode the line to UTF-8, cut it at the byte offset, decode, and take the length. The same fix was merged for SublimeLinter-xmllint and is proposed for SublimeLinter-cppcheck (byte columns as well). I can prepare a PR if you would like one.