Skip to content

Fix #833: check truncated UTF-8 in non-blocking Smile short Unicode decoding - #839

Merged
cowtowncoder merged 2 commits into
2.18from
tatu-claude/2.18/833-async-short-unicode-read-past
Oct 2, 2026
Merged

cowtowncoder merged 2 commits into
2.18from
tatu-claude/2.18/833-async-short-unicode-read-past

Conversation

@cowtowncoder

Copy link
Copy Markdown
Member

Fixes #833.

Problem

NonBlockingByteArrayParser._decodeShortUnicodeText() (short Unicode values and names) and _decodeLongUnicodeName() only checked bounds before the first byte of each character. A multi-byte UTF-8 lead byte near the end of the declared length caused continuation bytes to be read past the end of the String, resulting in:

  • ArrayIndexOutOfBoundsException when the String is at the end of the input buffer
  • bytes of following tokens silently decoded into the String
  • stale bytes from _inputCopy decoded when content was split across feeds

Fix

  • Check (inPtr + code) > end before decoding a multi-byte character, same as blocking SmileParser, and report the same errors:
    • "Truncated UTF-8 character in Short Unicode String value" / "... Short Unicode Name"
    • "Unexpected end-of-input in long field name"
  • Move _reportTruncatedUTF8InString() / _reportTruncatedUTF8InName() from SmileParser up to SmileParserBase (same protected signatures) so the non-blocking parser can use them.
  • Continuation-byte (10xxxxxx) validation is not added, to keep parity with the blocking short-text decoders.

Tests

New AsyncTruncatedUTF8Test: truncated 2/3/4-byte characters in tiny/small Unicode values, short names and long names. Each runs with feed sizes 1, 2, 3 and the whole document, with and without padding around fed chunks, and checks that blocking SmileParser reports the same error. Valid boundary cases are included. All 5 failure cases fail without the fix.

🤖 Generated with Claude Code

cowtowncoder and others added 2 commits October 2, 2026 12:39
…ecoding

Non-blocking `_decodeShortUnicodeText()` and `_decodeLongUnicodeName()` did not
verify that continuation bytes of a multi-byte UTF-8 character were within the
declared String/name length, reading past it (AIOOBE, or decoding bytes of
following tokens / stale buffer content). Now report the same errors as the
blocking `SmileParser`; error helpers moved up to `SmileParserBase`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cowtowncoder cowtowncoder self-assigned this Oct 2, 2026
@cowtowncoder
cowtowncoder merged commit 0481a68 into 2.18 Oct 2, 2026
4 checks passed
@cowtowncoder
cowtowncoder deleted the tatu-claude/2.18/833-async-short-unicode-read-past branch October 2, 2026 19:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant