Fix #833: check truncated UTF-8 in non-blocking Smile short Unicode decoding - #839
Merged
cowtowncoder merged 2 commits intoOct 2, 2026
Merged
Conversation
…ecoding Non-blocking `_decodeShortUnicodeText()` and `_decodeLongUnicodeName()` did not verify that continuation bytes of a multi-byte UTF-8 character were within the declared String/name length, reading past it (AIOOBE, or decoding bytes of following tokens / stale buffer content). Now report the same errors as the blocking `SmileParser`; error helpers moved up to `SmileParserBase`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cowtowncoder
deleted the
tatu-claude/2.18/833-async-short-unicode-read-past
branch
October 2, 2026 19:49
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #833.
Problem
NonBlockingByteArrayParser._decodeShortUnicodeText()(short Unicode values and names) and_decodeLongUnicodeName()only checked bounds before the first byte of each character. A multi-byte UTF-8 lead byte near the end of the declared length caused continuation bytes to be read past the end of the String, resulting in:ArrayIndexOutOfBoundsExceptionwhen the String is at the end of the input buffer_inputCopydecoded when content was split across feedsFix
(inPtr + code) > endbefore decoding a multi-byte character, same as blockingSmileParser, and report the same errors:_reportTruncatedUTF8InString()/_reportTruncatedUTF8InName()fromSmileParserup toSmileParserBase(sameprotectedsignatures) so the non-blocking parser can use them.10xxxxxx) validation is not added, to keep parity with the blocking short-text decoders.Tests
New
AsyncTruncatedUTF8Test: truncated 2/3/4-byte characters in tiny/small Unicode values, short names and long names. Each runs with feed sizes 1, 2, 3 and the whole document, with and without padding around fed chunks, and checks that blockingSmileParserreports the same error. Valid boundary cases are included. All 5 failure cases fail without the fix.🤖 Generated with Claude Code