From 8f9e72996d4b6b8810d2da45ce169e50920e2494 Mon Sep 17 00:00:00 2001 From: JohnnyT Date: Mon, 14 Sep 2026 04:56:39 -0600 Subject: [PATCH] Refuses an uppercase \U escape like \u The lexer's escape refusal matched lowercase ?u only, so "caf\U00e9" fell through the generic escape clause and lexed silently to "cafU00e9". Two spellings of one escape refused differently is the defect: the refusal clause in take_string/6 now guards on `u in [?u, ?U]` and returns the same error for both. Nothing else about escapes changes - `\x` still stands for itself. The language reference's "any other escaped character" rule names the two refused spellings, and the refusal example carries a doctest for each. Refs: px-2ms --- changelog.d/px-2ms.md | 4 ++++ docs/reference/language.md | 12 +++++++++--- lib/predicator/lexer.ex | 13 ++++++++----- test/predicator/lexer_test.exs | 13 +++++++++++++ 4 files changed, 34 insertions(+), 8 deletions(-) create mode 100644 changelog.d/px-2ms.md diff --git a/changelog.d/px-2ms.md b/changelog.d/px-2ms.md new file mode 100644 index 0000000..8f24e17 --- /dev/null +++ b/changelog.d/px-2ms.md @@ -0,0 +1,4 @@ +### Fixed + +- A `\U` escape in a string literal is refused with the same error as `\u`: + `"caf\U00e9"` no longer lexes silently to `cafU00e9`. diff --git a/docs/reference/language.md b/docs/reference/language.md index 6dd24e6..36d2919 100644 --- a/docs/reference/language.md +++ b/docs/reference/language.md @@ -45,8 +45,8 @@ recognized: All six are recognized in both quote styles - the quote character that does not delimit the literal needs no escape, but escaping it is still accepted. -Any other escaped character stands for itself and the backslash is dropped, -so `"a\qb"` is `aqb`. +Any other escaped character, apart from the refused `\u` and `\U` below, +stands for itself and the backslash is dropped, so `"a\qb"` is `aqb`. ```elixir iex> Predicator.evaluate(~S{"say \"hi\""}, %{}) @@ -60,12 +60,18 @@ iex> Predicator.evaluate(~S{"a\qb"}, %{}) ``` There is no numeric escape. A `\u` sequence is refused at parse time with an -error naming it, rather than decoding to the bare letter: +error naming it, rather than decoding to the bare letter. `\U` is refused the +same way, with the same error: both spellings of the letter are the refused +escape, so neither decodes silently. ```elixir iex> {:error, err} = Predicator.compile("\"caf\\u00e9\"") iex> err.message "Unsupported escape sequence \\u in string literal: predicator has no numeric escape; write the character itself (string literals are UTF-8)" + +iex> {:error, upper} = Predicator.compile("\"caf\\U00e9\"") +iex> upper.message +"Unsupported escape sequence \\u in string literal: predicator has no numeric escape; write the character itself (string literals are UTF-8)" ``` String literals are UTF-8 source text, so the character is written directly - diff --git a/lib/predicator/lexer.ex b/lib/predicator/lexer.ex index 0011a55..8467a4d 100644 --- a/lib/predicator/lexer.ex +++ b/lib/predicator/lexer.ex @@ -694,11 +694,14 @@ defmodule Predicator.Lexer do {:ok, acc, rest, count, line, col + 1} end - # `\uXXXX` is refused rather than decoded. A string literal is already UTF-8 - # source text, so the character can be written directly; a bug fix has no - # mandate to add grammar. The refusal is explicit so the escape never - # silently decodes to the bare letter instead. - defp take_string([?\\ | [?u | _rest]], _acc, _count, _quote_type, _line, _col) do + # `\uXXXX` is refused rather than decoded, in either spelling of the letter. + # A string literal is already UTF-8 source text, so the character can be + # written directly; a bug fix has no mandate to add grammar. The refusal is + # explicit so the escape never silently decodes to the bare letter instead, + # and `\U` is refused with `\u` because two spellings of one escape refused + # differently is the defect. + defp take_string([?\\ | [u | _rest]], _acc, _count, _quote_type, _line, _col) + when u in [?u, ?U] do {:error, "Unsupported escape sequence \\u in string literal: predicator has no " <> "numeric escape; write the character itself (string literals are UTF-8)"} diff --git a/test/predicator/lexer_test.exs b/test/predicator/lexer_test.exs index acb624b..f06092d 100644 --- a/test/predicator/lexer_test.exs +++ b/test/predicator/lexer_test.exs @@ -299,6 +299,19 @@ defmodule Predicator.LexerTest do assert message =~ "\\u" assert message =~ "string literal" end + + # Sabotage: drop the `?U` from the `when u in [?u, ?U]` guard on the + # refusal clause in `take_string/6` and this goes red - `\\U0041` falls + # through to the generic escape clause and lexes as "U0041". + test "refuses an uppercase \\U escape with the same error as \\u" do + assert {:error, message, 1, 1, _span} = Lexer.tokenize(~s("\\U0041")) + + assert message =~ "\\u" + assert message =~ "string literal" + + assert {:error, lower_message, 1, 1, _span} = Lexer.tokenize(~s("\\u0041")) + assert message == lower_message + end end describe "additional edge cases for coverage" do