diff --git a/changelog.d/px-2ms.md b/changelog.d/px-2ms.md new file mode 100644 index 0000000..8f24e17 --- /dev/null +++ b/changelog.d/px-2ms.md @@ -0,0 +1,4 @@ +### Fixed + +- A `\U` escape in a string literal is refused with the same error as `\u`: + `"caf\U00e9"` no longer lexes silently to `cafU00e9`. diff --git a/docs/reference/language.md b/docs/reference/language.md index 6dd24e6..36d2919 100644 --- a/docs/reference/language.md +++ b/docs/reference/language.md @@ -45,8 +45,8 @@ recognized: All six are recognized in both quote styles - the quote character that does not delimit the literal needs no escape, but escaping it is still accepted. -Any other escaped character stands for itself and the backslash is dropped, -so `"a\qb"` is `aqb`. +Any other escaped character, apart from the refused `\u` and `\U` below, +stands for itself and the backslash is dropped, so `"a\qb"` is `aqb`. ```elixir iex> Predicator.evaluate(~S{"say \"hi\""}, %{}) @@ -60,12 +60,18 @@ iex> Predicator.evaluate(~S{"a\qb"}, %{}) ``` There is no numeric escape. A `\u` sequence is refused at parse time with an -error naming it, rather than decoding to the bare letter: +error naming it, rather than decoding to the bare letter. `\U` is refused the +same way, with the same error: both spellings of the letter are the refused +escape, so neither decodes silently. ```elixir iex> {:error, err} = Predicator.compile("\"caf\\u00e9\"") iex> err.message "Unsupported escape sequence \\u in string literal: predicator has no numeric escape; write the character itself (string literals are UTF-8)" + +iex> {:error, upper} = Predicator.compile("\"caf\\U00e9\"") +iex> upper.message +"Unsupported escape sequence \\u in string literal: predicator has no numeric escape; write the character itself (string literals are UTF-8)" ``` String literals are UTF-8 source text, so the character is written directly - diff --git a/lib/predicator/lexer.ex b/lib/predicator/lexer.ex index 0011a55..8467a4d 100644 --- a/lib/predicator/lexer.ex +++ b/lib/predicator/lexer.ex @@ -694,11 +694,14 @@ defmodule Predicator.Lexer do {:ok, acc, rest, count, line, col + 1} end - # `\uXXXX` is refused rather than decoded. A string literal is already UTF-8 - # source text, so the character can be written directly; a bug fix has no - # mandate to add grammar. The refusal is explicit so the escape never - # silently decodes to the bare letter instead. - defp take_string([?\\ | [?u | _rest]], _acc, _count, _quote_type, _line, _col) do + # `\uXXXX` is refused rather than decoded, in either spelling of the letter. + # A string literal is already UTF-8 source text, so the character can be + # written directly; a bug fix has no mandate to add grammar. The refusal is + # explicit so the escape never silently decodes to the bare letter instead, + # and `\U` is refused with `\u` because two spellings of one escape refused + # differently is the defect. + defp take_string([?\\ | [u | _rest]], _acc, _count, _quote_type, _line, _col) + when u in [?u, ?U] do {:error, "Unsupported escape sequence \\u in string literal: predicator has no " <> "numeric escape; write the character itself (string literals are UTF-8)"} diff --git a/test/predicator/lexer_test.exs b/test/predicator/lexer_test.exs index acb624b..f06092d 100644 --- a/test/predicator/lexer_test.exs +++ b/test/predicator/lexer_test.exs @@ -299,6 +299,19 @@ defmodule Predicator.LexerTest do assert message =~ "\\u" assert message =~ "string literal" end + + # Sabotage: drop the `?U` from the `when u in [?u, ?U]` guard on the + # refusal clause in `take_string/6` and this goes red - `\\U0041` falls + # through to the generic escape clause and lexes as "U0041". + test "refuses an uppercase \\U escape with the same error as \\u" do + assert {:error, message, 1, 1, _span} = Lexer.tokenize(~s("\\U0041")) + + assert message =~ "\\u" + assert message =~ "string literal" + + assert {:error, lower_message, 1, 1, _span} = Lexer.tokenize(~s("\\u0041")) + assert message == lower_message + end end describe "additional edge cases for coverage" do