diff --git a/README.md b/README.md index 2f4a5c5..24f1998 100644 --- a/README.md +++ b/README.md @@ -267,12 +267,13 @@ bound is what makes that refusal a property of the source rather than of the machine that compiled it. `evaluate`, `execute` and `executeValue` each take that source text directly as -well, in place of the instruction list, and compile it before running it. The -string is compiled as an EXPRESSION at all three: a source that needs the -statement grammar is refused identically at every one of them. A caller with -a statement program to run compiles it with `compileProgram`, under -[Statement programs](#statement-programs) below, and passes the instruction -list. +well, in place of the instruction list, and compile it before running it, as +the reference implementation's three do. `evaluate` compiles the string as an +EXPRESSION, so a source that needs the statement grammar is refused there. +`execute` and `executeValue` compile it as a statement program, the +compilation `compileProgram` performs under +[Statement programs](#statement-programs) below, so the same source runs +there, and an expression's source runs as a program of one statement. ```ts import { evaluate, execute } from "@riddler/predicator"; @@ -291,9 +292,8 @@ if (!held.ok || held.value !== true) { // A source that does not compile comes back on the failing arm these three // already had, carrying the compiler's own refusal rather than a rewrapping of -// it. `execute` is no exception: it compiles an expression too, so an -// assignment is refused here exactly as it is at `evaluate`. -const assigned = execute("x = 1"); +// it. `evaluate` compiles an expression, so an assignment is refused there. +const assigned = evaluate("x = 1"); if (assigned.ok) { throw new Error("an assignment is not an expression"); @@ -310,6 +310,13 @@ if (assigned.error.reason !== "assignment_in_expression") { if (assigned.error.position.column !== 3) { throw new Error("the refusal points at the `=` the grammar had no room for"); } + +// `execute` compiles the same text as a statement program, and runs it. +const bound = execute("x = 1"); + +if (!bound.ok || bound.context.x !== 1) { + throw new Error("an assignment is a statement, and execute runs it"); +} ``` The cost is on the failing arm's type: it now admits a `ParseError` for every diff --git a/changelog.d/pts-0zns.md b/changelog.d/pts-0zns.md new file mode 100644 index 0000000..49c4ba2 --- /dev/null +++ b/changelog.d/pts-0zns.md @@ -0,0 +1,3 @@ +### Changed + +- `execute` and `executeValue` compile a source string as a statement program, as the reference's do, where they compiled it as an expression; `evaluate` still compiles an expression. Two effects follow: a source the expression grammar refused may now run (`execute("x = 1")` binds `x`), and a source either grammar refuses may answer the program grammar's refusal, with a different message or reason (`score 3` is refused as unexpected "after statement" rather than "after expression"). One more follows from the first: `executeValue` given an expression's source now answers that expression's value, where it answered undefined. A caller that wants a source refused unless it is an expression compiles it with `compile` and passes the instruction list. diff --git a/docs/adr/0004-the-compiler-surface.md b/docs/adr/0004-the-compiler-surface.md index 2fb32a2..9eb7ef5 100644 --- a/docs/adr/0004-the-compiler-surface.md +++ b/docs/adr/0004-the-compiler-surface.md @@ -1699,3 +1699,123 @@ go looking. The note on that amendment's acceptance says the type behind statement grammar says no name is exported for the result types of `compileProgramWithPositions` and `compileProgramWithSpans`; it is about those two results, and this change exports neither. + +## Amendment: `execute` and `executeValue` compile a source string as a statement program (2026-10-01) + +Status: proposed (2026-10-01) + +Recorded for `pts-0zns`, under the ruling that the two run entry points +compile a source string as the reference's do, a named host-visible change +(ruled by the operator, 2026-10-01). This entry is appended and removes no +line above. `src/` and `test/` are cited as the change carrying this entry +leaves them; the reference is cited at `v9.4.2`, read in a detached export of +predicator-ex whose `mix.exs` `@version` reads `9.4.2`. + +### What this amends + +The section headed "The three entry points take a source string" says that +`evaluate`, `execute` and `executeValue` each "compile it as an EXPRESSION", +that `execute("x = 1")` answers `assignment_in_expression` "exactly as +`evaluate("x = 1")` does", and that "A program-mode `execute(source)` that +compiles the statement grammar arrives with that grammar, not here." Those +statements are superseded for `execute` and `executeValue` by this entry and +still hold of `evaluate`. The section's second paragraph, on the failing arm +widening to admit `ParseError`, holds unchanged at all three. + +The statement-grammar amendment above names the same sentence in its +paragraph opening "One sentence nearby is NOT superseded", and says that +"Switching the two run entry points to program mode is a later change with +its own record." This entry is that record. That paragraph, and the first +paragraph of that amendment's "What this does not decide", were true of the +change they were written for; what `execute` and `executeValue` answer for a +source string from here on is stated below. + +### The decision + +**`execute` and `executeValue` compile a source string with the compilation +`compileProgram` performs; `evaluate` compiles one with the compilation +`compile` performs.** That is the reference's split at the tag: in +`lib/predicator.ex`, `evaluate/3`'s source clause hands its tokens to +`Parser.parse`, and `execute/3` runs through `execute_value/3`, whose source +clause hands them to `Parser.parse_program`. Here `programOf` in +`src/index.ts` takes the compiler as an argument; `evaluate` passes +`compile`, and `execute` and `executeValue` pass `compileProgram`. An +instruction list passed in place of a source runs exactly as before at all +three. + +An expression's source is a program of one bare expression statement, which +the program grammar compiles to the expression's own instruction list +followed by `["pop"]`. So for such a source `execute` answers the context or +the failing arm it answered before, and `executeValue` answers the same +context or failing arm with the expression's value beside it. + +What a host sees change, each answer run at `v9.4.2` through +`Predicator.execute_value/3` in the export and through this package after the +change: + +| at | source | before | after | +|---|---|---|---| +| `execute`, `executeValue` | `x = 1` | `assignment_in_expression` | runs, binding `x` to 1 | +| `execute`, `executeValue` | `score 3` | `trailing_token`, `Unexpected token number '3' after expression` | `trailing_token`, `Unexpected token number '3' after statement` | +| `execute`, `executeValue` | `if` | `statement_keyword`, at column 1 | `expected_primary`, at the end of input, column 3 | +| `executeValue` | `score > 1`, with `score` bound to 5 | the value absent | the value `true` | + +The first row is a source the expression grammar refused that now runs; the +second and third are a refusal that is now the program grammar's, with a +different message or reason; the fourth is the value an expression's source +now has at `executeValue`. A refusal at either entry point is the one +`compileProgram` answers for the same source, handed out unwrapped, and a +source either grammar refuses carries no context on the failing arm, as +before. + +**Evidence.** `test/index.test.ts` holds both halves. For every corpus case +that carries a source both grammars compile, it checks that the program is +the expression's list followed by `["pop"]`, that `execute` answers what the +expression's list answers, that `executeValue` answers the same context or +failing arm, and that `evaluate` is unchanged. It holds `execute` and +`executeValue` to four `program_refusal` rows of the compile transcript, read +through `compileTranscriptLines` - `stray-else/leading`, +`not-a-location/literal`, `missing-block/if-token` and +`after-statement/missing-separator`, under the prefix `program-refusal/` - +each message, position and span verbatim with its reason. The same four +sources run through `Predicator.execute_value/3` in the export answer the +same message, position and span as their rows. + +### What this does not decide + +It does not change `evaluate`, which compiles an expression, as the +reference's `evaluate/3` does. + +It adds no export and changes no signature. The failing arm of `execute` and +`executeValue` already admitted `ParseError`, and the three reasons only the +program grammar answers are already members of `ParseReason`. + +It does not carry a source location into a run. The reference runs a source +with its positions and segment positions tables (`execute_value_ast`), so an +evaluation error there can point into the source. Here an evaluation error's +`position` is the index of the failing instruction (`EvaluationError` in +`src/errors.ts`) and no entry point takes a positions table; the index is +into the list `compileProgram` answers for the same source, which +`compileProgramWithPositions` maps to a point. + +It does not hand back a context beside a refused source. The reference +answers its normalized input context beside a parse error; here a source that +did not compile carries no context on the failing arm, as it did before this +entry. + +It does not move the conformance registry's compiler claim or any corpus +expectation; `conformance/` is unchanged. + +### Consequences + +A host that keeps a statement program as text runs it directly, as a host of +the reference does, rather than compiling it first. + +A caller that relied on `execute` or `executeValue` refusing a source that is +not an expression no longer gets that refusal. It compiles the source with +`compile` and passes the instruction list, which is refused or run exactly as +before. + +The source form of `executeValue` now answers an expression's value, so an +expression's source answers the same value at `evaluate` and at +`executeValue` wherever `evaluate` answers one. diff --git a/src/compile.ts b/src/compile.ts index de970d1..4394194 100644 --- a/src/compile.ts +++ b/src/compile.ts @@ -185,8 +185,10 @@ function compileProgramAll(source: string): ProgramEmitResult { } /** - * Compiles a statement program into the instruction list `evaluate` - * consumes. + * Compiles a statement program into the instruction list `execute` and + * `executeValue` run. It is the compilation those two perform on a source + * string, so a program they are handed as text and the list this answers for + * the same text run alike. * * A program is one or more statements separated by `;`, with one trailing * `;` allowed. A statement is an assignment to a location - an identifier, diff --git a/src/index.ts b/src/index.ts index 6bd2d86..892f7ee 100644 --- a/src/index.ts +++ b/src/index.ts @@ -10,7 +10,7 @@ */ import { readDuration } from "./cast.js"; -import { compile } from "./compile.js"; +import { type CompileResult, compile, compileProgram } from "./compile.js"; import type { Ast } from "./decompile.js"; import type { ParseError } from "./errors.js"; import { @@ -82,11 +82,15 @@ export type ParseResult = /** * The program to run, from either accepted first argument. * - * A string is compiled as an EXPRESSION - the same compilation `compile` - * performs, from the same module, so the refusal a caller reads here is the - * one `compile` would have answered for that source, handed out unwrapped with - * its reason, message, position and span intact. A program is already what the - * evaluator wants and is passed on untouched. + * A string is compiled by the compiler the entry point names: `compile` at + * `evaluate`, which reads an expression, and `compileProgram` at `execute` and + * `executeValue`, which read a statement program, as the reference's + * `evaluate/3`, `execute/3` and `execute_value/3` do. Either way it is the + * same compilation the exported function performs, from the same module, so + * the refusal a caller reads here is the one that function would have + * answered for that source, handed out unwrapped with its reason, message, + * position and span intact. A program is already what the evaluator wants and + * is passed on untouched. * * It is not exported. What a caller holds is a result, and this shape exists * only so the three entry points share one answer to which of the two they @@ -94,11 +98,12 @@ export type ParseResult = */ function programOf( source: Program | string, + compiler: (text: string) => CompileResult, ): | { readonly ok: true; readonly program: Program } | { readonly ok: false; readonly error: ParseError } { if (typeof source !== "string") return { ok: true, program: source }; - const compiled = compile(source); + const compiled = compiler(source); if (!compiled.ok) return { ok: false, error: compiled.error }; return { ok: true, program: compiled.instructions }; } @@ -111,7 +116,10 @@ function programOf( * instruction list is, under the same context and the same options, so the two * accepted first arguments differ in what a caller stores rather than in what * the run does. A source that does not compile comes back on the failing arm - * below. + * below. The string is compiled as an EXPRESSION here and only here, as the + * reference's `evaluate/3` compiles one: `evaluate("x = 1")` answers + * `assignment_in_expression`, where `execute` and `executeValue` compile the + * same text as a statement program and run it. * * The context is normalized on the way in and the result is projected back to * plain host values on the way out, so a host writes and reads its own values @@ -147,7 +155,7 @@ export function evaluate( context?: unknown, options?: EvaluateOptions, ): EvaluateResult { - const resolved = programOf(program); + const resolved = programOf(program, compile); if (!resolved.ok) return { ok: false, error: resolved.error }; const outcome = evaluateToValue(resolved.program, context, options); return outcome.ok ? { ok: true, value: toHost(outcome.value) } : outcome; @@ -157,16 +165,13 @@ export function evaluate( * Runs a STATEMENT program and answers the context it halted with, from a * compiled instruction list or from source text. * - * THE SOURCE FORM COMPILES AN EXPRESSION, NOT A STATEMENT PROGRAM. A string - * here goes through the same expression compilation it goes through at - * `evaluate`, so a source that needs the statement grammar is refused with - * that grammar's own reason instead of running: `execute("x = 1")` answers - * `assignment_in_expression` and binds nothing, exactly as `evaluate("x = 1")` - * does. Compiling the statement grammar from source is not yet implemented and - * is not part of this entry point; it belongs with that grammar and arrives - * with it. Until then a caller with a statement program to run holds it as an - * instruction list and passes the list, which is what this entry point has - * always taken and what it still runs as a statement program. + * A string is compiled as a STATEMENT PROGRAM, by the compilation + * `compileProgram` performs, as the reference's `execute/3` compiles one: + * `execute("x = 1")` binds `x`, where `evaluate("x = 1")` refuses the same text + * with `assignment_in_expression`. An expression's source is a program of one + * bare expression statement, so it runs as the expression does and answers + * the context it was given. A source the program grammar refuses answers that + * grammar's `ParseError`, handed out unwrapped as `compileProgram` answers it. * * The mode is carried by the entry point rather than by the artifact: the same * instruction list runs here and at `evaluate`, and what differs is only what @@ -201,7 +206,7 @@ export function execute( context?: unknown, options?: EvaluateOptions, ): ExecuteResult { - const resolved = programOf(program); + const resolved = programOf(program, compileProgram); if (!resolved.ok) return { ok: false, error: resolved.error }; const outcome = executeToContext(resolved.program, context, options); if (outcome.ok) return { ok: true, context: projectContext(outcome.context) }; @@ -214,16 +219,10 @@ export function execute( * statement alongside the context, from a compiled instruction list or from * source text. * - * The source form compiles an expression, on the same terms as at `execute`: - * a source needing the statement grammar is refused rather than run, and - * compiling that grammar from source is not yet implemented here. That makes - * this the least useful of the three string forms, and deliberately so. The - * value below is the last expression STATEMENT's, and what the expression - * grammar compiles has no statement boundary to retain one, so a source - * argument answers the absence as its value however well it evaluates - a - * caller wanting an expression's value from its source text calls `evaluate`. - * The form is accepted here because it is accepted at all three, not because - * this is where it earns its keep. + * The source form compiles a statement program, on the same terms as at + * `execute`. An expression's source is a program of one bare expression + * statement, and so its value is that expression's: `executeValue("score > 1", + * { score: 5 })` answers `true`, as the reference's `execute_value/3` does. * * This is a host convenience rather than an instruction-set guarantee. The * value is what the statement boundary's `pop` discarded, retained rather than @@ -247,7 +246,7 @@ export function executeValue( context?: unknown, options?: EvaluateOptions, ): ExecuteValueResult { - const resolved = programOf(program); + const resolved = programOf(program, compileProgram); if (!resolved.ok) return { ok: false, error: resolved.error }; const outcome = executeToContext(resolved.program, context, options); if (outcome.ok) { diff --git a/test/index.test.ts b/test/index.test.ts index 98c26b6..0373969 100644 --- a/test/index.test.ts +++ b/test/index.test.ts @@ -5,7 +5,10 @@ // entry point does not have. import { describe, expect, it } from "vitest"; +import { loadCases } from "../scripts/lib/corpus.mjs"; import { + compile, + compileProgram, type EvaluateOptions, evaluate, execute, @@ -14,6 +17,8 @@ import { isaVersion, PDateTime, } from "../src/index.js"; +import { compileTranscriptLines } from "./conformance/compile-transcript.js"; +import { decodeCase } from "./conformance/runner.js"; describe("isaVersion", () => { // Sabotage: returning 5 instead of 6 turns this red. @@ -337,8 +342,12 @@ describe("a source string in place of a program", () => { // Sabotage: rewrapping the compiler's refusal in an EvaluationError, instead // of handing it out unchanged, turns the type, the reason and the message // assertions red together. It was run and reverted. - it("answers the compiler's own refusal on the failing arm, at all three", () => { - for (const outcome of [evaluate("x = 1"), execute("x = 1"), executeValue("x = 1")]) { + // + // `evaluate` is the one of the three that still compiles an EXPRESSION, as + // the reference's `evaluate/3` does, so an assignment is refused here and + // runs at the other two (the block below). + it("answers the expression compiler's own refusal at evaluate", () => { + for (const outcome of [evaluate("x = 1")]) { expect(outcome.ok).toBe(false); if (outcome.ok) continue; expect(outcome.error.type).toBe("ParseError"); @@ -366,21 +375,33 @@ describe("a source string in place of a program", () => { // context member at all, turns this red. A source that did not compile never // ran and so bound nothing. It was run and reverted. it("carries no context where the source did not compile", () => { - const refused = execute("x = 1", { visitor: { bucket: 12 } }); - expect(refused.ok).toBe(false); - if (refused.ok) return; - expect("context" in refused).toBe(false); + for (const refused of [ + execute("3 = renewals", { renewals: 1 }), + executeValue("3 = renewals", { renewals: 1 }), + ]) { + expect(refused.ok).toBe(false); + if (refused.ok) continue; + expect("context" in refused).toBe(false); + } }); - // Sabotage: retaining a value for the compiled expression - returning the - // stack top instead of the last expression statement's value - turns this - // red. The expression grammar emits no statement boundary, so there is no - // value to retain. It was run and reverted. - it("answers the absence at executeValue, whatever the expression evaluates to", () => { - expect(executeValue("score > 1", { score: 5 })).toEqual({ + // An expression's source is a program of one bare expression statement, so + // its value is the last expression statement's value. Before execute and + // executeValue compiled a program, this source answered the absence as its + // value here, because the expression grammar emitted no statement boundary. + // + // Sabotage: answering the absence as the value whenever the first argument + // is a string, or compiling it with `compile` again at executeValue, turns + // this red. Both were run and reverted. + it("answers an expression source's value at executeValue", () => { + expect(executeValue("loan.renewals > 1", { loan: { renewals: 2 } })).toEqual({ ok: true, - value: undefined, - context: { score: 5 }, + value: true, + context: { loan: { renewals: 2 } }, + }); + expect(execute("loan.renewals > 1", { loan: { renewals: 2 } })).toEqual({ + ok: true, + context: { loan: { renewals: 2 } }, }); }); @@ -401,3 +422,113 @@ describe("a source string in place of a program", () => { }); }); }); + +// execute and executeValue compile a source string as a statement PROGRAM, as +// the reference's execute/3 and execute_value/3 do (its source clause calls +// Parser.parse_program at v9.4.2); evaluate keeps compiling an expression. +describe("a source string at execute and executeValue is a statement program", () => { + // Sabotage: compiling the source with `compile` again at execute, or at + // executeValue, instead of `compileProgram`, turns this red with an + // assignment refused. Both were run and reverted. + it("runs a statement source the expression grammar refused", () => { + expect(execute("x = 1")).toEqual({ ok: true, context: { x: 1 } }); + expect(executeValue("x = 1")).toEqual({ ok: true, value: undefined, context: { x: 1 } }); + expect( + executeValue("if loan.overdue { fines = fines + 1 }; fines", { + loan: { overdue: true }, + fines: 2, + }), + ).toEqual({ ok: true, value: 3, context: { loan: { overdue: true }, fines: 3 } }); + expect(execute("patron.holds = patron.holds + 1", { patron: { holds: 1 } })).toEqual({ + ok: true, + context: { patron: { holds: 2 } }, + }); + }); + + // Sabotage: switching evaluate to `compileProgram` along with the other two + // turns this red: the assignment compiles and the run answers an empty stack + // rather than the grammar's refusal. It was run and reverted. + it("leaves evaluate compiling an expression", () => { + const refused = evaluate("status = 'late'"); + expect(refused.ok).toBe(false); + if (refused.ok) return; + expect(refused.error.type).toBe("ParseError"); + expect(refused.error.reason).toBe("assignment_in_expression"); + }); + + // The rows are the reference's own refusals, run at v9.4.2 into the compile + // transcript; each source below is refused by the program grammar with one + // of the three reasons only the program entry points answer, or with the + // statement grammar's wording of a family the union already names. + // + // Sabotage: replacing the refusal's message on its way out, or compiling + // with `compile` again at execute, turns this red, the second at the first + // row's reason. Both were run and reverted. + it("answers the program grammar's refusal, as its transcript row records it", () => { + const rows = new Map>(); + for (const line of compileTranscriptLines()) { + const row = JSON.parse(line) as Record; + if (row.kind === "program_refusal") rows.set(row.id as string, row); + } + const pinned: readonly (readonly [string, string])[] = [ + ["program-refusal/stray-else/leading", "unexpected_else"], + ["program-refusal/not-a-location/literal", "unassignable_location"], + ["program-refusal/missing-block/if-token", "expected_open_brace"], + ["program-refusal/after-statement/missing-separator", "trailing_token"], + ]; + for (const [id, reason] of pinned) { + const row = rows.get(id); + expect(row, id).toBeDefined(); + if (row === undefined) continue; + for (const outcome of [execute(row.source as string), executeValue(row.source as string)]) { + expect(outcome.ok, id).toBe(false); + if (outcome.ok) continue; + expect(outcome.error.type, id).toBe("ParseError"); + if (outcome.error.type !== "ParseError") continue; + expect(outcome.error.reason, id).toBe(reason); + expect(outcome.error.message, id).toBe(row.message); + expect(outcome.error.position, id).toEqual(row.position); + expect(outcome.error.span, id).toEqual(row.span); + } + } + }); + + // What answers as before. Every corpus case that carries a source and that + // both grammars compile is run three ways: at execute, the context or the + // failing arm the expression compilation answered (the old answer, by + // construction, since that is what execute ran); at executeValue, the same + // context or failing arm; and at evaluate, the source form answering what + // the expression's instruction list answers. + // + // Sabotage: switching evaluate to `compileProgram` turns this red at the + // first corpus case. It was run and reverted. This test pins the answers + // that must NOT move, so switching execute or executeValue back to `compile` + // leaves it green; the tests above are the ones that catch that. + it("answers an expression source's context and failing arm as before", () => { + let compared = 0; + for (const item of loadCases(9)) { + if (item.source === null) continue; + const expression = compile(item.source); + const program = compileProgram(item.source); + if (!expression.ok || !program.ok) continue; + expect(program.instructions, item.id).toEqual([...expression.instructions, ["pop"]]); + const context = decodeCase(item).context ?? {}; + expect(execute(item.source, context), item.id).toEqual( + execute(expression.instructions, context), + ); + const before = executeValue(expression.instructions, context); + const after = executeValue(item.source, context); + expect(after.ok, item.id).toBe(before.ok); + if (before.ok && after.ok) { + expect(after.context, item.id).toEqual(before.context); + } else { + expect(after, item.id).toEqual(before); + } + expect(evaluate(item.source, context), item.id).toEqual( + evaluate(expression.instructions, context), + ); + compared += 1; + } + expect(compared).toBeGreaterThan(200); + }); +});