Repository navigation
Ground image prompts in documentation and fail unaligned artwork - #169
Merged
Merged
Conversation
…work. Image elements now share the same fail-closed contract as subject-beat coverage: scene-spec-generate feeds source snippets into the LLM, image prompts must use documented terms, and image-generate wraps the Images API call with narration/source plus a validate gate. Co-authored-by: jmjava <jmjava@gmail.com>
The image stem was treated as an invented subject-beat label and failed coverage before the new prompt-alignment gate could run. Co-authored-by: jmjava <jmjava@gmail.com>
Prompt grounding only checked the caption. image-generate now OCRs the PNG for invented labels and vision-reviews it (OpenAI / Grok / Claude), retries once with the critique, and deletes a failing asset. validate adds image_asset_alignment (OCR on by default; vision opt-in). Co-authored-by: jmjava <jmjava@gmail.com>
* Add maintainability ratchets that fail when hotspots, dead code, or survivors grow. Keep the existing pull-request complexity diff. Hotspots fail only when a changed Python file is both complex and frequently changed. Vulture and mutmut compare to committed baselines and do not rewrite them. Co-authored-by: Cursor <cursoragent@cursor.com> * Run path_filters mutation against its own tests. The CI mutmut job was executing the full suite, including a narration test that calls the live API. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…eech-to-text. (#178) Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…ine. (#175) Co-authored-by: Cursor <cursoragent@cursor.com>
* Fail closed when scenes.py inlines or renames _TimedScene helpers. The stale-helper check returned no issues when none of the canonical helpers were top-level defs, so inlined or renamed helpers skipped staleness. Co-authored-by: Cursor <cursoragent@cursor.com> * Keep inlined-helper detection without raising cyclomatic complexity. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
* Fail closed when an ffprobe duration probe fails or returns an unusable duration. Co-authored-by: Cursor <cursoragent@cursor.com> * Keep the ffprobe failure path without raising cyclomatic complexity. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…t scores. Co-authored-by: Cursor <cursoragent@cursor.com>
The hotspot gate fails any edit to that file, so a filtered run marks its scores and compare_to_baseline scopes the missing-id check. Co-authored-by: Cursor <cursoragent@cursor.com>
Fail the benchmark when a baseline case disappears
…work. Image elements now share the same fail-closed contract as subject-beat coverage: scene-spec-generate feeds source snippets into the LLM, image prompts must use documented terms, and image-generate wraps the Images API call with narration/source plus a validate gate. Co-authored-by: jmjava <jmjava@gmail.com>
The image stem was treated as an invented subject-beat label and failed coverage before the new prompt-alignment gate could run. Co-authored-by: jmjava <jmjava@gmail.com>
Prompt grounding only checked the caption. image-generate now OCRs the PNG for invented labels and vision-reviews it (OpenAI / Grok / Claude), retries once with the critique, and deletes a failing asset. validate adds image_asset_alignment (OCR on by default; vision opt-in). Co-authored-by: jmjava <jmjava@gmail.com>
Move the new checks into quiet modules and thin wrappers so hotspot files match main and new functions stay within the existing CCN and NLOC limits. Co-authored-by: Cursor <cursoragent@cursor.com>
Keep the tree that moves alignment checks into quiet modules so the existing lint gates pass. Co-authored-by: Cursor <cursoragent@cursor.com>
jmjava
marked this pull request as ready for review
October 8, 2026 00:26
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Image quality was a caption-only hop:
scene-spec-generatecollected source snippets and never sent them, andimage-generateforwarded the authoredprompt:to the Images API with no check that it named anything from the docs — and no check that the PNG pixels matched either.This change closes both gaps the same way subject-beat coverage already works for box labels.
Prompt alignment (first commit)
scene-spec-generatenow includes--- SOURCE DOCUMENTATION ---in the user message (the snippets were collected and dropped).prompt:text must share documented terms with narration/source. Style-only prompts and invented subjects fail closed.image-generatewraps each Images API call with narration, source snippets, the box label, and an educational-diagram style prefix (image_generation.align_with_docs, default true).Pixel alignment (follow-up)
WidgetX orchestratorvs “checkout service”).chat_completion_with_image) answers PASS/FAIL against the same corpus. FAIL retries once with the critique appended to the prompt, then deletes the asset (image_generation.align_review, default true).validate/--pre-pushaddimage_asset_alignment(OCR on by default; vision opt-in viavalidation.image_asset_alignment.reviewso CI stays offline).Evidence
Scratch-bundle: invented OCR terms fail and the file is not kept; vision FAIL then PASS retries; validate OCR of matching terms passes.
pixel OCR fail, vision retry, validate pass
pytest tests/: 919 passed, 1 skipped.How to use it
Keep
manim_scene_generation.context.pathspointed at the real docs. Image prompts should name those terms. Opt out of vision (still keep prompt + OCR) withimage_generation.align_review: false. Opt out of all alignment withimage_generation.align_with_docs: false.To show artifacts inline, enable in settings.