Skip to content

fix(llm): distinguish failed chunks from model omissions - #3617

Open
GREYGROUPJP wants to merge 1 commit into
Graphify-Labs:v8from
GREYGROUPJP:fix/failed-chunk-diagnostics
Open

GREYGROUPJP wants to merge 1 commit into
Graphify-Labs:v8from
GREYGROUPJP:fix/failed-chunk-diagnostics

Conversation

@GREYGROUPJP

Copy link
Copy Markdown

Problem

When a semantic chunk raises before returning a usable result, the final reconciliation warning says the model returned a response but omitted the affected files. This contradicts the preceding chunk error and points operators toward model-output debugging instead of the dependency or provider failure that actually occurred.

Change

Track the files assigned to failed chunks and split uncovered files into two groups:

  • failed before returning a usable result; and
  • omitted from a successful model response.

Each group now receives an accurate warning. The existing omission wording remains unchanged for successful responses that omit files.

Fixes #3616.

Validation

  • Added a regression test for a missing optional dependency/backend exception.
  • Kept coverage for genuine model omissions and asserted the original warning remains.
  • pytest tests/ -q --tb=short: 5,643 passed, 96 skipped.
  • ruff check graphify/llm.py tests/test_chunking.py: passed.
  • git diff --check: passed.

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Formal verification. No changes could be formally verified in this run.


Graphify review — findings

Splits the corpus-parallel uncovered-files warning into two distinct messages: files whose semantic chunk raised before returning a result now get a "failed before returning a usable result" warning that points at the chunk errors above, while files the model genuinely returned but omitted keep the "returned a response but omitted them" message. Tracks which paths belonged to failed chunks via a new failed_files set so the two cases are classified correctly at reconciliation time.

No blocking issues surfaced. 2 lower-confidence candidates did not survive cross-model review.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 913 functions depend on the 270 functions this change touches.

Health — this change adds coupling hotspots:

  • new: deduplicate_entities() — 77 callers, 24 callees
  • new: build_merge() — 76 callers, 14 callees
  • new: extract_files_direct() — 17 callers, 20 callees
  • new: build() — 52 callers, 6 callees
  • new: extract_corpus_parallel() — 27 callers, 11 callees
  • new: _call_claude_cli() — 33 callers, 9 callees
  • new: dispatch_command() — 2 callers, 125 callees
  • new: _extract_with_adaptive_retry() — 22 callers, 10 callees
  • …and 16 more — each is listed as a finding

Verification — 913 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 578 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

41 of 286 test file(s) selected (14%) via static blast radius.

  • tests/test_backend_env_isolation.py — impact
  • tests/test_backend_extras.py — impact
  • tests/test_build.py — impact
  • tests/test_build_merge_dedup_scope.py — impact
  • tests/test_build_merge_hyperedges_and_prune.py — impact
  • tests/test_build_merge_shrink_guard.py — impact
  • tests/test_carried_hyperedge_remap.py — impact
  • tests/test_charmap_encoding.py — impact
  • tests/test_chunking.py — impact, changed-test
  • tests/test_claude_cli_backend.py — impact
  • tests/test_corrupt_graph_json.py — impact
  • tests/test_cross_extension_reexport_self_cycle.py — impact
  • tests/test_dedup.py — impact
  • tests/test_dedup_remaps_hyperedges.py — impact
  • tests/test_dedup_survivor_richness.py — impact
  • tests/test_evidence_binding.py — impact
  • tests/test_file_slice.py — impact
  • tests/test_global_graph.py — impact
  • tests/test_go_qualified_resolution.py — impact
  • tests/test_hyperedge_member_shapes.py — impact
  • tests/test_image_vision.py — impact
  • tests/test_injection_sentinel_coverage.py — impact
  • tests/test_issue_3472_source_file_collision.py — impact
  • tests/test_label_retry.py — impact
  • tests/test_labeling.py — impact
  • tests/test_llm_backends.py — impact
  • tests/test_llm_parser.py — impact
  • tests/test_llm_parser_reasoning.py — impact
  • tests/test_no_dedup_flag.py — impact
  • tests/test_non_string_node_ids.py — impact
  • tests/test_ollama.py — impact
  • tests/test_ollama_retry_cap.py — impact
  • tests/test_oversized_document_slicing.py — impact
  • tests/test_partial_cache.py — impact
  • tests/test_pdf_slicing.py — impact
  • tests/test_pdf_token_estimate.py — impact
  • tests/test_provider_registry.py — impact
  • tests/test_prs.py — impact
  • tests/test_prune_sweeps_orphans.py — impact
  • tests/test_semantic_fragment_sanitize.py — impact
  • tests/test_unverified_semantic_shrink.py — impact

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

Could not verify: Could not verify extract\_corpus\_parallel.

The verifier did not have enough to check extract\_corpus\_parallel, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: parameter `root` is annotated `Path` — outside the synthesizable primitive/collection set

· 24 more finding(s) on lines outside this diff (see the check run).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

extract_corpus_parallel misreports failed chunks as successful model omissions

1 participant