Skip to content

feat(client): grounded judge context and per-judge diagnostics - #84

Draft
apucacao wants to merge 1 commit into
alexis/judge-output-format-ignoredfrom
alexis/grounded-judge-context
Draft

apucacao wants to merge 1 commit into
alexis/judge-output-format-ignoredfrom
alexis/grounded-judge-context

Conversation

@apucacao

@apucacao apucacao commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Why

A judge is a second AI call that grades the first one. Three things were wrong with how it ran. This PR fixes all three.

  1. The judge was grading blind. It saw the user request and the model's answer, but not the data the run actually looked at. So it had to guess. It now receives that evidence as context.

  2. A broken judge could hold up a good answer. Judges run after the primary call is finished and billed. Python already caught a judge's errors and logged them, but nothing bounded a judge that hung, so a stuck judge held the response the caller already paid for. Each judge is now isolated and bounded by judge_timeout_ms: its failure is recorded, not raised, and a malformed judgeConfiguration entry no longer escapes either.

  3. Failures were invisible. A judge that was disabled, timed out, or errored simply vanished from the results, with no way to tell which. Every skip or failure now leaves a JudgeDiagnostic behind.

What changed

  • config(...) gains judge_context (a lazy callback, resolved once after the primary call) and judge_timeout_ms.
  • ProviderResponse, the stream done event, and JudgeTask gain judge_context; those plus the graph response and the final graph().stream() done event gain judge_diagnostics.
  • The evidence block is one more part of judge_scoring.build_message_history, after the answer (and the tool trajectory) and before the formatting instructions, so run_judges and run_judge on a JudgeTask still show a judge the same conversation. A callback that returns None supplies no evidence.
  • The context is validated as JSON under 64 KiB, then injected into each judge's message_history inside UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN / ..._END. It never reaches the primary model and is never recorded in telemetry. A judge's own prompt is a LaunchDarkly AI Config and stays swappable at runtime; the evidence rides in a fixed slot that prompt cannot displace.

Breaking change

run_judges and build_judge_tasks now return RunJudgesResult { judge_results, judge_diagnostics } and BuildJudgeTasksResult { judge_tasks, judge_diagnostics, judge_context } instead of the bare value. The per-entry shape inside is unchanged. This is a clean break on purpose: a wrapper that returned only results would reintroduce the silent-failure hole point 3 closes.

Stacked

Based on #86 (fix(judges): a judge config's outputFormat must not reach the provider), not on main, because both change judges.py. Review #86 first; its diff is small. GitHub retargets this PR to main automatically when #86 merges.

🤖 Generated with Claude Code

@apucacao
apucacao force-pushed the alexis/grounded-judge-context branch from 1279bfd to 8a96052 Compare September 11, 2026 10:17
@apucacao
apucacao changed the base branch from main to alexis/judge-output-format-ignored September 11, 2026 14:09
@apucacao
apucacao force-pushed the alexis/grounded-judge-context branch from 8a96052 to c5635ab Compare September 11, 2026 14:09
@apucacao
apucacao force-pushed the alexis/judge-output-format-ignored branch from 95c088d to 627f1fa Compare October 2, 2026 15:07
@apucacao
apucacao force-pushed the alexis/grounded-judge-context branch from c5635ab to 6d15c9d Compare October 2, 2026 15:21
@apucacao

apucacao commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/client/src/launchdarkly_ai_server/judges.py
Comment thread packages/client/src/launchdarkly_ai_server/judges.py
@apucacao
apucacao force-pushed the alexis/grounded-judge-context branch from 6d15c9d to 4ed75a2 Compare October 2, 2026 15:32
@apucacao

apucacao commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/client/src/launchdarkly_ai_server/judges.py
Callers can now hand judges evidence about what actually happened during a
request. config() takes a lazy judge_context callback because the value does
not exist yet when config() is called: the caller's tools fill it while the
primary handler runs. The SDK resolves it exactly once, right after the
primary handler succeeds and before output-format parsing, so every request
with a callback freezes the same snapshot whether or not a judge is sampled,
and whether or not skip_judges is set, on invoke() and stream() alike.

A resolved context must be acyclic JSON of at most 64 KiB encoded. It is
returned unchanged on ProviderResponse.judge_context and reaches a judge only
through that judge's message_history variable, between the
UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN and UNTRUSTED_ACTUATOR_EVIDENCE_END lines.
It never reaches the primary model, the track data, or a span. The block is
one more part of judge_scoring.build_message_history, after the answer and
before the formatting instructions, so the inline path and run_judge on a
JudgeTask still show a judge the same conversation, trajectory included. With
no callback the judge prompt is byte-identical to before.

Judges are now isolated from each other and from the primary result. Config
lookup, provider call, parse and tracking each sit behind their own boundary,
bounded by judge_timeout_ms, and a failure produces one JudgeDiagnostic
instead of discarding work that already succeeded. A judge that beats the
clock and then fails to track keeps its result. A judge that misses the clock
has its late completion consumed silently, so it can neither mutate results
nor emit the score metric. Diagnostics carry codes only, never exception text,
and the same codes are shared with the TypeScript SDK.

graph().invoke() and the final graph().stream() done event forward the
graph-level judge's diagnostics. Graph nodes do not receive a caller judge
context in v1.

run_judges and build_judge_tasks now return result objects carrying both the
results and the diagnostics, which is a breaking change for direct callers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

BREAKING CHANGE: run_judges and build_judge_tasks now return a result
object (RunJudgesResult with judge_results/judge_diagnostics, and
BuildJudgeTasksResult with judge_tasks/judge_diagnostics/judge_context)
instead of the bare value. The per-entry shape inside judge_results is
unchanged. Both are exported from the package.
@apucacao
apucacao force-pushed the alexis/grounded-judge-context branch from 4ed75a2 to 827218c Compare October 2, 2026 15:46
@apucacao

apucacao commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 827218c. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant