Verification-gated skill routing and self-improvement harness for Hermes-style agent skills
-
Updated
Jul 21, 2026 - Python
Verification-gated skill routing and self-improvement harness for Hermes-style agent skills
The deterministic merge gate for AI-generated agent capability changes — a local-first, static Tool-Use Readiness review for MCP, OpenAPI, and SDK tool surfaces. Open-source CLI + GitHub Action.
Black-box-first skill suite for manual test design, risk-based prioritization, minimal test case synthesis, and release gate decisions. / ブラックボックス前提で、手動テストの観点出し・リスク優先度付け・最小ケース化・リリースゲート判定まで行う skills
AI-powered release gate for SAP Transport Requests. CLI collects evidence from ADT; SKILL guides LLMs to review code, config, interfaces, and release risk — online or offline.
Risk-based release gate for CRA (Article 14), NIS2 and DORA compliance with SBOM, KEV correlation, and auditable evidence.
Human-in-the-loop AI release gate: Gemini 3 on Vertex AI reasons over Arize Phoenix telemetry to approve, block, or guardrail a version promotion, with idempotent gated writes. Google Cloud Rapid Agent Hackathon.
Deterministic, evidence-backed QA runtime for Codex, Claude Code, and Cursor. Versioned artifact contracts, Playwright browser driver, redacted evidence, reproducible release gates. Never calls an LLM.
ML release-gate checks for drift, performance regression, latency, and JSON/Markdown deployment reports.
QA Architecture Blueprint — Python/Playwright/Pytest, Docker-first CI Gate Chain, Release-Readiness Governance, Agentic QA Workflow Foundation
Validate, pressure-test, and release-gate SKILL.md packages for OpenAI and portable agent runtimes.
Run your intentic agent from CI — block on an agent-judged release gate, or wake the agent with the event payload
A closed-loop safety harness for agentic LLMs: stress-test → regression → release-gate → incident-replay.
Evidence-backed release audits for AI agent harnesses, covering loop correctness, tool safety, context integrity, recovery, and production readiness. 为 AI Agent 运行底座做有证据的上线审查,验证循环、工具、上下文、故障恢复与生产安全。
Evidence-first product quality checks for web apps
Offline Python CLI (nine subcommands, no network) evaluating whether transgender and nonbinary patients' identity and clinical-context data survives registration, EHR, HL7/FHIR, and lab boundaries, emitting a deterministic receipt that states its own limits. Synthetic fixtures only; the service is a plan, not clinically governed or validated.
Side-effect governance, approvals, crash-safe idempotency, reconciliation, and audit for AI agents.
Claude Code QA agent with memory, plus a read-only MCP server (verdict-qa-mcp on PyPI): baseline → delta runs (NEW/REGRESSED), a fact harness so the model judges while the system measures, signed run history, flaky quarantine with expiry, and an exit-code release gate for CI. Ships six scored eval fixtures — misses published.
Deterministic, evidence-bound release gate for AI systems.
Trace-first release gate for coding-agent skills. Scores baseline vs. candidate traces across a 7-dimension rubric and emits a PASS / HOLD / INVESTIGATE Skill Delta Report.
PreFlight — Agent Service Release Gate
To associate your repository with the release-gate topic, visit your repo's landing page and select "manage topics."