Skip to content

OCE report: add novelty noise gate and redesign the Broker report around it, Fixes AB#3733390 - #460

Merged
Shahzaib (shahzaibj) merged 3 commits into
masterfrom
shjameel-microsoft-oce-noise-gate
Aug 25, 2026
Merged

Shahzaib (shahzaibj) merged 3 commits into
masterfrom
shjameel-microsoft-oce-noise-gate

Conversation

@shahzaibj

@shahzaibj Shahzaib (shahzaibj) commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Why

The 60-day trend section of the weekly Broker report had grown into a catalog. Every code with a 60-day regression got a full row and a chart, whether or not anything had changed recently — so each week the on-call engineer was asked to re-triage the same long-standing known regressions. Meanwhile the section that actually matters, "things that need attention this week", had no charts at all.

That's backwards. The report was generating noise where nothing had changed and withholding evidence where something had.

What this does

Adds assets/scripts/classify-novelty.js — a novelty classifier that reads a bucket-trends.js --json= sidecar and labels each key against its own 7-week baseline:

Label Meaning
NEW absent (or negligible) in the baseline, present now
ACCELERATING already elevated, and materially worse this week
ONGOING elevated but flat — a known regression, not news
VOLATILE swings wide enough that this week isn't distinguishable from noise
RECOVERY / IMPROVING moving the right way
STABLE nothing to say

It also clusters related codes into families, so a single upstream failure doesn't consume six attention slots.

The classifier is the noise gate. Its attention set (NEW + ACCELERATING), plus at most 2 wins, is all that renders visibly with charts. Everything else collapses into a fold — still one click away, never deleted, but no longer competing for the reader's attention.

The Broker report template is redesigned around that gate:

  • Attention rows carry a mandatory inline sparkline, so a claim and its evidence sit on the same line.
  • The 60-day section becomes a detector, not a catalog — capped at 6 charts.
  • The attention list is capped at 8 visible rows.

validate-report.ps1 gains hard checks 13–18 so a future run can't quietly regress the gate: row-body specificity, no flat top row, no suppressed-ratio chip, mandatory .item-spark on attention rows, the 8-row cap, and the 6-chart cap.

What this is not

This does not change what the classifier is fed. It grades Sunday-aligned calendar weeks, which is what it did before. That window is misaligned with the rolling 7-day window the report displays — a real bug, but a separate one, fixed in the follow-up PR so it can be reviewed on its own evidence.

Check 12 in validate-report.ps1 is deliberately left as a reserved gap. The prose refers to checks by number, and the Authenticator profile (next PR) fills that slot — keeping the numbering stable across the stack means the follow-up's validator diff is a pure insertion rather than a renumbering.

Verification

validate-report.ps1 -Path oncall-wow-report-2026-08-18.html — all hard checks pass, exit 0, including 13–18:

[OK] All 4 visible attention row(s) carry an inline sparkline
[OK] Section 2 attention list is short (4 visible row(s))
[OK] 60-day section is a detector, not a catalog (0 visible chart(s), cap 6)
[OK] Section 2 row bodies are row-specific
[OK] No VOLATILE/RECOVERY row headlines a WoW percentage

Stack

This is 1 of 3. Each PR is reviewable in isolation:

  1. this PR — novelty noise gate + Broker report redesign → master
  2. OCE report: add Authenticator app telemetry and turn the skill into a router, Fixes AB#3731627 #461 — Authenticator app report + router-ify the skill → this branch
  3. OCE report: align the noise gate to the report's rolling window, Fixes AB#3731628 #462 — align the noise gate to the report's rolling window → PR 2's branch

Splitting this way keeps the Broker-behaviour changes separate from the purely-additive Authenticator support, so neither has to be reviewed through the other.

Fixes AB#3733390

…und it

The weekly Broker report had become a browsing exercise rather than a triage
tool. The 60-day section rendered 38 charts, ~93% of which duplicated rows in
the error tables below it, while the "needs attention this week" section --
the part an on-call engineer actually reads first -- carried 13 volume-ranked
rows and zero charts. A flat-but-huge code led the list; the genuinely new
ipc_* family sat at positions #6/#9/#10.

Root cause: bucket-trends.js reports what MOVED, but nothing decided whether a
movement was NEWS. Ranking by device count is not a proxy for novelty.

This change adds classify-novelty.js, which labels every series against its own
7-week baseline (NEW / ACCELERATING / ONGOING / VOLATILE / RECOVERY / IMPROVING
/ STABLE) and emits an `attention` set = NEW + ACCELERATING. That set, plus at
most 2 wins, is all that renders visibly with charts; everything still-elevated
collapses into a fold with its weeksElevated count. The ACCELERATING/ONGOING
split is the whole fix: only "still getting worse" earns a second look.

The 60-day section becomes a slow-burn DETECTOR -- it charts only what it
promotes (rows rising on 60d and absent from the attention section, typically
0-3, often zero) and folds the full classification with no chart column.
Sections 6/7 keep a per-row sparkline as a deliberate exemption: they are
lookup tables, not a browsing section.

Also guards two measured false positives: a WoW % off an anomalous prior week
(429 headlined at +397.8% while sitting 94.5% BELOW its own 60-day median), and
a slow drift in block means labelling a flat, WoW-negative series ACCELERATING.

validate-report.ps1 gains checks 13-18 to enforce all of this (row-body
specificity, no flat top row, no suppressed-ratio chip, mandatory .item-spark,
<=8 visible rows, <=6 charts in the 60-day section). Check 12 is left reserved.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@shahzaibj
Shahzaib (shahzaibj) requested a review from a team as a code owner August 20, 2026 01:27
@github-actions github-actions Bot changed the title OCE report: add novelty noise gate and redesign the Broker report around it OCE report: add novelty noise gate and redesign the Broker report around it, Fixes AB#3733390 Aug 20, 2026
@github-actions

Copy link
Copy Markdown

✅ Work item link check complete. Description contains link AB#3733390 to an Azure Boards work item.

@github-actions

Copy link
Copy Markdown

✅ Work item link check complete. Description contains link AB#3733390 to an Azure Boards work item.

Shahzaib (shahzaibj) and others added 2 commits August 25, 2026 15:04
classify-novelty.js emitted the `attention` array sorted purely by
`b.current - a.current`, so the JSON sidecar the report renders from ranked
by device volume. That contradicts the rule the noise gate exists to enforce
(SKILL.md: Section 2 leads with NEW, never with the highest-volume row), and
the docstring immediately above the emit claims this list is what keeps
attention "from drifting back into a volume-ranked dump".

The console output was already correct -- it loops over ORDER -- so only the
sidecar was affected, which is the artifact that actually matters.

Observed on the 2026-08-22 run: state_mismatch, the week's only NEW code,
ranked 4th of 4 behind three ACCELERATING ipc_* codes purely because it had
the fewest devices (10,226 vs 138,288). With this change it sorts first.

No existing check would have caught it: check 14 only warns when the top row
is a flat mover, and an ACCELERATING code clears that by construction. It also
compounds with the 8-row cap, which drops the lowest-volume tail -- exactly
where NEW codes land, since a genuinely new code is low-volume relative to one
that has been climbing for weeks.

Sort by ORDER index first, keeping volume as the within-label tiebreak.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@shahzaibj
Shahzaib (shahzaibj) merged commit ddb1c3d into master Aug 25, 2026
2 checks passed
@shahzaibj
Shahzaib (shahzaibj) deleted the shjameel-microsoft-oce-noise-gate branch August 25, 2026 23:56
Shahzaib (shahzaibj) added a commit that referenced this pull request Aug 28, 2026
… router, Fixes AB#3731627 (#461)

## Why

On-call engineers had Broker telemetry in a weekly report and
Authenticator app health in a Kusto dashboard nobody opened during a
rotation. The ask was explicit: **one slash command, both reports** —
not two commands, and not one merged document.

## What this does

`SKILL.md` becomes a thin **router**. It resolves the reporting window
*once*, picks a mode, and dispatches. All Broker analysis moves
**verbatim** into `assets/playbooks/broker.md`; the new
`assets/playbooks/authapp.md` is its Authenticator counterpart.

| Mode | Produces |
|---|---|
| `both` *(default)* | both reports **+** `oce-index-<curEnd>.html` |
| `broker` | `oncall-wow-report-<curEnd>.html` |
| `authapp` | `authapp-wow-report-<curEnd>.html` |

### Two reports, not one

The two apps have different owners, different triage ladders and
different escalation paths. A merged report forces every reader through
the half they don't own. The `both` mode instead emits two standalone
reports plus a one-page index digest linking them.

### The playbooks are never read into one context

In `both` mode they run as **parallel sub-agents**. This isn't only
about wall-clock — their Kusto conventions are **mutually
incompatible**:

- Broker: HLL device counting; `sum(countDevices)` is **actively
wrong**.
- Authenticator: `sum(SucceededDCount)` is the **correct** idiom.

Interleaving them in one context risks writing one app's numbers under
the other app's rules. The router says so explicitly, and the shared
hard-rules section calls out that app-specific rules are never
interchangeable.

### Authenticator coverage

Scenario funnels (Passkey / Entra MFA / Entra PSI / MSA NGC+SA),
error-reason decomposition, abandonment, Broker API responsiveness,
version share, and an optional App Center crash layer (`--skip-crashes`,
since it needs a secret).

### Plumbing

`bootstrap-report.ps1`, `run-kql.ps1`, `validate-report.ps1` and
`find-suspect-prs.ps1` all gain `-App broker|authapp`.
`validate-report.ps1` also gains the Authenticator check profile — which
fills the check **12** slot deliberately reserved in the previous PR, so
this validator diff is a pure insertion with no renumbering.
`find-suspect-prs.ps1` gains `-Repos` so it can scan the authenticator
repo instead of broker/common.

New `build-index.ps1` reads the headline KPI tiles out of both finished
reports and emits the digest. It is a **digest, not an analysis** — a
cross-app finding gets written into *both* reports, and the index just
links them.

## Scope

This PR is **purely additive Authenticator support plus the router
refactor**. It contains no Broker behaviour changes — those are in the
parent PR, and the window-alignment fix is in the child. That separation
is the whole point of the split.

`SKILL.md` shrinks substantially because its Broker content **moves** to
`assets/playbooks/broker.md` rather than being deleted.

## Verification

Full E2E run in default `both` mode: both reports generated, **both
validators pass**, index built. Reports land in
`%USERPROFILE%\android-oce-reports\` — outside the workspace, so they
can't be committed by accident.

## Stack

This is **2 of 3**:

1. #460 — novelty noise gate + Broker report redesign → `master`
2. **this PR** — Authenticator app report + router → PR 1's branch
3. #462 — align the noise gate to the report's rolling window → this
branch

Fixes
[AB#3731627](https://identitydivision.visualstudio.com/fac9d424-53d2-45c0-91b5-ef6ba7a6bf26/_workitems/edit/3731627)

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Shahzaib (shahzaibj) added a commit that referenced this pull request Sep 22, 2026
…s AB#3731628 (#462)

## The bug

The novelty classifier graded **Sunday-aligned calendar weeks**. The
report displays a **rolling 7-day window**. Those two windows drift
apart by up to six days — so anything that turned in the last ~6 days,
*exactly the period an on-call engineer cares about most*, was
structurally invisible to the noise gate.

Measured on the **2026-08-01** run: the gate's "current" week was `07/19
→ 07/26` against a report window of `07/25 → 08/01`. **One day of
overlap.**

| Code | Report showed | Classifier saw | Verdict printed |
|---|---|---|---|
| `authorization_pending` | **+63.2%** | −37.1% | *ONGOING — do not
re-triage* |
| `expired_token` | **+26.7%** | −51.0% | *ONGOING — do not re-triage* |

Both were filed "don't look at this" **directly beneath their own rising
numbers**. On the Authenticator side the same defect surfaced
differently: red scoreboard pills (rolling-derived) sitting above an
empty *"Needs attention"* section (calendar-derived) — the exact
confusion reported.

## The fix

Bucket the 60-day trend and the sparklines with:

```kql
bin_at(<TIME>, 7d, datetime(<TREND_END>))
```

instead of `startofweek(<TIME>)`.

The final bucket then **is** the report's displayed window, every bucket
is a complete 7 days, and **classifier WoW == displayed WoW by
construction** — not by convention, and not something a future change
can quietly break.

`--include-partial-end` and `TREND_CLASS_END` become obsolete and are
removed. `bucket-trends.js` now **warns if `--end` is omitted**, because
its partial-end auto-drop heuristic (`if (!endArg …)`) would otherwise
silently discard a genuinely complete final bucket — under rolling
alignment a real 70% collapse could be thrown away as "looks partial".

Because 60 isn't a multiple of 7, the **oldest** bucket is the partial
one — the safe end to be partial on — and `--start` drops it, leaving 8
complete weeks.

## This is not "more alerts"

A/B on real data:

- attention set went **4 → 5** keys
- both mis-filed codes promoted to `ACCELERATING`
- `access_denied` correctly **demoted** — it was actually **−53.2%**, a
false positive the calendar window had been surfacing

Alignment **adds real signal and removes phantom signal**. A controlled
A/B over the affected window confirmed the classifier's WoW input
matched the displayed WoW on **8 of 8** sampled codes after the change,
versus 0 of 8 before.

> ⚠️ One correction to earlier framing, worth stating plainly for
reviewers: the *"promoted both to ACCELERATING and demoted
`access_denied`"* result is specific to the **2026-08-01** window, and
the prose now says so. On other windows the same fix produces different
(still correct) label changes. The claim being made here is about
**input correctness**, not about any one code's verdict.

## Also: reconciling red pills

Adds `validate-report.ps1` check **19**. Scoreboard tables colour a row
from its **own rolling delta**; the attention section is populated from
the classifier's **novelty** verdict. Those answer different questions,
so a row can legitimately be red in the table *and* legitimately absent
from attention — but a reader who sees that mismatch unexplained
concludes the report is broken.

Precedent: `Passkey WebAuthN Registration` shipped carrying `tag-bad`
(−1.27 pts, worst delta in its table) directly above the words *"Quiet
week — 0 NEW or ACCELERATING"*. Both statements were true — the scenario
peaks at ~732 bad-outcome devices, below the 1,000-device classification
floor, so it is **structurally excluded** and can never appear in
attention however sharply it moves.

Every `tag-bad`/`tag-warn` row must now be **either** promoted into
attention **or** named in a muted `.reconcile-note` giving the reason,
tested in order: (1) below the classification floor, (2) within its own
normal band, (3) ONGOING and flat. Check 19 hard-fails an unreconciled
pill.

This closes the "red pill above an empty attention section" confusion at
the **report** level, independently of window alignment.

## Stack

This is **3 of 3**:

1. #460 — novelty noise gate + Broker report redesign → `master`
2. #461 — Authenticator app report + router → PR 1's branch
3. **this PR** — rolling-window alignment + check 19 → PR 2's branch

Fixes
[AB#3731628](https://identitydivision.visualstudio.com/fac9d424-53d2-45c0-91b5-ef6ba7a6bf26/_workitems/edit/3731628)

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants