Skip to content

Commit 40e0937

Browse files
committed
docs: bring P-033 (probabilistic-data-structures) in from main
Completes the rebase — all P-026..P-033 targets now exist in this branch's own tree (copied verbatim from main), so docs/proposals/README.md's index no longer has dead links when this branch is viewed standalone.
1 parent 8fc13d4 commit 40e0937

1 file changed

Lines changed: 273 additions & 0 deletions

File tree

Lines changed: 273 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,273 @@
1+
# P-033 — In-process sketches and bitmap indexes for legacy .NET diagnostics
2+
3+
- **Status:** draft. Imported from a design discussion and normalized into the
4+
proposal series (the pasted original suggested the then-taken number P-028).
5+
Note on scope: the subject is **not** the OwnLang analyzer itself but the
6+
legacy .NET Framework / WPF desktop application the audit targets (see
7+
[`audit/README.md`](../../audit/README.md)). The proposed module would ship
8+
as instrumentation guidance for audited legacy apps — e.g. alongside the
9+
runtime harnesses in [`audit/runtime/`](../../audit/runtime/README.md) — not
10+
as part of the analyzer.
11+
12+
## Summary
13+
14+
The legacy .NET application under audit (the WPF/.NET Framework desktop app targeted by `audit/README.md`) should gain a small, dependency-light module for compact runtime diagnostics and fast set operations using classic probabilistic and compressed data structures:
15+
16+
- bitsets / roaring-style bitmap indexes;
17+
- Top-K / heavy-hitter counters;
18+
- Count-Min Sketch for approximate frequencies;
19+
- t-digest or DDSketch-style latency summaries;
20+
- optional Bloom/Cuckoo filters for import and lookup pre-checks;
21+
- optional SimHash for grouping similar errors.
22+
23+
The goal is not to turn a legacy desktop .NET application into a fake distributed analytics platform. That would be architecture cosplay, and nobody needs that circus. The goal is narrower: improve local diagnostics, filtering, dirty tracking, and performance visibility without requiring Redis, Valkey, Kafka, or some other infrastructure animal.
24+
25+
## Problem
26+
27+
The legacy .NET application under audit has several known pain points:
28+
29+
- large legacy WPF/.NET Framework surface;
30+
- heavy dictionaries and reference data;
31+
- expensive recalculation paths;
32+
- memory-sensitive UI workflows;
33+
- difficult-to-debug performance spikes;
34+
- repeated validation and import scenarios;
35+
- need for better local evidence before changing architecture.
36+
37+
Current code can observe some issues, but it likely lacks compact, queryable runtime summaries:
38+
39+
- which operations are actually slow at p95/p99;
40+
- which validations fail most often;
41+
- which dictionary/reference entries are hot;
42+
- which rows/documents are affected by a recalculation;
43+
- which errors are effectively the same root cause;
44+
- which imports contain duplicates or obviously invalid references.
45+
46+
Without compact summaries, developers either over-log, under-measure, or guess. Guessing is not engineering. It is astrology with stack traces.
47+
48+
## Proposed solution
49+
50+
Add an internal module to the audited application, tentatively named
51+
`Own.Diagnostics.Sketches`.
52+
53+
The module should expose simple interfaces, not leak implementation details into business logic.
54+
55+
Example conceptual interfaces:
56+
57+
```csharp
58+
public interface ILatencySketch
59+
{
60+
void Record(long elapsedMilliseconds);
61+
LatencySnapshot Snapshot();
62+
}
63+
64+
public interface IHeavyHitters<T>
65+
{
66+
void Add(T item, long weight = 1);
67+
IReadOnlyList<HeavyHittersEntry<T>> Top(int count);
68+
}
69+
70+
public interface IApproxFrequency<T>
71+
{
72+
void Add(T item, long count = 1);
73+
long Estimate(T item);
74+
}
75+
76+
public interface IBitmapIndex
77+
{
78+
void Add(int id);
79+
void Remove(int id);
80+
bool Contains(int id);
81+
// Non-mutating: each set operation returns a new index; the receiver
82+
// and `other` are never modified (no aliasing surprises for callers).
83+
IBitmapIndex And(IBitmapIndex other);
84+
IBitmapIndex Or(IBitmapIndex other);
85+
IBitmapIndex Except(IBitmapIndex other);
86+
}
87+
```
88+
89+
The first implementation may be deliberately boring:
90+
91+
- `BitArray` / custom packed bitset for dense ids;
92+
- `HashSet<int>` fallback for sparse ids;
93+
- simple Space-Saving Top-K;
94+
- simple Count-Min Sketch;
95+
- latency sketch adapter with an initially simple histogram implementation.
96+
97+
The point is to introduce the model safely before chasing cleverness. Cleverness without containment is how a “small optimization” becomes a haunted subsystem.
98+
99+
## Candidate use cases
100+
101+
### 1. Dirty tracking and affected-row calculation
102+
103+
Use bitmap indexes to represent sets such as:
104+
105+
- rows with validation errors;
106+
- rows affected by changed customs rate;
107+
- rows requiring recalculation;
108+
- rows visible after current filter;
109+
- rows already processed;
110+
- rows excluded by user action.
111+
112+
Instead of scanning large collections repeatedly, compute set operations:
113+
114+
```text
115+
RowsToRecalculate =
116+
AffectedByRateChange
117+
AND CurrentDeclarationRows
118+
AND NOT AlreadyRecalculated
119+
```
120+
121+
This is especially suitable when ids are stable integer indexes within a document/import/session.
122+
123+
### 2. Validation and import diagnostics
124+
125+
Use Top-K and Count-Min Sketch to track:
126+
127+
- most frequent validation errors;
128+
- most frequent invalid TNVED codes;
129+
- most frequent import normalization problems;
130+
- most frequently missing reference data;
131+
- most common user correction patterns.
132+
133+
This helps answer:
134+
135+
Which 20 validation problems actually hurt users most?
136+
137+
Not “which validation problems look important in a meeting”, because apparently humans needed a database to learn humility.
138+
139+
### 3. Performance telemetry
140+
141+
Use latency sketches to record p50/p90/p95/p99 for operations such as:
142+
143+
- opening large WPF forms;
144+
- loading reference dictionaries;
145+
- graph 47 recalculation;
146+
- report generation;
147+
- import parsing;
148+
- SQL query wrappers;
149+
- UI filtering.
150+
151+
The output should be local and cheap:
152+
153+
```text
154+
Operation: LoadTnvedTree
155+
Count: 143
156+
p50: 120 ms
157+
p95: 2.4 s
158+
p99: 8.1 s
159+
Max: 9.6 s
160+
```
161+
162+
Average latency alone should be treated as suspicious. Averages hide pain like a rug hides broken glass.
163+
164+
### 4. Error grouping
165+
166+
Use SimHash-like fingerprints to group similar:
167+
168+
- exception messages;
169+
- stack traces;
170+
- validation failure clusters;
171+
- SQL error patterns.
172+
173+
This can later connect to the existing idea of error ids, hidden stack traces, and build-aware deobfuscation.
174+
175+
## Scope
176+
177+
### MVP
178+
179+
The MVP should include:
180+
181+
1. `ILatencySketch`
182+
2. `IHeavyHitters<T>`
183+
3. `IBitmapIndex`
184+
4. one local diagnostic sink:
185+
- JSON file;
186+
- text report;
187+
- or debug window export.
188+
5. instrumentation examples for 2–3 real operations.
189+
190+
Suggested first targets:
191+
192+
- dictionary/reference loading;
193+
- graph 47 recalculation;
194+
- validation/import flow.
195+
196+
### Phase 2
197+
198+
Add:
199+
200+
- Count-Min Sketch;
201+
- Bloom filter for import pre-checks;
202+
- SimHash grouping;
203+
- optional compact binary export;
204+
- analyzer/test coverage for misuse.
205+
206+
### Phase 3
207+
208+
Integrate with OwnAudit or 007 by exporting normalized evidence
209+
(OwnAudit's `docs/sketch-based-evidence.md` already anticipates ingesting
210+
these runtime diagnostic exports; the 007-side run evidence is specified in
211+
007's `docs/sketch-aware-evidence.md`):
212+
213+
```json
214+
{
215+
"schema": "own.sketches.v1",
216+
"source": "own.diagnostics.sketches",
217+
"operation": "LoadTnvedTree",
218+
"latency": {
219+
"p50_ms": 120,
220+
"p95_ms": 2400,
221+
"p99_ms": 8100
222+
},
223+
"top_errors": [],
224+
"affected_sets": []
225+
}
226+
```
227+
228+
## Non-goals
229+
230+
This proposal explicitly does not include:
231+
232+
- adding Redis/Valkey as a runtime dependency;
233+
- replacing SQL Server;
234+
- changing business rules;
235+
- introducing approximate answers into critical legal/business decisions;
236+
- using Bloom/HLL/Count-Min for authorization, licensing, billing, or correctness checks;
237+
- rewriting existing WPF flows around sketches.
238+
239+
Approximate structures may support diagnostics and optimization. They must not become the source of truth for business decisions. Works fine?! A cart with three wheels “works fine” too.
240+
241+
## Safety rules
242+
243+
1. Every approximate structure must expose its error model in docs.
244+
2. Approximate values must be named as estimates.
245+
3. Exact fallback must exist where correctness matters.
246+
4. Sketches must be resettable and exportable.
247+
5. No global mutable singleton dumping random metrics from everywhere.
248+
6. No business logic may depend on false-positive behavior.
249+
250+
## Acceptance criteria
251+
252+
The proposal is successful when:
253+
254+
- a developer can instrument an operation in fewer than 10 lines;
255+
- bitmap indexes can represent affected row sets and combine them efficiently;
256+
- p95/p99 latency is visible for selected operations;
257+
- Top-K diagnostics identify frequent validation/import issues;
258+
- exported evidence can be consumed later by OwnAudit;
259+
- no new infrastructure is required;
260+
- no correctness-sensitive path relies only on probabilistic results.
261+
262+
## Expected benefit
263+
264+
The audited legacy application gets a practical local observability and set-processing layer:
265+
266+
- fewer full scans;
267+
- better dirty tracking;
268+
- better recalculation targeting;
269+
- better import diagnostics;
270+
- clearer performance evidence;
271+
- less guessing before refactoring.
272+
273+
This is not highload cosplay. It is a small internal toolbox for making the old codebase confess where it hurts.

0 commit comments

Comments
 (0)