Calibration and conformal prediction for tabular models. Original C11 numerical engine, JavaScript/WASM and Python frontends. Probability calibration and array APIs for split conformal intervals and class sets, plus owning conformal classifiers and regressors, CV+/jackknife+, and joint regression regions are implemented, including simultaneous multilabel sets, recall risk control and selective prediction. Empirical distributions and split conformal predictive systems, owning predictive-system regressors, proper scores and calibration diagnostics are implemented. Native model-output adapters and broader ensemble integration remain planned.
RiskController fits a matrix of bounded per-observation losses, one column per
candidate action. Declare candidates and their preference/testing order before
using the calibration data. It supports two different guarantees:
method: 'crc': conformal risk control for expected loss. Each observation's losses must decrease along the candidate order. SupplyfallbackLossBound, a known pointwise bound on the final action, no greater thantargetRisk. The bound must hold for every possible observation; empirical zero is insufficient.method: 'ltt': Learn-Then-Test controls the probability of selecting an action whose population risk exceedstargetRisk.confidencedefaults to 0.95.bound: 'hoeffding_bentkus'accepts bounded fractional losses;binomialrequires losses equal to zero ormaximumLoss. Bonferroni tests the fixed family;correction: 'fixed_sequence'tests in reverse candidate order and stops at the first failure. No restart or reordering using calibration outcomes is permitted.
CRC uses exchangeability and monotone bounded losses; LTT assumes IID calibration
observations and a family fixed independently of them. selection.guarantee
records these assumptions. LTT can return a null selection if no candidate is
certified and no safe fallback was supplied. Individual fixed-sequence upper
bounds are not simultaneous bounds for every candidate; only the selection
procedure has the stated familywise guarantee. See the original
CRC and
Learn-Then-Test constructions.
RecallController supplies a concrete bounded loss: the fraction of true labels
missed by thresholding independent multilabel probabilities. An observation with
no true labels has loss zero. Its strictly decreasing thresholds end at zero,
which includes every label and provides the safe fallback. Threshold ties are
included. Its default grid is 1, 0.99, …, 0; custom grids must be frozen before
calibration. predict returns a binary label matrix, not a conformal region of
possible label vectors.
const { RecallController } = require('@wlearn/uncertainty')
const recall = await RecallController.create({ targetRisk: .1 })
recall.fit(calibrationProbabilities, calibrationLabels)
const selectedLabels = recall.predict(testProbabilities)
const decision = recall.selection // threshold, fallback and guaranteefrom wlearn_uncertainty import RecallController
recall = RecallController(target_risk=.1)
recall.fit(calibration_probabilities, calibration_labels)
selected_labels = recall.predict(test_probabilities)
decision = recall.selectionBoth controllers support WLRN save/load, transactional refitting and diagnostics.
Expected missed-label loss, high-probability population-risk control and
simultaneous label-vector coverage are distinct quantities. Evaluate label-set
size together with recall on untouched data. Small calibration sets may require
the full-label fallback.
SelectiveController accepts a fixed, increasing grid of finite confidence-score
thresholds. Scores need not be probabilities; larger means more confident.
fit(scores, errors) takes one finite score and one binary error (0 or 1) per
held-out observation. Both the predictor and the score function must already be
fixed. accept(scores) returns a byte mask, with threshold ties accepted.
const { SelectiveController } = require('@wlearn/uncertainty')
const selective = await SelectiveController.create({
thresholds: [.5, .6, .7, .8, .9], targetRisk: .05, confidence: .95
})
selective.fit(calibrationConfidence, calibrationErrors)
const accepted = selective.accept(testConfidence)Python uses SelectiveController(thresholds=[...], target_risk=.05, confidence=.95) with the same fit and accept methods. selection reports the
chosen threshold or null/None with abstention everywhere when none is certified.
The controlled quantity is population error conditional on acceptance.
Exact binomial tests use errors and counts only within each accepted subset.
Bonferroni correction is the default; reverse fixed_sequence testing stops at
the first failure, including an empty accepted subset. Conditional error need
not improve monotonically with the score threshold. Inspect acceptance rates,
accepted counts, empirical conditional errors and upper bounds in diagnostics.
Zero accepted observations give no risk estimate; abstention everywhere carries
no claim that an undefined conditional error is zero.
This fixed-family construction uses the conditional-error/binomial formulation of selective classification and simultaneous testing. It does not choose its grid from calibration scores. Do not substitute an all-row average that counts rejected predictions as zero errors: that measures a different loss. The assumptions and confidence apply to the selected population risk, not a promise about every future finite test sample.
EmpiricalDistribution consumes a core Prediction containing finite samples
with [row, draw, target] layout. It provides weighted marginal CDFs, inverted-CDF
quantiles and seeded resampling. Zero-weight atoms do not affect support extrema.
Weights are nonnegative relative draw weights, shared across rows. Omit them for
uniform weights. CDF queries have [row, target] shape and include equality ties;
quantile levels must strictly increase in [0,1], with endpoints returning the
positive-mass support extrema.
const { EmpiricalDistribution } = require('@wlearn/uncertainty')
const distribution = await EmpiricalDistribution.create()
distribution.fit({
rows: 1, sampleCount: 3, sampleKind: 'outcome',
samples: [10, 20, 30], sampleWeights: [1, 2, 1]
})
const cdf = distribution.cdf([[20]]) // 0.75
const quantiles = distribution.quantiles([.1, .5, .9])
const draws = distribution.sample(20, { seed: 7 })from wlearn.prediction import create_prediction
from wlearn_uncertainty import EmpiricalDistribution
distribution = EmpiricalDistribution().fit(create_prediction(
rows=1, sample_count=3, sample_kind='outcome',
samples=[10, 20, 30], sample_weights=[1, 2, 1]))
cdf = distribution.cdf([[20]])
quantiles = distribution.quantiles([.1, .5, .9])
draws = distribution.sample(20, seed=7)Declare sampleDependence: 'joint' (sample_dependence in Python) only when
target columns within each draw belong together. Joint resampling reuses one
draw index across targets. Undeclared or explicit marginal samples are resampled
independently by target and carry no joint-distribution claim. sampleKind: 'mean'
preserves uncertainty about the regression function; it does not become an
outcome distribution by resampling. CDF and quantile outputs retain this distinction
in metadata. An empirical distribution has no conformal calibration guarantee.
prediction returns a copy of the stored samples, normalized weights and axes.
WLRN roundtrips preserve these, including target names. A fixed seed and matching
rowOffset/row_offset make resampling stable across chunks with the same draw
count and target axis. Prediction rows are bound to this distribution object;
fit a new object on samples for a different query batch.
ConformalPredictiveSystem calibrates signed residuals. The default residual
method uses y - prediction; normalized divides by a positive scale supplied
for every row at calibration and inference. Freeze the predictor, scale function
and optional groupLabels before calibration. Python uses group_labels.
const { ConformalPredictiveSystem } = require('@wlearn/uncertainty')
const cps = await ConformalPredictiveSystem.create()
cps.fit(calibrationPredictions, calibrationTargets)
const ranks = cps.cdf(testPredictions, testTargets, { seed: 9 })
const intervals = cps.predictInterval(testPredictions, [.8, .9, .95])
const quantiles = cps.predictQuantiles(testPredictions, [.1, .5, .9], { bound: 'upper' })Python provides cdf, predict_interval and predict_quantiles, with seed=9
and bound='upper' as keyword arguments. cdf returns lower, upper and an
optional randomized vector. With n calibration scores, the bounds are
count(score < query)/(n+1) and (count(score <= query)+1)/(n+1). An independent
uniform randomizer selects a rank between them, including ties, as in the
split predictive-system construction.
Omit both seed and tau for deterministic rank envelopes. Supply a seed for
reproducible pseudorandom ranks and rowOffset/row_offset for chunk consistency,
or supply tau directly as a number/vector in [0,1]. These options are mutually
exclusive. Exact uniform-rank validity assumes exchangeability and an independent
uniform randomizer; a fixed supplied tau is not exactly uniform. A randomized
rank function has different tails from an ordinary finite empirical CDF. It is
not fed directly into CRPS as a proper CDF.
Quantiles require an explicit lower or upper envelope. Their finite-sample
uncertainty includes infinite extreme quantiles: level zero is -Infinity and
the upper level-one quantile is Infinity. Equal-tail intervals use the lower
and upper envelopes. Interval/set/region methods accept a single coverage number
or an increasing array of coverage levels. Sparse or unseen fixed groups retain conservative support;
an empty stratum returns a rank envelope [0,1] and an unbounded interval.
Diagnostics include calibration counts by stratum. WLRN artifacts preserve
signed score banks and method parameters, and failed refits retain prior state.
ConformalPredictiveRegressor owns a point estimator (which may itself be a
Pipeline) and an optional separate scale estimator. It shares fitting, invalidation
and nested persistence with ConformalRegressor. Training and calibration remain
separate operations:
const { ConformalPredictiveRegressor } = require('@wlearn/uncertainty')
const model = await ConformalPredictiveRegressor.create({ estimator: fittedPipeline })
model.calibrate(calibrationX, calibrationY)
const intervals = model.predictInterval(testX, [.9])
const ranks = model.predictCDF(testX, testY, { seed: 19 })For a normalized system, declare calibration: { method: 'normalized' } and
scaleEstimator, then call fitScale(developmentX, developmentY) before
calibration. Its positive scale floor defaults to 1e-12 and is persisted.
Optional row IDs detect overlap between declared training/development/calibration
rows; without IDs, the caller is responsible for their separation. fit,
fitScale and setParams invalidate calibration. Python uses
ConformalPredictiveRegressor(estimator, ...), fit_scale, predict_cdf,
predict_quantiles and predict_interval.
PredictionMetrics is stateless. JavaScript construction initializes WASM;
subsequent calls are synchronous. Python uses the same C kernels:
const { PredictionMetrics } = require('@wlearn/uncertainty')
const metrics = await PredictionMetrics.create()
const crps = metrics.score('crps', outcomePrediction, testY)
const losses = metrics.losses('pinball', quantilePrediction, testY)
const objective = metrics.measure('pinball', { levels: [.1, .5, .9] })from wlearn_uncertainty import PredictionMetrics
metrics = PredictionMetrics()
crps = metrics.score('crps', outcome_prediction, test_y)
objective = metrics.measure('pinball', levels=[.1, .5, .9])All scores are losses to minimize. losses returns a row matrix: columns flatten
target,level for pinball/interval loss, one column per target for CRPS, and one
column for energy/variogram loss. score averages columns equally and accepts
relative row weights (sampleWeight / sample_weight). These are distinct from
Prediction.sampleWeights, which weight predictive draws. A zero-weight row is
ignored even when its loss is infinite. Unbounded or empty intervals receive
infinite interval loss; endpoints with zero pinball coefficient contribute zero.
Pinball scores evaluate declared quantile levels. The interval score targets
central equal-tail intervals at each declared coverage; evaluating an asymmetric
interval with this score does not change its original coverage interpretation.
CRPS and energy follow Gneiting and Raftery.
Variogram loss uses the sum over all ordered target pairs with unit pair weights
and exponent power (default 0.5, strictly between 0 and 2), following
Scheuerer and Hamill.
Target units matter for joint scores; choose any rescaling independently of
held-out evaluation. Variogram loss is proper but not strictly proper.
Sample scores evaluate the supplied weighted empirical distribution, without a
Monte Carlo bias correction. CRPS sorts each margin; exact energy uses quadratic
work in the number of draws. All require sampleKind: 'outcome'. Multivariate
energy and variogram additionally require sampleDependence: 'joint'. Mean
posterior draws and undeclared dependence are rejected, not silently reinterpreted.
measure returns a core Measure definition, without registering a global name.
Quantile/interval objectives require fixed levels and reject predictions with
different axes. Pass predictionOptions / prediction_options to request options
such as { bound: 'upper' } for conformal quantile envelopes. Direct scores may
be infinite; core estimator scoring rejects nonfinite objectives for selection.
reliability(probabilities, binaryOutcomes) takes matching matrices of binary
events: one-vs-rest class columns, positive multilabel columns, or top-label
confidence paired with correctness. Preserve class order when constructing the
events. Each column reports fixed-bin counts, mean probability, observed frequency,
Brier loss, binary log loss, ECE and maximum calibration error. Empty-bin means
are null/None; impossible observed events give infinite log loss. Optional
edges must increase from zero to one; bins include their left edge, and the
last bin also includes one. Diagnostics are unweighted; do not present them as
weighted-population estimates.
empiricalPIT / empirical_pit returns the left/right empirical CDF at each
observed outcome, optionally randomized within atom jumps using tau or a seed.
Its [row,target] matrices can be passed to pitDiagnostics / pit_diagnostics.
Conformal predictive-system randomized ranks can likewise be supplied explicitly
as columns. PIT diagnostics report a histogram, mean, population variance, KS
distance and Cramer–von Mises statistic against uniformity. They provide no IID
p-values: test ranks sharing a fitted calibration sample need not be independent.
A fixed midpoint randomizer is not a uniform-rank guarantee. ECE and PIT summaries
are descriptive diagnostics, not substitutes for proper scores or coverage tests.
JavaScript: @wlearn/uncertainty. Python: wlearn-uncertainty, imported as
wlearn_uncertainty. Both depend on the wlearn core; no Polygrad, SciPy or BLAS
runtime dependency.
const { ProbabilityCalibrator } = require('@wlearn/uncertainty')
const calibration = await ProbabilityCalibrator.create({ method: 'isotonic' })
calibration.fit([.1, .2, .4, .6, .8, .9], [0, 1, 0, 1, 1, 1])
const probabilities = calibration.transform([.15, .7]) // DenseMatrix, two columns
const bytes = calibration.save() // WLRNfrom wlearn_uncertainty import ProbabilityCalibrator
calibration = ProbabilityCalibrator(method='isotonic')
calibration.fit([.1, .2, .4, .6, .8, .9], [0, 1, 0, 1, 1, 1])
probabilities = calibration.transform([.15, .7]) # NumPy matrix
calibration.save('calibration.wlrn')These small arrays demonstrate the API, not calibration quality. Fit maps on held-out predictions, then evaluate on separate untouched observations.
| Method | Construction | Input |
|---|---|---|
temperature |
One nonnegative inverse temperature, cross-entropy fit | Probabilities or logits |
sigmoid |
Platt scaling with smoothed targets | Probabilities or scores |
isotonic |
Weighted pooled-adjacent-violators, linear interpolation, clipped extrapolation | Probabilities or scores |
beta |
Logistic map of log probability and negative log complement; nonnegative slopes | Probabilities |
venn_abers |
Isotonic fits for both hypothetical labels | Probabilities or scores, unit weights |
input defaults to probabilities; logits and scores require explicit
selection. Temperature uses clipped log probabilities when supplied probabilities.
Beta and log conversion clip at 1e-15. l2 defaults to zero and regularizes slope
parameters when requested. tolerance and maxIterations control optimization;
Python spells the latter max_iterations. Inspect diagnostics for convergence,
iteration count, objective and knot count. An iteration limit does not become a
successful-convergence claim. Temperature searches a bounded scaled coefficient;
separable data can have no finite maximum-likelihood temperature.
Classification uses explicit class columns. Supply classes for nonstandard
labels or classes absent from the calibration sample. Binary one-column inputs
refer to the positive class; two-column non-temperature fits use the second
class. Multiclass non-temperature calibration fits one-vs-rest maps and normalizes
rows; all-zero rows become uniform. This reduction does not transfer binary
Venn–Abers validity to the normalized multiclass point probabilities.
For independent labels use task: 'multilabel' and matrix targets, optionally
targetNames (target_names in Python). Each column is calibrated separately;
columns are not normalized across labels. Temperature uses a binary map per label.
These independent maps do not define a joint label distribution.
predictPairs / predict_pairs preserves Venn–Abers' two hypothetical-label
outputs. JavaScript uses [row, column, hypothetical label] flat data; Python
returns an array with those axes. The point probability uses the log-loss minimax
formula. Pairs are not confidence intervals; multiclass normalization and later
ensembling have distinct semantics.
const { CalibratedClassifier } = require('@wlearn/uncertainty')
// predictor implements wlearn's lifecycle; it may already be a fitted Pipeline.
const model = await CalibratedClassifier.create({
estimator: predictor, calibration: { method: 'temperature' }
})
// If the predictor is not already fitted:
await model.fit(XTrain, yTrain)
await model.calibrate(XCalibration, yCalibration)
const probabilities = await model.predictProba(XTest)Python uses CalibratedClassifier(predictor, calibration={'method': 'temperature'}),
fit, calibrate and predict_proba. Ordinary predictions select the largest
calibrated class probability; multilabel predictions threshold each column at 0.5.
The wrapper owns the transferred predictor. Do not mutate it externally. Refitting or changing predictor parameters invalidates calibration before mutation. Failed recalibration preserves valid existing maps; failed predictor fitting requires a successful refit. Updating parameters can recover a failed fit. Predictions stay synchronous for synchronous JavaScript predictors and become Promises when needed.
fit and calibrate accept rowIds (row_ids); the wrapper rejects known overlap.
A prefitted predictor can supply trainingRowIds (training_row_ids) at creation.
Raw arrays without identities cannot prove their training history. Training weights
are not conformal importance weights. Normal calibration accepts row weights;
Venn–Abers accepts only unit weights.
Call dispose() when replacing many objects or finishing long-running workflows.
Wrappers dispose their children. WLRN stores numeric map blobs and nested predictor
artifacts, class order and supplied training identities. Python save(path=None)
returns bytes and optionally writes a file; loaders accept bytes or paths.
CrossVennAbersClassifier accepts an estimator specification and fits its own
complementary folds. JavaScript uses
await CrossVennAbersClassifier.create({ estimator: ['model', ModelClass, params], cv: 5 });
Python uses CrossVennAbersClassifier(('model', ModelClass, params), cv=5).
Call fit(X, y) once; each fold predictor calibrates on its held-out observations.
The default fold assignment is independent of labels. Explicit folds must cover
each row once and train on every other row. Rare-class folds retain the global
class axis; constant training folds require no classifier backend.
Point probabilities use the published geometric log-loss aggregation of fold
Venn pairs. predictFoldPairs / predict_fold_pairs exposes the separate outputs
as [row, fold, column, hypothetical label]. This aggregation is not an interval
or a general multiclass calibration guarantee. Multilabel columns remain
independent. Point inference is chunked (chunkSize / chunk_size); explicit
pair outputs are bounded and may require smaller caller batches. Training weights
affect fold model fitting, while held-out Venn observations have unit weights.
Saved models retain all fold predictors and maps. Loading permits inference;
refitting requires a new estimator specification through setParams / set_params.
const { IntervalCalibrator, SetCalibrator } = require('@wlearn/uncertainty')
const intervals = await IntervalCalibrator.create({ method: 'absolute' })
intervals.fit(calibrationPredictions, calibrationTargets)
const result = intervals.predictInterval(testPredictions, [.8, .9, .95])
// result.interval: [row, target=0, coverage, lower/upper], contiguous
// result.metadata.uncertainty: coverage scope and assumptions
const sets = await SetCalibrator.create({ method: 'aps', classes: [0, 1, 2] })
sets.fit(calibrationProbabilities, calibrationLabels)
const classification = sets.predictSet(testProbabilities, [.9])
// classification.sets: [row, coverage, class], uint8 membershipPython exposes IntervalCalibrator and SetCalibrator with predict_interval
and predict_set. Results use wlearn's Prediction object: flat .interval or
.sets, explicit .rows, .coverage_levels and class/target metadata. Reshape
using those axes; do not infer interval meaning from an unlabeled matrix.
Interval methods are absolute, normalized, cqr and asymmetric_cqr.
Normalized fitting and inference require a positive scale value per row from a
separately fitted, frozen scale model. CQR takes two ordered prediction columns;
crossing inputs reject. Negative CQR corrections are retained. Empty resulting
intervals use [Infinity, -Infinity]; unsupported ranks use [-Infinity, Infinity].
Asymmetric CQR allocates the lower-tail error fraction with lowerTailFraction
(lower_tail_fraction), default 0.5. Tail allocation is fixed before calibration.
Set methods are lac, aps and raps. RAPS accepts nonnegative penalty and
regularizedAfter (regularized_after); choose them on separate development data.
Default scores are deterministic. Class-probability ties follow declared class
order. Randomized APS/RAPS require randomized: true and an explicit seed.
Calibration and prediction use separate reproducible random streams. For chunked
randomized inference, pass the batch's starting rowOffset (row_offset) so the
result matches one complete call; default offset is zero.
Both APIs accept a fixed groupLabels (group_labels) axis at construction and
one groups label per row at fitting/inference. Groups declared but absent during
calibration, and unseen prediction groups, receive conservative outputs; they are
never pooled. Undeclared calibration groups reject. classConditional: true
(class_conditional) calibrates class-specific strata, crossed with groups when
provided. diagnostics retains calibration counts, including the empty fallback
stratum. Sparse strata can make outputs uninformative.
Coverage statements are marginal over calibration/test draws within the specified stratum, assuming exchangeability and a predictor/auxiliary models fixed before calibration. They are not pointwise guarantees or guarantees under arbitrary distribution shift. Array inputs cannot establish their own training history. Ordinary training weights are not accepted as conformal calibration weights.
References: CQR, APS/RAPS, jackknife+/CV+ ranks and bounds.
ConformalClassifier owns an estimator and exposes the same set calibration:
await ConformalClassifier.create({ estimator: predictor, calibration: { method: 'aps' } }).
Call fit(XTrain, yTrain) if needed, then calibrate(XCalibration, yCalibration)
and predictSet(XTest, [.9]). Python uses ConformalClassifier(predictor, calibration=...)
and predict_set. Ordinary predict and predictProba/predict_proba retain the
underlying classifier's outputs; set construction does not alter its probabilities.
Row identities, invalidation and ownership follow the calibrated-classifier rules
above. Nested artifacts retain the predictor and the set-calibration state.
ConformalRegressor similarly wraps a scalar regressor or Pipeline:
await ConformalRegressor.create({ estimator: predictor }). Its default is absolute
residual split conformal. Fit the predictor, calibrate on held-out observations,
then call predictInterval(XTest, [.8, .9, .95]) (predict_interval in Python).
Ordinary predictions continue to come from the supplied primary estimator.
For normalized conformal, supply scaleEstimator (scale_estimator) and
calibration: { method: 'normalized' }. Call fitScale(XDevelopment, yDevelopment)
(fit_scale) on separate development observations: it fits the scale model to
absolute primary-model residuals. Its predictions are clamped in C to a positive
minimumScale (minimum_scale, default 1e-12 in target units), fixed before
calibration. This handles zero or negative scale-model predictions, including
constant targets. Set the floor to a scale appropriate to the target units.
Prefitted scale models are also supported. Supply scaleTrainingRowIds
(scale_training_row_ids) for their known training history. Calibration checks
both primary and auxiliary identities when available.
CQR accepts either a primary model implementing the structured predictQuantiles
contract, with two fixed quantileLevels (default [.05, .95]), or explicit
lowerEstimator and upperEstimator models (lower_estimator, upper_estimator).
Configure those models' quantile objectives yourself; uncertainty contains no
model-family parameter logic. The primary model may also be the lower or upper
model: aliased roles fit/dispose once and retain their identity in WLRN. In that
case ordinary point predictions remain those of the chosen primary model.
The scale predictor must be a separate model because it fits a different target.
src/: canonical C11 algorithms and checked buffer ABI.js/: asynchronous WASM construction, synchronous numerical calls, model ownership, class/target axes, WLRN and browser bundles.py/: matching native extension and orchestration; no Python numerical reimplementation.test/: C, JS/Python lifecycle, reference and interoperability checks.
The core defines Prediction, errors and WLRN. Uncertainty does not depend on AutoML or a particular model package. Array APIs work with external prediction producers.
make test JOBS=4 builds C and runs native checks. npm run build --prefix js
synchronizes canonical C and builds WASM. Install the declared JS dependencies,
then run npm test --prefix js, npm run test:types --prefix js and
npm run build:browser --prefix js. Python tests use pytest test/; C development
can set UNCERTAINTY_LIB_PATH to a built library. Generated C package copies are
checked by node js/scripts/sync-csrc.js --check.
The C API checks dimensions and caps each workspace buffer at 256 MiB; WASM can grow up to 1 GiB. No parallel numerical kernels are used. Isotonic state stores unique sorted scores and fitted probabilities. Venn state stores sorted score groups, counts and positive counts; prediction currently performs PAV per query and hypothetical label. Scaling measurements and further optimization are pending.
References: temperature scaling, beta calibration, Venn–Abers, scikit-learn calibration conventions. Implementation is original; reference packages are test-only dependencies. Apache-2.0; WASM/runtime and embedded core notices accompany the npm distribution.
CVPlusRegressor fits complementary fold predictors and retains each held-out
absolute residual with its originating model. Ordinary point predictions average
fold predictions. Interval prediction uses the CV+ order statistics, not a pooled
residual radius around that average.
const { CVPlusRegressor } = require('@wlearn/uncertainty')
const { RFModel } = require('@wlearn/rf')
const intervals = await CVPlusRegressor.create({
estimator: ['rf', RFModel, { task: 'regression', seed: 42 }], cv: 5
})
await intervals.fit(X, y)
const prediction = intervals.predictInterval(Xtest, [.9, .95])Python uses CVPlusRegressor(('rf', RFModel, {'seed': 42}), cv=5),
fit(X, y) and predict_interval(Xtest, [.9, .95]). Supply a factory for any
compatible scalar regressor, including a complete Pipeline so preparation fits
inside each fold. Use method: 'jackknife_plus' (Python method='jackknife_plus')
and omit cv to fit one model per omitted observation. Its training complements
are temporary; its artifact stores implicit membership. Model storage and fitting
still grow with the number of observations. chunkSize / chunk_size bounds
inference batches. Model fitting weights are sliced by fold; residual ranks are
unweighted. Loaded predictors need a new estimator specification before refitting.
Requested coverage levels specify the published endpoint ranks. They are not
the worst-case guarantee: jackknife+ guarantees at least 1 - 2α when the requested
level is 1 - α, assuming exchangeable observations and a permutation-symmetric
fitting algorithm. Equal-fold CV+ has an additional finite-sample correction.
The returned metadata.uncertainty.coverageBound reports that bound separately.
Custom or unequal folds return a null bound and an explicit unverified status;
no rows are dropped to make folds equal. These distinctions follow
Barber et al., jackknife+ and CV+ Theorems 1 and 4.
Tuning or preprocessing outside the retained fold predictors can violate the
assumptions. Row-dependent training weights require exchangeability of the weighted
observations and a symmetric fitting procedure as well.
RegionCalibrator consumes matrices of predictions and targets. Rectangles use
one conformal score per row: the maximum absolute residual divided by its fixed
target scale. Ellipsoids use the Mahalanobis residual norm under a fixed positive
definite covariance. The calibrated radius controls joint target coverage.
It does not fit a separate interval for each target.
Use scales or covariance to provide a shape chosen independently of calibration.
Without either, the shape is identity. Alternatively, call fitShape / fit_shape
on separate development predictions and targets before fit on calibration data.
The helper computes RMS scales for rectangles or centered sample covariance for
ellipsoids. shrinkage (default0.01) multiplies off-diagonal covariances by
1 - shrinkage; minimumScale / minimum_scale (default1e-12 in target units)
adds its square to covariance diagonals or floors rectangle scales. Singular
results fail rather than triggering an undocumented regularization retry.
const { RegionCalibrator } = require('@wlearn/uncertainty')
const regions = await RegionCalibrator.create({ method: 'ellipsoid' })
regions.fitShape(developmentPredictions, yDevelopment, { rowIds: developmentIds })
regions.fit(calibrationPredictions, yCalibration, { rowIds: calibrationIds })
const joint = regions.predictRegion(testPredictions, [.9, .95])Python has fit_shape, fit, and predict_region. Rectangles return the existing
Prediction.interval layout [row,target,coverage,bound]; ellipsoids return
centers, a shared precision matrix, and radii. metadata.uncertainty identifies
joint coverage and contains log volumes [row,coverage]. A zero-radius region
has log volume -Infinity; insufficient calibration support yields an unbounded
region. Group calibration uses the same fixed groups and conservative unseen-group
behavior as scalar intervals. The implementation uses a global frozen shape;
Messoudi et al. also study
instance-dependent shapes, which this API does not implement.
ConformalMultiOutputRegressor owns a compatible multioutput predictor, including
@wlearn/ensemble's MultiOutputRegressor. Its sequence is fit, optional
fitShape, then calibrate on separate rows; inference is predictRegion.
Python uses ConformalMultiOutputRegressor with the corresponding snake_case
methods. predictInterval is available for rectangles. Refitting the predictor
invalidates its shape and calibration; a failed calibration preserves its prior
state. Known training/development/calibration row identities are checked for overlap.
The nested WLRN artifact preserves the predictor, frozen shape and calibration.
MultiLabelSetCalibrator consumes independent positive-state probabilities
[row,label]; rows need not sum to one. It calibrates LAC, APS or RAPS binary
scores for each label. Fixed allocationWeights / allocation_weights distribute
the total error budget across labels; the default is equal allocation. Weights
are relative nonnegative shares with a positive total. A zero share assigns zero
error to that label and retains both states. The allocation must be chosen before
calibration; updating it invalidates the fitted map.
const { MultiLabelSetCalibrator } = require('@wlearn/uncertainty')
const sets = await MultiLabelSetCalibrator.create({ allocationWeights: [1, 2, 1] })
sets.fit(calibrationProbabilities, calibrationLabels)
const result = sets.predictSet(testProbabilities, [.9, .95])The compact Prediction.sets layout is [row,coverage,label,state], where states
are zero and one. The represented label-vector set is the Cartesian product of
those allowed states; no label vectors are enumerated. The union bound yields
whole-vector coverage under the stated split-conformal assumptions, without
assuming independence between labels. Metadata keeps that joint bound separate
from the allocated per-label coverage levels. This is not an expected-recall
guarantee. Sparse or unseen groups and insufficient calibration counts retain
both states as needed. Randomized APS/RAPS use an explicit seed and row-label
stream; rowOffset / row_offset preserves chunk equivalence.
For an owned model or Pipeline, use the existing ConformalClassifier with
task: 'multilabel' (Python task='multilabel'). Its fit, calibrate and
predictSet lifecycle is the same as for a classifier; probability outputs keep
the label matrix axis and classes() returns null/None. It composes with the
ensemble package's MultiLabelClassifier and persists all heads and calibration
in its nested WLRN artifact.