Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

uncertainty

Calibration and conformal prediction for tabular models. Original C11 numerical engine, JavaScript/WASM and Python frontends. Probability calibration and array APIs for split conformal intervals and class sets, plus owning conformal classifiers and regressors, CV+/jackknife+, and joint regression regions are implemented, including simultaneous multilabel sets, recall risk control and selective prediction. Empirical distributions and split conformal predictive systems, owning predictive-system regressors, proper scores and calibration diagnostics are implemented. Native model-output adapters and broader ensemble integration remain planned.

Risk-controlled decisions

RiskController fits a matrix of bounded per-observation losses, one column per candidate action. Declare candidates and their preference/testing order before using the calibration data. It supports two different guarantees:

  • method: 'crc': conformal risk control for expected loss. Each observation's losses must decrease along the candidate order. Supply fallbackLossBound, a known pointwise bound on the final action, no greater than targetRisk. The bound must hold for every possible observation; empirical zero is insufficient.
  • method: 'ltt': Learn-Then-Test controls the probability of selecting an action whose population risk exceeds targetRisk. confidence defaults to 0.95. bound: 'hoeffding_bentkus' accepts bounded fractional losses; binomial requires losses equal to zero or maximumLoss. Bonferroni tests the fixed family; correction: 'fixed_sequence' tests in reverse candidate order and stops at the first failure. No restart or reordering using calibration outcomes is permitted.

CRC uses exchangeability and monotone bounded losses; LTT assumes IID calibration observations and a family fixed independently of them. selection.guarantee records these assumptions. LTT can return a null selection if no candidate is certified and no safe fallback was supplied. Individual fixed-sequence upper bounds are not simultaneous bounds for every candidate; only the selection procedure has the stated familywise guarantee. See the original CRC and Learn-Then-Test constructions.

RecallController supplies a concrete bounded loss: the fraction of true labels missed by thresholding independent multilabel probabilities. An observation with no true labels has loss zero. Its strictly decreasing thresholds end at zero, which includes every label and provides the safe fallback. Threshold ties are included. Its default grid is 1, 0.99, …, 0; custom grids must be frozen before calibration. predict returns a binary label matrix, not a conformal region of possible label vectors.

const { RecallController } = require('@wlearn/uncertainty')
const recall = await RecallController.create({ targetRisk: .1 })
recall.fit(calibrationProbabilities, calibrationLabels)
const selectedLabels = recall.predict(testProbabilities)
const decision = recall.selection // threshold, fallback and guarantee
from wlearn_uncertainty import RecallController
recall = RecallController(target_risk=.1)
recall.fit(calibration_probabilities, calibration_labels)
selected_labels = recall.predict(test_probabilities)
decision = recall.selection

Both controllers support WLRN save/load, transactional refitting and diagnostics. Expected missed-label loss, high-probability population-risk control and simultaneous label-vector coverage are distinct quantities. Evaluate label-set size together with recall on untouched data. Small calibration sets may require the full-label fallback.

Selective prediction

SelectiveController accepts a fixed, increasing grid of finite confidence-score thresholds. Scores need not be probabilities; larger means more confident. fit(scores, errors) takes one finite score and one binary error (0 or 1) per held-out observation. Both the predictor and the score function must already be fixed. accept(scores) returns a byte mask, with threshold ties accepted.

const { SelectiveController } = require('@wlearn/uncertainty')
const selective = await SelectiveController.create({
  thresholds: [.5, .6, .7, .8, .9], targetRisk: .05, confidence: .95
})
selective.fit(calibrationConfidence, calibrationErrors)
const accepted = selective.accept(testConfidence)

Python uses SelectiveController(thresholds=[...], target_risk=.05, confidence=.95) with the same fit and accept methods. selection reports the chosen threshold or null/None with abstention everywhere when none is certified.

The controlled quantity is population error conditional on acceptance. Exact binomial tests use errors and counts only within each accepted subset. Bonferroni correction is the default; reverse fixed_sequence testing stops at the first failure, including an empty accepted subset. Conditional error need not improve monotonically with the score threshold. Inspect acceptance rates, accepted counts, empirical conditional errors and upper bounds in diagnostics. Zero accepted observations give no risk estimate; abstention everywhere carries no claim that an undefined conditional error is zero.

This fixed-family construction uses the conditional-error/binomial formulation of selective classification and simultaneous testing. It does not choose its grid from calibration scores. Do not substitute an all-row average that counts rejected predictions as zero errors: that measures a different loss. The assumptions and confidence apply to the selected population risk, not a promise about every future finite test sample.

Empirical predictive distributions

EmpiricalDistribution consumes a core Prediction containing finite samples with [row, draw, target] layout. It provides weighted marginal CDFs, inverted-CDF quantiles and seeded resampling. Zero-weight atoms do not affect support extrema. Weights are nonnegative relative draw weights, shared across rows. Omit them for uniform weights. CDF queries have [row, target] shape and include equality ties; quantile levels must strictly increase in [0,1], with endpoints returning the positive-mass support extrema.

const { EmpiricalDistribution } = require('@wlearn/uncertainty')
const distribution = await EmpiricalDistribution.create()
distribution.fit({
  rows: 1, sampleCount: 3, sampleKind: 'outcome',
  samples: [10, 20, 30], sampleWeights: [1, 2, 1]
})
const cdf = distribution.cdf([[20]])             // 0.75
const quantiles = distribution.quantiles([.1, .5, .9])
const draws = distribution.sample(20, { seed: 7 })
from wlearn.prediction import create_prediction
from wlearn_uncertainty import EmpiricalDistribution
distribution = EmpiricalDistribution().fit(create_prediction(
    rows=1, sample_count=3, sample_kind='outcome',
    samples=[10, 20, 30], sample_weights=[1, 2, 1]))
cdf = distribution.cdf([[20]])
quantiles = distribution.quantiles([.1, .5, .9])
draws = distribution.sample(20, seed=7)

Declare sampleDependence: 'joint' (sample_dependence in Python) only when target columns within each draw belong together. Joint resampling reuses one draw index across targets. Undeclared or explicit marginal samples are resampled independently by target and carry no joint-distribution claim. sampleKind: 'mean' preserves uncertainty about the regression function; it does not become an outcome distribution by resampling. CDF and quantile outputs retain this distinction in metadata. An empirical distribution has no conformal calibration guarantee.

prediction returns a copy of the stored samples, normalized weights and axes. WLRN roundtrips preserve these, including target names. A fixed seed and matching rowOffset/row_offset make resampling stable across chunks with the same draw count and target axis. Prediction rows are bound to this distribution object; fit a new object on samples for a different query batch.

Conformal predictive systems

ConformalPredictiveSystem calibrates signed residuals. The default residual method uses y - prediction; normalized divides by a positive scale supplied for every row at calibration and inference. Freeze the predictor, scale function and optional groupLabels before calibration. Python uses group_labels.

const { ConformalPredictiveSystem } = require('@wlearn/uncertainty')
const cps = await ConformalPredictiveSystem.create()
cps.fit(calibrationPredictions, calibrationTargets)
const ranks = cps.cdf(testPredictions, testTargets, { seed: 9 })
const intervals = cps.predictInterval(testPredictions, [.8, .9, .95])
const quantiles = cps.predictQuantiles(testPredictions, [.1, .5, .9], { bound: 'upper' })

Python provides cdf, predict_interval and predict_quantiles, with seed=9 and bound='upper' as keyword arguments. cdf returns lower, upper and an optional randomized vector. With n calibration scores, the bounds are count(score < query)/(n+1) and (count(score <= query)+1)/(n+1). An independent uniform randomizer selects a rank between them, including ties, as in the split predictive-system construction.

Omit both seed and tau for deterministic rank envelopes. Supply a seed for reproducible pseudorandom ranks and rowOffset/row_offset for chunk consistency, or supply tau directly as a number/vector in [0,1]. These options are mutually exclusive. Exact uniform-rank validity assumes exchangeability and an independent uniform randomizer; a fixed supplied tau is not exactly uniform. A randomized rank function has different tails from an ordinary finite empirical CDF. It is not fed directly into CRPS as a proper CDF.

Quantiles require an explicit lower or upper envelope. Their finite-sample uncertainty includes infinite extreme quantiles: level zero is -Infinity and the upper level-one quantile is Infinity. Equal-tail intervals use the lower and upper envelopes. Interval/set/region methods accept a single coverage number or an increasing array of coverage levels. Sparse or unseen fixed groups retain conservative support; an empty stratum returns a rank envelope [0,1] and an unbounded interval. Diagnostics include calibration counts by stratum. WLRN artifacts preserve signed score banks and method parameters, and failed refits retain prior state.

ConformalPredictiveRegressor owns a point estimator (which may itself be a Pipeline) and an optional separate scale estimator. It shares fitting, invalidation and nested persistence with ConformalRegressor. Training and calibration remain separate operations:

const { ConformalPredictiveRegressor } = require('@wlearn/uncertainty')
const model = await ConformalPredictiveRegressor.create({ estimator: fittedPipeline })
model.calibrate(calibrationX, calibrationY)
const intervals = model.predictInterval(testX, [.9])
const ranks = model.predictCDF(testX, testY, { seed: 19 })

For a normalized system, declare calibration: { method: 'normalized' } and scaleEstimator, then call fitScale(developmentX, developmentY) before calibration. Its positive scale floor defaults to 1e-12 and is persisted. Optional row IDs detect overlap between declared training/development/calibration rows; without IDs, the caller is responsible for their separation. fit, fitScale and setParams invalidate calibration. Python uses ConformalPredictiveRegressor(estimator, ...), fit_scale, predict_cdf, predict_quantiles and predict_interval.

Proper scores and calibration diagnostics

PredictionMetrics is stateless. JavaScript construction initializes WASM; subsequent calls are synchronous. Python uses the same C kernels:

const { PredictionMetrics } = require('@wlearn/uncertainty')
const metrics = await PredictionMetrics.create()
const crps = metrics.score('crps', outcomePrediction, testY)
const losses = metrics.losses('pinball', quantilePrediction, testY)
const objective = metrics.measure('pinball', { levels: [.1, .5, .9] })
from wlearn_uncertainty import PredictionMetrics

metrics = PredictionMetrics()
crps = metrics.score('crps', outcome_prediction, test_y)
objective = metrics.measure('pinball', levels=[.1, .5, .9])

All scores are losses to minimize. losses returns a row matrix: columns flatten target,level for pinball/interval loss, one column per target for CRPS, and one column for energy/variogram loss. score averages columns equally and accepts relative row weights (sampleWeight / sample_weight). These are distinct from Prediction.sampleWeights, which weight predictive draws. A zero-weight row is ignored even when its loss is infinite. Unbounded or empty intervals receive infinite interval loss; endpoints with zero pinball coefficient contribute zero.

Pinball scores evaluate declared quantile levels. The interval score targets central equal-tail intervals at each declared coverage; evaluating an asymmetric interval with this score does not change its original coverage interpretation. CRPS and energy follow Gneiting and Raftery. Variogram loss uses the sum over all ordered target pairs with unit pair weights and exponent power (default 0.5, strictly between 0 and 2), following Scheuerer and Hamill. Target units matter for joint scores; choose any rescaling independently of held-out evaluation. Variogram loss is proper but not strictly proper.

Sample scores evaluate the supplied weighted empirical distribution, without a Monte Carlo bias correction. CRPS sorts each margin; exact energy uses quadratic work in the number of draws. All require sampleKind: 'outcome'. Multivariate energy and variogram additionally require sampleDependence: 'joint'. Mean posterior draws and undeclared dependence are rejected, not silently reinterpreted.

measure returns a core Measure definition, without registering a global name. Quantile/interval objectives require fixed levels and reject predictions with different axes. Pass predictionOptions / prediction_options to request options such as { bound: 'upper' } for conformal quantile envelopes. Direct scores may be infinite; core estimator scoring rejects nonfinite objectives for selection.

reliability(probabilities, binaryOutcomes) takes matching matrices of binary events: one-vs-rest class columns, positive multilabel columns, or top-label confidence paired with correctness. Preserve class order when constructing the events. Each column reports fixed-bin counts, mean probability, observed frequency, Brier loss, binary log loss, ECE and maximum calibration error. Empty-bin means are null/None; impossible observed events give infinite log loss. Optional edges must increase from zero to one; bins include their left edge, and the last bin also includes one. Diagnostics are unweighted; do not present them as weighted-population estimates.

empiricalPIT / empirical_pit returns the left/right empirical CDF at each observed outcome, optionally randomized within atom jumps using tau or a seed. Its [row,target] matrices can be passed to pitDiagnostics / pit_diagnostics. Conformal predictive-system randomized ranks can likewise be supplied explicitly as columns. PIT diagnostics report a histogram, mean, population variance, KS distance and Cramer–von Mises statistic against uniformity. They provide no IID p-values: test ranks sharing a fitted calibration sample need not be independent. A fixed midpoint randomizer is not a uniform-rank guarantee. ECE and PIT summaries are descriptive diagnostics, not substitutes for proper scores or coverage tests.

Probability calibration

JavaScript: @wlearn/uncertainty. Python: wlearn-uncertainty, imported as wlearn_uncertainty. Both depend on the wlearn core; no Polygrad, SciPy or BLAS runtime dependency.

const { ProbabilityCalibrator } = require('@wlearn/uncertainty')
const calibration = await ProbabilityCalibrator.create({ method: 'isotonic' })
calibration.fit([.1, .2, .4, .6, .8, .9], [0, 1, 0, 1, 1, 1])
const probabilities = calibration.transform([.15, .7]) // DenseMatrix, two columns
const bytes = calibration.save()                       // WLRN
from wlearn_uncertainty import ProbabilityCalibrator
calibration = ProbabilityCalibrator(method='isotonic')
calibration.fit([.1, .2, .4, .6, .8, .9], [0, 1, 0, 1, 1, 1])
probabilities = calibration.transform([.15, .7])  # NumPy matrix
calibration.save('calibration.wlrn')

These small arrays demonstrate the API, not calibration quality. Fit maps on held-out predictions, then evaluate on separate untouched observations.

Method Construction Input
temperature One nonnegative inverse temperature, cross-entropy fit Probabilities or logits
sigmoid Platt scaling with smoothed targets Probabilities or scores
isotonic Weighted pooled-adjacent-violators, linear interpolation, clipped extrapolation Probabilities or scores
beta Logistic map of log probability and negative log complement; nonnegative slopes Probabilities
venn_abers Isotonic fits for both hypothetical labels Probabilities or scores, unit weights

input defaults to probabilities; logits and scores require explicit selection. Temperature uses clipped log probabilities when supplied probabilities. Beta and log conversion clip at 1e-15. l2 defaults to zero and regularizes slope parameters when requested. tolerance and maxIterations control optimization; Python spells the latter max_iterations. Inspect diagnostics for convergence, iteration count, objective and knot count. An iteration limit does not become a successful-convergence claim. Temperature searches a bounded scaled coefficient; separable data can have no finite maximum-likelihood temperature.

Classification uses explicit class columns. Supply classes for nonstandard labels or classes absent from the calibration sample. Binary one-column inputs refer to the positive class; two-column non-temperature fits use the second class. Multiclass non-temperature calibration fits one-vs-rest maps and normalizes rows; all-zero rows become uniform. This reduction does not transfer binary Venn–Abers validity to the normalized multiclass point probabilities.

For independent labels use task: 'multilabel' and matrix targets, optionally targetNames (target_names in Python). Each column is calibrated separately; columns are not normalized across labels. Temperature uses a binary map per label. These independent maps do not define a joint label distribution.

predictPairs / predict_pairs preserves Venn–Abers' two hypothetical-label outputs. JavaScript uses [row, column, hypothetical label] flat data; Python returns an array with those axes. The point probability uses the log-loss minimax formula. Pairs are not confidence intervals; multiclass normalization and later ensembling have distinct semantics.

Owning an estimator

const { CalibratedClassifier } = require('@wlearn/uncertainty')
// predictor implements wlearn's lifecycle; it may already be a fitted Pipeline.
const model = await CalibratedClassifier.create({
  estimator: predictor, calibration: { method: 'temperature' }
})
// If the predictor is not already fitted:
await model.fit(XTrain, yTrain)
await model.calibrate(XCalibration, yCalibration)
const probabilities = await model.predictProba(XTest)

Python uses CalibratedClassifier(predictor, calibration={'method': 'temperature'}), fit, calibrate and predict_proba. Ordinary predictions select the largest calibrated class probability; multilabel predictions threshold each column at 0.5.

The wrapper owns the transferred predictor. Do not mutate it externally. Refitting or changing predictor parameters invalidates calibration before mutation. Failed recalibration preserves valid existing maps; failed predictor fitting requires a successful refit. Updating parameters can recover a failed fit. Predictions stay synchronous for synchronous JavaScript predictors and become Promises when needed.

fit and calibrate accept rowIds (row_ids); the wrapper rejects known overlap. A prefitted predictor can supply trainingRowIds (training_row_ids) at creation. Raw arrays without identities cannot prove their training history. Training weights are not conformal importance weights. Normal calibration accepts row weights; Venn–Abers accepts only unit weights.

Call dispose() when replacing many objects or finishing long-running workflows. Wrappers dispose their children. WLRN stores numeric map blobs and nested predictor artifacts, class order and supplied training identities. Python save(path=None) returns bytes and optionally writes a file; loaders accept bytes or paths.

Cross Venn–Abers

CrossVennAbersClassifier accepts an estimator specification and fits its own complementary folds. JavaScript uses await CrossVennAbersClassifier.create({ estimator: ['model', ModelClass, params], cv: 5 }); Python uses CrossVennAbersClassifier(('model', ModelClass, params), cv=5). Call fit(X, y) once; each fold predictor calibrates on its held-out observations. The default fold assignment is independent of labels. Explicit folds must cover each row once and train on every other row. Rare-class folds retain the global class axis; constant training folds require no classifier backend.

Point probabilities use the published geometric log-loss aggregation of fold Venn pairs. predictFoldPairs / predict_fold_pairs exposes the separate outputs as [row, fold, column, hypothetical label]. This aggregation is not an interval or a general multiclass calibration guarantee. Multilabel columns remain independent. Point inference is chunked (chunkSize / chunk_size); explicit pair outputs are bounded and may require smaller caller batches. Training weights affect fold model fitting, while held-out Venn observations have unit weights. Saved models retain all fold predictors and maps. Loading permits inference; refitting requires a new estimator specification through setParams / set_params.

Conformal prediction arrays

const { IntervalCalibrator, SetCalibrator } = require('@wlearn/uncertainty')
const intervals = await IntervalCalibrator.create({ method: 'absolute' })
intervals.fit(calibrationPredictions, calibrationTargets)
const result = intervals.predictInterval(testPredictions, [.8, .9, .95])
// result.interval: [row, target=0, coverage, lower/upper], contiguous
// result.metadata.uncertainty: coverage scope and assumptions

const sets = await SetCalibrator.create({ method: 'aps', classes: [0, 1, 2] })
sets.fit(calibrationProbabilities, calibrationLabels)
const classification = sets.predictSet(testProbabilities, [.9])
// classification.sets: [row, coverage, class], uint8 membership

Python exposes IntervalCalibrator and SetCalibrator with predict_interval and predict_set. Results use wlearn's Prediction object: flat .interval or .sets, explicit .rows, .coverage_levels and class/target metadata. Reshape using those axes; do not infer interval meaning from an unlabeled matrix.

Interval methods are absolute, normalized, cqr and asymmetric_cqr. Normalized fitting and inference require a positive scale value per row from a separately fitted, frozen scale model. CQR takes two ordered prediction columns; crossing inputs reject. Negative CQR corrections are retained. Empty resulting intervals use [Infinity, -Infinity]; unsupported ranks use [-Infinity, Infinity]. Asymmetric CQR allocates the lower-tail error fraction with lowerTailFraction (lower_tail_fraction), default 0.5. Tail allocation is fixed before calibration.

Set methods are lac, aps and raps. RAPS accepts nonnegative penalty and regularizedAfter (regularized_after); choose them on separate development data. Default scores are deterministic. Class-probability ties follow declared class order. Randomized APS/RAPS require randomized: true and an explicit seed. Calibration and prediction use separate reproducible random streams. For chunked randomized inference, pass the batch's starting rowOffset (row_offset) so the result matches one complete call; default offset is zero.

Both APIs accept a fixed groupLabels (group_labels) axis at construction and one groups label per row at fitting/inference. Groups declared but absent during calibration, and unseen prediction groups, receive conservative outputs; they are never pooled. Undeclared calibration groups reject. classConditional: true (class_conditional) calibrates class-specific strata, crossed with groups when provided. diagnostics retains calibration counts, including the empty fallback stratum. Sparse strata can make outputs uninformative.

Coverage statements are marginal over calibration/test draws within the specified stratum, assuming exchangeability and a predictor/auxiliary models fixed before calibration. They are not pointwise guarantees or guarantees under arbitrary distribution shift. Array inputs cannot establish their own training history. Ordinary training weights are not accepted as conformal calibration weights.

References: CQR, APS/RAPS, jackknife+/CV+ ranks and bounds.

ConformalClassifier owns an estimator and exposes the same set calibration: await ConformalClassifier.create({ estimator: predictor, calibration: { method: 'aps' } }). Call fit(XTrain, yTrain) if needed, then calibrate(XCalibration, yCalibration) and predictSet(XTest, [.9]). Python uses ConformalClassifier(predictor, calibration=...) and predict_set. Ordinary predict and predictProba/predict_proba retain the underlying classifier's outputs; set construction does not alter its probabilities. Row identities, invalidation and ownership follow the calibrated-classifier rules above. Nested artifacts retain the predictor and the set-calibration state.

ConformalRegressor similarly wraps a scalar regressor or Pipeline: await ConformalRegressor.create({ estimator: predictor }). Its default is absolute residual split conformal. Fit the predictor, calibrate on held-out observations, then call predictInterval(XTest, [.8, .9, .95]) (predict_interval in Python). Ordinary predictions continue to come from the supplied primary estimator.

For normalized conformal, supply scaleEstimator (scale_estimator) and calibration: { method: 'normalized' }. Call fitScale(XDevelopment, yDevelopment) (fit_scale) on separate development observations: it fits the scale model to absolute primary-model residuals. Its predictions are clamped in C to a positive minimumScale (minimum_scale, default 1e-12 in target units), fixed before calibration. This handles zero or negative scale-model predictions, including constant targets. Set the floor to a scale appropriate to the target units. Prefitted scale models are also supported. Supply scaleTrainingRowIds (scale_training_row_ids) for their known training history. Calibration checks both primary and auxiliary identities when available.

CQR accepts either a primary model implementing the structured predictQuantiles contract, with two fixed quantileLevels (default [.05, .95]), or explicit lowerEstimator and upperEstimator models (lower_estimator, upper_estimator). Configure those models' quantile objectives yourself; uncertainty contains no model-family parameter logic. The primary model may also be the lower or upper model: aliased roles fit/dispose once and retain their identity in WLRN. In that case ordinary point predictions remain those of the chosen primary model. The scale predictor must be a separate model because it fits a different target.

Architecture and development

  • src/: canonical C11 algorithms and checked buffer ABI.
  • js/: asynchronous WASM construction, synchronous numerical calls, model ownership, class/target axes, WLRN and browser bundles.
  • py/: matching native extension and orchestration; no Python numerical reimplementation.
  • test/: C, JS/Python lifecycle, reference and interoperability checks.

The core defines Prediction, errors and WLRN. Uncertainty does not depend on AutoML or a particular model package. Array APIs work with external prediction producers.

make test JOBS=4 builds C and runs native checks. npm run build --prefix js synchronizes canonical C and builds WASM. Install the declared JS dependencies, then run npm test --prefix js, npm run test:types --prefix js and npm run build:browser --prefix js. Python tests use pytest test/; C development can set UNCERTAINTY_LIB_PATH to a built library. Generated C package copies are checked by node js/scripts/sync-csrc.js --check.

The C API checks dimensions and caps each workspace buffer at 256 MiB; WASM can grow up to 1 GiB. No parallel numerical kernels are used. Isotonic state stores unique sorted scores and fitted probabilities. Venn state stores sorted score groups, counts and positive counts; prediction currently performs PAV per query and hypothetical label. Scaling measurements and further optimization are pending.

References: temperature scaling, beta calibration, Venn–Abers, scikit-learn calibration conventions. Implementation is original; reference packages are test-only dependencies. Apache-2.0; WASM/runtime and embedded core notices accompany the npm distribution.

CV+ and jackknife+

CVPlusRegressor fits complementary fold predictors and retains each held-out absolute residual with its originating model. Ordinary point predictions average fold predictions. Interval prediction uses the CV+ order statistics, not a pooled residual radius around that average.

const { CVPlusRegressor } = require('@wlearn/uncertainty')
const { RFModel } = require('@wlearn/rf')
const intervals = await CVPlusRegressor.create({
  estimator: ['rf', RFModel, { task: 'regression', seed: 42 }], cv: 5
})
await intervals.fit(X, y)
const prediction = intervals.predictInterval(Xtest, [.9, .95])

Python uses CVPlusRegressor(('rf', RFModel, {'seed': 42}), cv=5), fit(X, y) and predict_interval(Xtest, [.9, .95]). Supply a factory for any compatible scalar regressor, including a complete Pipeline so preparation fits inside each fold. Use method: 'jackknife_plus' (Python method='jackknife_plus') and omit cv to fit one model per omitted observation. Its training complements are temporary; its artifact stores implicit membership. Model storage and fitting still grow with the number of observations. chunkSize / chunk_size bounds inference batches. Model fitting weights are sliced by fold; residual ranks are unweighted. Loaded predictors need a new estimator specification before refitting.

Requested coverage levels specify the published endpoint ranks. They are not the worst-case guarantee: jackknife+ guarantees at least 1 - 2α when the requested level is 1 - α, assuming exchangeable observations and a permutation-symmetric fitting algorithm. Equal-fold CV+ has an additional finite-sample correction. The returned metadata.uncertainty.coverageBound reports that bound separately. Custom or unequal folds return a null bound and an explicit unverified status; no rows are dropped to make folds equal. These distinctions follow Barber et al., jackknife+ and CV+ Theorems 1 and 4. Tuning or preprocessing outside the retained fold predictors can violate the assumptions. Row-dependent training weights require exchangeability of the weighted observations and a symmetric fitting procedure as well.

Joint regression regions

RegionCalibrator consumes matrices of predictions and targets. Rectangles use one conformal score per row: the maximum absolute residual divided by its fixed target scale. Ellipsoids use the Mahalanobis residual norm under a fixed positive definite covariance. The calibrated radius controls joint target coverage. It does not fit a separate interval for each target.

Use scales or covariance to provide a shape chosen independently of calibration. Without either, the shape is identity. Alternatively, call fitShape / fit_shape on separate development predictions and targets before fit on calibration data. The helper computes RMS scales for rectangles or centered sample covariance for ellipsoids. shrinkage (default0.01) multiplies off-diagonal covariances by 1 - shrinkage; minimumScale / minimum_scale (default1e-12 in target units) adds its square to covariance diagonals or floors rectangle scales. Singular results fail rather than triggering an undocumented regularization retry.

const { RegionCalibrator } = require('@wlearn/uncertainty')
const regions = await RegionCalibrator.create({ method: 'ellipsoid' })
regions.fitShape(developmentPredictions, yDevelopment, { rowIds: developmentIds })
regions.fit(calibrationPredictions, yCalibration, { rowIds: calibrationIds })
const joint = regions.predictRegion(testPredictions, [.9, .95])

Python has fit_shape, fit, and predict_region. Rectangles return the existing Prediction.interval layout [row,target,coverage,bound]; ellipsoids return centers, a shared precision matrix, and radii. metadata.uncertainty identifies joint coverage and contains log volumes [row,coverage]. A zero-radius region has log volume -Infinity; insufficient calibration support yields an unbounded region. Group calibration uses the same fixed groups and conservative unseen-group behavior as scalar intervals. The implementation uses a global frozen shape; Messoudi et al. also study instance-dependent shapes, which this API does not implement.

ConformalMultiOutputRegressor owns a compatible multioutput predictor, including @wlearn/ensemble's MultiOutputRegressor. Its sequence is fit, optional fitShape, then calibrate on separate rows; inference is predictRegion. Python uses ConformalMultiOutputRegressor with the corresponding snake_case methods. predictInterval is available for rectangles. Refitting the predictor invalidates its shape and calibration; a failed calibration preserves its prior state. Known training/development/calibration row identities are checked for overlap. The nested WLRN artifact preserves the predictor, frozen shape and calibration.

Simultaneous multilabel sets

MultiLabelSetCalibrator consumes independent positive-state probabilities [row,label]; rows need not sum to one. It calibrates LAC, APS or RAPS binary scores for each label. Fixed allocationWeights / allocation_weights distribute the total error budget across labels; the default is equal allocation. Weights are relative nonnegative shares with a positive total. A zero share assigns zero error to that label and retains both states. The allocation must be chosen before calibration; updating it invalidates the fitted map.

const { MultiLabelSetCalibrator } = require('@wlearn/uncertainty')
const sets = await MultiLabelSetCalibrator.create({ allocationWeights: [1, 2, 1] })
sets.fit(calibrationProbabilities, calibrationLabels)
const result = sets.predictSet(testProbabilities, [.9, .95])

The compact Prediction.sets layout is [row,coverage,label,state], where states are zero and one. The represented label-vector set is the Cartesian product of those allowed states; no label vectors are enumerated. The union bound yields whole-vector coverage under the stated split-conformal assumptions, without assuming independence between labels. Metadata keeps that joint bound separate from the allocated per-label coverage levels. This is not an expected-recall guarantee. Sparse or unseen groups and insufficient calibration counts retain both states as needed. Randomized APS/RAPS use an explicit seed and row-label stream; rowOffset / row_offset preserves chunk equivalence.

For an owned model or Pipeline, use the existing ConformalClassifier with task: 'multilabel' (Python task='multilabel'). Its fit, calibrate and predictSet lifecycle is the same as for a classifier; probability outputs keep the label matrix axis and classes() returns null/None. It composes with the ensemble package's MultiLabelClassifier and persists all heads and calibration in its nested WLRN artifact.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages