Skip to content

Add calibration metrics and patient-clustered bootstrap CIs - #1268

Merged
jhnwu3 merged 1 commit into
sunlabuiuc:masterfrom
solarsys:feat/calibration-metrics-bootstrap
Oct 8, 2026
Merged

jhnwu3 merged 1 commit into
sunlabuiuc:masterfrom
solarsys:feat/calibration-metrics-bootstrap

Conversation

@solarsys

@solarsys solarsys commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Problem

Clinical prediction papers, and reporting guidelines such as TRIPOD+AI, expect:

  • calibration measures: Brier score, observed/expected ratio, calibration slope and intercept;
  • confidence intervals for every metric, using a patient-level (cluster) bootstrap, because samples from one patient are correlated;
  • paired differences when comparing models.

binary_metrics_fn had discrimination metrics and ECE only, and PyHealth had no interval utilities, so downstream projects wrote these by hand.

Changes

Calibration in binary_metrics_fn (new metric names):

Name Definition Ideal
brier Mean squared error of the probabilities (sklearn.metrics.brier_score_loss) 0
oe_ratio Observed / expected events, sum(y_true) / sum(y_prob) 1
calibration_slope b in the logistic recalibration y ~ a + b·logit(p); below 1 = overconfident 1
calibration_intercept Calibration-in-the-large: a in y ~ a + offset(logit(p)); above 0 = risks underestimated 0

The intercept follows the calibration hierarchy (Van Calster et al.): it is fitted with the slope fixed at 1. The logistic fits are a small unpenalised Newton solve in NumPy, so there's no new dependency.

New pyhealth.metrics.bootstrap (exported from pyhealth.metrics):

bootstrap_ci(y_true, y_prob, "pr_auc", groups=patient_ids, n_boot=1000, seed=0, alpha=0.05)
# {'estimate': …, 'lower': …, 'upper': …, 'n_boot': …, 'n_skipped': …}
paired_bootstrap_diff(y_true, y_prob_a, y_prob_b, "brier", groups=patient_ids)
  • Percentile intervals.
  • groups= resamples whole patients.
  • The paired version scores both models on identical resamples.
  • Resamples containing a single class are skipped and counted in n_skipped.
  • Results are deterministic given seed.
  • metric is any binary_metrics_fn name or a callable.

Trainer.evaluate(..., ci=True) is left for a follow-up.

Tests, docs, example

  • New tests/core/test_calibration_bootstrap.py (11 tests):
    • brier matches scikit-learn;
    • oe_ratio is exactly observed/expected;
    • a calibrated synthetic model gives slope ≈ 1, intercept ≈ 0, O/E ≈ 1;
    • an overconfident one gives slope ≈ 0.5;
    • an under-predicting one gives intercept ≈ +1;
    • the slope equals scikit-learn's unpenalised LogisticRegression to 4 decimal places;
    • the bootstrap is deterministic and its interval contains the estimate;
    • whole patients are resampled (checked inside a callable metric);
    • single-class resamples are skipped and counted;
    • the paired difference of a model with itself is exactly 0.
  • New docs page pyhealth.metrics.bootstrap, and the new metric names in the binary_metrics_fn docstring. The >>> examples were run as doctests.
  • New examples/calibration_and_bootstrap_ci.py: two models with identical ROC-AUC and PR-AUC, where one is clearly miscalibrated (slope 0.49 [0.33, 0.65], O/E 2.5). It also shows the paired Brier difference, whose interval excludes 0.
  • Full core suite: Ran 1408 tests … OK (skipped=76). tools/check_pr_rules.py passes.

🤖 Generated with Claude Code

binary_metrics_fn had discrimination metrics and ECE, but none of the
calibration measures TRIPOD+AI asks for, and PyHealth had no way to put
confidence intervals on a metric or on the difference between models.

- binary_metrics_fn: "brier" (sklearn), "oe_ratio" (observed/expected
  events), "calibration_slope" (b in y ~ a + b*logit(p)) and
  "calibration_intercept" (calibration-in-the-large: a with logit(p) as
  an offset). The logistic fits are a small unpenalised Newton solve in
  NumPy; no new dependency.
- New pyhealth.metrics.bootstrap: bootstrap_ci and paired_bootstrap_diff.
  Percentile intervals; `groups=` resamples whole patients (cluster
  bootstrap); the paired difference scores both models on identical
  resamples; single-class resamples are skipped and counted; results are
  deterministic given `seed`. Metric = a binary_metrics_fn name or a
  callable. Exported from pyhealth.metrics.
- tests/core/test_calibration_bootstrap.py: brier vs sklearn, O/E, a
  calibrated model (slope ~1, intercept ~0, O/E ~1), an overconfident one
  (slope ~0.5), an under-predicting one (intercept ~+1), slope equal to
  sklearn's unpenalised logistic regression to 4 places; bootstrap
  determinism, whole-patient resampling, skipped single-class resamples,
  identical resamples in the paired difference.
- docs: pyhealth.metrics.bootstrap page; metric names in the docstring.
- examples/calibration_and_bootstrap_ci.py.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

@jhnwu3 jhnwu3 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@jhnwu3
jhnwu3 merged commit 91f6449 into sunlabuiuc:master Oct 8, 2026
2 checks passed
@solarsys
solarsys deleted the feat/calibration-metrics-bootstrap branch October 9, 2026 04:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants