QUANTIX BIO · STUDIES
PYTHON 3.9 · NUMPY 2.0 LAST RUN 2026-08-07 RESEARCH USE ONLY
← BACK TO QUANTIX BIO
QUANTIXMIND / BIO / STUDIES 01–19 · JUL TO AUG 2026

The
Studies,
In Full

Each study below is broken down to its parts: the question, the cohort, the method, the numbers it produced, the figures exactly as the code drew them, and the limitations a reviewer would raise. Studies designed and run by Karanvir Singh, July and August 2026. Dataset citations are listed with each study and in the provenance table at the end.

ILLUSTRATION · DNA MOTIF rotational · 12s period
19completed studies
160k+records, beats, molecules
11manuscripts drafted
01 keptresult that went the wrong way
⚠ BINDING

Nothing in this work is FDA cleared, CE marked, or clinically validated, and nothing here is medical advice. These are methods studies on public research datasets, research use only. All performance numbers are computed out of fold, on data the models never touched during fitting.

ACQUIREpublic datasets, provenance logged
ƒ
MODELmethods written in numpy
VALIDATEout of fold, bootstrapped
FIGUREdrawn by our own plotter
PUBLISHresults, code, limitations
SEC/00

Manuscripts

tiered: journal track · technical reports · demonstrations

PAPER 1 · TIER 1, JOURNAL TRACK · 8 DATASETS

Post hoc recalibration can make clinical risk models worse

Platt scaling, beta calibration and isotonic regression, applied to holdout sets of 25 to 200 observations across eight clinical datasets, harm calibration more often than they help below roughly 100 observations or 25 minority events.

72%repeats harmed at n=25
8datasets
50repeats per cell
READ THE FULL PAPER →
PAPER 2 · TIER 3, METHODOLOGY DEMONSTRATION · SIMULATION

Evaluating early warning systems by lead time per unit of alarm burden

Two ward alerting systems with identical detection rates differ two fold in false alarms per patient day. The proposed operating curve makes the difference visible; single threshold metrics hide it.

½the alarm burden
6.6 hmedian lead time
READ THE FULL PAPER →
PAPER 3 · METHODS STUDY · 8 DATASETS

When abstention helps: error concentration under selective prediction

Letting a model defer its least confident cases nearly eliminates error on separable tasks, where the bottom fifth of confidence holds 90 percent of errors, but captures only a third of errors on the hard tasks where clinical models actually live.

90%errors caught, easy tasks
32%errors caught, hard tasks
READ THE FULL PAPER →
PAPER 4 · APPLIED STUDY · REAL ICU DATA

Sepsis early warning across two hospital systems

The lead time per alarm burden framework on real hourly ICU vitals from the PhysioNet 2019 cohort: derived on hospital system A, externally validated on hospital system B, quantifying what threshold scores cost in false alarms per patient day.

2hospital systems
hourlyreal vitals
READ THE FULL PAPER →
PAPER 5 · SIGNAL PROCESSING · 109,494 BEATS

A dependency free QRS detector on the complete MIT-BIH database

Pan and Tompkins from the 1985 paper in two hundred lines of numpy, with hand written binary format readers, evaluated beat by beat on all 48 records: 98.6 percent sensitivity, and a census of exactly which pathologies defeat the classical pipeline.

98.6%gross sensitivity
48records, fixed params
READ THE FULL PAPER →
PAPER 6 · TIER 1, JOURNAL TRACK · 45,412 ADULTS

Discrimination survives, calibration decays: twenty years of NHANES

A frozen 2002 diabetes score carried through 2018: AUC flat, calibration error up fourfold, tracking prevalence at correlation 0.96. The prospective recalibration rule from Paper 1 beats always and never recalibrating over the full two decades.

0.96drift vs prevalence corr
10survey cycles
calibration decay
READ THE FULL PAPER →
PAPER 7 · CARDIOLOGY · 108,534 BEATS

Ventricular ectopy defeats RR interval AF screening

Beat level, leave one record out: interval irregularity finds AF with AUC 0.935 in ectopy free rhythm and 0.826 where ectopy is present, which is half of all windows in this ambulatory population. The wearable screening premise, stratified.

0.935AUC, clean rhythm
0.826AUC, with ectopy
50%windows with ectopy
READ THE FULL PAPER →
PAPER 8 · SIGNAL PROCESSING · INTER PATIENT

Honest ECG beat classification: nine features, disjoint patients

Under the strict de Chazal patient split with thresholds frozen before testing: ventricular beats at 92 percent sensitivity, and the famous supraventricular failure diagnosed precisely as an inter patient calibration problem, not a capacity one.

92.4%V sensitivity
0.986V ranking AUC
READ THE FULL PAPER →
PAPER 9 · TIER 1, JOURNAL TRACK · HEALTH EQUITY

Aggregate monitoring masks sixfold subgroup calibration disparity

Paper 6's drift decomposed by race, sex and age: by 2018 the frozen model is six times more miscalibrated for Mexican American than white participants, along an axis the model never sees, with a survey weighted replication.

subgroup disparity
10 yrearlier subgroup alarm
READ THE FULL PAPER →
PAPER 10 · HEALTH EQUITY · 14,080 ADULTS

Removing the race column does not remove the race information

Routine clinical measurements reconstruct self identified race and ethnicity at AUC up to 0.76, socioeconomic variables add almost nothing on top, and the race blind score still distributes flags unequally at clinical thresholds. Measured, not assumed, with the social category framing stated first.

0.76reconstruction AUC
9 ptsflag rate gap
p<.005score leakage
READ THE FULL PAPER →
PAPER 11 · TIER 1, JOURNAL TRACK · TWO HOSPITALS

Covariate shift is the wrong worry

The feature family that fingerprints its hospital at AUC 0.99 transports its sepsis signal intact; the family a site classifier cannot distinguish collapses and inverts externally. Practice pattern features fail across sites; physiology travels.

0.99site AUC, transports fine
−0.33transport share, inverts
READ THE FULL PAPER →
SEC/01

Interactive Bench

the actual result data, live

Recalibration explorer

Drag the calibration set size and watch when recalibration stops hurting. Bars are mean test ECE over 50 repeated splits from Paper 1's experiment; lower is better.

25
isotonic vs none

Abstention explorer

Choose how many cases the model may defer to a clinician. Numbers are means over 20 repeated cross validations from Paper 3's experiment.

80%
accuracy on answered cases
accuracy, no abstention
cases deferred

Outbreak simulator

The 1978 boarding school influenza outbreak (dots are the real daily counts, 763 boys). Drag transmission and recovery and watch the epidemic respond; the fitted values from Study 05 are marked.

1.67 0.44
3.77reproduction number R0
peak boys in bed
final attack rate
SEC/02

Study Index

jump to a study

QB/01 · CARDIOLOGY

Calibrated, abstaining coronary risk

Calibration, abstention, and what deferring uncertain cases buys.

REAL · n=303NEGATIVE RESULT KEPT
QB/02 · ONCOLOGY

Decision curve analysis of breast masses

Past AUC, to the question of whether the model helps at all.

REAL · n=569
QB/03 · SURVIVAL

Advanced lung cancer prognosis

Cox regression from the partial likelihood up, proven against the literature.

REAL · n=228VALIDATED
QB/04 · DISCOVERY

Aqueous solubility from five descriptors

Most of a landmark QSAR model, in numbers anyone can audit.

REAL · n=1128
QB/05 · EPIDEMIOLOGY

SIR dynamics of a 1978 outbreak

A fully observed epidemic reduced to two parameters.

REAL · N=763
QB/06 · PATIENT SAFETY

Alarm fatigue vs warning time

The tradeoff nurses feel, measured where ground truth is exact.

SYNTHETIC BY DESIGN
QB/07 · CALIBRATION

When does recalibration hurt?

Eight datasets, four strategies, and the sample size where helping begins. Basis of Paper 1.

REAL · 8 DATASETSPAPER 1
QB/08 · SELECTIVE PREDICTION

How much does abstention buy?

Errors concentrate in low confidence only when the task is nearly separable. Basis of Paper 3.

REAL · 8 DATASETSPAPER 3
SEC/03

The Studies

eight foundational dossiers; studies 09 to 14 are documented inside the manuscripts above

QB/01

Calibrated, abstaining coronary disease risk

UCI CLEVELAND · 303 PATIENTS5 FOLD CVNEGATIVE RESULT
The question

A risk model in front of a clinician must have literal probabilities, and it must know when to defer. Does isotonic recalibration plus selective prediction deliver that on a small real cohort, and what does abstention buy per case deferred?

The cohort

303 real patients from the Cleveland Clinic: 13 clinical features from age and chest pain type through fluoroscopy vessel count. Outcome is angiographically confirmed coronary disease, prevalence 46 percent. Six missing values, median imputed.

The method

L2 regularized logistic regression solved by Newton's method. Isotonic calibration by pool adjacent violators, fit on a held out quarter of each training fold so it never sees test data. Abstention inside a widening band around 0.5. Bootstrap confidence intervals, 2000 resamples.

What we found

Recalibration backfired: with roughly 55 calibration points per fold, the isotonic step function overfit and calibration error doubled. The raw logistic model was already nearly calibrated. Meanwhile abstention worked exactly as designed: declining the most uncertain 43 percent of cases lifted accuracy on the answered cases from 82 to 95 percent. We kept both findings.

Limitations. Single site, a 1980s cohort, n of 303. These are methods results on a research dataset; nothing transfers to any modern population as clinical performance.
$ python3 experiments/exp01_calibrated_heart_risk/run.py
Results · out of fold
AUC, 95% bootstrap CI0.897 · [0.860, 0.930]
Brier score0.126
Calibration error, raw0.052
Calibration error, after isotonic0.099 ▲ worse
Accuracy, full coverage82.2%
Accuracy at 57% coverage90.8%
Accuracy at 34% coverage95.2%
-0 0.2 0.4 0.6 0.8 1 -0 0.2 0.4 0.6 0.8 1 ROC, 5 fold cross validated (n=303) False positive rate True positive rate logistic AUC 0.897 (95% CI 0.860 to 0.930)
QB/01 · A — receiver operating characteristic
0.4 0.5 0.6 0.7 0.8 0.9 1 0.825 0.85 0.875 0.9 0.925 0.95 Selective prediction: accuracy vs coverage Coverage (fraction of cases answered) Accuracy on answered cases abstain inside band around 0.5
QB/01 · B — selective prediction, accuracy vs coverage
QB/02

Decision curve analysis for breast mass classification

WDBC · 569 SAMPLES5 FOLD CV
The question

Discrimination is not benefit. Across the range of threshold probabilities a clinician might hold, does acting on the model beat the two default strategies, biopsy everyone and biopsy no one? That is the question decision curve analysis answers, and the one AUC cannot.

The cohort

569 real fine needle aspirate samples from the Wisconsin Diagnostic Breast Cancer study: 30 morphometric features of cell nuclei as mean, standard error and worst value. 37 percent malignant.

The method

Regularized logistic regression, every number out of fold. Net benefit at threshold t weighs true positives against false positives at the odds t over one minus t, the standard formulation in which the threshold itself encodes how a clinician trades a missed cancer against an unnecessary biopsy.

What we found

The model dominates both defaults at every plotted threshold. At a working threshold of ten percent, the margin over biopsying everyone is 5.8 net true positives per hundred patients. The heaviest coefficients are worst texture, radius error, worst radius and worst concave points: large, irregular, deeply concave nuclei, which is precisely the pathologist's account of malignancy. The model's reasoning is open and it agrees with the field.

Limitations. A single early 1990s dataset from one laboratory. Near ceiling AUC on WDBC is well known; it says the dataset is separable, not that the problem is solved.
$ python3 experiments/exp02_breast_cancer_decision_curve/run.py
Results · out of fold
AUC, 95% bootstrap CI0.995 · [0.988, 0.999]
Brier score0.022
Net benefit at t = 0.100.360
Net benefit, biopsy everyone0.303
Margin per 100 patients+5.8 net TP
-0 0.1 0.2 0.3 0.4 0.5 0.6 -0.6 -0.4 -0.2 -5.6e-17 0.2 0.4 Decision curve analysis, 5 fold out of fold (n=569) Threshold probability Net benefit logistic model biopsy everyone biopsy no one
QB/02 · A — decision curve, net benefit vs threshold
-0 2 4 6 8 0.55 0.6 0.65 0.7 0.75 Largest standardized coefficients log odds per SD
QB/02 · B — largest standardized coefficients
QB/03

Survival analysis of advanced lung cancer

NCCTG TRIAL · 228 PATIENTSVALIDATED VS LITERATURE
The question

How much prognostic signal do sex and ECOG performance status carry in advanced lung cancer, and can survival machinery written from scratch, Kaplan Meier, the log rank test, Cox proportional hazards, be shown correct against the field's reference implementation?

The cohort

228 real patients from the North Central Cancer Treatment Group trial, 164 deaths observed, the remainder censored. One patient dropped for missing performance status.

The method

Kaplan Meier with Greenwood variance. Two group log rank. Cox proportional hazards fit by Newton's method on the Breslow partial likelihood, standard errors from the inverse information matrix, discrimination by Harrell's concordance.

What we found

The hazard ratios reproduce the published reference values to two decimals: 0.58 for female sex, 1.59 per ECOG point, age contributing little once performance status is in the model. Median survival 270 days for men against 426 for women. The agreement is the finding: an auditable implementation, validated externally, which is the exact property the wider Quantix Bio suite claims.

Limitations. A single trial cohort predating modern systemic therapy; absolute survival times do not transfer to current patients. Concordance of 0.637 is honest: three baseline covariates only explain so much.
$ python3 experiments/exp03_lung_survival/run.py
Results
HR, female sex · 95% CI0.58 · [0.41, 0.80]
HR, per ECOG point · 95% CI1.59 · [1.27, 1.98]
HR, age per decade · 95% CI1.12 · [0.93, 1.34]
Log rank by sexχ² 10.0 · p = .0016
Median survival, men / women270 / 426 days
Harrell's concordance0.637
-0 200 400 600 800 1.0e+03 -0 0.2 0.4 0.6 0.8 1 Overall survival by sex, NCCTG lung cancer (n=227) Days from enrolment Survival probability male female
QB/03 · A — Kaplan Meier by sex, Greenwood bands
-0 200 400 600 800 1.0e+03 -0 0.2 0.4 0.6 0.8 1 Overall survival by ECOG performance status Days from enrolment Survival probability ECOG 0 (n=63) ECOG 1 (n=113) ECOG 2 to 3 (n=51)
QB/03 · B — Kaplan Meier by ECOG performance status
QB/04

Aqueous solubility from five physicochemical descriptors

ESOL · 1,128 MOLECULES10 FOLD CV
The question

Solubility gates oral bioavailability, which makes predicting it the first question asked of any drug candidate. How much of Delaney's landmark ESOL model can a fully transparent ridge regression on five descriptors recover, and which descriptors carry the signal?

The data

1,128 small molecules with experimentally measured aqueous log solubility, with molecular weight, hydrogen bond donors, ring count, rotatable bonds and polar surface area precomputed.

The method

Ridge regression, ten fold cross validation, every metric out of fold, compared against the ESOL reference predictions shipped with the dataset and against a predict the mean baseline.

What we found

Five interpretable numbers recover most of the reference model: 1.19 log units of error against ESOL's 0.91, from a baseline of 2.10. The physics reads correctly, heavier and more ring laden molecules dissolve worse, polar surface dissolves better. The remaining 0.28 log units localize precisely what the missing lipophilicity term is worth, since logP requires atom level structure our descriptor set does not carry.

Limitations. Descriptors taken as given rather than computed from structure. Linear, no interactions. Solubility measurements themselves carry roughly half a log unit of experimental noise, which bounds any model.
$ python3 experiments/exp04_esol_solubility/run.py
Results · out of fold
Baseline, predict the meanRMSE 2.10
This model, five descriptorsRMSE 1.19 · R² 0.68
ESOL reference modelRMSE 0.91
Effect, molecular weight−1.34 logS / SD
Effect, polar surface area+1.14 logS / SD
-10 -7.5 -5 -2.5 0 2.5 -10 -7.5 -5 -2.5 0 2.5 Measured vs predicted log solubility, 10 fold CV (n=1128) Predicted logS Measured logS
QB/04 · A — measured vs predicted log solubility
-0 1 2 3 4 -1 -0.5 0 0.5 1 1.5 Descriptor effects on solubility descriptor index (see results.json) logS change per SD
QB/04 · B — descriptor effects per standard deviation
QB/05

SIR dynamics of the 1978 boarding school influenza

BMJ 1978 · N=763 · CLOSED POPULATION
The question

Can a mechanistic SIR model, integrated with a hand written Runge Kutta solver and fitted by least squares, recover the transmission dynamics of a real and essentially completely observed outbreak, and what reproduction number does it imply?

The data

In January 1978, influenza A swept a boys' boarding school in northern England. 763 boys, a closed population, and an infirmary that counted every bed every day for fourteen days. Among the cleanest outbreak records in existence.

The method

The SIR ordinary differential equations, integrated by fourth order Runge Kutta at 0.05 day steps from a single infectious boy. Transmission and recovery rates fitted by least squares over a refined grid. Final attack rate from the final size equation.

What we found

Two fitted parameters reproduce the entire epidemic within seventeen boys RMS on a peak of 298, giving a reproduction number of 3.77 and an infectious period of 2.26 days, matching published analyses of this outbreak. The counterfactual runs make "flatten the curve" a computation: a forty percent transmission reduction roughly halves the peak and delays it by days.

Limitations. Bedridden is a proxy for infectious. Deterministic SIR ignores stochastic early dynamics. A closed boarding school is the most favorable possible setting; this reproduction number does not transfer to open populations.
$ python3 experiments/exp05_sir_influenza/run.py
Results
Basic reproduction number R₀3.77
Transmission rate β1.67 / day
Recovery rate γ0.44 / day
Infectious period2.26 days
Peak, model vs observed292 vs 298 boys
Fit error17.2 boys RMS
Predicted final attack rate97.5%
-0 2.5 5 7.5 10 12.5 15 -0 50 100 150 200 250 300 SIR fit, 1978 boarding school influenza (N=763) Day of outbreak Boys confined to bed observed SIR fit, R0=3.77
QB/05 · A — SIR fit to daily bedridden counts
-0 2.5 5 7.5 10 12.5 15 -0 50 100 150 200 250 300 Counterfactual transmission reduction Day of outbreak Boys in bed fitted beta beta reduced 20% beta reduced 40%
QB/05 · B — counterfactual transmission reductions
QB/06

Alarm fatigue versus lead time in ward early warning

SYNTHETIC BY DESIGNMETHODOLOGY STUDY
The question

Ward early warning systems fail in practice through alarm fatigue: staff silence what cries wolf. The evaluation that matters is therefore not AUC but the operating curve of warning hours against false alerts per patient day. How does a trend aware model compare to a standard threshold score on that curve?

Why synthetic, stated plainly

Every vital sign trajectory here is generated by the study's own code, because simulation is the only setting where the true deterioration time is known exactly and lead time can be measured without label noise. No clinical performance is claimed and none should be inferred. The deliverable is the evaluation method, which transfers to real wards; the numbers do not.

The method

400 train and 400 test simulated patients, 72 hours of hourly vitals, fifteen percent deteriorating toward a sepsis like pattern. A NEWS2 style aggregate score against logistic regression on current vitals plus four hour slopes, both swept across their thresholds; each threshold yields median lead time, detection fraction, and false alert episodes per patient day.

What we found

At matched alarm burden, the trend aware model reaches similar warning time at roughly half the false alerts per day. The general lesson stands regardless of the simulator: comparing alert systems at one threshold is meaningless, and the denominator staff actually experience is per patient day.

Limitations. The generator is simplistic: monotone deterioration, Gaussian noise, no missingness, no interventions bending trajectories. Real validation needs real data under governance, which is exactly why this study stops here.
$ python3 experiments/exp06_early_warning_tradeoff/run.py
Results · synthetic cohort
Threshold score at 0.27 false / day6.6 h median lead
Trend model, similar lead time≈ 0.13 false / day
Alarm burden at equal warning≈ ½
Events detected, both systems100%
-0 0.25 0.5 0.75 1 1.25 5 10 15 20 25 Lead time vs alarm burden (synthetic ward cohort) False alerts per patient day Median lead time, hours NEWS2 style threshold logistic + 4h trends
QB/06 · A — median lead time vs false alerts per patient day
-0 0.25 0.5 0.75 1 1.25 0.94 0.95 0.96 0.97 0.98 0.99 1 Detection rate vs alarm burden (synthetic) False alerts per patient day Fraction of events detected NEWS2 style threshold logistic + 4h trends
QB/06 · B — detection fraction vs alarm burden
QB/07

When does post hoc recalibration hurt?

8 DATASETS · 50 REPEATS PER CELLBASIS OF PAPER 1
The question

Recalibration maps are routinely fit on whatever holdout data happens to exist. How many observations do Platt scaling, beta calibration and isotonic regression need before they help a well specified model more often than they hurt it?

The method

Eight public clinical datasets, repeated random partitions into training, calibration and test sets, calibration set sizes 25 to 200, four strategies, 50 repeats per cell, with the minority event count of every calibration set recorded. All metrics on the untouched test partition.

What we found

At 25 observations every method harmed calibration on every dataset, isotonic worst (72 percent of repeats harmed pooled, 96 percent in the worst cell). Harm shrank monotonically, crossed zero near 100 observations, and turned to benefit near 200. In the more portable unit, harm below roughly 25 minority events and benefit above roughly 45. Flexibility ordered the harm exactly: Platt, then beta, then isotonic.

Limitations. One base model class, chosen because it is approximately calibrated by construction; badly miscalibrated learners shift the break even downward. Full limitations in the paper.
$ python3 experiments/exp07_calibration_sample_size/run.py · theory: exp09_scaling_theory/run.py
Results
Pooled ΔECE, isotonic, n=25+0.035 · 72% harmed
Worst cell (diabetes, n=25)ECE 0.071 → 0.131 · 96%
Pooled ΔECE, all methods, n=100≈ 0 · coin flip
Pooled ΔECE, isotonic, n=200−0.003 · 42% harmed
Break even, in minority events≈ 25 to 45
50 100 150 200 -0 0.01 0.02 0.03 0.04 Recalibration effect vs calibration set size (8 datasets) Calibration set size Mean change in ECE vs uncalibrated Platt beta isotonic no effect
QB/07 · A — pooled recalibration effect vs calibration size
-0 0.2 0.4 0.6 0.8 1 0.2 0.4 0.6 0.8 Worst cell: pima, 25 calibration obs (pooled over 50 repeats) Predicted probability Observed event rate uncalibrated, ECE 0.026 isotonic, ECE 0.099
QB/07 · B — worst cell reliability diagram
QB/08

How much accuracy does abstention buy?

8 DATASETS · 20 REPEATSBASIS OF PAPER 3
The question

Selective prediction assumes a model's errors live where its confidence is low, so deferring uncertain cases to a clinician removes them. How strongly does that premise actually hold across clinical prediction tasks?

The method

Out of fold logistic predictions on the eight datasets, confidence scored as distance from 0.5, coverage swept from 100 down to 50 percent, 20 repeated cross validations. Primary quantity: the share of all errors sitting in the least confident fifth of cases.

What we found

Two regimes. On the two nearly separable tasks the least confident fifth held about nine tenths of all errors, so a 20 percent deferral budget effectively eliminated machine error. On the six hard tasks it held only a third, and even deferring half of all cases left selective accuracy at 0.82 to 0.93. Base accuracy predicted error concentration almost perfectly: abstention is a property of the task before it is a property of the model.

Limitations. One model class and one confidence score; ensembles or conformal methods may concentrate error better on hard tasks, and measuring that is the direct extension. Full limitations in the paper.
$ python3 experiments/exp08_selective_prediction/run.py
Results
Errors in bottom 20%, breast tasks88 to 90%
Errors in bottom 20%, hard tasks32 to 44%
Accuracy at 80% coverage, WDBC0.978 → 0.997
Accuracy at 80% coverage, liver0.718 → 0.761
Uninformative baseline20%
0.5 0.6 0.7 0.8 0.9 1 -0 0.025 0.05 0.075 0.1 0.125 0.15 Selective accuracy gain vs coverage (8 datasets, 20 repeats) Coverage (fraction of cases answered) Accuracy gain over full coverage cleveland wdbc pima transfusion haberman wisc_original ilpd saheart
QB/08 · A — selective accuracy gain vs coverage
-0 2 4 6 8 0.2 0.4 0.6 0.8 1 Error concentration in the least confident 20% of cases dataset index (see results.json order) Share of all errors captured uninformative confidence (0.20)
QB/08 · B — error concentration per dataset
SEC/04

Methods Implemented

each written directly in numpy

InstrumentImplementationUsed in
logistic_fitL2 regularized logistic regression, solved by Newton's method on the full HessianQB/01 · 02 · 06
isotonic_fitPool adjacent violators, monotone calibration map with held out fittingQB/01
auc · roc_curveMann Whitney statistic with tie correction; full operating curveQB/01 · 02
ece · reliabilityExpected calibration error and reliability binningQB/01
bootstrap_ciPercentile bootstrap, 2000 resamples, any statisticQB/01 · 02
kaplan_meierProduct limit estimator with Greenwood variance bandsQB/03
logrankTwo group log rank test, hypergeometric varianceQB/03
cox_fitProportional hazards by Newton on the Breslow partial likelihood, SEs from inverse informationQB/03
concordanceHarrell's C over comparable pairsQB/03
linreg_fitRidge regression by normal equationsQB/04
rk4Fixed step fourth order Runge Kutta integratorQB/05
net_benefitDecision curve analysis, Vickers and Elkin weightingQB/02
svgplotThe lab's own figure engine: every plot on this page, zero plotting dependenciesALL
SEC/05

Data Provenance

where every record came from

DatasetOriginContentsCitation
CLEVELANDUCI ML Repository303 patients, Cleveland Clinic, angiographic outcomesDetrano et al, 1989
WDBCUCI ML Repository569 fine needle aspirates, 30 nuclear featuresStreet, Wolberg, Mangasarian, 1993
NCCTG LUNGR survival package228 advanced lung cancer patients, time to deathLoprinzi et al, JCO 1994
ESOLDeepChem mirror1,128 measured aqueous solubilities with descriptorsDelaney, 2004
BOARDING SCHOOLBMJ 1978;1:58714 daily bedridden counts, 763 boys, closed populationAnonymous, BMJ 1978
WARD COHORTGenerated in study800 synthetic trajectories, ground truth exactQB/06, this lab

All datasets are fully deidentified public research data. No protected health information is used anywhere in this work.