Uncertainty reduction over time LIVE METRICS

Every model update is re-validated against held-out labs before it ships, and this page regenerates from the accepted-checkpoint history on every deploy — it cannot go stale. The 90% prediction-interval width is the product; watching it shrink is watching the model earn trust. Latest checkpoint b034288 (2026-07-24): n=173 · MAE 5.34 MPa · R² 0.744 · PI ×/÷1.85 cross-study · ← strength predictor · print settings

90% prediction-interval width factor (×/÷)

1.351.892.432.973.5107-215b60fff07-21420761107-21c2feb1607-2112cc8da07-24b034288publish bar 1.51.851.64
cross-study (publish this) · within-CV
Cross-study is the honest number for a material from a lab we have never seen; within-CV assumes the corpus. The publish bar ×/÷1.5 is where the beta label drops. The 2026-07-21 spike is the true-Z ingest — the model told the truth about upright-print scatter before it learned to explain it.

Cross-validated MAE vs family-mean baseline (MPa)

0.02.75.48.010.707-215b60fff07-21420761107-21c2feb1607-2112cc8da07-24b0342885.37.0
model CV MAE · family-mean(LOO) baseline
The baseline predicts every coupon with its family's mean strength (leave-one-out). The model must beat it at every checkpoint — when the gap narrows it is usually because new, harder data raised both lines (the corpus hardened, e.g. the ×2-strength PA-CF ingest at "live").

Training rows (published measured coupons with stated density)

0479314018707-215b60fff07-21420761107-21c2feb1607-2112cc8da07-24b034288173
n_train
Only rows whose load-bearing fraction is known (stated infill or measured part density) are allowed to train — metadata completeness, not row count, drives accuracy.

Remaining targeted-research gap score (lower is better)

015130245360407-25search07-25evidence07-25search520
remaining priority score
The score is each named uncertainty term's variance (sigma²) multiplied by its controlled-evidence gap. It changes only when a source meets the acquisition brief and is deterministically normalized; search volume alone cannot improve it. Current qualifying controlled sources: 2.

How to read this. Each x-axis point is an accepted model checkpoint (git short-sha of accepted_metrics.json; "live" = the current deployed model). The research-gap series on this same page is separate from model performance: searches are recorded but remain flat unless qualified controlled evidence reduces a named gap. A model checkpoint is only accepted after the regression harness passes: determinism, physics monotonicity, honesty floors (beats the baseline, calibrated coverage) and a metric drift gate.

Widths can go UP honestly. New data that widens the interval (new orientation axes, new families, harder specimens) is progress too — the interval is measured, not promised.

Model + data: open source · predictions CC BY 4.0 · research preview, AS-IS, not for safety-relevant design — see Terms.