Every model update is re-validated against held-out labs before it
ships, and this page regenerates from the accepted-checkpoint history on every
deploy — it cannot go stale. The 90% prediction-interval width is the product;
watching it shrink is watching the model earn trust.
Latest checkpoint b034288 (2026-07-24): n=173 · MAE 5.34 MPa · R² 0.744 · PI ×/÷1.85 cross-study · ← strength predictor · print settings
90% prediction-interval width factor (×/÷)
cross-study (publish this) · within-CV
Cross-study is the honest number for a material from a lab we have
never seen; within-CV assumes the corpus. The publish bar ×/÷1.5 is where
the beta label drops. The 2026-07-21 spike is the true-Z ingest — the model told
the truth about upright-print scatter before it learned to explain it.
Cross-validated MAE vs family-mean baseline (MPa)
model CV MAE · family-mean(LOO) baseline
The baseline predicts every coupon with its family's mean strength
(leave-one-out). The model must beat it at every checkpoint — when the gap narrows
it is usually because new, harder data raised both lines (the corpus hardened,
e.g. the ×2-strength PA-CF ingest at "live").
Training rows (published measured coupons with stated density)
n_train
Only rows whose load-bearing fraction is known (stated
infill or measured part density) are allowed to train — metadata completeness,
not row count, drives accuracy.
Remaining targeted-research gap score (lower is better)
remaining priority score
The score is each named uncertainty term's variance
(sigma²) multiplied by its controlled-evidence gap. It changes only when a
source meets the acquisition brief and is deterministically normalized; search
volume alone cannot improve it. Current qualifying controlled sources:
2.
How to read this. Each x-axis point is an accepted model checkpoint
(git short-sha of accepted_metrics.json; "live" = the current
deployed model). The research-gap series on this same page is separate from model
performance: searches are recorded but remain flat unless qualified controlled
evidence reduces a named gap. A model checkpoint is only accepted after the regression harness
passes: determinism, physics monotonicity, honesty floors (beats the baseline,
calibrated coverage) and a metric drift gate.
Widths can go UP honestly. New data that widens the interval
(new orientation axes, new families, harder specimens) is progress too — the
interval is measured, not promised.
Model + data: open
source · predictions CC BY 4.0 · research preview, AS-IS, not for
safety-relevant design — see Terms.