Skip to main content
Track Record · Case Study

Can You Predict How Someone Will Age?

AI is not the right answer to every health question. Across thirteen analyses of NHANES, CHNS, and HRS, we asked when a simple clinical or population model is enough and when a learned model adds useful signal. The work mixes cross-sectional assessments, mortality-linked NHANES follow-up, and HRS trajectories extending to 30 years; it is not one longitudinal cohort. Question-specific samples range from 1,907 to 39,839 people. The study groups three questions with domain-led or simple models, six with richer learned models, two with hybrids, and two where the added SDOH or sleep features contribute little incremental discrimination.

By Michael Key · ORCID

1,90739,839
Sample per Analysis
3
Datasets
30 yr
Longitudinal Tracking
3–6–2–2
Domain–ML–Hybrid–Null
NHANES and CHNS provide cross-sectional measurements; HRS provides longitudinal follow-up.

The Data

Three public-use health studies span two countries and three decades. NHANES (USA) and CHNS (China) provide detailed biomarker snapshots; linked NHANES mortality files add later outcomes for two questions. HRS follows participants every two years for up to 30 years. The longitudinal questions can be compared with later recorded outcomes; the cross-sectional questions cannot.

Datasets
3 public health studies
Analytic Samples
1,90739,839 per question
Countries
USA + China
Timespan
30 years (1992–2022)
Biomarkers
26 blood measures
Questions
13 investigations
How to Read This Study

For each question, we predict which method should win — and why — before running any models. The validation target follows the question. Five HRS analyses and two mortality-linked NHANES analyses use later recorded outcomes; six cross-sectional analyses use same-wave measurements, labels, or within-cohort criteria. The method tier was assigned before model results were generated, but this was an informal directional check, not a formal preregistered hypothesis test. The resulting study grouping is 3–6–2–2 across domain-led or simple models, richer learned models, hybrids, and cases where added data contributes little.

The Study

Thirteen Questions, Four Methods

We asked thirteen questions about aging, health prediction, and intervention, including mortality. Each uses its own eligible sample, endpoint, analytical method, and fidelity level. The right method depends on what the question is actually asking.

#QuestionDecision TypeBest MethodData
Domain-Led or Simple Models
4US norms in China?Transfer ValidationDomain-guided recalibrationNHANES+CHNS
11Physical independence?Decline PredictionLogistic regressionHRS 30yr
1324-month mortality?Survival PredictionComparison inconclusiveNHANES
ML Adds Genuine Value
2Diabetes risk?Risk ScreeningGradientBoostingNHANES
5Individual inflammation?Biomarker PredictionGradient BoostingNHANES
6Health in 10 years?Trajectory ForecastML on early wavesHRS 30yr
7Heart disease risk?Disease PredictionML vs published RRNHANES
8Lifestyle interactions?Interaction DiscoveryNeural net + SHAPNHANES
14Cognitive decline?Regression ForecastGradientBoostingHRS 30yr
More Data Doesn’t Always Help
15Wealth beyond health?Feature EvaluationHealth+SDOH MLNHANES
16Sleep predicts decline?Feature EvaluationHealth+Sleep MLHRS 30yr
Hybrid: Encode + Learn
9Weight trajectory?Long-Horizon ForecastPhysics + residualHRS 30yr
10Biological age?Health AssessmentKDM + MLNHANES

Methodology note: Methods and validation designs differ by question. Classification analyses generally compare a domain baseline with logistic regression and/or GradientBoosting; regression and trajectory analyses use their own splits and metrics. Confidence intervals, calibration checks, and bootstrap designs are reported where the individual analysis generated them, not as one shared protocol across all thirteen questions. Sample sizes range from 1,907 for Q10 biological age to 39,839 for Q2 diabetes classification. Q7 includes 38,033 people, Q15 includes 28,636, and Q13 uses a separate 3,578-person NHANES mortality-linked sample. See each investigation page for its design and provenance.

The Investigations

Each Question Gets the Method It Needs

Each tier groups questions by which method won.

Domain-Led or Simple Models
ML Adds Genuine Value
Investigation 2

Diabetes Risk

“Does this person have diabetes given their biomarker panel?”
39,839 NHANES adults (2005–2018); 6,433 diabetic by ADA criteria (16.15% prevalence). ADA-aligned published score (AUC 0.784 [0.7780.790]) vs GradientBoosting on the full biomarker panel (AUC 0.851 [0.8460.856]). +6.7 AUC points, non-overlapping CIs. Gap is largest in Age 60+ (+8.8 points). A hybrid model that adds the domain score to ML yields AUC 0.851 — adds nothing.
ADA-aligned Score vs GradientBoosting ML Wins non-overlap CI
Investigation 5

Individual Inflammation

“What are this person’s inflammatory markers given their lifestyle?”
CRP varies widely across individuals. GradientBoosting more than doubles explained variance on the log scale (R²=0.259 vs 0.114), although most variation remains unexplained. Population averages are insufficient for accurate individual prediction here.
Age-Sex Norms vs Gradient Boosting ML Wins 2x
Investigation 6

Health Trajectory

“How will this person’s health change over 10 years?”
7,864 people tracked 8+ waves. ML beats population curves by 1519% RMSE improvement, with gains growing at longer horizons. The trajectory IS the signal.
Population Curve vs ML on Early Waves ML Wins +18%
Investigation 7

Heart Disease Prediction

“Does this person have heart disease given their biomarker panel?”
38,033 NHANES adults (2005–2018); 4,193 with self-reported CVD (11.0% prevalence). A Framingham-derived score reaches AUC 0.786 [0.7790.793], versus 0.843 [0.8380.849] for GradientBoosting on the richer panel (permutation p=0.0005). This is classification of current self-reported CVD, not prospective event prediction.
Framingham Score vs GradientBoosting ML Higher (p=0.0005)
Investigation 8

Lifestyle Interactions

“Which lifestyle and metabolic factors show non-additive associations with CRP?”
Three-tier comparison: additive R²=0.208, +published interactions R²=0.218 (+5%), GBM R²=0.245 (+12%). In this cross-sectional analysis, obesity×diabetes is associated with +2.1 mg/L CRP beyond the additive fit, and sedentary×diabetes with +1.8 mg/L. These are candidate non-additive associations, not causal effects; residual confounding may explain part or all of them.
Three-Tier: Additive → +Interactions → GBM Candidate Nonlinearities
Investigation 14

Cognitive Decline

“Can we predict who will experience cognitive decline?”
5,212 people tracked across 11 waves of cognitive testing. ML beats age-education curves by 18% at 10 years (RMSE 4.38 vs 5.33). Individual cognitive slope, depression, and cardiovascular risk improve prediction beyond population averages.
Age-Education Curve vs GradientBoosting ML Wins +18% RMSE
More Data Doesn’t Always Help
Hybrid: Encode What You Know, Learn the Rest

The Honest Finding

Income and sleep are the surprises. The NHANES mortality gradient across income quintiles is 2.3×, yet the tested SDOH variables add only +0.003 AUC after measured health and biomarkers. Sleep quality likewise adds little incremental discrimination. These are predictive comparisons, not causal explanations.

Study grouping: three domain-led or simple, six richer learned models, two hybrid, and two limited incremental value.
The Fidelity Lesson

No Single Method Works

Each axis represents one of the thirteen questions. The radius shows where the best-performing approach in that analysis landed. Cross-population transfer used domain knowledge to identify what needed recalibration, then learned the new magnitudes. The biological-age analysis achieved its best held-out correlation with the hybrid model. SDOH and sleep (rose) sit at the bottom because, in these comparisons, adding them to the health models yielded little incremental discrimination. The shape is irregular. That’s the point.

Thirteen axes, four tiers. Teal = domain knowledge. Blue = ML needed. Rose = more data doesn’t help. Gold = hybrid.

Limitations

What We Measured and What We Didn't

Data Sources

Public-use Data, Local Reproduction

NHANES, CHNS, and HRS are public-use research sources under their respective access, registration, and citation terms; raw records are not redistributed here. A code-backed local reproduction package now covers all thirteen published investigations, but the website does not yet provide a public source repository or one-click reproduction path. Until that package is released, the study is not independently reproducible from this page alone.

  • Primary Dataset — USA National Health and Nutrition Examination Survey (NHANES) 2017–2018. Centers for Disease Control and Prevention (CDC). 9,254 Americans with full biomarker panels + lifestyle surveys. Cross-sectional. CDC NHANES →
  • Cross-Population Dataset — China China Health and Nutrition Survey (CHNS) 2009. University of North Carolina at Chapel Hill. 9,549 Chinese adults with 26 fasting blood biomarkers. Enables cross-population transfer validation. UNC CHNS →
  • Longitudinal Dataset — 30 Years Health and Retirement Study (HRS) RAND Longitudinal File, waves 1992–2022. University of Michigan. The source cohort supports five longitudinal investigations, each with its own analytic sample. Access requires HRS registration and acceptance of its conditions of use. HRS conditions of use →
  • Biological Age Reference Klemera, P. & Doubal, S. (2006). “A new approach to the concept and computation of biological age.” Mechanisms of Ageing and Development, 127(3), 240–248. Defines the KDM biological age estimator used in Investigation 10.
  • Frailty Criteria Fried, L.P. et al. (2001). “Frailty in older adults: Evidence for a phenotype.” J. Gerontology, 56A(3), M146–M156. Defines the five-criterion frailty index used as domain baseline in Investigation 11.