Can You Predict How Someone Will Age?
AI is not the right answer to every health question. Across thirteen analyses of NHANES, CHNS, and HRS, we asked when a simple clinical or population model is enough and when a learned model adds useful signal. The work mixes cross-sectional assessments, mortality-linked NHANES follow-up, and HRS trajectories extending to 30 years; it is not one longitudinal cohort. Question-specific samples range from 1,907 to 39,839 people. The study groups three questions with domain-led or simple models, six with richer learned models, two with hybrids, and two where the added SDOH or sleep features contribute little incremental discrimination.
By Michael Key · ORCID
The Data
Three public-use health studies span two countries and three decades. NHANES (USA) and CHNS (China) provide detailed biomarker snapshots; linked NHANES mortality files add later outcomes for two questions. HRS follows participants every two years for up to 30 years. The longitudinal questions can be compared with later recorded outcomes; the cross-sectional questions cannot.
For each question, we predict which method should win — and why — before running any models. The validation target follows the question. Five HRS analyses and two mortality-linked NHANES analyses use later recorded outcomes; six cross-sectional analyses use same-wave measurements, labels, or within-cohort criteria. The method tier was assigned before model results were generated, but this was an informal directional check, not a formal preregistered hypothesis test. The resulting study grouping is 3–6–2–2 across domain-led or simple models, richer learned models, hybrids, and cases where added data contributes little.
Three publicly available datasets form the foundation for all thirteen investigations. NHANES gives biomarker depth. CHNS lets us test whether American patterns hold in China. HRS tracks people long enough to see if predictions came true.
Counting note: 9,254 NHANES participants and 9,549 CHNS participants describe the cross-population comparison. HRS and later NHANES questions use different, overlapping analytic samples, so the page does not add them into a study-wide participant total. Sample sizes are reported investigation by investigation.
Thirteen Questions, Four Methods
We asked thirteen questions about aging, health prediction, and intervention, including mortality. Each uses its own eligible sample, endpoint, analytical method, and fidelity level. The right method depends on what the question is actually asking.
| # | Question | Decision Type | Best Method | Data |
|---|---|---|---|---|
| Domain-Led or Simple Models | ||||
| 4 | US norms in China? | Transfer Validation | Domain-guided recalibration | NHANES+CHNS |
| 11 | Physical independence? | Decline Prediction | Logistic regression | HRS 30yr |
| 13 | 24-month mortality? | Survival Prediction | Comparison inconclusive | NHANES |
| ML Adds Genuine Value | ||||
| 2 | Diabetes risk? | Risk Screening | GradientBoosting | NHANES |
| 5 | Individual inflammation? | Biomarker Prediction | Gradient Boosting | NHANES |
| 6 | Health in 10 years? | Trajectory Forecast | ML on early waves | HRS 30yr |
| 7 | Heart disease risk? | Disease Prediction | ML vs published RR | NHANES |
| 8 | Lifestyle interactions? | Interaction Discovery | Neural net + SHAP | NHANES |
| 14 | Cognitive decline? | Regression Forecast | GradientBoosting | HRS 30yr |
| More Data Doesn’t Always Help | ||||
| 15 | Wealth beyond health? | Feature Evaluation | Health+SDOH ML | NHANES |
| 16 | Sleep predicts decline? | Feature Evaluation | Health+Sleep ML | HRS 30yr |
| Hybrid: Encode + Learn | ||||
| 9 | Weight trajectory? | Long-Horizon Forecast | Physics + residual | HRS 30yr |
| 10 | Biological age? | Health Assessment | KDM + ML | NHANES |
Methodology note: Methods and validation designs differ by question. Classification analyses generally compare a domain baseline with logistic regression and/or GradientBoosting; regression and trajectory analyses use their own splits and metrics. Confidence intervals, calibration checks, and bootstrap designs are reported where the individual analysis generated them, not as one shared protocol across all thirteen questions. Sample sizes range from 1,907 for Q10 biological age to 39,839 for Q2 diabetes classification. Q7 includes 38,033 people, Q15 includes 28,636, and Q13 uses a separate 3,578-person NHANES mortality-linked sample. See each investigation page for its design and provenance.
Each Question Gets the Method It Needs
Each tier groups questions by which method won.
Cross-Population Transfer
Functional Decline
Short-Term Mortality (NHANES)
Diabetes Risk
Individual Inflammation
Health Trajectory
Heart Disease Prediction
Lifestyle Interactions
Cognitive Decline
Wealth & Longevity (NHANES)
Sleep & Health
BMI Trajectory
Biological Age
No Single Method Works
Each axis represents one of the thirteen questions. The radius shows where the best-performing approach in that analysis landed. Cross-population transfer used domain knowledge to identify what needed recalibration, then learned the new magnitudes. The biological-age analysis achieved its best held-out correlation with the hybrid model. SDOH and sleep (rose) sit at the bottom because, in these comparisons, adding them to the health models yielded little incremental discrimination. The shape is irregular. That’s the point.
Thirteen axes, four tiers. Teal = domain knowledge. Blue = ML needed. Rose = more data doesn’t help. Gold = hybrid.
What We Measured and What We Didn't
Self-reported data in HRS. Five of thirteen investigations use the Health and Retirement Study (Q6, Q9, Q11, Q14, and Q16), where health status (1–5 scale), disease onset, and BMI are self-reported. Self-reported BMI is typically underestimated by 1–2 kg/m², and disease onset dates reflect when a diagnosis was received, not necessarily when the condition began. NHANES provides measured biomarkers; its survey measurements are cross-sectional, while separate mortality linkage supports Q13 and Q15. It does not provide repeated biomarker trajectories for these analyses. This asymmetry means the HRS trajectories carry more self-report measurement noise than the NHANES biomarker analyses.
No external validation cohort. All comparisons remain internal to their source datasets, using cross-validation, held-out splits, or split-half checks as specified by the individual investigation. A true external validation — testing HRS-trained models on ELSA (English Longitudinal Study of Ageing) or SHARE (European equivalent) — would strengthen generalizability claims. This is deferred pending data access agreements.
Cross-sectional vs. longitudinal design. Q2, Q4, Q5, Q7, Q8, and Q10 use cross-sectional measurements or current-condition labels. They can evaluate within-wave classification, association, or transfer, but not future outcomes. Q13 and Q15 use NHANES linked mortality follow-up. Q6, Q9, Q11, Q14, and Q16 use longitudinal HRS outcomes. These three designs answer different questions and should not be described as one common validation scheme.
Survivorship bias in HRS. Participants must survive to each follow-up wave to contribute data. People who die between waves are lost, and those who drop out tend to be sicker. This biases longitudinal analyses toward healthier survivors, potentially underestimating the true predictive power of health decline indicators.
Model complexity ceiling. All ML models are GradientBoosting ensembles or logistic regression. Deep learning, survival models (Cox), or time-series architectures might perform differently. We deliberately chose interpretable models to keep the focus on whether ML helps, not on squeezing marginal gains from architecture search.
Explore the Data
Dig into the data behind the findings.
The Biology Paradox
Three methods predict biological age. Step through to see which one loses — and why more computation made it worse.
Bio-Age Calculator
Enter your biomarkers and see your estimated biological age. Compare your position against 1,907 NHANES participants.
Prediction Duel
Four investigations with direct domain-vs-ML comparisons. ROC curves, calibration, reclassification, and feature importance heatmap.
The Data Paradox
Income predicts a 2.3× mortality gap across NHANES quintiles. Sleep separates healthy from declining. But adding this data to models changes almost nothing. Why?
Public-use Data, Local Reproduction
NHANES, CHNS, and HRS are public-use research sources under their respective access, registration, and citation terms; raw records are not redistributed here. A code-backed local reproduction package now covers all thirteen published investigations, but the website does not yet provide a public source repository or one-click reproduction path. Until that package is released, the study is not independently reproducible from this page alone.
- Primary Dataset — USA National Health and Nutrition Examination Survey (NHANES) 2017–2018. Centers for Disease Control and Prevention (CDC). 9,254 Americans with full biomarker panels + lifestyle surveys. Cross-sectional. CDC NHANES →
- Cross-Population Dataset — China China Health and Nutrition Survey (CHNS) 2009. University of North Carolina at Chapel Hill. 9,549 Chinese adults with 26 fasting blood biomarkers. Enables cross-population transfer validation. UNC CHNS →
- Longitudinal Dataset — 30 Years Health and Retirement Study (HRS) RAND Longitudinal File, waves 1992–2022. University of Michigan. The source cohort supports five longitudinal investigations, each with its own analytic sample. Access requires HRS registration and acceptance of its conditions of use. HRS conditions of use →
- Biological Age Reference Klemera, P. & Doubal, S. (2006). “A new approach to the concept and computation of biological age.” Mechanisms of Ageing and Development, 127(3), 240–248. Defines the KDM biological age estimator used in Investigation 10.
- Frailty Criteria Fried, L.P. et al. (2001). “Frailty in older adults: Evidence for a phenotype.” J. Gerontology, 56A(3), M146–M156. Defines the five-criterion frailty index used as domain baseline in Investigation 11.