Skip to main content
The Mission

Aiming AI at the diseases we can’t yet predict.

Independent research on Alzheimer’s and ALS, judged against lab measurements, public benchmarks, and held-out outcomes.

This work is personal — why →

Start where the result can be checked.

There’s a corner of this science where a laptop is enough to make a useful, verifiable contribution. Scientists have already measured how mutations change protein aggregation, and much of that data is public. Those measurements give a model a clear test: if its predictions are wrong, the data shows it. The experiments already exist, so a careful solo effort can build and evaluate models without pretending to replace a wet lab.

AI lets one researcher compare more candidate models, test more variations, and examine more failure cases than would otherwise be practical. It supplies speed and range. It does not decide what counts as evidence.

I choose the question, decide how much detail it needs, design the validation, and interpret the result. A claim survives only if it agrees with an external check: a lab measurement, a public benchmark, or a held-out outcome.

How Analysis Driven Modeling works →

Two completed studies, each with a held-out check.

The molecular study is evaluated against held-out lab measurements. The HRS study uses a later time-period holdout within the same registered public-release cohort; it is internal validation, not independent-cohort validation. Every displayed number is regenerated from the recorded analysis.

Study 01 · Molecular aggregation & variant effect
0.82

Reproduced the field’s model

A recent, strong model (CANYA) predicts protein aggregation from sequence. Rerun on its own held-out data, it scored an AUROC of 0.82 — matching the paper’s 0.809. The pipeline is faithful, and there’s a concrete baseline to improve on.

~6x

Corrected its calibration

A model can rank well and still misstate how its scores relate to observed outcomes. This one’s calibration error was 0.068; a standard recalibration cut it roughly six-fold, to 0.011, with its ranking preserved. Calibrated uncertainty is the point, not a footnote.

0.62vs 0.20

The Aβ-specific result

Alzheimer’s has a handful of inherited mutations that cause the early-onset form. Trained on Aβ double-mutant measurements and tested on 468 held-out single-mutant outcomes, a purpose-built model’s correlation with the lab measurements was 0.62 — versus 0.20 for the general model applied cold. On separating the inherited mutations from the other measured variants, it reached about the same discrimination as the assay itself (~0.90). The familial substitutions were represented in the double-mutant training data; their single-mutant outcomes were not.

Then a wall worth hitting

Pointing what works on Alzheimer’s Aβ at the ALS protein TDP-43 didn’t just fail — the measured relationships ran in opposite directions across the Aβ nucleation and TDP-43 toxicity assays. The tested ESM-2 representation did not rescue the transfer. Because the assays measure different phenotypes, the study cannot separate protein biology from assay mismatch. That is a useful boundary, and a genuine open problem.

What I will not claim
  • I didn’t “beat” the measurement — matching ~0.90 is within its noise. “Matched the ceiling,” not “surpassed it.”
  • The headline rests on only 8 inherited mutations, so the result I actually trust is the 0.62 vs 0.20 correlation across 468 points.
  • It’s not a fair fight, and I’ll say so: my model was trained on Aβ data; the general one wasn’t. The honest claim is “a protein-specific model far outperforms a transferred general one,” not “my architecture is better.”
  • The method is standard. The contribution is applying it carefully, transparently, and in the open — not inventing something new.
  • The data and the comprehensive molecular picture belong to the Lehner/Bolognesi lab. The 2025 CANYA work is Thompson et al.; the energetic analysis is Arutyunyan et al.; and the familial-mutation discrimination result is the lab’s earlier Seuma et al. finding. This is a calibrated reproduction and extension on their foundation, not a competing discovery.
— from the build log, “The molecular track”
Study 03 · HRS cognitive-decline validation

The simple model earned its place — and the cohort exposed where the decision remains unsafe.

A minimal cognition model produced well-calibrated two-year probabilities to the Langa-Weir silver label. More complex models added small, statistically detectable gains on this holdout. Their point estimates fell below the study’s operating threshold, but the uncertainty crossed it, so the stop is provisional rather than proof of equivalence. Competing death risk and subgroup performance mattered more to the decision than another fraction of pooled AUROC.

What the HRS validation added
Temporal holdout 5,022 Person-periods in the 2020→2022 test wave, including 195 incident events.
Calibration 0.89 Calibration slope to the silver label on the out-of-time holdout.
Complexity gain +0.0196 Point estimate just below the 0.02 operating threshold; its interval crosses the threshold and equivalence was not established.
Competing risk 14% Relative overstatement of CIND incidence when death is censored rather than modeled.

How the mission is organized.

The research is at the center. The surrounding pages explain why it exists, how it is done, and what remains before any result can reach families.

Live

The Work

Completed evidence, active studies, deferred questions, and the next gate for each. Research is the canonical record of what has—and has not—been established.

See the research →
Live

Motivation

Why this exists — two parents, two diseases, and a modeler’s conviction that the right tools, rigorously applied, can contribute something that holds up.

Read the why →
Live

Methods

The right-fidelity discipline underneath all of it — match the model to the decision, quantify what you don’t know, never ship false certainty.

The methodology →
Live

Build Log

The decisions, false starts, and changes of direction behind the formal study pages.

Read the build log →

Useful scrutiny is welcome.

Right now this is an independent, one-person research effort led by me, Michael Key. The most useful help is specific: finding a flaw, checking whether a result would matter in practice, providing a lawful route to external validation, or strengthening the work’s governance.

Review the methods

If you work in biostatistics, machine learning, molecular aggregation, or clinical prediction, I welcome a direct challenge to the design, analysis, or interpretation. The full code package is available on request while its public release review is completed.

See what can be checked →

Test whether it travels

Clinicians, trialists, outcomes researchers, and cohort custodians can help identify where a forecast would be unsafe or irrelevant and whether a result survives in another population. The current need is interpretation and external validation, not endorsement.

michael@rightfidelity.ai →

Provide a lawful data path

Data providers and research groups can help establish a governed route to representative longitudinal outcomes. Repository terms and participant protections govern every step.

Read the data boundaries →

About the researcher. I developed Analysis Driven Modeling over a 24-year career in modeling, simulation, and analysis. The background page gives the context; Vision explains the longer horizon.