Molecular Aggregation & Variant Effects.
A rigorous, calibrated reproduction and extension of amyloid-β variant-effect modeling. Every displayed result traces to script-written output from checksum-locked public data. Every caveat kept in the record.
By Michael Key · ORCID
The data and assays belong primarily to the Lehner / Bolognesi lab: the Aβ and TDP-43 deep-mutational-scanning atlases (GSE151147 / GSE128165), the familial-mutation analysis (Seuma et al., eLife 2021), the CANYA random-peptide nucleation model (Thompson et al., Science Advances 2025), and the energetic / thermodynamic analysis of Aβ aggregation and measured couplings (Arutyunyan et al., Science Advances 2025). The assay’s discrimination of known familial Alzheimer’s-disease mutations is the source lab’s published finding; the result below is reproduced on that foundation, not discovered here. This study’s contribution is not a new measurement or a standalone discovery. It is a calibrated reproduction and extension focused on where an Aβ model is informative and where its uncertainty fails. Everything here concerns in-vitro nucleation or toxicity assays: not therapy, not a clinical claim, and not a statement about any person.
What this means. An Aβ-specific model tracked this assay better than a model transferred from random peptides, and calibrated intervals exposed where the simple model had never learned a substitution effect.
What it does not mean. It does not predict Alzheimer’s disease in a person, identify a treatment, or establish a model that transfers between Aβ nucleation and TDP-43 toxicity.
A calibrated Aβ model that knows its own uncertainty.
The foundation: reproducing the field’s model
Before building anything, I reproduced CANYA — a leading published aggregation model — on its own held-out data: AUROC 0.82 on 7,626 random peptides (paper reports 0.809). The pipeline is faithful and there’s a concrete baseline to improve on.
That faithful reproduction also revealed something important: CANYA ranks well but its probability estimates were materially miscalibrated. Its raw calibration error was 0.068. Isotonic recalibration cut it ~6x, to 0.011, with ranking unchanged. This is a standard technique — not a breakthrough — and the improvement is measured on this held-out set, not established in a new assay or population.
The Aβ-specific model
Trained on 14,015 Aβ double-mutants and tested on 468 held-out single-mutant outcomes, a from-scratch additive model predicts measured nucleation at Spearman 0.62 versus 0.20 for zero-shot CANYA on the same points. All 8 dominant familial-Alzheimer’s single-mutant outcomes were held out from fitting, but their substitution features were represented among the double-mutant training examples. This is an outcome holdout, not a substitution-level cold start. In that eight-positive familial / extreme-assay evaluation — the source lab’s published test, reproduced here rather than discovered — the model reaches AUROC 0.92, within the noise of the measured assay’s own 0.90 discrimination and above zero-shot CANYA (0.76).
- The familial AUROC has a training-advantage component. My model trained on Aβ data; CANYA was applied cold. The honest claim is that the Aβ-specific model performs better than a transferred random-peptide model on this assay, not that its architecture is generally better. Restricting the comparison to variants with represented substitution features gives a lower familial AUROC; the load-bearing result is the large-n Spearman across 468 points, not the 8-positive AUROC.
- The 0.92 and measured 0.90 values are indistinguishable at this sample size. This is not evidence that a model surpassed the measurement. The assay’s familial discrimination is the source lab’s result (Seuma et al., eLife 2021), reproduced here.
- The familial AUROC rests on 8 positive examples. A small-n result. The Spearman across 468 is the number to trust.
- No model cures anything. Cures need labs and trials. This is a contribution to mechanistic understanding of in-vitro nucleation.
Calibrated uncertainty on the Aβ model
Conformal prediction intervals exposed something accuracy metrics cannot see: a structural blind spot affecting nearly half the test variants. These are substitutions never seen among the training doubles, so the position-additive model has no learned effect for them. The naïve way of reporting uncertainty was badly wrong.
A single marginal interval over-covered variants with represented substitutions and under-covered unseen substitutions. Group-conditional intervals brought both groups to their nominal target coverage on this held-out set, with wider intervals for the unseen group. That is an internal calibration result; it does not guarantee coverage in another assay or population.
One load-bearing exception remains. All eight familial substitutions are in the represented group, but their effects extend beyond that group’s interval, with both misses above it. The appropriate output is a sharp interval for broader variant-effect screening plus a separately disclosed wider upper band for the familial extreme — not one interval everywhere.
A stronger representation does not solve the cross-assay transfer barrier.
A natural hope: one model spanning the misfolding diseases. I tested it directly. Naïve transfer between the Aβ nucleation assay and the TDP-43 toxicity assay fails — backwards. Transfer Spearman: -0.19. A hydrophobic substitution correlates with harm in Aβ nucleation (+0.30) but with protection in TDP-43 toxicity (-0.45). The opposing measured relationships are consistent with known biology, but the assays also measure different phenotypes.
The sharper question: is that barrier merely a weakness of crude features? I re-ran the transfer with a strong protein-language-model representation. Under matched five-fold cross-validation, the pLM was substantially stronger within each assay than the crude biophysical baseline. Yet it transferred no better — near zero for Aβ→TDP-43, strongly negative for TDP-43→Aβ. This tested representation was not enough to rescue transfer; that does not prove the barrier is purely biological.
The study used 1,196 TDP-43 single-mutant measurements (GSE128165) as the transfer target.
The two assays measure different phenotypes: Aβ nucleation vs TDP-43 toxicity. I cannot isolate protein-specific biology from the phenotype mismatch without a shared-assay control — data I do not have. I can say the failure is not fixed by the stronger representation tested here; I cannot rule out every representation or attribute the failure to protein biology alone. A clean test would need two proteins assayed for the same phenotype.
The suppressor question — a go/no-go.
The boldest disease-framed question the atlas can answer directly: for each familial mutation, does a second mutation make it aggregate less than expected? This is a suppressor — not a therapy (you cannot add a second mutation to a person), but a mechanistic signal. The approach needs no model: epistasis is measured(double) minus the sum of individual effects, straight from the published scores. Using the established epistasis method from the Lehner/Bolognesi lab’s own published library.
Verdict: PARTIAL — at its thinnest, one replicate away from a clean negative. With two real confounds handled — a global sub-additive baseline (suppression is the atlas-wide norm, so every value must be offset-corrected) and a measurement ceiling sitting on the most aggregating familial mutations — the signal is enhancer-dominated. Four of the six well-covered familial mutations show no suppression at all. One non-ceiling lead emerged: fragile, resting on two discordant replicates, losing significance on either alone.
The honest kept-in-the-record result: I checked rigorously, and the public atlas does not support a familial-suppressor map. Doing it properly would require the lab’s full thermodynamic modeling, which sits above this study’s scope and data.
The tested escalations don’t fix the tail.
Between the model and the suppressor check, I tested whether selected higher-fidelity models close the Aβ model’s remaining weakness: systematic under-prediction of the most extreme-nucleating variants. The tested escalations do not.
Protein-language-model head. A pLM-based model (ESM-2, 8M parameters) scores higher than the additive model on overall correlation — but primarily by gaining useful predictions for the cold-start variants the additive model is blind to. Where the additive model can learn, it remains competitive. And the pLM’s cold-start gain is the result most exposed to the model’s pretraining familiarity with this exact protein sequence.
Scaling the pLM four-fold (8M → 35M parameters) reduces tail under-prediction but regresses cold-start performance — no clean win on both simultaneously.
Asymmetric (tail-aware) training loss. Reduces under-prediction, but turns out to be global debiasing, not a tail fix — it lifts bulk residuals as much as tail residuals, and a trivial one-number offset matches it.
The supported conclusion is narrower: neither tested ESM-2 scale nor the tested tail-aware objective resolves the pathogenic-tail error while preserving the other gains. Structure-aware inputs and greater coverage of the assay’s extreme tail are plausible next rungs, not requirements established by these experiments. The honest cheap fix for the model’s global low-bias is a one-number debias folded into the calibration layer.
What this contributes, honestly stated.
- A rigorously reproduced, calibrated Aβ variant-effect model that distinguishes represented from unseen substitutions and reports where its interval coverage fails at the familial extreme.
- Two controlled negatives: the tested stronger representation does not overcome the Aβ-nucleation / TDP-43-toxicity transfer barrier, and a bigger model or tail-aware loss does not fix the extreme-nucleation underprediction.
- A PARTIAL on the suppressor question: one fragile lead, not a map. The public atlas does not support a familial-suppressor claim.
- Every claim carries its caveat, every displayed result traces to script-written output, and the negatives are kept in the record.
The molecular track closes here. The Lehner/Bolognesi lab supplied the data and the main mechanistic foundation. My contribution is narrower: a calibrated extension that exposes where the simple model is blind, plus a controlled cross-assay result whose phenotype mismatch remains unresolved. These are bounded inputs to the field, not a clinical model.
Every displayed result has a trace.
Primary data. GSE268261 (CANYA evaluation set), GSE151147 (Aβ atlas), and GSE128165 (TDP-43 atlas). The study locks the retrieved inputs by sha256 checksum.
Primary papers. Seuma et al. 2021 for the Aβ atlas and familial-mutation result; Thompson et al. 2025 for CANYA; Arutyunyan et al. 2025 for the energetic / thermodynamic analysis.
Reproduction. Each investigation regenerates script-written results, and the reproduction harness runs it twice to check byte-identical output. The code, environment records, research protocol, and per-investigation provenance are maintained together. They are available for review on request while a separate public-release review is pending; this page therefore does not claim that every artifact is independently downloadable today.
Public artifacts. The molecular build log records the study’s progression and corrections. A reviewed release package is still pending.