MAGELLAN LONGEVITY

The Longevity Evidence Gap Report 2026

We took the 130 products in the Magellan Longevity catalogue, the 2,353 unique PubMed-indexed citations behind them and the 314 structured study write-ups we maintain, and asked one question of the whole market at once: how strong is the evidence, really? Every number below is computed from published data by a script we ship with the report. Where a number was not computable, we say so and drop the claim.

Original researchOpen dataset (CC BY 4.0)2,353 citations analysed
Published 2026-08-11 · Version 1.0 · Author Gabriel Radu, DO — physiatrist (NY licence 275110, NPI 1376861765) · Corpus 1984–2026 · Reuse free with attribution and a link

The seven numbers

11 of 130
carry a Strong evidence grade
That is 8.5% of the catalogue. 74 are Moderate, 20 Emerging, 4 Preliminary.
31.8%
of citations are randomized
749 of 2,353 unique citations carry a randomized-trial signal. The rest is cohort, mechanistic and review literature.
21
median citations per product
Interquartile range 18–24; the thinnest product still carries 14.
2020
median year of the evidence
23.1% of the corpus was published in 2014 or earlier; 43.2% since 2021.
61 of 144
randomized write-ups report a null
42.4% of the randomized evidence we summarise reports at least one explicitly non-significant result.
7 of 130
have no randomized citation at all
4 of those 7 are graded as measurement instruments, where a randomized trial is not the relevant design.
r = 0.18
star rating vs evidence grade
95% CI -0.02 to 0.38 (n = 90). Amazon ratings tell you almost nothing about evidence.
0 of 130
rest on animal data alone
Every product in the catalogue has at least one citation carrying a human-research signal. Only 90 of 2,353 individual citations (3.8%) are preclinical-only.
Contents
  1. Why this report exists
  2. Methodology
  3. The grading rubric
  4. Finding 1 — the shape of the market
  5. Finding 2 — how deep the citations go
  6. Finding 3 — how old the evidence is
  7. Finding 4 — how much of it is randomized
  8. Finding 5 — how often the trials came back null
  9. Finding 6 — which categories are best and worst supported
  10. Finding 7 — price and star ratings do not track evidence
  11. Limitations
  12. Claims we could not compute, and dropped
  13. How to reproduce this
  14. How to cite this

Why this report exists

The longevity market is sold on mechanism. A compound activates a pathway in a cell line, the pathway is implicated in aging, and by the time the story reaches a product label it has quietly become a benefit. Nobody publishes the denominator — how much of the shelf actually has randomized human evidence behind it, how old that evidence is, and how often it came back negative.

Magellan Longevity has, for its own storefront, been forced to answer that question product by product: every item we list carries an evidence grade and links to its primary citations, and the grade is set independently of what the product pays in affiliate commission. That grading work leaves behind a dataset — 130 graded products, 2,353 unique PubMed-indexed citations, 314 structured study write-ups — that can be read sideways, as a description of the market rather than of any one product.

That is what this report is. It is not a ranking, and it is not an accusation: a Moderate grade is a normal, respectable place for a supplement to sit. It is an attempt to put an honest number on how much of the evidence base is randomized, how recent it is, and how much of it is null — and to publish the method so that anyone who disagrees can rerun it and argue with the code rather than with us.

Methodology

Corpus. The unit of analysis is the shipped catalogue: 130 products across 7 categories, mapped to 117 unique compound or device monographs. Each monograph carries a hand-assembled citation list; across the catalogue those lists resolve to 2,353 unique studies, of which 2,353 (100.0%) carry a PubMed ID and 2,022 (85.9%) carry a DOI. Publication years span 1984–2026. A separate layer of 314 structured write-ups — each with an explicit Question, Methods, Results, Conclusion and Limitations field — covers the most load-bearing studies and is used for the design and null-result analyses.

Grades. Grades are not recomputed for this report. The script reads the same three objects the live site reads — the tier definitions, the product-specific grade map and the compound-level grade map — out of the site shell and applies the site’s own precedence rule: a product-specific curated grade wins; failing that, diagnostics and wearables are graded as instruments; failing that, the compound-level curated grade applies; and only if none of those exist does an automatic citation-count fallback fire. Across this catalogue the fallback never fires. 26 of 130 grades come from a product-specific curated entry and 104 from a compound-level curated entry; 0 fall through to the instrument rule and 0 to the automatic citation-count fallback. Every grade in this report is therefore human editorial judgement, not an algorithm dressed up as one.

Design classification. Individual citations carry a title, a curated one-sentence excerpt from the abstract and — where a write-up exists — a Methods paragraph. Those three fields are concatenated and matched against three published regular expressions. This is a keyword classifier, not a human read of every full text, and its error profile is described under Limitations.

RANDOMIZED  /randomi[sz]ed|randomi[sz]ation|randomly (assigned|allocated)|
            placebo-controlled|double-blind|triple-blind|single-blind|
            crossover (trial|study|design)|RCTs?/i

HUMAN       /patients?|participants?|subjects?|volunteers?|humans?|men|women|
            adults?|children|cohort|population-based|NHANES|UK Biobank|
            case-control|cross-sectional|clinical trial|meta-analy|
            systematic review|epidemiolog/i

PRECLINICAL /mice|mouse|murine|rats?|rodents?|in vitro|cell (line|culture|s)|
            C. elegans|caenorhabditis|drosophila|zebrafish|yeast|preclinical|
            animal model|fibroblasts?|myotubes?|ex vivo/i

Null-result detection. Within the 314 structured write-ups, the Results and Conclusion fields are matched against 12 patterns for explicit statements of a non-significant finding (no significant difference, did not improve, failed to meet, not superior to placebo, nonsignificant, and eight more, all listed in the shipped script). A write-up counts as randomized if its title, Question or Methods matches the RANDOMIZED pattern.

Statistics. Medians and interquartile ranges are reported for skewed counts. Correlations are Pearson’s r with a Fisher z 95% confidence interval, plus Spearman’s ρ because the grade ladder is ordinal. Grades are mapped Strong = 4, Moderate = 3, Emerging = 2, Preliminary = 1; measurement tools and the refuted grade sit off that ladder and are excluded from correlations. No p-values are reported — this is a census of one catalogue, not a sample from a population, so the confidence intervals describe estimation noise in the mapping, not a hypothesis test.

Rounding. Percentages are given to one decimal place using ordinary half-up rounding, and never nudged toward a more striking number. Every percentage in this report is printed next to its raw numerator and denominator so the rounding can be checked.

The grading rubric

Six tiers. The first four are a strength ladder; the last two are not on it. Measurement tool exists because grading a blood-pressure cuff on "treatment effect" is a category error — an instrument is graded on how accurately it measures and on whether the marker predicts anything. Tested — did not work exists because "we looked and found nothing" is a different, stronger statement than "nobody has looked yet", and collapsing the two would flatter the shelf.

Table 1. The six evidence tiers and the number of catalogue products in each.
TierLetterWhat it meansProductsShare
Strong evidenceAConsistent human evidence, including randomized trials, on outcomes that matter.118.5%
Moderate evidenceBReal human evidence, but mixed, limited to surrogate endpoints, or small in scale.7456.9%
Emerging evidenceCEarly human signal on top of a solid mechanism — promising, not settled.2015.4%
PreliminaryC-Mechanism or very early human data only. Buy it as an experiment, not a plan.43.1%
Measurement toolBGraded on measurement accuracy and on whether the marker predicts anything — not on a treatment effect.2015.4%
Tested — did not workDAdequately tested and did not work. Not thin evidence — evidence pointing at nothing.10.8%

Finding 1 — the shape of the market

Only 11 of 130 products (8.5%) earn the top grade. The catalogue’s centre of gravity is Moderate — 74 products, 56.9% — which is the honest resting place for most of this market: real human data, but mixed, small, or measured on a surrogate endpoint rather than on an outcome anyone lives or dies by.

Evidence grade distribution across 130 productsStrong evidence: 11; Moderate evidence: 74; Emerging evidence: 20; Preliminary: 4; Measurement tool: 20; Tested — did not work: 1019375674Strong evidenceStrong evidence: 11 (8.5%)11 (8.5%)Moderate evidenceModerate evidence: 74 (56.9%)74 (56.9%)Emerging evidenceEmerging evidence: 20 (15.4%)20 (15.4%)PreliminaryPreliminary: 4 (3.1%)4 (3.1%)Measurement toolMeasurement tool: 20 (15.4%)20 (15.4%)Tested — did not workTested — did not work: 1 (0.8%)1 (0.8%)
Figure 1. Grade distribution, n = 130 products. The Strong tier is highlighted; all other bars share one neutral tone because the categories are not a magnitude scale. Every bar is labelled with its count and share — colour carries no information the label does not.

The products at the top of the ladder are not the ones the category is marketed on. They are, in full: Bright Light Therapy Lamp, Lutein and Zeaxanthin, Voltaren Arthritis Pain Gel (Diclofenac 1%), Creatine Monohydrate, Betaine (TMG), Daily Mineral SPF Sunscreen, Validated Home BP Monitor, Balance and Strength Trainer, Azelaic Acid Gel, 24h Ambulatory BP Monitor, Hand Grip Dynamometer Pro. Sunscreen, creatine, a light box, a grip dynamometer, a validated blood-pressure cuff, lutein and zeaxanthin, a balance trainer, betaine, azelaic acid, an ambulatory BP monitor and topical diclofenac. None of them is a NAD⁺ precursor, a senolytic or a "cellular rejuvenation" formula.

At the other end sit 4 Preliminary-grade products (Copper Tripeptide (GHK-Cu) Serum, Saliva Cortisol Stress Test, Allergen Alert Mini Food Lab, Phosphatidylcholine (PC) Complex) and exactly one product graded Tested — did not work: SYB Clear Blue Light Glasses. That single entry is the most useful cell in the whole table, because it is the one place where the literature is not thin but adequate — and adequate literature pointing at nothing is a finding, not a gap. Related background: blue light.

Finding 2 — how deep the citations go

Median 21 citations per product (IQR 18–24; range 14–104). No product ships with fewer than 14. That is a floor, not a boast: it means the interesting variation in this market is not whether a compound has published literature — essentially all of them do — but what kind of literature it is. Citation count is the weakest measure in this report, and we lead with it only in order to retire it.

Distribution of citations per product14–17: 20 products; 18–20: 42 products; 21–24: 40 products; 25–29: 12 products; 30–49: 8 products; 50–104: 8 products01121324214–17: 202014–1718–20: 424218–2021–24: 404021–2425–29: 121225–2930–49: 8830–4950–104: 8850–104
Figure 2. Number of products (y) by size of their linked citation list (x), n = 130. The long right tail is driven by shared monographs — see the caveat below.
Caveat you should apply to Figure 2. 130 products map to 117 unique monographs, so citation lists are shared. The largest single monograph carries 104 citations and is shared by 8 wearable devices; those devices do not each have 104 device-specific studies. Read per-product citation counts as "depth of literature on the underlying compound or device class", never as "studies on this SKU". Table 5 reports the category figures on unique compounds instead, which removes this inflation.

Finding 3 — how old the evidence is

Median publication year 2020. But a median hides the shape: 543 citations (23.1%) were published in 2014 or earlier, and 1,016 (43.2%) since 2021. Roughly a quarter of the evidence base behind the modern longevity shelf is more than a decade old.

Publication years of the citation corpus1984–1999: 34; 2000–2009: 224; 2010–2014: 285; 2015–2019: 621; 2020–2022: 597; 2023–2026: 59201553114666211984–1999: 34341984–19991.4%2000–2009: 2242242000–20099.5%2010–2014: 2852852010–201412.1%2015–2019: 6216212015–201926.4%2020–2022: 5975972020–202225.4%2023–2026: 5925922023–202625.2%
Figure 3. Publication year of all 2,353 unique citations behind the catalogue. Bins are uneven in width — read the counts, not the silhouette.

The counter-intuitive part is what happens when you split evidence age by grade. The Strong tier rests on the oldest literature in the catalogue — pooled median citation year 2018. The newest citations sit behind measurement tools (2021), with both the Moderate and Emerging tiers at 2020.

Median citation year by evidence gradeStrong evidence: 2018; Moderate evidence: 2020; Emerging evidence: 2020; Preliminary: 2019; Measurement tool: 2021; Tested — did not work: 2020201620202023Strong evidence (n=11)Strong evidence (n=11): 20182018Moderate evidence (n=74)Moderate evidence (n=74): 20202020Emerging evidence (n=20)Emerging evidence (n=20): 20202020Preliminary (n=4)Preliminary (n=4): 20192019Measurement tool (n=20)Measurement tool (n=20): 20212021Tested — did not work (n=1)Tested — did not work (n=1): 20202020
Figure 4. Pooled median publication year of every citation attached to the products in each grade. The horizontal axis is truncated to 2016–2023 to make a 3-year spread legible; the whole corpus spans 1984–2026. Dots, not bars, because the axis does not start at zero.

The likeliest reading is boring and worth saying plainly: things get graded Strong once the literature has matured and stopped moving, so the citation list stops accumulating new entries. A rising median year is a sign of an active question, not a settled one. It is not evidence that the Strong grades are stale — but it does mean "backed by the newest science" is close to an anti-signal in this market.

Table 2. Grade, citation depth, evidence age, price and consumer rating by tier.
TierProductsMedian citationsPooled median citation yearMedian priceMean star rating
Strong evidence11202018$44.004.55 (n=11)
Moderate evidence74212020$25.874.53 (n=60)
Emerging evidence2020.52020$32.984.46 (n=17)
Preliminary419.52019$80.974.35 (n=2)
Measurement tool20222021$134.004.43 (n=15)
Tested — did not work1222020$38.004.00 (n=1)

Finding 4 — how much of it is randomized

749 of 2,353 citations (31.8%) carry a randomized-trial signal. Put the other way: about two thirds of the literature holding up the longevity shelf is observational, mechanistic or narrative-review work. That is not a scandal — mechanism is how you decide what to trial — but it is the number that should sit next to any sentence beginning "studies show".

At the product level the picture is better than the citation level, because randomized evidence is concentrated where it matters: 123 of 130 products (94.6%) have at least one randomized citation. The 7 that do not are listed in full below, because a report that only prints aggregates is asking to be trusted rather than checked.

Table 3. Every catalogue product with no randomized-trial citation (n = 7 of 130). The last column is derived, not editorial: it states only whether the grade came from the instrument rule.
ProductCategoryGradeCitationsOf those, human-signalHow it is graded
Wearable Pulse OximeterAir, Water & Sleep SystemsMeasurement tool2212Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials
Shower Dechlorination FilterAir, Water & Sleep SystemsEmerging evidence2211Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence
Smart Humidity MonitorAir, Water & Sleep SystemsModerate evidence187Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence
Grip-Strength DynamometerDiagnostics & Epigenetic TestsMeasurement tool1612Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials
Hand Grip Dynamometer ProAir, Water & Sleep SystemsStrong evidence2119Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence
Fasting Insulin / HOMA-IR KitDiagnostics & Epigenetic TestsMeasurement tool1613Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials
Cystatin C / eGFR Kidney TestDiagnostics & Epigenetic TestsMeasurement tool1811Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials

4 of those 7 are graded as instruments, where the absence of a randomized trial is expected rather than damning: you validate a home kidney test or a grip dynamometer against a reference method and against prognostic cohorts, not against placebo. See grip strength for the kind of literature that supports them. The remaining 3 sit on the strength ladder, which means their grades rest entirely on non-randomized human evidence — and one of them, Hand Grip Dynamometer Pro, holds a Strong grade on that basis: every one of its 21 citations carries a human-research signal and none carries a randomized one. That is the most attackable cell in this report, and it is printed rather than smoothed over — whether prognostic and validation evidence can carry a top grade is a legitimate argument to have with our rubric.

The mirror-image finding is the reassuring one. Zero of 130 products rest on animal or in-vitro data alone — every product has at least one citation carrying a human-research signal — and only 90 of 2,353 individual citations (3.8%) are preclinical with no human signal anywhere in their text. The mouse work is present in this corpus as supporting mechanism, not as the load-bearing evidence. Background reading: senolytics, autophagy, NAD⁺.

Finding 5 — how often the trials came back null

Of the 314 studies we maintain full structured write-ups for, 144 describe randomized evidence (individual trials plus meta-analyses and systematic reviews of trials). Of those, 61 (42.4%) report at least one explicitly null or non-significant result in their Results or Conclusion.

Read that claim precisely. It says "reports at least one null result", not "the trial failed". Many of these nulls are on secondary outcomes, on one dose arm of a dose-ranging trial, or on a subgroup — a trial can hit its main endpoint and still contribute a null here. We deliberately did not claim a primary-endpoint failure rate; see Claims we could not compute for why.

We hand-audited a deterministic sample of these matches — every 2nd match in alphabetical key order, 25 write-ups — against the underlying Results text on 2026-08-11. All 25 of 25 were genuine statements of a null or non-significant finding on at least one reported outcome. The classifier’s error, in other words, is almost entirely in the other direction: trials that reported a null without using any of our trigger phrases are counted as non-null, so 42.4% is a floor.

This is the number the industry does not print. A shelf where four in ten of the randomized write-ups contain a null is not a broken shelf — it is what an honest evidence base looks like when nobody has filtered it. The filtering is the problem everywhere else.

Table 4. Study design mix across the 314 structured write-ups.
Design (first matching rule wins)Write-upsShare
Meta-analysis or systematic review8627.4%
Randomized trial (not a review)9831.2%
Other human study (cohort, cross-sectional, validation)8426.8%
Preclinical (animal or in vitro)113.5%
Unclassified by the rules above3511.1%

The unclassified row is published rather than redistributed. It is mostly narrative reviews and mechanistic papers whose Methods text names no study population.

Finding 6 — which categories are best and worst supported

Two of the 7 categories — Diagnostics & Epigenetic Tests and Biometric Wearables & Sensors — consist mostly of measurement instruments, which sit off the strength ladder by design. They score zero Strong-or-Moderate products by construction, not by weakness of evidence, so ranking them against supplement categories would be a category error. They are excluded from the comparison below and reported on their own terms in Table 5.

Among the 5 categories that are graded on the strength ladder, the best-supported is Recovery, Sauna & Light Therapy (11 of 13 products at Strong or Moderate, 84.6%) and the worst-supported is Air, Water & Sleep Systems (8 of 15, 53.3%).

Table 5. Evidence profile by catalogue category. Citation depth is computed on unique compounds or device classes, not on products, so shared monographs are counted once.
CategoryProductsUnique compoundsGraded as instrumentsMedian citations per compoundMedian citation yearProducts with ≥1 randomized citationStrong / Strong+Moderate
Nutraceuticals & Cellular Energizers65600212020100%3 / 53
Diagnostics & Epigenetic Tests1010918201870%0 / 0
Air, Water & Sleep Systems1514121202073%4 / 8
Biometric Wearables & Sensors1141021.52021100%0 / 0
Recovery, Sauna & Light Therapy13130182021100%1 / 11
Skin & Topical Longevity11110192015100%2 / 9
Topical Pain Relief550182016–2017100%1 / 4

That "worst" label needs a footnote of its own, because Air, Water & Sleep Systems is also the category with the most Strong grades in the catalogue (4 of the 11). It is bimodal rather than weak: a few well-tested interventions sitting beside hardware whose randomized evidence is about the intervention class — filtered air, filtered water — and not about the particular machine. A single summary statistic hides that, which is why the full per-product rows are published as a CSV.

The honest summary is that Nutraceuticals & Cellular Energizers — the largest category, 65 products — has universal randomized coverage at the compound level and only 3 products at Strong. Depth of literature and strength of conclusion are close to independent in this market, which is exactly why citation counts should never be used as a proxy for evidence quality. Compare the compound pages for creatine and NMN to see the difference in kind rather than in volume.

Finding 7 — price and star ratings do not track evidence

If the market were efficient at pricing evidence, better-evidenced products would cost more, or at least be rated higher. Neither holds.

Mean Amazon star rating by evidence gradeStrong evidence: 4.55; Moderate evidence: 4.53; Emerging evidence: 4.46; Preliminary: 4.354.04.44.8Strong evidence (n=11)Strong evidence (n=11): 4.55★4.55★Moderate evidence (n=60)Moderate evidence (n=60): 4.53★4.53★Emerging evidence (n=17)Emerging evidence (n=17): 4.46★4.46★Preliminary (n=2)Preliminary (n=2): 4.35★4.35★
Figure 5. Mean star rating for products on each rung of the evidence ladder, n = 90. The axis is truncated to 4.0–4.8 of a possible 0–5 in order to make the differences visible at all — which is the point: the entire spread from the best-evidenced tier to the worst is 0.20 stars. Dots on a truncated axis, never bars.

The practical reading: a shopper sorting by star rating is sorting on nothing that has to do with whether the product works. The 4 Preliminary-grade products carry a median price of $80.97, against $25.87 for the 74 Moderate-grade products — the thinnest evidence in the catalogue is attached to some of its most expensive items. This is the single most defensible reason an evidence grade has to be published separately from the commercial signals, and the reason ours is set independently of affiliate commission.

Limitations

  1. This is one catalogue, not the market. 130 products chosen by one editorial team. Products were included because they were judged worth grading, which means genuinely evidence-free products are under-represented — a shelf assembled by someone else would almost certainly grade worse, not better. Every percentage here should be read as "of this catalogue", never "of the supplement industry".
  2. The design classifier reads titles and excerpts, not full texts. A randomized trial whose title, curated abstract excerpt and Methods paragraph never use a trigger word is counted as non-randomized. The bias therefore runs toward under-counting randomized evidence, which makes 31.8% a floor rather than an estimate. The same logic applies in reverse to the preclinical count: a mouse study whose abstract mentions "patients" in its framing will not be flagged preclinical-only, so 3.8% is also a floor.
  3. Grades are editorial judgements. They are consistent, documented and applied by one credentialed reviewer, but they are not the output of a formal GRADE or Cochrane risk-of-bias assessment and should not be cited as if they were. A different reviewer working from the same citations would produce a different distribution.
  4. Citation lists are curated, not exhaustive. They are the studies we judged load-bearing for a consumer decision. They are not a systematic search, they carry the selection preferences of whoever built them, and citation counts are therefore a measure of editorial attention as much as of literature volume.
  5. Monographs are shared across products. 130 products map to 117 monographs, so per-product citation counts and per-category medians computed on products are inflated for the shared cases. Table 5 avoids this by computing on unique compounds; Figure 2 does not, and says so.
  6. The null-result analysis covers 314 write-ups, not all 2,353 citations. The write-ups skew toward the studies that mattered most to a grading decision, which plausibly over-samples contested compounds — exactly where nulls live. Treat 42.4% as a property of that curated subset.
  7. Prices and star ratings are point-in-time snapshots from the catalogue and move constantly. The correlations in Finding 7 describe the snapshot in this dataset version, not a stable market parameter.
  8. No causal claim is made anywhere in this report. "Strong-graded products have older citations" is a description of a dataset. It is not a claim about why, and we offer the maturation explanation as a hypothesis, not a result.

Claims we could not compute, and dropped

Three numbers were on the original outline for this report and are not in it, because the data cannot support them. They are listed here rather than quietly omitted, since which questions a dataset cannot answer is part of describing it honestly.

How to reproduce this

Every figure above is emitted by a single script from three public files. Nothing is hand-entered, so a disagreement with any number in this report is a disagreement with code you can run.

Computed figures — JSONPer-product rows — CSVSource: data.js (catalogue, monographs, citations)Source: study-writeups.js (314 structured write-ups)
# the three inputs
curl -O https://magellanlongevity.com/data.js              # catalogue, monographs and 2,756 citation records
#                                          (2,353 of them are referenced by the shipped catalogue)
curl -O https://magellanlongevity.com/study-writeups.js    # 314 structured study write-ups
curl -O https://magellanlongevity.com/index.html           # carries TIER, EVID and EVID_PROD, the grading maps

# the analysis
node tools/gen_evidence_gap_report.js

# the outputs
reports/evidence-gap-2026.html
reports/evidence-gap-2026.json
reports/evidence-gap-2026-products.csv

The CSV carries one row per product: id, name, category, grade, citation count, median citation year, oldest and newest citation year, randomized-citation count, preclinical-only citation count, price and star rating. It is the file to start from if you want to recompute any figure independently, disagree with a grade, or test the classifier against your own read of the abstracts.

Licence. The computed figures, tables and charts in this report are released under CC BY 4.0 — reuse them anywhere, including commercially, with attribution and a link to this page. The underlying study abstracts belong to their publishers and are excerpted here under fair use for commentary; the licence does not extend to them.

How to cite this

Suggested citation:

Radu G. The Longevity Evidence Gap Report 2026. Magellan Longevity; 2026-08-11. Version 1.0. Available at: https://magellanlongevity.com/reports/evidence-gap-2026.html
@techreport{radu2026evidencegap,
  author      = {Radu, Gabriel},
  title       = {The Longevity Evidence Gap Report 2026},
  institution = {Magellan Longevity},
  year        = {2026},
  month       = {8},
  type        = {Research report},
  number      = {MAG-EGR-2026-1.0},
  url         = {https://magellanlongevity.com/reports/evidence-gap-2026.html}
}

Journalists and researchers: a press kit with the five most quotable findings, each written with its exact method and sample size, is maintained alongside this report. Corrections are welcome and will be published with a dated changelog on this page rather than edited in silently.

About the author

GR

Gabriel Radu, DO is a physiatrist — a physician specialising in physical medicine and rehabilitation — licensed in New York (licence 275110, NPI 1376861765) and the founder of Magellan Longevity. He designed the evidence rubric used here and assigned every curated grade in the dataset.

Editorial firewall. Magellan Longevity earns Amazon affiliate commission on some products in this catalogue. Evidence grades are assigned from the published literature before and independently of any commercial consideration, and no brand has ever paid for inclusion or for a grade. The clearest test of that policy is in this dataset: the correlation between price and evidence grade is -0.09, and the one product graded Tested — did not work is still listed, still linked, and still carries its commission.

Educational information, not medical advice. Nothing here is intended to diagnose, treat, cure, or prevent any disease. This report describes the strength of published evidence behind categories of products; it is not a recommendation for any individual. Talk to your physician before starting any supplement or device, especially if you are pregnant, nursing, or taking medication.

Read the underlying evidence

Evidence-based longevity supplements

The graded guide this dataset is drawn from.

Senolytics

The compound class where the preclinical-vs-human question in Finding 4 is most live.

NMN

A Moderate grade on 34 citations — the pattern Finding 6 describes.

Creatine

One of the 11 Strong grades, and the least glamorous.

Wearables & sensors

The shared-monograph caveat in Figure 2, in practice.

The research map

Every study in the corpus, by pathway.

All Magellan reports

Methods, datasets and prior versions.

How we grade (in the app)

The rubric as it appears to shoppers.

This report is a static page. The same grades, citations and study write-ups are browsable interactively in the Magellan app — open the grading methodology in the app →