We took the 130 products in the Magellan Longevity catalogue, the 2,353 unique PubMed-indexed citations behind them and the 314 structured study write-ups we maintain, and asked one question of the whole market at once: how strong is the evidence, really? Every number below is computed from published data by a script we ship with the report. Where a number was not computable, we say so and drop the claim.
The longevity market is sold on mechanism. A compound activates a pathway in a cell line, the pathway is implicated in aging, and by the time the story reaches a product label it has quietly become a benefit. Nobody publishes the denominator — how much of the shelf actually has randomized human evidence behind it, how old that evidence is, and how often it came back negative.
Magellan Longevity has, for its own storefront, been forced to answer that question product by product: every item we list carries an evidence grade and links to its primary citations, and the grade is set independently of what the product pays in affiliate commission. That grading work leaves behind a dataset — 130 graded products, 2,353 unique PubMed-indexed citations, 314 structured study write-ups — that can be read sideways, as a description of the market rather than of any one product.
That is what this report is. It is not a ranking, and it is not an accusation: a Moderate grade is a normal, respectable place for a supplement to sit. It is an attempt to put an honest number on how much of the evidence base is randomized, how recent it is, and how much of it is null — and to publish the method so that anyone who disagrees can rerun it and argue with the code rather than with us.
Corpus. The unit of analysis is the shipped catalogue: 130 products across 7 categories, mapped to 117 unique compound or device monographs. Each monograph carries a hand-assembled citation list; across the catalogue those lists resolve to 2,353 unique studies, of which 2,353 (100.0%) carry a PubMed ID and 2,022 (85.9%) carry a DOI. Publication years span 1984–2026. A separate layer of 314 structured write-ups — each with an explicit Question, Methods, Results, Conclusion and Limitations field — covers the most load-bearing studies and is used for the design and null-result analyses.
Grades. Grades are not recomputed for this report. The script reads the same three objects the live site reads — the tier definitions, the product-specific grade map and the compound-level grade map — out of the site shell and applies the site’s own precedence rule: a product-specific curated grade wins; failing that, diagnostics and wearables are graded as instruments; failing that, the compound-level curated grade applies; and only if none of those exist does an automatic citation-count fallback fire. Across this catalogue the fallback never fires. 26 of 130 grades come from a product-specific curated entry and 104 from a compound-level curated entry; 0 fall through to the instrument rule and 0 to the automatic citation-count fallback. Every grade in this report is therefore human editorial judgement, not an algorithm dressed up as one.
Design classification. Individual citations carry a title, a curated one-sentence excerpt from the abstract and — where a write-up exists — a Methods paragraph. Those three fields are concatenated and matched against three published regular expressions. This is a keyword classifier, not a human read of every full text, and its error profile is described under Limitations.
RANDOMIZED /randomi[sz]ed|randomi[sz]ation|randomly (assigned|allocated)|
placebo-controlled|double-blind|triple-blind|single-blind|
crossover (trial|study|design)|RCTs?/i
HUMAN /patients?|participants?|subjects?|volunteers?|humans?|men|women|
adults?|children|cohort|population-based|NHANES|UK Biobank|
case-control|cross-sectional|clinical trial|meta-analy|
systematic review|epidemiolog/i
PRECLINICAL /mice|mouse|murine|rats?|rodents?|in vitro|cell (line|culture|s)|
C. elegans|caenorhabditis|drosophila|zebrafish|yeast|preclinical|
animal model|fibroblasts?|myotubes?|ex vivo/iNull-result detection. Within the 314 structured write-ups, the Results and Conclusion fields are matched against 12 patterns for explicit statements of a non-significant finding (no significant difference, did not improve, failed to meet, not superior to placebo, nonsignificant, and eight more, all listed in the shipped script). A write-up counts as randomized if its title, Question or Methods matches the RANDOMIZED pattern.
Statistics. Medians and interquartile ranges are reported for skewed counts. Correlations are Pearson’s r with a Fisher z 95% confidence interval, plus Spearman’s ρ because the grade ladder is ordinal. Grades are mapped Strong = 4, Moderate = 3, Emerging = 2, Preliminary = 1; measurement tools and the refuted grade sit off that ladder and are excluded from correlations. No p-values are reported — this is a census of one catalogue, not a sample from a population, so the confidence intervals describe estimation noise in the mapping, not a hypothesis test.
Rounding. Percentages are given to one decimal place using ordinary half-up rounding, and never nudged toward a more striking number. Every percentage in this report is printed next to its raw numerator and denominator so the rounding can be checked.
Six tiers. The first four are a strength ladder; the last two are not on it. Measurement tool exists because grading a blood-pressure cuff on "treatment effect" is a category error — an instrument is graded on how accurately it measures and on whether the marker predicts anything. Tested — did not work exists because "we looked and found nothing" is a different, stronger statement than "nobody has looked yet", and collapsing the two would flatter the shelf.
| Tier | Letter | What it means | Products | Share |
|---|---|---|---|---|
| Strong evidence | A | Consistent human evidence, including randomized trials, on outcomes that matter. | 11 | 8.5% |
| Moderate evidence | B | Real human evidence, but mixed, limited to surrogate endpoints, or small in scale. | 74 | 56.9% |
| Emerging evidence | C | Early human signal on top of a solid mechanism — promising, not settled. | 20 | 15.4% |
| Preliminary | C- | Mechanism or very early human data only. Buy it as an experiment, not a plan. | 4 | 3.1% |
| Measurement tool | B | Graded on measurement accuracy and on whether the marker predicts anything — not on a treatment effect. | 20 | 15.4% |
| Tested — did not work | D | Adequately tested and did not work. Not thin evidence — evidence pointing at nothing. | 1 | 0.8% |
Only 11 of 130 products (8.5%) earn the top grade. The catalogue’s centre of gravity is Moderate — 74 products, 56.9% — which is the honest resting place for most of this market: real human data, but mixed, small, or measured on a surrogate endpoint rather than on an outcome anyone lives or dies by.
The products at the top of the ladder are not the ones the category is marketed on. They are, in full: Bright Light Therapy Lamp, Lutein and Zeaxanthin, Voltaren Arthritis Pain Gel (Diclofenac 1%), Creatine Monohydrate, Betaine (TMG), Daily Mineral SPF Sunscreen, Validated Home BP Monitor, Balance and Strength Trainer, Azelaic Acid Gel, 24h Ambulatory BP Monitor, Hand Grip Dynamometer Pro. Sunscreen, creatine, a light box, a grip dynamometer, a validated blood-pressure cuff, lutein and zeaxanthin, a balance trainer, betaine, azelaic acid, an ambulatory BP monitor and topical diclofenac. None of them is a NAD⁺ precursor, a senolytic or a "cellular rejuvenation" formula.
At the other end sit 4 Preliminary-grade products (Copper Tripeptide (GHK-Cu) Serum, Saliva Cortisol Stress Test, Allergen Alert Mini Food Lab, Phosphatidylcholine (PC) Complex) and exactly one product graded Tested — did not work: SYB Clear Blue Light Glasses. That single entry is the most useful cell in the whole table, because it is the one place where the literature is not thin but adequate — and adequate literature pointing at nothing is a finding, not a gap. Related background: blue light.
Median 21 citations per product (IQR 18–24; range 14–104). No product ships with fewer than 14. That is a floor, not a boast: it means the interesting variation in this market is not whether a compound has published literature — essentially all of them do — but what kind of literature it is. Citation count is the weakest measure in this report, and we lead with it only in order to retire it.
Median publication year 2020. But a median hides the shape: 543 citations (23.1%) were published in 2014 or earlier, and 1,016 (43.2%) since 2021. Roughly a quarter of the evidence base behind the modern longevity shelf is more than a decade old.
The counter-intuitive part is what happens when you split evidence age by grade. The Strong tier rests on the oldest literature in the catalogue — pooled median citation year 2018. The newest citations sit behind measurement tools (2021), with both the Moderate and Emerging tiers at 2020.
The likeliest reading is boring and worth saying plainly: things get graded Strong once the literature has matured and stopped moving, so the citation list stops accumulating new entries. A rising median year is a sign of an active question, not a settled one. It is not evidence that the Strong grades are stale — but it does mean "backed by the newest science" is close to an anti-signal in this market.
| Tier | Products | Median citations | Pooled median citation year | Median price | Mean star rating |
|---|---|---|---|---|---|
| Strong evidence | 11 | 20 | 2018 | $44.00 | 4.55 (n=11) |
| Moderate evidence | 74 | 21 | 2020 | $25.87 | 4.53 (n=60) |
| Emerging evidence | 20 | 20.5 | 2020 | $32.98 | 4.46 (n=17) |
| Preliminary | 4 | 19.5 | 2019 | $80.97 | 4.35 (n=2) |
| Measurement tool | 20 | 22 | 2021 | $134.00 | 4.43 (n=15) |
| Tested — did not work | 1 | 22 | 2020 | $38.00 | 4.00 (n=1) |
749 of 2,353 citations (31.8%) carry a randomized-trial signal. Put the other way: about two thirds of the literature holding up the longevity shelf is observational, mechanistic or narrative-review work. That is not a scandal — mechanism is how you decide what to trial — but it is the number that should sit next to any sentence beginning "studies show".
At the product level the picture is better than the citation level, because randomized evidence is concentrated where it matters: 123 of 130 products (94.6%) have at least one randomized citation. The 7 that do not are listed in full below, because a report that only prints aggregates is asking to be trusted rather than checked.
| Product | Category | Grade | Citations | Of those, human-signal | How it is graded |
|---|---|---|---|---|---|
| Wearable Pulse Oximeter | Air, Water & Sleep Systems | Measurement tool | 22 | 12 | Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials |
| Shower Dechlorination Filter | Air, Water & Sleep Systems | Emerging evidence | 22 | 11 | Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence |
| Smart Humidity Monitor | Air, Water & Sleep Systems | Moderate evidence | 18 | 7 | Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence |
| Grip-Strength Dynamometer | Diagnostics & Epigenetic Tests | Measurement tool | 16 | 12 | Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials |
| Hand Grip Dynamometer Pro | Air, Water & Sleep Systems | Strong evidence | 21 | 19 | Graded on the strength ladder, so this grade rests entirely on non-randomized human evidence |
| Fasting Insulin / HOMA-IR Kit | Diagnostics & Epigenetic Tests | Measurement tool | 16 | 13 | Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials |
| Cystatin C / eGFR Kidney Test | Diagnostics & Epigenetic Tests | Measurement tool | 18 | 11 | Graded as a measurement instrument — validation and prognostic designs apply, not randomized trials |
4 of those 7 are graded as instruments, where the absence of a randomized trial is expected rather than damning: you validate a home kidney test or a grip dynamometer against a reference method and against prognostic cohorts, not against placebo. See grip strength for the kind of literature that supports them. The remaining 3 sit on the strength ladder, which means their grades rest entirely on non-randomized human evidence — and one of them, Hand Grip Dynamometer Pro, holds a Strong grade on that basis: every one of its 21 citations carries a human-research signal and none carries a randomized one. That is the most attackable cell in this report, and it is printed rather than smoothed over — whether prognostic and validation evidence can carry a top grade is a legitimate argument to have with our rubric.
The mirror-image finding is the reassuring one. Zero of 130 products rest on animal or in-vitro data alone — every product has at least one citation carrying a human-research signal — and only 90 of 2,353 individual citations (3.8%) are preclinical with no human signal anywhere in their text. The mouse work is present in this corpus as supporting mechanism, not as the load-bearing evidence. Background reading: senolytics, autophagy, NAD⁺.
Of the 314 studies we maintain full structured write-ups for, 144 describe randomized evidence (individual trials plus meta-analyses and systematic reviews of trials). Of those, 61 (42.4%) report at least one explicitly null or non-significant result in their Results or Conclusion.
We hand-audited a deterministic sample of these matches — every 2nd match in alphabetical key order, 25 write-ups — against the underlying Results text on 2026-08-11. All 25 of 25 were genuine statements of a null or non-significant finding on at least one reported outcome. The classifier’s error, in other words, is almost entirely in the other direction: trials that reported a null without using any of our trigger phrases are counted as non-null, so 42.4% is a floor.
This is the number the industry does not print. A shelf where four in ten of the randomized write-ups contain a null is not a broken shelf — it is what an honest evidence base looks like when nobody has filtered it. The filtering is the problem everywhere else.
| Design (first matching rule wins) | Write-ups | Share |
|---|---|---|
| Meta-analysis or systematic review | 86 | 27.4% |
| Randomized trial (not a review) | 98 | 31.2% |
| Other human study (cohort, cross-sectional, validation) | 84 | 26.8% |
| Preclinical (animal or in vitro) | 11 | 3.5% |
| Unclassified by the rules above | 35 | 11.1% |
The unclassified row is published rather than redistributed. It is mostly narrative reviews and mechanistic papers whose Methods text names no study population.
Two of the 7 categories — Diagnostics & Epigenetic Tests and Biometric Wearables & Sensors — consist mostly of measurement instruments, which sit off the strength ladder by design. They score zero Strong-or-Moderate products by construction, not by weakness of evidence, so ranking them against supplement categories would be a category error. They are excluded from the comparison below and reported on their own terms in Table 5.
Among the 5 categories that are graded on the strength ladder, the best-supported is Recovery, Sauna & Light Therapy (11 of 13 products at Strong or Moderate, 84.6%) and the worst-supported is Air, Water & Sleep Systems (8 of 15, 53.3%).
| Category | Products | Unique compounds | Graded as instruments | Median citations per compound | Median citation year | Products with ≥1 randomized citation | Strong / Strong+Moderate |
|---|---|---|---|---|---|---|---|
| Nutraceuticals & Cellular Energizers | 65 | 60 | 0 | 21 | 2020 | 100% | 3 / 53 |
| Diagnostics & Epigenetic Tests | 10 | 10 | 9 | 18 | 2018 | 70% | 0 / 0 |
| Air, Water & Sleep Systems | 15 | 14 | 1 | 21 | 2020 | 73% | 4 / 8 |
| Biometric Wearables & Sensors | 11 | 4 | 10 | 21.5 | 2021 | 100% | 0 / 0 |
| Recovery, Sauna & Light Therapy | 13 | 13 | 0 | 18 | 2021 | 100% | 1 / 11 |
| Skin & Topical Longevity | 11 | 11 | 0 | 19 | 2015 | 100% | 2 / 9 |
| Topical Pain Relief | 5 | 5 | 0 | 18 | 2016–2017 | 100% | 1 / 4 |
That "worst" label needs a footnote of its own, because Air, Water & Sleep Systems is also the category with the most Strong grades in the catalogue (4 of the 11). It is bimodal rather than weak: a few well-tested interventions sitting beside hardware whose randomized evidence is about the intervention class — filtered air, filtered water — and not about the particular machine. A single summary statistic hides that, which is why the full per-product rows are published as a CSV.
The honest summary is that Nutraceuticals & Cellular Energizers — the largest category, 65 products — has universal randomized coverage at the compound level and only 3 products at Strong. Depth of literature and strength of conclusion are close to independent in this market, which is exactly why citation counts should never be used as a proxy for evidence quality. Compare the compound pages for creatine and NMN to see the difference in kind rather than in volume.
If the market were efficient at pricing evidence, better-evidenced products would cost more, or at least be rated higher. Neither holds.
The practical reading: a shopper sorting by star rating is sorting on nothing that has to do with whether the product works. The 4 Preliminary-grade products carry a median price of $80.97, against $25.87 for the 74 Moderate-grade products — the thinnest evidence in the catalogue is attached to some of its most expensive items. This is the single most defensible reason an evidence grade has to be published separately from the commercial signals, and the reason ours is set independently of affiliate commission.
Three numbers were on the original outline for this report and are not in it, because the data cannot support them. They are listed here rather than quietly omitted, since which questions a dataset cannot answer is part of describing it honestly.
Every figure above is emitted by a single script from three public files. Nothing is hand-entered, so a disagreement with any number in this report is a disagreement with code you can run.
# the three inputs
curl -O https://magellanlongevity.com/data.js # catalogue, monographs and 2,756 citation records
# (2,353 of them are referenced by the shipped catalogue)
curl -O https://magellanlongevity.com/study-writeups.js # 314 structured study write-ups
curl -O https://magellanlongevity.com/index.html # carries TIER, EVID and EVID_PROD, the grading maps
# the analysis
node tools/gen_evidence_gap_report.js
# the outputs
reports/evidence-gap-2026.html
reports/evidence-gap-2026.json
reports/evidence-gap-2026-products.csvThe CSV carries one row per product: id, name, category, grade, citation count, median citation year, oldest and newest citation year, randomized-citation count, preclinical-only citation count, price and star rating. It is the file to start from if you want to recompute any figure independently, disagree with a grade, or test the classifier against your own read of the abstracts.
Licence. The computed figures, tables and charts in this report are released under CC BY 4.0 — reuse them anywhere, including commercially, with attribution and a link to this page. The underlying study abstracts belong to their publishers and are excerpted here under fair use for commentary; the licence does not extend to them.
Suggested citation:
@techreport{radu2026evidencegap,
author = {Radu, Gabriel},
title = {The Longevity Evidence Gap Report 2026},
institution = {Magellan Longevity},
year = {2026},
month = {8},
type = {Research report},
number = {MAG-EGR-2026-1.0},
url = {https://magellanlongevity.com/reports/evidence-gap-2026.html}
}Journalists and researchers: a press kit with the five most quotable findings, each written with its exact method and sample size, is maintained alongside this report. Corrections are welcome and will be published with a dated changelog on this page rather than edited in silently.
The graded guide this dataset is drawn from.
The compound class where the preclinical-vs-human question in Finding 4 is most live.
A Moderate grade on 34 citations — the pattern Finding 6 describes.
One of the 11 Strong grades, and the least glamorous.
The shared-monograph caveat in Figure 2, in practice.
Every study in the corpus, by pathway.
Methods, datasets and prior versions.
The rubric as it appears to shoppers.
This report is a static page. The same grades, citations and study write-ups are browsable interactively in the Magellan app — open the grading methodology in the app →