MAGELLAN LONGEVITY
HomeAbout › How we grade

How we grade evidence

Written and medically reviewed by , physiatrist · Last reviewed · How we grade

Magellan exists to close the gap between “here’s a study” and “here’s what’s actually worth buying.” This page is the complete rubric: the six grades, what evidence counts toward each, the exact order in which a grade is decided, a worked example at every level, and the specific things that get a product downgraded.

The rubric is applied by Gabriel Radu, DO, a physiatrist. It is applied before any commercial relationship exists, and the editorial policy explains why a commission cannot move a grade.

On this pageThe six gradesWhere the catalogue sits todayWhat counts as evidenceHow a grade is decidedWorked examplesWhat causes a downgradeWhat we never do

The six grades

Ingestibles and interventions are graded on the strength of human evidence for the active compound. Diagnostics and wearables are graded on how well the thing they measure predicts health — not on a treatment effect. The grade always answers a specific question: strong evidence for what, measured how.

Strong evidence

Strong evidence

Consistent randomized human trials — or large, convergent human cohort data, sometimes backed by genetic (Mendelian) evidence — show a real benefit. That benefit is usually a specific, measurable outcome (blood pressure, blood glucose, joint pain, muscle strength), not proof of a longer lifespan. Each product page and caveat states exactly what the strong evidence is for.

Moderate evidence

Moderate evidence

Several human studies point the same way, but with real limits — sample size, duration, industry funding, or mixed results across trials. This is the most common grade in the catalogue, and it is an honest one: it means “probably something, at a size worth arguing about”.

Emerging evidence

Emerging evidence

The mechanism is plausible and early human data exists, but it is small, short, or not yet independently replicated. Expect the grade to move in either direction as trials report.

Preliminary

Preliminary

Mostly mechanistic or animal data, or very few human studies. Interesting, not proven. A product can sit here for years with a large and enthusiastic literature behind it if none of that literature is a controlled human trial.

Measurement tool

Measurement tool

A test or sensor. Graded as an instrument on two questions: how accurately does it measure what it claims to measure, against what reference standard — and how well does that marker predict health? A measurement only adds value if you act on it, so every instrument page says what you would do differently with the number.

Tested — did not work

Tested — did not work

The bottom of the scale, and the most useful grade on the site. It is not for products with thin evidence — those are Preliminary. It is for products with adequate evidence pointing at nothing. Keeping a product in the catalogue at this grade, with the trials that sank it quoted on the page, is more useful to a reader than quietly deleting it.

Where the catalogue sits today

The distribution below is generated from the live catalogue, not written by hand. It is deliberately unflattering: most things are Moderate, few are Strong, and one has been tested and failed.

GradeProducts
Strong evidence11 of 130
Moderate evidence74 of 130
Emerging evidence20 of 130
Preliminary4 of 130
Measurement tool20 of 130
Tested — did not work1 of 130

What counts as evidence, in order

  1. Systematic reviews and meta-analyses of randomized human trials — including Cochrane reviews. The strongest single input, and the one most likely to produce a downgrade rather than an upgrade.
  2. Randomized controlled trials in humans, read at the primary endpoint. A trial that missed its primary endpoint and reported a positive secondary one is treated as a miss, and the page says so.
  3. Large prospective cohorts, especially where several independent cohorts converge. Cohorts establish association, not causation, and the page says which one it is.
  4. Mendelian randomization and other genetic evidence, which can lift a cohort association toward causal.
  5. Mechanistic and animal work, which explains how something might work and can never on its own establish that it works in people.
  6. Regulatory status — an FDA clearance for a device feature, or a prescription approval — as a fact about validation, not as proof of benefit.

Effect statements printed on product pages are quoted from a cited study rather than paraphrased, precisely so that a stronger claim cannot creep in during the rewrite.

How a grade is decided

Grades are resolved in a fixed order, and the more specific rule always wins:

  1. Product-level curated grade (26 products today). Used where the evidence genuinely differs between products that share an ingredient or a category — nine wearables share one compound entry, but an FDA-cleared ECG feature and an unvalidated proprietary “readiness” score are not the same claim.
  2. Instrument default. A diagnostic or wearable with no product-level entry is graded as a Measurement tool.
  3. Compound-level curated grade (104 products today). A physician-assigned grade for the active compound, with a dose shown only where a verified standard range exists.
  4. Provisional grade, derived conservatively from the peer-reviewed studies linked to the page: three or more linked studies → Moderate, two → Emerging, one or none → Preliminary. A provisional grade can never reach Strong. It stays capped until a human reviews it.

Every product in the catalogue today is graded by one of the first three rules; the provisional path exists so that nothing new can appear on the site with an unearned grade.

A worked example at every grade

These are live pages, with the caveat text quoted exactly as it is published.

Strong evidence

Creatine Monohydrate

Nutraceuticals & Cellular Energizers · 23 cited studies · grade set at the compound level

Why this grade: Dozens of randomized trials and meta-analyses converge on the same functional outcomes — strength, power and lean mass with resistance training — with a long safety record. Note what the grade is not: it is not evidence for lifespan.

The caveat published on the page: “One of the most studied supplements there is: dozens of randomized trials and meta-analyses show reliable gains in strength, power and lean muscle alongside resistance training, and it is well tolerated in healthy people. The strong evidence is for muscle and exercise performance; cognitive, bone and healthy-aging benefits are promising but less settled. It supports training — it is not a treatment.”

Dose shown: 3–5 g/day (creatine monohydrate); an optional ~20 g/day loading week works faster but isn’t required

Moderate evidence

Omega-3 EPA/DHA Fish Oil

Nutraceuticals & Cellular Energizers · 22 cited studies · grade set at the compound level

Why this grade: The biochemical effect is real and dose-dependent, but two large randomized trials failed to show the outcome people actually buy it for, and a safety signal exists at higher doses. Real effect, contested importance — that is what Moderate means.

The caveat published on the page: “Standard over-the-counter doses did not prevent cardiovascular events or cancer in large trials (VITAL, STRENGTH) and meta-analyses, though triglyceride lowering is real. Doses above 1 g/day carry a genuine atrial-fibrillation signal - about 13% higher incident risk in a large cohort of healthy users.”

Dose shown: Typical studied range 0.5-4 g/day EPA+DHA; triglyceride and blood-pressure effects are dose-dependent, clearest around 2-3 g/day. Event-reduction evidence is for prescription 4 g/day purified EPA in high-risk patients.

Emerging evidence

Magnesium L-Threonate

Nutraceuticals & Cellular Energizers · 19 cited studies · grade set at the compound level

Why this grade: Small, short, industry-funded randomized trials that agree with the rodent mechanism, with no independent replication and the key mechanistic step unconfirmed in humans.

The caveat published on the page: “A few small, short, industry-funded randomized trials report cognitive and sleep benefits consistent with rodent mechanism data, but there is no independent replication and human brain-magnesium elevation has not been directly confirmed. It provides little elemental magnesium, so it is a cognitive-support candidate, not a way to correct magnesium deficiency.”

Dose shown: 1-2 g/day in trials (about 144 mg elemental magnesium per 2 g)

Preliminary

Copper Tripeptide (GHK-Cu) Serum

Skin & Topical Longevity · 21 cited studies · grade set at the compound level

Why this grade: Decades of laboratory and animal data with essentially no rigorous human trials for the claimed use — the exact profile the Preliminary grade exists to describe.

The caveat published on the page: “Decades of laboratory and animal data show wound-healing and collagen-stimulating effects, but rigorous human trials for skin aging are essentially absent — the best controlled study found no objective benefit, only higher patient satisfaction. Skin penetration is a known limitation that newer liposomal formulations are trying to solve.”

Measurement tool

Oura Ring Gen 4

Biometric Wearables & Sensors · 104 cited studies · grade set at the product level

Why this grade: Graded as an instrument: excellent against an ECG reference for cardiac metrics, softer on sleep staging, and with proprietary composite scores that have no published validation at all. The grade describes the measurement, not a benefit.

The caveat published on the page: “For duration and cardiac metrics this is the best-validated device on this list: against an ECG reference across 536 nights the Gen 4 gave nocturnal resting heart rate within about 2% and HRV within about 6%, the best of five wearables tested, and against multi-night polysomnography the Gen 3 did not differ significantly from PSG for total sleep time, wake after sleep onset, light sleep or deep sleep. Sleep staging is still the soft spot everywhere, including here: per-stage sensitivity ran about 76-80% in a laboratory comparison, wake-detection specificity was 0.41 in a week of unrestricted home sleep, and the most favourable of those studies has an author who sits on Oura's medical advisory board. Readiness and Cardiovascular Age are proprietary composites with no published validation against any reference standard and no FDA clearance for any feature. Nobody has shown that acting on a Readiness score changes sleep, fitness or illness risk, and for a subset of people nightly scoring makes sleep worse rather than better - the pattern sleep physicians named orthosomnia.”

Dose shown: Wear it nightly for at least two weeks before reading anything into a number; single nights are noise. The defensible outputs are sleep duration, nocturnal resting heart rate and the HRV trend. Readiness is a composite, not a measurement.

Tested — did not work

SYB Clear Blue Light Glasses

Air, Water & Sleep Systems · 22 cited studies · grade set at the product level

Why this grade: Adequate evidence pointing at nothing. A 2023 Cochrane review pooled 17 randomized trials and found no reliable benefit — which is why this sits at the bottom of the scale rather than in the “promising” tier.

The caveat published on the page: “These have been tested properly and they did not work. That distinction matters: this is not thin evidence, it is adequate evidence pointing at nothing, which is why it is graded at the bottom rather than as something promising. The 2023 Cochrane review pooled 17 randomized trials and concluded that blue-light filtering lenses may make no difference to eye strain with computer use over short follow-up, probably make little or no difference to visual acuity, and have indeterminate effects on sleep — three trials reported improvement, three reported none. Cochrane also recorded the reported harms, infrequent but real, including headache, discomfort, lower mood and increased depressive symptoms; there is no evidence at all on macular health, contrast sensitivity or melatonin, because no trial measured them.”

Dose shown: No protocol, because no wearing schedule has been shown to produce a benefit. If the goal is screen comfort, the treatable causes are dry eye, an uncorrected refractive error and not blinking — an eye exam is the better $38. If the goal is sleep, dim the lights and cut screen time in the hour before bed, which is where the evidence actually points.

What causes a downgrade

Grades move down more often than up. Any one of the following is enough to trigger a review, and several are enough to move a grade:

  1. A large randomized trial reports null on the outcome the grade was resting on. Two large trials failing to reduce cardiovascular events is why the marine omega-3 entry sits at Moderate rather than Strong despite an unambiguous effect on triglycerides.
  2. A systematic review pools the trials and finds nothing. A Cochrane review of 17 randomized trials is what moved blue-light-filtering lenses to “Tested — did not work”.
  3. The trial missed its primary endpoint. A positive secondary outcome after a missed primary is hypothesis-generating, and the caveat says so — as it does on the urolithin A page.
  4. The evidence turns out to be industry-funded and unreplicated. Manufacturer-funded trials are not disqualified, but a grade will not exceed Emerging on manufacturer-funded data alone.
  5. The benefit requires a co-intervention the product does not include. Home blood-pressure monitoring lowers blood pressure when it is paired with clinician-guided titration; the device on its own does little, and the caveat states it.
  6. A safety signal appears — for example the atrial-fibrillation signal at higher omega-3 doses. A safety finding is published on the page whether or not it changes the grade.
  7. The measurement cannot be validated. For instruments, a proprietary score with no published comparison against a reference standard is described as unvalidated, no matter how confident the app looks.

Every grade change is dated and published in the evidence changelog. Nothing is edited silently.

What we never do

Questions about the rubric

How many evidence grades does Magellan use?

Six. Strong evidence, Moderate evidence, Emerging evidence, Preliminary, Measurement tool, and Tested — did not work. The first four rank the strength of human evidence for an active compound; Measurement tool is used for tests and sensors, which are graded on what they measure rather than on a treatment effect; the last one is reserved for interventions that have been tested properly and failed.

Does an affiliate commission change an evidence grade?

No. Grades are assigned from the human and mechanistic literature before any commercial link exists, and no brand can pay for a listing, a grade, or placement. One product in the catalogue currently carries the grade "Tested — did not work" and still carries its honest write-up.

What happens when a product has very few studies?

It is capped. Without a curated review, a product with three or more linked studies can rise no higher than Moderate, two studies gives Emerging, and one or none gives Preliminary. A provisional grade can never reach Strong.

Why are diagnostics graded differently from supplements?

Because a test does not treat anything. A diagnostic or wearable is graded on how accurately it measures what it claims to measure and on how well that marker predicts health outcomes — and a measurement only adds value if you act on it.

Scope. Magellan Longevity publishes general wellness and educational information. It is not medical advice, it does not create a doctor–patient relationship, and nothing here is intended to diagnose, treat, cure, or prevent any disease. Talk to your own clinician before starting or stopping anything, especially if you are pregnant, nursing, or taking medication.

Related: who assigns these grades · the change log · the supplement guide · Diagnostics & Epigenetic Tests · the methodology page in the store →