How we grade evidence
Every evidence label you see on this site comes from the system below. It exists because the standard way of grading health claims has a bias built into it, and that bias quietly distorts what people believe about their options.
The problem with one scale
Most evidence scales are a single ladder. Randomized controlled trials sit at the top, and everything else is ranked by how closely it resembles one. A practice with eight hundred years of documented clinical use lands near the bottom, in the same band as something a stranger posted on a forum.
That ladder answers a question nobody asked. It measures how closely evidence resembles a particular Western method, then presents the answer as if it measured whether something works. Those are not the same question.
The cost is specific and it is not hypothetical. When the FDA reviewed a longevity peptide in 2026, it stated in its own briefing document that it could not evaluate two of the key studies because they were published in Russian and it could not locate English versions. On a single ladder, that renders as a low grade, and the reader concludes the compound does not work. What actually happened is that nobody read the papers. A language barrier became a scientific verdict.
Two questions, kept separate
So we ask two questions instead of one, and we never collapse them into a single score.
Corroboration
How much independent confirmation a claim has. This is the only axis we rank, because convergence across separate investigations genuinely is a measure of confidence.
- Established
- Findings converge across multiple independent investigations. Independence matters: a large body of work from a single group does not reach this level.
- Emerging
- Real investigation exists, but findings have not yet converged across independent sources. This describes the state of the research, not a judgment that the approach does not work.
Provenance
What kind of investigation produced the evidence. These are peers, not tiers. We do not rank a tradition above another, and none of them is treated as the default against which others fall short.
- Clinical
- Human clinical research. Country of origin does not affect this label: Russian, American, and European clinical programs are the same kind of evidence.
- Traditional
- Classical lineage with documented long-term use. Long use is strong evidence about tolerability and weaker evidence about efficacy.
- Preclinical
- Animal and laboratory research. Establishes mechanism and plausibility; does not establish an effect in humans.
- Field
- Practitioner observation and recorded case series. Real patients and named authors, without a control arm.
Contrary finding
Reserved for research that tested something properly and found it did not work. We apply it rarely and deliberately, because a label like this is only meaningful if it is never used to mean "we did not look".
- Did not replicate
- Adequately tested and failed to show the expected effect. Applied only to genuine negative results, never to an absence of research.
What we will not do
Score the absence of a trial
If nobody has run the study, that is a fact about the research landscape, not a property of the thing being studied. It gets described, never graded. You will never see a low mark here that actually means “underfunded” or “not translated”.
Rank one tradition above another
Provenance labels carry their own colours, deliberately kept off the green-to-amber scale, because they are categories rather than grades. Russian clinical work and American clinical work carry the same label. Neither is the default that the other is measured against.
Put numbers on a badge
A label that says “five studies” is really claiming that five is all that exists, which is a claim about the entire world literature that we cannot verify and that stops being true the moment somebody publishes. Specific studies, sample sizes, and citations appear in our detailed compound reports, where they are dated and scoped to exactly what we reviewed.
What this does not mean
Refusing to rank traditions is not the same as treating all claims as equal. Long use tells you a great deal about whether something is tolerated and considerably less about whether it does what people believe it does. Centuries of practitioners would have noticed patients being harmed. They would not necessarily have noticed a placebo response. We hold both halves of that at once.
Corroboration is the one axis we do rank, and independence is what it measures. A large body of work from a single research group does not reach the same level as findings that separate teams arrived at on their own. That standard applies everywhere, without exception for the traditions we find most compelling.
Evidence labels describe research. They are not medical advice and not a recommendation to use anything. If you want help thinking through what any of this means for your situation specifically, that is what a discovery call is for.