In this note

That gap is not a rounding artefact or a disagreement about fees. It is the visible edge of a structural problem, and it applies to both of the labels that onchain currently offers for identifying a good trader. One is a leaderboard rank. The other is a smart money tag. Below is our own data on each, and on the question neither of them asks.

The first label: leaderboard rank

Start with a concession, because it matters and because most critiques of leaderboards get it wrong.

The leaderboard formula is not naive. It does not mistake your deposits for profit. It divides realised and unrealised PnL by the account value you started the window with, plus the maximum net deposits you made during it, with a floor to stop microscopic accounts producing absurd percentages. Flows are already handled.

That is worth stating plainly because it is checkable, and because the interesting failure is not the one people assume.

Recovering the formula

Hyperliquid no longer publishes this formula. The leaderboard documentation page returns a 404 and there is no leaderboard section in the documentation corpus served for machine reading. So we recovered it empirically.

On 7 August 2026 we pulled the live leaderboard payload, 41,381 rows, and back-solved the denominator each row implies by dividing its reported PnL by its reported ROI. If the denominator is the account value at the start of the window, the implied figure should equal current account value minus window PnL. If it is the current account value, it should equal account value itself.

WindowImplied denominator ÷ (account value − PnL)Within 2%
Daymedian 1.000085.0%
Weekmedian 1.000084.8%
Monthmedian 1.000044.2%
All timemedian 1.000035.4%

Measured against current account value instead, the median falls to 0.98 and only 2.9% to 44.5% of rows land within 2%.

The denominator is the value at the start of the window. And the fit degrades exactly as the window lengthens: 36.5% of all-time rows imply a denominator larger than their start value, which is what deposits sitting inside a denominator look like when that denominator never moves.

That last part is the whole problem.

One denominator, one window

The formula knows how much capital showed up. It cannot know when.

In the case above, the strategy earned $67,390 of genuine trading profit while its book was small. Then $5,815,749 of deposits arrived. The losses came afterwards, on the much larger book.

Divide the full-window profit by one static denominator and you get +0.47%, which is positive and therefore reads as a good fortnight. Weight each day by the capital that was actually exposed on that day, and chain those daily returns together, and the same fortnight is −5.07%.

Account value, meanwhile, finished the window up 67.94%, which is the figure most people eyeball first. The strategy later breached our drawdown limit and came out of the Index.

How often this happens

If this were one unusual wallet it would be an anecdote.

Of the 887 wallets in our current screening universe, 307 carry flows large enough that we refuse to read a return off the top line at all. That is 34.6% of the visible field.

A third of the wallets you can see on a leaderboard cannot have a return inferred from their headline numbers without knowing when their capital arrived. The rank is not lying to you. It is answering a narrower question than the one you are asking.

The second label: smart money

So people reach for the other tag.

In practice, smart money means one thing: this address was early to a token that worked.

Across 924 wallets we screened in July, the median wallet earned 52.0% of its realised PnL from a single token. 283 of them, 30.6%, earned more than 80% from one. And 106, 11.5%, earned 30% or more of everything they had ever made from a single trade.

That is not a track record. It is one observation, dressed as a credential.

There is a structural point underneath the statistical one. The tag attaches to an address. Not to a person, not to a process, not to a risk policy. Addresses change hands, and nothing about the label survives that transfer or even registers it.

The question neither label asks

Both labels share a defect that is easy to miss because it is a defect of omission.

Each one scores a single window and implies the result continues. Neither is ever re-tested. There is no second look, no fresh evidence, no mechanism by which a label expires when the thing it described stops being true.

We re-test, because it is the only part of this that is falsifiable.

We select 30 Hyperliquid strategies against a published ruleset, then re-apply that same ruleset to fresh evidence a fortnight later. At the most recent reselection, 13 of the 30 strategies still qualified. Seventeen that had qualified two weeks earlier no longer did.

If that is the turnover on a ruleset built specifically to look for repeatability, it is worth asking what the turnover is on a label that never checks at all.

How we measure instead

The method is deliberately unexciting, and it is the only reason the −5.07% exists.

We align Hyperliquid’s own account value history and PnL history by timestamp. The gap between them is external flow: whatever moved the account value that trading did not. Each interval return is that interval’s PnL divided by the capital genuinely at risk during it, with positive flow added to prior NAV before the capital base is sized. Then the interval returns are chained.

The result answers a different question from the leaderboard’s, and a more useful one for anyone deciding whether to put money next to a strategy: what would a dollar alongside this have done, given when it was actually exposed.

Limits

The flow series is inferred from Hyperliquid’s official account value and PnL histories rather than from a deposit ledger. We label it a research proxy, not audited flow accounting, and we would rather say so than let someone discover it.

The 924-wallet figures come from our July screening chain and describe that cohort, not the current universe. The 887-wallet figure is current generation, from 5 August.

The 13-of-30 result is selection turnover under our own rules re-applied to fresh evidence. It is not a formal persistence test, and we are not presenting it as one.

What comes next

We ranked those same 924 wallets two ways: by dollars made, and by return on the capital actually at risk.

The overlap between the two lists is smaller than almost anyone would guess. That one gets its own piece, because the objection it invites, that dollars and rates are simply different questions, deserves a proper answer rather than a footnote.

Index inclusion does not constitute an O.P.E.N. Profile, human-reviewed Strategy Passport assessment, investment recommendation, institutional approval or conclusion that a strategy is accessible or investable.

OSIQ 30 is a public research benchmark and is not currently investable.

Cite this research note.

OpenStrategy Research. “Onchain Has Two Ways to Tell You a Trader Is Good. Neither One Is Ever Re-Tested.” OpenStrategy, 7 August 2026. https://openstrategy.trade/research/measurement/two-labels-never-re-tested/