Home / Steward Tips / Platforms & Tools
Platforms & Tools

The Twin Test: Why Every Strategy Has to Beat Its Own Random Twin

September 19, 2026 · 4 min read · Part of Platforms & Tools

Waterfall 101 · Step 0Also lesson 10 of 12 in Platforms & Tools, part of Waterfall 101 — free.Read it in the course →

Here is a question that is uncomfortable to ask about any strategy you love: what if the entry is doing nothing? What if the good results are coming from the stop, the target and the instrument, and the entry rule you are so proud of is just choosing a random moment to start? It is uncomfortable because, in our lab, the answer has been "yes, exactly that" far more often than not. The twin test is how we find out.

What the twin is

Take the strategy exactly as it is. Keep the instruments, the sizing, the stop, the target, the time-of-day rules, the costs. Then replace the entry with a coin flip, at the same frequency the real rule fires. That is the twin. It is not a straw man; it gets every advantage the real strategy gets except the one thing the strategy claims to be good at, which is choosing *when*. Run the twin many times, because each run is one random draw, and you get a distribution of what "no skill at all" looks like for this exact structure.

What the comparison tells you

Now put the real strategy against that distribution. If it sits in the middle of the twin's results, the entry is decoration. If it sits at the edge, the entry is doing something, and you can say how much. The twin turns "it made money in the backtest" into "it did better than chance would have, by this margin, over this many years", which is the only sentence about a strategy that is worth saying out loud.

Two details matter and both bit us. First, the twin has to be fair. A random walk drifts, and if the twin's synthetic prices are built carelessly the twin will lose money on its own for reasons that have nothing to do with skill, which flatters the real strategy. We had to correct our own twin for exactly this. Second, the twin has to see the same fills. If the real strategy is getting filled at a level the market gapped through, and the twin is not, the comparison is broken. Every "edge" we later withdrew had one of those two problems hiding in it.

Why win rate cannot substitute for this

A high win rate is the most seductive number in trading and it is nearly meaningless without the twin. A strategy with a tight target and a wide stop will win most of its trades whether the entry is skilled or random; that is arithmetic, not edge. The twin has the same target and stop, so it has the same win rate, and only the *difference* between the two tells you anything. We put this at the centre of everything we teach about tools because it is the difference between an edge you can explain and a story you tell yourself.

What it did to our own work

It killed most of it. Ideas we had run for weeks, ideas with beautiful curves, ideas that felt obviously right, went into the twin test and came out indistinguishable from chance. Conduit is one we have not yet been able to eliminate. Several refinements to it we eliminated immediately, and we wrote those up, which is the point. A test that never kills anything is not a test.

How to run your own

You do not need our harness. Take your strategy, keep everything but the entry, replace the entry with a random moment at roughly the same frequency, and run that fifty times on as much history as you can get. If your real results are not clearly better than the best of those fifty, that is worth knowing before any money is involved, and a smooth curve on its own is not evidence. We build these habits on a demonstration account first, where the lesson costs nothing.

Kingdom Portfolios LLC is an independent education publisher. It is not registered with the CFTC or the NFA, is not a Commodity Trading Advisor, does not offer or manage investments, and does not trade anyone else's capital. Nothing here is investment advice, a signal, an offer, or a solicitation.

No performance results of any kind are presented in this note and none should be inferred. Where this note refers to our lab, our harness or our measurement, it means simulated research on historical data. Hypothetical performance results have many inherent limitations: they are prepared with the benefit of hindsight, they do not involve financial risk, no hypothetical record can completely account for the impact of financial risk in actual trading, and no representation is made that any account will or is likely to achieve results similar to anything described. Past performance is not indicative of future results.

The algorithms described trade only Kingdom Portfolios' own accounts. They are not for sale, not licensed, and not run for anyone else. Trading forex and futures carries substantial risk of loss, including loss of the whole account. Education only.

Common Questions

How many years of data is enough?

More than the last good one. Our own rule is two decades where the data exists, because a strategy that only works in one regime will look excellent for a year and then fail, and a year is not long enough to tell those apart.

Does the twin test prove a strategy works?

No test proves that. The twin test disproves the ones that do not, which is most of them, and tells you how much margin the survivors have over chance. That margin is what you then watch decay in the forward record.

Keep Building

Start the Free Curriculum

Foundations of Stewardship Trading walks you from the basics to disciplined scaling, grade by grade, no hype, education only.

Enter the SchoolTry the Free Demo Challenge

Education only. This article is general financial education, not investment, legal, or tax advice and not a recommendation to buy, sell, or trade any asset. Kingdom Portfolios does not manage money, accept investor funds, or guarantee any result. Trading involves substantial risk of loss. Consult your own licensed professionals before making decisions.

Keep Reading