ZEME Lab: Why Most Signals That Look Good Are Not
A demonstration of the evaluation method behind signal-based systems, run on synthetic data generated in your browser. No market data, no track record, and no advice.
A demonstration of the evaluation method we use on signal-based systems, run on synthetic data generated in your browser. No market data, no ZeMe output and no track record is shown here. This is educational and is not investment advice or a recommendation of any kind.
Why most signals that look good are not
The single most expensive mistake in this field is testing a rule on the same data used to find it. Below, a signal is discovered on synthetic history and then tested two ways: in-sample, and walk-forward on data it never saw. Run it a few times.
This is the variable nobody reports. The more rules you try, the better your best one looks, even when none of them work.
ZeMe is a production AI system BvLogic built and operates, in a domain that punishes wishful measurement faster than any other.
This page shows none of its results, and that is deliberate.
Why there is no chart that goes up#
Every vendor in this field has one. A chart that goes up proves nothing, because you cannot tell from looking at it how many variants were tried before that one was kept, or whether the test data was ever genuinely held back.
What is worth evaluating in a supplier is whether they know why their own backtest might be lying to them. So the demo shows that instead, on data generated in front of you.
What the demo does#
It generates a random series with no signal in it whatsoever. It then searches for the best-performing rule on the first half, and tests that rule on the second half, which the rule has never seen.
The in-sample result looks good. The walk-forward result does not, because there was never anything to find.
Run the 200-iteration version. Mean in-sample return comes back strongly positive and mean walk-forward at approximately zero, on pure noise. Around half of the discovered rules are positive out of sample by coin flip.
The three lessons that follow#
The number of variants you tried is part of the result. Searching twenty rules and keeping the best does not find a better edge in noise, it finds a better-looking accident. This is the figure nobody reports, and without it an in-sample return is uninterpretable: it measures how hard you searched.
A single walk-forward pass is evidence, not proof. Roughly half the rules above survive the holdout by chance. One passing test on real data means considerably less than people assume, which is why serious evaluation uses several folds and reports the worst one.
In-sample results routinely overstate by a factor. If you are shown a backtest with no walk-forward, you are looking at the artifact this demo produces. The correct response is to ask for the holdout rather than to discount the number by instinct.
Why this generalises well beyond markets#
Anything that makes a prediction and then acts on it has this problem. Churn models, fraud scoring, demand forecasting, lead scoring and predictive maintenance can all be fitted to history, and all of them look excellent until they meet data nobody has seen.
The discipline is identical. Hold data out properly, count the variants, treat one passing test as weak evidence, and report the fold that performed worst rather than the one that performed best.
What is not on this page#
No market data, no ZeMe output, no track record, no performance claim, and nothing that constitutes investment advice or a recommendation. Synthetic data, generated in your browser, demonstrating a method.
If you want to see the product itself it is at ZeMe. If you are building something that predicts and acts, the method above is the conversation worth having.
What else is coming for ZEME Lab
Experiment Ready
What we tried, and what it showed.
Diagram Not yet
How it is put together.
Worked Example Not yet
A run, in full.
FAQ Not yet
What people ask about this one.