LemmaRequest early access

Methodology

This page is written for the reader who assumes any given backtest is wrong until shown otherwise. It describes the process Lemma uses to generate and validate trade hypotheses, the ways that process can fail, and the limits of backtesting itself. There is no marketing framing here. This is the same standard we hold our own research to.

Failure mode taxonomy

Every hypothesis is evaluated against a fixed list of known ways a backtest can mislead. This list is generic by design. It describes the categories of failure, not the specific parameters or thresholds used internally.

Overfitting to sample

Parameters tuned until a backtest looks good will fit the noise in that specific sample. Every hypothesis here is tested on a fixed recent window with no out of sample holdout, so this risk is disclosed explicitly in every dossier's validation transparency section rather than presented as solved.

Regime dependence

An edge that only exists under a specific volatility or liquidity regime will look robust across a sample dominated by that regime and fail the moment conditions shift. Every dossier states the exact data window tested and flags limited regime coverage directly, rather than claiming coverage it doesn't have.

Multiple testing / data snooping

Testing enough variants of an idea will eventually produce one that looks statistically significant by chance. Every hypothesis carries the cost of the search that produced it, not just the cost of the one test that happened to work.

Cost model blindness

Fees, slippage, and funding are easy to underestimate, especially at size or in thinner markets. Cost assumptions are made explicit and stress tested rather than left as a footnote.

Redundancy

A 'new' signal that is highly correlated with a known factor is not adding anything. It's a repackaged exposure with extra steps. Candidates are checked against a library of known signals before they're taken seriously.

Concentration

A strong headline return driven by a handful of trades is fragile, not robust. Performance attribution is checked at the trade level, not just the portfolio level.

Look ahead bias and data leakage

Information that would not have been available at decision time (restated data, survivorship filtered universes, mistimed timestamps) can quietly leak into a backtest and inflate results. Data pipelines are built to prevent this rather than to catch it after the fact.

Why preregistration and stopping rules matter

Acceptance criteria, kill criteria, and evaluation windows are fixed before a hypothesis is tested, not after the results come in. Without this, it is trivial to rationalize a marginal result (extend the sample a little, drop an inconvenient outlier, adjust a threshold) after seeing whether it helps. Preregistration removes that degree of freedom. If a hypothesis needs its stopping rule changed after the fact to survive, it fails.

Why Lemma publishes rejections

A track record built only from ideas that worked is a survivorship biased sample of one firm's own output. The rejection rate is part of the evidence: a process that rejects most of what it generates is a process that is actually applying a standard, rather than approving whatever it produces. The Risk Officer's role exists specifically to keep that rejection rate honest. It reports outside the research chain and is evaluated on catching bad ideas, not on approving them.

The honest limits of backtesting

No amount of validation converts a backtest into a guarantee. The limitations below are structural, not something a better process removes:

  • Regime dependence. Markets are not stationary. A relationship that held for the backtest period can weaken or invert as market structure, participants, or macro conditions change.
  • Multiple testing across the field. Even with correction applied internally, any sufficiently large population of researchers testing ideas will produce some that look good by chance alone.
  • Cost model uncertainty. Modeled fees, slippage, and funding are estimates. Real execution costs shift with venue liquidity, position size, and market conditions in ways a backtest can only approximate.
  • Capacity and market impact. A strategy's historical performance does not account for the price impact of capital actually deployed against it at scale.

Hypothetical, backtested performance has inherent limitations and does not guarantee future results. Every figure Lemma presents (in a dossier, on this site, or anywhere else) should be read with that in mind.

Request early access

Lemma is in early access. We are accepting a limited number of professional users while we refine the platform. All features are free in exchange for your feedback.

Asset classes of interest

Early access is granted in the order requests are received.