RESEARCH AND PRACTICE
Point-in-time crypto backtests: avoiding survivor and availability bias
Reconstruct the market universe and information actually available at each decision, including removed tokens, listing changes and delayed data.

Historical prices are only one part of a defensible backtest. The test also needs to know which instruments could be considered at the time and when each input became available. Starting with today's surviving markets and filling in their past can answer a much easier question than the one a live scanner faced. A point-in-time dataset preserves that changing decision environment instead of quietly replacing it with the benefit of hindsight.
Store universe membership as a dated fact
Keep snapshots or a dated history of eligible instruments, including venue, product type, quote asset and trading status. Current instrument APIs are useful for current constraints, but their present response is not automatically an archive of previous listings and removals. Record additions, suspensions, symbol changes and the provenance of each historical status. Do not identify an asset by ticker alone when migrations or reused names can create ambiguity. The backtest should be able to explain why an instrument was eligible on a particular date and why another was absent.
A hypothetical survivor-only result
Imagine that 20 markets were eligible at the beginning of a study. Ten remain in the current catalogue and average a 20% return; ten later disappear from that catalogue and average a 60% loss over the specified holding interval. An equally weighted calculation over only the current survivors reports a 20% gain. The same simple calculation over all 20 starting markets gives a 20% loss. The numbers are hypothetical, and an actual investable portfolio requires additional execution assumptions. Their purpose is to show that changing membership can reverse an apparent result before any model is involved.
Separate event dates from availability dates
An economic observation, token-supply figure or exchange status can describe an earlier period while being published later. A feature may be usable only after publication, delivery and processing, not at the date printed as its reference period. Store both the economic timestamp and the earliest supported availability timestamp. Revised data should be versioned when they can change past feature values. If the original release cannot be recovered, disclose the limitation or restrict the experiment. Moving the revised value backward into history changes what the simulated decision-maker knew.
Preserve inconvenient outcomes and missing intervals
A removed market may have incomplete terminal prices or restricted withdrawals, making its economic outcome difficult to estimate. Do not silently delete it from the sample or replace the missing final value with a convenient zero return. Specify the disposition, show sensitivity to defensible alternatives and separate unknown outcomes from measured losses. Similarly, record data outages during volatile periods. If those observations are systematically excluded, the remaining sample may describe only the easiest trading conditions. Coverage by date and instrument should accompany any performance table.
Audit the joins before tuning a model
Choose a small set of historical decision times and reconstruct the eligible universe, market constraints and input versions by hand. Test that no join selects a record published after the decision simply because its reference date is earlier. Include a listing, a suspension, a renamed symbol and a delayed observation among the test cases. Save the dataset version with the experiment. Research on backtest overfitting explains another risk—selecting a strategy after trying many alternatives—but solving that problem does not repair a universe or feature table that already contains future information.
Publish an evidence ledger with the result
Document the source of historical membership, the coverage of removed assets, the rule for delayed data and every unresolved gap. Compare the full point-in-time sample with the survivor-only sample as a diagnostic, not as competing results from which to choose the more attractive one. Keep parameter selection, calibration and final evaluation periods separate. The most useful conclusion may be that the available history cannot support the original claim. A smaller, reproducible experiment with explicit limits is more credible than a long equity curve whose decision universe cannot be reconstructed.
Sources and example scope
Sources support the definitions and mechanisms. Numerical scenarios are hypothetical teaching examples, not live prices, forecasts or reported HOSTuvo returns. Images are editorial illustrations.