RESEARCH AND PRACTICE
Walk-forward crypto validation when forecast horizons overlap
Use chronological evaluation with mature labels, purged overlaps and separate calibration so repeated crypto forecasts do not leak future outcomes.

Sorting observations by their issue time is necessary for a realistic forecast test, but it may not be sufficient. A forecast can be issued before the test period while its outcome extends into that period. Repeated predictions can also describe nearly the same market move. A useful walk-forward design therefore tracks both the decision timeline and the outcome timeline, and explains exactly when information is allowed to influence the next model.
Write an information timeline for every observation
Retain at least the issue time, the latest availability time of the input data, the start of evaluation, the end of evaluation and the time the label was actually settled. A model fitted at a particular cutoff can use only labels known by that cutoff. The same condition applies to threshold selection, score polarity, calibration and any rule deciding whether a research model is eligible. An old issue timestamp does not make a later outcome historical knowledge. This distinction matters particularly when short- and long-horizon predictions coexist in one table.
Separate fitting, calibration and the final test
Use an earlier interval to fit the predictor, a later permitted interval to choose calibration or decision settings, and a still later interval for evaluation. Keep the final interval untouched by selection decisions. Repeat the process through time under a predeclared schedule, preserving the exact model and settings used for each forecast. The scikit-learn TimeSeriesSplit interface provides chronological splits and a gap measured in samples. It does not automatically understand variable-length market outcomes or convert that sample count into minutes; those requirements belong to the experiment's own data contract.
A hypothetical overlap across the split boundary
A forecast issued at 10:00 has a 90-minute outcome ending at 11:30. It cannot supply a known training label for a model deployed at 10:30, even though its issue time precedes the new test observation. A 240-minute forecast from the same issue time remains unresolved until 14:00 under a matching horizon convention. Remove or defer training records whose required outcome information was unavailable at fitting time. If evaluation begins at a later full-bar entry, use that actual end time. A rounded bucket identifier is not a substitute for the label's information interval.
Purge overlap without pretending samples are independent
Near a split boundary, inspect the intervals used to construct labels and remove prohibited intersections according to the written protocol. Where an embargo is used, define its direction and purpose; a gap is not a universal cure for every leakage path. Within the test set, several forecasts from one evolving move may still be legitimate live observations but strongly dependent. Report both the number of forecasts and a meaningful grouping such as days or market events. Resampling uncertainty by independent-looking rows can otherwise make a heavily overlapping stream appear more certain than it is.
Audit transformations and model selection too
Feature scaling, missing-value treatment and feature selection must be fitted only from the allowed history for each split. A transformation learned from the entire dataset can leak information even when the prediction model itself is trained chronologically. Keep a record of the alternatives tried and the rule used to select among them. Bailey and colleagues' work on backtest overfitting addresses the danger of selecting from many historical experiments; a clean final timeline does not erase repeated tuning on the same supposed holdout. Reserve evidence that can still challenge the preferred configuration.
Replay the sequence that a user would have seen
For each test issue, load the model state that existed then, calculate the forecast from available inputs, and freeze the emitted result. Settle it only after the required future data become complete. Review performance by horizon and market condition alongside missing-data coverage and a simple benchmark. If later calibration changes the interpretation of an earlier score, store that as a retrospective analysis rather than overwriting the original output. The central question is whether the procedure works as a chronological decision process, not whether a hindsight transformation can make its history look attractive.
Sources and example scope
Sources support the definitions and mechanisms. Numerical scenarios are hypothetical teaching examples, not live prices, forecasts or reported HOSTuvo returns. Images are editorial illustrations.