RESEARCH AND PRACTICE
Calibrated crypto probabilities: reliability, Brier score and base rates
Evaluate whether a crypto probability means what it says using a defined target, reliability groups, proper scoring rules and honest out-of-sample evidence.

A score of 70 is not automatically a 70% probability. A probability needs a specific event, a horizon and evidence connecting previously issued numbers with subsequently observed outcomes. Even a model that ranks opportunities well can overstate or understate their likelihood. This article explains a practical probability review using hypothetical examples; none of the percentages below describe measured HOSTuvo performance.
Define the event before discussing confidence
Specify whether success means a positive return at the horizon, touching an upper threshold first, or exceeding estimated costs. State the entry reference, observation window and treatment of ambiguous paths. These targets can produce different labels for the same price sequence. Keep invalid or incomplete outcomes out of the binary label calculation while reporting their frequency separately. A probability is conditional on the population to which it applies, so an estimate derived from liquid markets should not silently become a promise about a different universe or a newly added low-liquidity instrument.
Compare stated probabilities with observed frequencies
Group genuinely out-of-sample predictions into predeclared probability ranges and compare the mean prediction with the observed event frequency in each range. Show the count in every group. If a hypothetical set of 100 forecasts averages 70% and 70 events occur, that group is consistent with the stated rate, but the single result does not prove future reliability. Small groups can fluctuate widely, and overlapping forecasts reduce the amount of independent evidence. Examine whether calibration differs across horizons and conditions rather than hiding every pattern in one reassuring overall percentage.
Use a scoring rule without mistaking it for the whole answer
For a binary event, the Brier loss is the average squared difference between the forecast probability and the outcome coded as zero or one. Two hypothetical forecasts of 0.8, followed by one event and one non-event, have losses of 0.04 and 0.64, averaging 0.34. Forecasting 0.5 for those same two outcomes gives 0.25. This tiny example illustrates the penalty for confident errors, not a statistically meaningful model comparison. As the scikit-learn documentation explains, Brier loss reflects discrimination as well as calibration, so a lower score alone does not establish better reliability.
Choose a baseline that respects the event frequency
Suppose a separate hypothetical population contains 20 events among 100 observations. A constant probability of 0.2 produces an average Brier loss of 0.16, while a constant 0.5 produces 0.25. Always predicting the negative class would achieve 80% classification accuracy in that population, yet that headline says little about the usefulness of an event warning. Estimate any baseline frequency from information permitted at the evaluation time, not the full future test set. Compare identical observations and targets so a change in the event mix cannot masquerade as a modelling improvement.
Fit calibration without consuming the evidence used to judge it
A calibrator learns a mapping from a model's raw output to probabilities. Training it on observations the original predictor already fitted can make the mapping overly confident. For market data, use chronology and label maturity as well as separation between fitting and evaluation. Keep the calibration version attached to every emitted forecast. If too few outcomes support a fine-grained mapping, use fewer claims and disclose the limitation; adding a flexible curve does not create information. Check that the relevant positive and negative outcomes are represented in the calibration sample.
Turn the review into a restrained interface
Display a percentage only when its target and evidence support a probabilistic interpretation. Otherwise, label the number as a score or ranking and explain what it orders. Pair probability reporting with sample size, coverage, the evaluation period and uncertainty where estimable. Preserve the originally published value so later corrections do not rewrite the evidence. Calibration also does not establish economic attractiveness: a correctly estimated event can still be too small, too costly or too difficult to execute. Keep probability quality, candidate selection and executable economics as separate questions.
Sources and example scope
Sources support the definitions and mechanisms. Numerical scenarios are hypothetical teaching examples, not live prices, forecasts or reported HOSTuvo returns. Images are editorial illustrations.