KUNKAFA FREE FORECAST TOOLS

Probability Calibration Calculator by Kunkafa

Compare forecast probabilities with recorded outcomes. Calculate a binary Brier score and inspect a reliability diagram using your own data. Everything runs locally in your browser.

Enter forecast and outcome pairs

Use decimal probabilities from 0 to 1 (0.7 means 70%). Outcome 1 means the specified event happened; 0 means it did not. One pair per line, up to 10,000 rows. The header is optional.

The starting rows are invented for illustration. They are not Kunkafa performance data.

Forecast reliability results

Forecast pairs
10
Binary Brier score
0.1450
Reliability diagramEach purple point is one nonempty probability bucket. The diagonal represents equal predicted probability and observed frequency. The table gives exact bucket values and counts.0%100%0%100%Mean predicted probabilityObserved event frequency
Buckets with fewer than 20 rows are marked “small sample” as a practical warning, not a statistical significance test.
Five fixed buckets: lower bound included; upper bound excluded, except 100% is included in the final bucket.
Probability rangeCountMean forecastObserved rate
0–20%1Small sample10.0%0.0%
20–40%2Small sample25.0%50.0%
40–60%2Small sample45.0%50.0%
60–80%2Small sample65.0%50.0%
80–100%3Small sample90.0%100.0%

Lower Brier scores mean lower average squared probability error on this dataset. A low score alone does not prove calibration, profitable trades or out-of-sample skill.

What does probability calibration mean?

If forecasts assigned about a 70% chance to an event, that event should occur about 70% of the time across many comparable forecasts. It does not mean every individual 70% forecast will succeed.

Each point groups forecasts into a fixed probability range and compares their mean probability with their observed event frequency. Points near the diagonal suggest agreement within those buckets. Empty buckets are omitted from the plot; small buckets can move substantially with one extra outcome.

How the binary Brier score is calculated

For each row, square the difference between the probability and the binary outcome, then average those squared differences. The binary score here ranges from 0 (perfect probabilities for these outcomes) to 1 (confidently wrong for every outcome). A 50% probability contributes 0.25 regardless of whether the event occurs.

Brier score measures more than calibration. Event frequency, discrimination and the dataset affect it. Compare models on the same events and dates, against an appropriate baseline, rather than comparing unrelated scores.

Use comparable, completed forecasts

Choose one precise event definition and duration. Record the probability before the outcome, include failures and wait until each outcome can be resolved. Overlapping market forecasts may not be independent. This tool does not validate your labels, remove selection bias or calculate confidence intervals.

Kunkafa publishes market forecasts as predicted changes, direction and probabilities. Use this calculator to understand probability evidence; use Kunkafa’s published forecasting results for the site’s own evidence and definitions.

Kunkafa forecasting methodology · Kunkafa risk/reward calculator