AI CALIBRATION DIAGNOSTIC
How accurate is the LLM's confidence? Predicted probability vs actual outcome.
Overall Brier Score
—
0 = oracle · 0.25 = coin flip · 1 = worst
Mean Predicted vs Mean Actual
—
—
RELIABILITY DIAGRAM
Diagonal = perfectly calibrated. Below = overconfident. Above = underconfident. Bubble size = sample count.
PER-BUCKET BREAKDOWN (predicted-probability-of-winning)
Predicted bucket N Mean predicted Actual WR Gap (pp) Net PnL
BY CONFIDENCE LEVEL
Confidence N Mean pred Actual WR Gap PnL
BY ENTRY-PRICE BUCKET
Entry bucket N Mean pred Actual WR Gap PnL