Resolution accuracy
Every settled market gets scored against the price the market was charging an hour before resolution. We publish the receipts: how often the favorite won, how confident the markets were, and which platforms call it best.
Events scored
338,222
Settled markets with a usable pre-settlement snapshot.
Favorite hit rate
59.4%
How often the pre-settlement favorite was the actual winner.
Brier score
0.2089
Mean squared probability error. Lower is more accurate. 0.25 is a coin flip.
16.4% better than random guess
Calibration
Markets that priced an outcome inside each 10pp band — how often did it actually happen? In-line means the observed rate landed inside the bucket (well calibrated).
Accuracy over time
Daily favorite hit rate and Brier score across the whole settled corpus, tracked snapshot over snapshot.
Premium: accuracy over time
Members see today’s numbers. Premium unlocks the full daily trend line — 77 day(s) of resolution history and counting.
By platform
How each platform’s pre-settlement pricing held up against the actual results, counted only on markets that at least two platforms priced. Scored against a whole book this would measure which categories a venue lists rather than how well it prices them — crypto resolves close to a coin flip and dominates some books entirely.
| Platform | Contested markets | Favorite hit rate | Brier score |
|---|---|---|---|
| Polymarket | 37,770 | 72.9% | 0.1610 |
| Polymarket US | 23,809 | 76.5% | 0.1539 |
| Kalshi | 18,268 | 73.0% | 0.1557 |
| Azuro | 16,836 | 70.9% | 0.1680 |
| Limitless | 3,069 | 71.5% | 0.1848 |
| Gemini | 1,572 | 82.6% | 0.1017 |
By category
Some markets are easier to call than others. Categories with more settled markets carry more weight. Click a category for the calibration breakdown.
| Category | Events scored | Favorite hit rate | Brier score |
|---|---|---|---|
| Crypto | 207,705 | 50.2% | 0.2478 |
| Sports | 111,034 | 71.9% | 0.1639 |
| Tennis | 58,147 | 70.7% | 0.1844 |
| Soccer | 19,234 | 74.0% | 0.1282 |
| Tech & Science | 10,318 | 90.7% | 0.0175 |
| Esports | 8,576 | 81.4% | 0.1335 |
| Baseball | 4,560 | 68.8% | 0.1375 |
| Markets | 3,404 | 82.9% | 0.0492 |
| Basketball | 1,915 | 71.4% | 0.1241 |
| American Football | 1,722 | 78.2% | 0.1013 |
| Cricket | 1,392 | 90.3% | 0.0675 |
| Politics | 1,363 | 90.8% | 0.0374 |
| World | 1,265 | 86.3% | 0.0621 |
| Entertainment | 1,129 | 78.1% | 0.0713 |
| Golf | 1,069 | 53.9% | 0.1492 |
| MMA | 569 | 68.4% | 0.1888 |
| Hockey | 427 | 68.8% | 0.1517 |
How we score this
- Corpus: Settled markets with both a settlement timestamp and a winning outcome. Voided and canceled markets are excluded.
- Snapshot: For each market we read the price an hour before settlement. The one-hour buffer keeps the settlement spike — where the eventual winner pre-prints to near 1.0 in the final minutes — out of the scoring. This instant is relative to settlement, not to kickoff: for a sports fixture it usually lands after the game has been played, while the venue is still waiting to resolve. Read these figures as a record of how markets priced outcomes, not as a forecast you could have placed a bet on.
- Favorite: The outcome with the highest cross-platform consensus price at snapshot time. Ties broken by stable canonical id. The "favorite hit rate" is the share of markets where the favorite was the actual winner.
- Brier score: Per market, the mean squared distance between the snapshot probability and the realized result (1 for the winner, 0 for losers). Overall Brier is the mean across markets. 0.0 is perfect; 0.25 is a coin flip on a binary market. Each market’s prices are normalized to sum to 1 before scoring: venues quote different things — an executable ask on some, a book midpoint on others — so scoring the raw quote would grade a venue on its pricing convention rather than its judgement. Markets where an outcome went unquoted carry no Brier, because a partial book cannot be normalized.
- Per-platform: Each platform is scored against its own pricing, normalized the same way, and only on markets where it had an opinion on the eventual winning outcome — so neither coverage gaps nor a venue’s quoting convention penalize accuracy. The comparison counts only markets at least two platforms priced: measured across a whole book it would rank which categories a venue lists rather than how well it prices them, since some question types resolve close to a coin flip and make up most of some books. Platforms with few contested markets are published but marked provisional rather than hidden.
- Calibration: Each canonical outcome’s normalized probability is binned into one of ten 10pp-wide buckets. For each bucket we publish the observed hit rate — the share of outcomes in that bucket that actually won — measured against the average probability the bucket predicted rather than the bucket’s midpoint, since predictions do not sit at the middle of their bin. A well-calibrated market lands on what it predicted; persistent skew up or down points at systematic over- or under-confidence.