Back to Research
A probability curve with a highlighted tail

The short version

Prediction markets are supposed to make a surprise expensive to hold and cheap to see coming. On Polymarket, most of the time, they do. This is a look at the 540 times they didn't, and at what a market built to price a whole distribution would and wouldn't have changed about it.

247,367Markets screened
540Resolved against a 95%+ price
$120MWrong-side exposure · cumulative volume, an upper bound
164In the bracket cohort
  • The sample is every resolved Polymarket market that closed between January 2024 and April 2026 carrying at least $10,000 in event volume. Of those, 540 traded past 95% confidence and then resolved the other way.
  • Twelve geopolitical markets sit outside the core sample. The Khamenei succession market carries $129.6 million by itself, more than the other 540 combined, and the twelve together carry $195.5 million against their $120.4 million. Left in, every category and cohort comparison in the piece becomes a report on one cluster of events.
  • Bracket markets are heavily over-represented. They cut a number into mutually exclusive ranges and trade each range as its own contract, and they are 8.9% of the sample, 20% of the markets that reach 95% confidence, and 30% of the breaks, worth $36.0 million.
  • Brackets resolve against a confident price at 1.4 to 1.8 times the rate of other high-confidence markets. The spread is how settled-price entries are counted in the denominator; 1.4 is the least favourable way we counted them.
  • We sort the 40 largest breaks into four kinds to make them tractable. Three describe something about the world. The fourth, structural, describes the contract, and it is the one that shows what a different market design would and would not change.

What counts as a bracket market

A bracket market takes a question that is really about a number and cuts that number into mutually exclusive ranges, each range becoming its own yes-or-no contract, with exactly one resolving YES. "What price will Bitcoin close at in January?" becomes "between $90,000 and $95,000," then "between $95,000 and $100,000," and so on. "Republican Senate seats after the 2026 midterms?" is the same shape, eleven contracts for one number. On Polymarket these trade as negRisk multi-outcome events.

Two structures look similar and are excluded, because neither forms a partition.

Three shapes that all look like a range market

Only the first partitions the outcome. The other two are excluded from the bracket cohort.

A bracket partitions a number into non-overlapping ranges with exactly one paying. A threshold ladder nests, so several legs pay together. A touch market is path-dependent, so every strike the price crosses pays. BRACKET - ranges partition the number, one pays PAYS $1 $90k $95k $100k $105k $110k $115k THRESHOLD LADDER - legs nest, several pay together above $180 above $190 above $200 outcome $205 all three pay TOUCH - path-dependent, every strike crossed pays $90k ✓ $85k ✓ $80k
Figure 1. A bracket forces one choice among exclusive ranges. A ladder and a touch market both let several legs resolve YES, so neither creates the forced-bucket problem.
  • Threshold ladders. "Solana above $190 on August 13" sits beside "above $180" and "above $200." The legs are nested. Above $190 implies above $180, and several resolve YES together.
  • Barrier and touch markets. "Will Bitcoin dip to $80,000 in January?" is path-dependent. Every strike the price touches resolves YES.

Brackets are the negRisk events whose legs parse as non-overlapping numeric intervals. That is 22,077 markets across 2,656 events, 8.9% of the sample, of which 164 resolved against a confident price carrying $36.0 million.

The underlying outcome is never naturally binary. It is a single value that could land anywhere along a range, and the structure makes a trader pick which bucket catches it.


How often confident prices are wrong

Prices near certainty are usually right, and slightly more often right than the price implies, the favorite-longshot bias found across betting markets generally. Bürgi, Deng and Whelan document it directly in Kalshi contracts. Alex McCullough's Polymarket dashboard points the other way, toward mild overestimation of event probabilities. The two disagree on the direction of the bias and both leave misses above 95% rare, which is what makes the ones that occur worth reading one at a time.

functionSPACE's work on Polymarket order flow points at something that would push in the same direction. Across 28,793 on-chain trades, demand tracked how cheap a token was rather than which side of the question it sat on. Small trades clustered where tokens were cheap, and within a price bucket every trade size behaved alike. Apply that to a near-certain market and the dismissed leg is the cheap one. A contract the information would put at a cent attracts buying for being a cent, and sits at four or five instead. Some of what we count here as a market at 95% was a market at 99% with a bid under its tail. That inflates the number of markets crossing our line, and it puts traders into the exposure figure who were buying cheapness rather than expressing a view. We have not measured how large the effect is on this sample.

Scoring rules ask a different question from the one here. Clinton and Huang applied log-loss and Brier metrics to election markets on Polymarket, Kalshi and PredictIt, both of which penalise a confident miss far harder than a close one. Nothing here revises that work. A Brier score describes how a platform performed across a population of markets. It does not say what happened inside any one of them, and that is the subject here.


Four kinds of break

A price at 95% is wrong one time in twenty by construction, and a good share of the 540 are exactly that. What we want out of the list is the subset where something else was going on, so the sort has to separate a market that met its own odds from one that was mispriced for a reason.

We ask one question of each break.

What would have had to be different for the confident price to be right?

The answers land in four places: who knew what, when the fact came into existence, how the question was settled, and what the contract was able to express. This is our model for reasoning about why a break happened, not a taxonomy anyone else uses.

  • Information-asymmetric. One side held information the price did not reflect.
  • Exogenous shock. New information surfaced after the price had formed.
  • Resolution dispute. The oracle went against the public evidence.
  • Structural. The price was right about the world and wrong about the contract, because the contract could not express what the trader believed.

The first three describe the world. The fourth describes the contract, and it is the only one that can be identified from structure without reading the case.

Kind Of the 40 In core sample Core exposure
Structural2424$45.6M
Information-asymmetric76$8.8M
Exogenous shock72$2.6M
Resolution dispute22$10.3M
Total4034$67.3M

Five of the seven exogenous shocks are the held-out geopolitical markets, carrying $191.5 million between them. Left in, one kind of event arriving from outside any price would account for most of the money in the study, which is why they are held out.

Structural cases run wider than the bracket cohort. Fourteen of the 24 are brackets. The other ten are threshold ladders, touch markets and subscription-style thresholds on a continuous quantity. They have the same problem without forming a partition. A trader holds a view about a number, and the contract only accepts a view about a line. The bracket cohort is the part that can be identified automatically rather than by reading.

The four kinds describe those 40 markets. They are not a partition of all 540.


What the breaks look like

Against the 63,453 markets that reached 95% confidence, the 540 core breaks are an aggregate rate of 0.85%.

The rate has been falling. Breaks per 1,000 markets resolved ran between 5 and 8 through mid-2024 and sit near 1.1 by early 2026, while the number of markets Polymarket resolves each month grew from a few hundred to tens of thousands.

Breaks against a growing denominator

Monthly, March 2024 to April 2026, by the month a market actually closed.

Breaks per 1,000 markets resolved fell from around 8 in early 2024 to around 1.1 by early 2026, while monthly resolved markets grew from a few hundred to over 40,000. BREAKS PER 1,000 MARKETS RESOLVED THAT MONTH 0 3 6 9 MARKETS RESOLVED THAT MONTH 0 20k 40k 2024-03 2024-06 2024-09 2024-12 2025-03 2025-06 2025-09 2025-12 2026-03
Figure 2. The denominator is every market resolved that month, not every market that reached 95%, which the free API does not let us count monthly.

The fall sits almost entirely outside the bracket cohort. Non-bracket markets went from 2.98 breaks per 1,000 resolved in the first half of the window to 1.52 in the second. Brackets went from 9.47 to 7.12, and month to month they show no downward trend at all. One caveat holds the reading open: the denominator is every market resolved, not every market that reached 95%, so some part of the fall is Polymarket listing more markets that never approach a confident price.

A market that reaches 95% confidence is not one kind of object. A two-sided binary gets there because traders formed a strong view. A leg of an eleven-way partition gets there because ten of its eleven legs cannot win, and the structure supplies that near-certainty before anyone trades. Bracket legs reach 95% at 58.6% against 22.9% for everything else, 2.6 times as often. Brackets manufacture confident prices in volume, and a price that was cheap to reach is cheap to be wrong about.

Median confidence right before the break was 97.3%. 316 of the 540 sit inside negRisk multi-outcome events where only one branch resolves YES, and the rest are standalone binaries. The 540 came out of 495 separate events, and in 8% of those, more than one market broke at once.

85% of breaks were confident NOs that resolved YES. Most of that skew is mechanical. In a negRisk event with N legs one resolves YES and N−1 resolve NO, so the sample holds far more legs priced under five cents than over 95, and 59% of the core sample is negRisk. It is a property of how the sample is built before it is anything about behaviour.

Where the exposure sits

Core sample, 540 markets. Exposure is cumulative volume, an upper bound.

Wrong-side exposure by category. Crypto $46.6M across 157 markets, Other $25.4M across 206, Tech $23.5M across 30, Politics $18.6M across 74, Entertainment $3.1M across 13, Sports $2.5M across 52, Geopolitics $0.4M across 2, Macro $0.4M across 6. $0 $10M $20M $30M $40M $50M Crypto Other Tech Politics Entertainment Sports Geopolitics Macro $46.6M · 157 $25.4M · 206 $23.5M · 30 $18.6M · 74 $3.1M · 13 $2.5M · 52 $0.4M · 2 $0.4M · 6
Figure 3. Wrong-side exposure by category, with the number of broken markets after each bar. Geopolitics is small here because the twelve largest geopolitical breaks are held out.

Tech is almost entirely the weekly Musk tweet-count brackets, which is why 30 markets carry more than Politics carries across 74.


Bracket compression

Somebody decides where the lines go before trading opens. Where the ranges start, how wide they are, how many there are. A trader who thinks Bitcoin closes around $114,000 holds a view about a spread of prices, and the contract asks which single $500 band it lands in. Covering three bands means three trades in three thin order books, so most traders pick one.

A functionSPACE census of 18,863 Polymarket multi-market events found the top five contracts take around 90% of event volume however many exist, and that continuous events spread somewhat more evenly than categorical ones as the count grows. Brackets are the better-behaved half of that finding rather than an exception to it. Neither that census nor this dataset can see where positions actually sat across strikes at the moment a market broke, which needs per-strike depth.

158 of the 164 bracket breaks, 96%, landed in the exact bucket the market had dismissed. 26 of those never traded above five cents.

The weekly Musk tweet-count market makes the point without any volatility story attached to it. It runs on an integer count of posts with a fully public history, and it still breaks at a median confidence of 98.2%. Those 24 markets carry $23.0 million of the cohort's $36.0 million. The rest is Ethereum price bands, $GME market-cap ranges, Musk net-worth ranges, box-office spreads and daily temperature bands for London, Miami and New York.

Brackets resolve against a confident price about half again as often as everything else. Some markets in the sample never traded at all across the measurement window because the outcome was already decided, so their price records a result rather than a forecast, and they sit mostly in the comparison group. Counting those four different ways moves the ratio, and we do not have the trade timestamps to settle which count is right.

How much more often brackets break

Bracket break rate divided by the rate for everything else.

Across four ways of counting settled prices the ratio runs 1.41 to 1.83 with a primary estimate of 1.49. The event-cluster bootstrap gives 1.51, 95% interval 1.24 to 1.82. Counting each partition once rather than per leg gives 2.34. All sit above 1. 1.0 1.5 2.0 2.5 1.0, no difference FOUR WAYS OF COUNTING SETTLED PRICES 1.41 to 1.83 EVENT-CLUSTER BOOTSTRAP, 95% INTERVAL 1.24 to 1.82 COUNTING EACH PARTITION ONCE, NOT PER LEG 2.34
Figure 4. Filled dots are point estimates. The bottom row counts an eleven-leg partition as one observation instead of eleven.

Bracket width, count and placement are set at creation, under the same uncertainty the traders face and with less information than the traders will have by the time it matters. A $500 band centred on consensus is a contract that pays nothing for being approximately right. That cause sits in the contract rather than in the world, which is what separates it from the other three kinds.


The other three kinds

Information-asymmetric

The 2025 Nobel Peace Prize market was 99.3% confident María Corina Machado would not win, on $2.2 million. She won. The committee seals its deliberations for fifty years, so the decisive information was never available to trade on. The 2025 papal conclave has the same shape and broke three markets at once. "Will Robert Francis Prevost be the next pope?" sat at 1.5 cents on $1.4 million of volume, alongside markets on the next pope's continent and country. A conclave is sealed by design.

Exogenous shock

"Megaquake before August?" sat at 99.4% against on $1.4 million and resolved YES when an M8.8 earthquake struck off Kamchatka on 29 July 2025, the largest anywhere since 2011. Nothing preceded it that a price could have reflected.

Resolution dispute

The Trump-Ukraine minerals market traded at 98.2% NO with $6.7 million behind it and resolved YES. Public reporting before resolution showed no deal signed. It resolved YES through UMA's token-weighted dispute process. XO Labs has documented this case as the canonical "whale tax" failure and we defer to their account. It belongs to the oracle layer rather than the contract.

Most of crypto's losses sit outside the bracket cohort. Threshold ladders and touch markets account for $39.5 million across 99 breaks, against $7.1 million across 58 crypto brackets. Volatility explains the ladders.

What separates the four kinds is what a trader could have known beforehand.

KindSignal available before the break
StructuralFully public. Price feeds, options-implied densities, historical tweet counts.
Information-asymmetricPartial. The decisive fact was sealed.
Exogenous shockNone. Nothing preceded the event.
Resolution disputeStrong, pointed the other way, and was overridden.

Only the first row describes a market where the information was there and the contract could not hold it.


What one market across the range would change

Eleven bands means eleven order books. Liquidity splits eleven ways before a trade happens and stays split, because nothing moves depth from the band nobody wants to the band that ends up paying. Cutting the same range into twenty $250 bands rather than eleven $500 bands halves the depth behind each one, so precision and liquidity trade against each other at creation time.

The alternative is one market covering the whole range, with a single pool of collateral behind it and every bin priced off the same cost function. The idea is not new. Goldman Sachs and Deutsche Bank launched Economic Derivatives in 2002, listing a ladder of digital ranges over payrolls, CPI and initial claims, for the reason anyone would give now: buying one range is equivalent to shorting all the others, so the market clears without an opposing buyer at that strike. Gurkaynak and Wolfers found the resulting densities well calibrated and slightly better than the consensus survey. It cleared parimutuel in scheduled auctions, so a trader did not know the price until the auction closed, and it was dealer-intermediated and institutional.

functionSPACE puts the same idea under continuous trading, and the difference from the 2002 version matters. Every bin is quoted continuously from the cost function, and each share pays $1 on the realised bin, fixed at the moment you trade. Nothing is pooled and divided among winners at settlement.

In that design, bin count stops costing depth. The market maker's worst-case loss does not grow with the number of bins, so cutting a range finer does not thin the market. That is the term forcing a bracket creator to compromise between precision and having a book at all.

A trader also buys a shape in one transaction. "Somewhere around $114,000, probably a little under" is expressible directly at one cost, rather than approximated by three trades across three spreads.

And being approximately right starts paying approximately. In a bracket, the trader who put everything on the $113,500 to $114,000 band and watched it close at $114,100 collects nothing, and the payoff treats being $600 out the same as being $30,000 out. Holding shares across neighbouring bins pays on whichever one lands. Against the 158 of 164 that landed in the bucket the market had dismissed, most of those markets were not badly wrong about the number. They were one band out, against a payoff with a cliff in it.

The 26 bracket markets that never traded above five cents are the sharper version. A bin with no book has no price, so nothing in the market said anything about that outcome until it happened. A cost function quotes every bin whether or not anyone has traded it.

Three limits. A histogram still has bins, so discretisation does not disappear and a trader can still be wrong about the shape rather than the band. A quoted tail is not a correct tail, and a price that exists because the cost function exists is not the same as a price someone put money behind. And whether this design would have turned these 164 breaks into smaller losses rather than different ones needs the counterfactual run against these markets.

KindWhat the design reaches
StructuralThe cause directly. Depth stops being split across bins and the payoff stops having a cliff.
Information-asymmetricThe cost, not the cause. No design unseals a conclave. A trader holding a distribution loses less when the unlikely outcome lands than one who sold it at two cents.
Exogenous shockNothing about the event. Where the outcome is a number, a trader holding shares across the range has some exposure wherever the shock lands rather than none. Where the question is binary, as the megaquake was, nothing.
Resolution disputeNothing about the oracle. Where the outcome is a number, a contested settlement moves a spread of positions partially instead of wiping a single bucket. Where the question is binary, as the minerals deal was, it stays all or nothing.

The split that matters as much as the four kinds is whether the underlying outcome is a number. Strikes get cut crudely for a good reason, because a hundred liquid strikes on one question is not achievable in separate books, and that crudeness is what makes a break all-or-nothing whatever caused it. Pricing the whole range from one cost function makes the payoff graded, so a trader who was roughly right keeps something back in every case, including the two nothing in the design was ever going to prevent. Where the question is genuinely binary, a deal signed or not, an earthquake before August or not, there is no distribution to hold and none of this applies.


Appendix: how the sample was built

A $10,000 floor on event volume

Below that a price is one or two trades. We are looking for markets that held a view, and a thin book does not hold one.

Three price points per market

The final price before resolution, the price nearest 24 hours out, and the price nearest 7 days out. A single closing price catches markets that touched 95% for an hour on the way to resolving.

12-hour granularity

That is what the free CLOB endpoint returns for resolved markets. A market that went from even money to resolution inside its last 12 hours never appears here, so every count is a lower bound.

95% as the line

Raising it thins the sample to fewer than 80 markets at 99%. Lowering it stops being a study of confident prices. 95% leaves a set large enough to sort into kinds and small enough to read.

Exposure is cumulative market volume

An upper bound on what was at stake, not a loss figure. Volume counts both sides of the book and counts trades made earlier at lower confidence. Separating the trader who bought at 97 cents from the one who bought at 60 and sold at 90 needs per-trade fills, which the free API does not return.

Twelve geopolitical markets held out

The Khamenei succession market carries $129.6 million, more than the other 540 markets put together. The twelve carry $195.5 million against the core sample's $120.4 million. Left in, every table here would be a table about one cluster of events in early 2026.

Settled prices out of the denominator

A championship leg dies the moment the team is eliminated, and a "by December 31" market can be decided in March. The contract keeps its listed end date, but the price has stopped forecasting and started recording. Those markets sit at 99% or 1% and can essentially never resolve against themselves, so counting them as high-confidence markets pushes the base rate down. They fall unevenly across the comparison: 24.5% of non-bracket high-confidence markets are settled this way against 7.2% of brackets. We identify them as markets whose price did not move across the measurement window.

Identifying brackets

An event qualifies on the negRisk flag together with legs that parse as non-overlapping numeric intervals. Both are required. Keyword matching on question text alone pulls in threshold ladders and touch markets, which carry negRisk = False, and misses genuine partitions whose legs read as "between $A and $B." The classifier and the frozen market list are published, so the cohort can be re-derived without an API call.

Classifying the 40

Hand-labelled, taking the primary cause where more than one applied. The annotated file carries event IDs so it joins to the dataset, and a status column, and only verified rows are used.

What is inside the "other" category

"Other" is the largest bucket by count and the least informative, because it is Polymarket's topic label rather than anything about the contract. A crude pass over the question text of all 206 splits it five ways.

What is actually in "other"

206 core breaks, grouped by the semantic structure of the question.

The 206 Other-category breaks split into weather brackets 54, field winner 48, sports spreads and matchups 39, numeric range or threshold 38, and dated yes/no events 27. 206 markets GROUP MARKETS EXPOSURE Weather brackets 54 · 26% $1.6M Field winner 48 · 23% $11.0M Sports spreads and matchups 39 · 19% $0.8M Numeric range or threshold 38 · 18% $6.3M Dated yes/no event 27 · 13% $5.8M Already inside the bracket cohort 61
Figure A1. Groups assigned by pattern-matching the question text, so the edges are rough. Weather brackets and numeric ranges are the same shape as the bracket cohort, and 61 of the 206 are already counted in it.

Weather brackets are the biggest group and the cheapest, 54 markets carrying $1.6 million between them, about $29,000 each. Field-winner markets carry most of the money, $11.0 million across 48. The 39 sports spreads and matchups are markets whose category field Polymarket left blank. And 61 of the 206 already sit inside the bracket cohort, mostly the daily temperature bands for London, Miami and New York, so "other" overlaps the structure we are tracking rather than sitting beside it.

What the free API does not give us

No per-strike volume or order book depth, so we can see which bin resolved but not how positions were spread across bins. No per-trade fills, so exposure stays a bound. No last-traded timestamps, so the settled-price correction is approximate and the break rates carry a range rather than a point.


Limitations

  • Price history comes at 12-hour granularity, so a market that went from even money to resolution inside its last 12 hours never appears. Every count here is a lower bound.
  • Exposure is cumulative market volume. It is an upper bound on what was at stake and not a measure of realised losses.
  • The bracket-versus-rest ratio depends on how settled prices are counted, and across the four ways we counted them it runs from 1.41 to 1.83. We do not have the trade timestamps to settle which is right.
  • The concentration step in the bracket argument is inferred from the structure. Nothing here measures where positions actually sat across strikes when a market broke.
  • The fall in break rate over time is measured per market resolved, not per market that reached 95%, so it mixes any change in pricing with a change in what Polymarket lists.
  • Category labels are inferred from titles wherever Polymarket left the field blank.
  • The four kinds were assigned by hand across the 40 largest breaks, and some calls took judgement, particularly separating an exogenous shock from a structural cause on crypto markets. The annotated file records each decision.
  • Breaks cluster by event and by recurring event family, so significance is reported at event level and through an event-cluster bootstrap. A single-market chi-square would overstate it.
  • Markets where the dismissed side sat at exactly five cents are kept. Dropping them leaves 516 markets and changes nothing.

None of these changes the shape of the result. The pattern that repeats across Polymarket's confident misses is bracket compression on numerical outcomes, 30% of the breaks and 30% of the exposure, at a rate 1.4 to 1.8 times the rest of the market. It is also the one of the four kinds where the cause sits in the contract rather than in the world.


References

  • Bürgi, Constantin, Deng, and Karl Whelan. "Makers and Takers: The Economics of the Kalshi Prediction Market." UCD Working Paper 2025/19, September 2025.
  • Clinton, Joshua, and TzuFeng Huang. "Prediction Markets? The Accuracy and Efficiency of $2.4 Billion in the 2024 Presidential Election." Vanderbilt University, SocArXiv preprint, 2025. doi.org/10.31235/osf.io/d5yx2_v3
  • McCullough, Alex. "How Accurate Is Polymarket", Dune dashboard.
  • Gurkaynak, Refet S., and Justin Wolfers. "Macroeconomic Derivatives: An Initial Analysis of Market-Based Macro Forecasts, Uncertainty, and Risk." NBER Working Paper 11929, 2007.
  • XO Labs. "Resolving Prediction Markets: An AI-Driven Layered Approach," May 2026, and their account of the Trump-Ukraine minerals settlement.
  • functionSPACE. "The Real Cost of Liquidity Discretisation," and its numeric slice, on the census of 18,863 Polymarket multi-market events. functionspace.dev/research
  • functionSPACE. "The YES Bias Might Not Exist," on 28,793 on-chain Polymarket trades. functionspace.dev/research
  • functionSPACE. Probability Markets white paper, mechanism v0.4.1.
  • Data: Polymarket Gamma API and CLOB API. The pipeline, the frozen 247,367-market list, the denominator aggregates and the annotated top-40 classification file are at github.com/OJdominus/polymarket-high-confidence-breaks-main.