Where crossing is free, the mispricing is gone too: a strategy screen on Polymarket
A systematic screen of arbitrage, calibration, momentum and lead-lag strategies across ~2,400 Polymarket markets. Almost everything died at the taker fee — and when we found the 14% of markets where Polymarket waives that fee entirely, the mispricings were not there either.
Prediction markets look inefficient from the outside. Retail-dominated flow, thin books, contracts whose fair value is a probability that anyone can reason about, and a set of hard no-arbitrage constraints — YES and NO must sum to a dollar, a ladder of thresholds must be monotone — that are trivially checkable and frequently violated on screen.
That combination invites a systematic screen, so we ran one on Polymarket: complete-set arbitrage, ladder relative value, calibration and favourite–longshot, momentum, lead–lag between related markets, catalyst convergence, carry, and classic two-sided market making. Roughly 2,400 markets, several independent samples, on-chain fee tracing and a book recorder. Everything below is measured on Polymarket only; no second venue is involved.
Almost all of it is dead, and the interesting part is not that it died. It is what killed it, and what happened when we found the segment of the venue where that cause is absent.
The cost structure decides everything
Polymarket charges takers and pays makers nothing. That is visible on chain rather than inferred from documentation: in a single fill, the two OrderFilled events carry different fees — maker fee = 0, taker fee = 1.012150 — and the collateral ledger agrees, with the taker short by notional plus fee and the maker short by notional alone.
The fee is not linear in notional. It is shaped:
| Quantity | Value |
|---|---|
| Fee formula | theta × shares × p × (1−p) |
| theta (per market) | 0.04 – 0.07 |
| Taker cost mid-book | 250 – 500 bps |
| Taker cost at p > 0.9 | 15 – 70 bps |
| Measured folded round-trip spread | ~200 bps (median 2.00c) |
Read the last two rows together. A single marketable leg costs more than the entire round-trip spread. The venue's economics are that resting is free and subsidised, and crossing is taxed beyond any edge in this study.
What that does to a strategy list
Complete-set arbitrage is the cleanest illustration. Across three independent samples taken hours apart on two machines (1,755 / 2,416 / 926 markets), ask(YES) + ask(NO) had a minimum of 100.10c and a median of 101.00c, with zero observations below par in 2,416 markets. The floor sits exactly one tick above a dollar.
The natural reading is that competition has picked it clean. The fee structure gives a better one: the arbitrage is not competed away, it is uneconomic to take. Capturing a 200 bps spread requires two marketable legs at 250–500 bps each. Nobody needs to be fast here. The trade does not exist at any speed.
The same arithmetic disposes of most of the list:
| Strategy | Measured effect | Verdict |
|---|---|---|
| Ladder monotonicity / convexity | Violations in 66% of samples, median persistence 300s, max 3.4h — but median resting edge 0.20c | Frequent, persistent, and smaller than the fee |
| Lead–lag between related markets on Polymarket | Real lead within negRisk families — multi-outcome events whose legs must sum to one — at t = 5.9–9.1 over ~200k obs, capped at 0.068c against a 1.25c crossing cost | 18× too small; standalone markets with structurally independent books show a flat null |
| Catalyst convergence | Within-market paired test: −1.00c mean, −0.26c median, trimmed t ≈ −2.8, sign test p = 0.25 | 0.44× the 2.27c round trip |
| Time-value / carry | Residuals +0.000 to +0.011 against noise floors 0.005–0.026 across 1h–7d | Clean null — no detectable discount |
| Oracle-latency race | Winning the race means taking | Fee-killed by construction |
Note the shape of the lead–lag result, because it is a useful template. To be precise about what is being compared: this is one Polymarket market leading another Polymarket market, not one venue leading another. The pairs are legs of the same negRisk family — a multi-outcome event, such as a field of election candidates, where exactly one leg resolves YES so the legs must sum to a dollar.
Within those families the lead is real and highly significant — and also arithmetic: it is the sum-to-one constraint re-asserting itself after one leg moves, not information propagating between books. The control is to re-run on standalone markets, which have no mechanical link to each other, and there the effect is 0.0007c at t = 0.06. A large t-statistic on a mechanically-linked family is not evidence of a signal.
One null that died to a control, not to a fee
Calibration deserves separate mention because it failed for a different and more embarrassing reason.
The initial measurement showed longshots systematically underpriced — realized minus quoted of +0.04 to +0.08 in the 0.05–0.40 band, the reverse of the classic favourite–longshot bias. It survived a parametric bootstrap. It survived removing a look-ahead confound that had accounted for several times the raw effect. It survived clustering the bootstrap by event rather than by market.
It did not survive checking what the labels meant. 35.5% of resolved markets are not Yes/No contracts. They are head-to-head or over/under pairs, where there is no YES token at all and the convention of treating the first token as YES is meaningless. Roughly a third of all outcome labels were therefore randomised — and randomised labels regress realized outcomes toward the base rate, which lifts exactly the low-price bins where the effect appeared.
Restricted to genuine Yes/No markets with deadline resolution, the bootstrap returns p = 0.09 — consistent with perfect calibration. What remained was small, non-monotone in price, and driven by single thin bins. A real behavioural bias would be smooth and monotone in price. This was a metadata artefact wearing the shape of one.
Then the premise broke
Everything above was computed against a single global fee. That was wrong.
Polymarket exposes a fee schedule per market, and the parameters vary: rate in {0.04, 0.05, 0.07}, with a rebate rate in {0.15, 0.20, 0.25} that was not in the original tracing at all. More importantly:
84 of 599 sampled markets — 14% — have fees disabled entirely. And they are not small. Those markets carried $5.98M of 24-hour volume against $16.24M for the fee-enabled set: 27% of sampled volume trades with no taker fee at all. They cluster in geopolitics — the promoted, heavily-trafficked questions.
This is a premise error rather than a measurement error, and it is potentially fatal to the whole screen. Every “dead — fee-killed” verdict was computed against a cost of 250–500 bps that, in a quarter of the venue by volume, is zero. Several of those strategies died by a margin smaller than the fee they were charged. They had to be re-run.
The re-run, and the actual finding
Restricted to the fee-free segment:
| Test | Fee-free | Fee-enabled |
|---|---|---|
| Complete set below par | 0 violations (90 markets, 113,578 samples) | 1 violation, 1.00c, lasting 60s (584 markets) |
| Sum-to-one violations (of 2,271 already measured) | 0 | all 2,271 |
| Monotonicity violations (of 6,255 already measured) | 0 | all 6,255 |
Not one of the 8,526 measured mispricings lies in the segment where crossing is free.
So the verdicts stand, but for a materially better reason than the one originally given. The fee was never what stood between us and these trades. Where the fee is absent, the mispricing is absent too.
Why that is not a coincidence
Fees and mispricing are inversely related because both track how contested a market is.
The fee-free markets are the promoted ones — the questions the venue is pushing, carrying the most volume and the most attention. That is precisely where the sharpest participants are and where the books are tightest. Waiving the fee is not a gift to arbitrageurs; it is a subsidy that attracts the competition that removes the edge. The remaining venue-wide inefficiency lives where nobody is looking, and it is taxed at 250–500 bps to touch.
This is the same conclusion we reached testing a zero-fee perpetual DEX, by a different route: if your strategies die net of fees, a venue with no fees should resurrect them, and none came back. Cost is rarely the binding constraint on a dead edge. It is the thing that gets blamed because it is the thing that is easy to measure.
What the book is actually doing
The screen produced one result that has nothing to do with fees and is the most useful thing in it.
Two measurements looked contradictory. Conditioning on trades, the mid continued in the direction of the trade — ordinary adverse selection. Conditioning on all mid moves, prices reversed significantly at every horizon (t = −4.5 to −4.9, 529 markets, ~72k observations), and the reversal was monotone increasing in move size, which rules out bid-ask bounce: a bounce artefact lives entirely in the smallest bucket, and this did the opposite. Roughly 20% of a one-cent move retraced within five minutes.
Both cannot be true of the same population, so we split every mid move on the one-minute grid by whether a public trade occurred in the same minute:
| Class | Horizon | Continuation (cents) | t | Observations |
|---|---|---|---|---|
| Trade-initiated | 5 min | +0.187 | +2.37 | 4,651 |
| Trade-initiated | 20 min | −0.093 | −1.03 | 4,506 |
| Quote-driven | 5 min | −0.129 | −4.56 | 28,763 |
| Quote-driven | 20 min | −0.140 | −3.94 | 28,196 |
Both findings were correct. They conditioned on different populations, and the populations move in opposite directions with similar magnitude — which is why pooling them nets to approximately nothing.
Three things are visible only in the split:
- 86% of mid movement has no trade behind it — 29,070 quote-driven observations against 4,700 trade-initiated. Most apparent price discovery on this venue is makers repricing each other, not information arriving.
- Information is impounded in about ten minutes, then overshoots. Trade-initiated continuation peaks at five minutes, holds through ten, and has turned negative by twenty.
- Neither class is crossable. At 0.13–0.19c against a 2.00c folded round trip, these are quoting parameters, not strategies.
That last point is the honest limit of the result. It does not produce a trade. It produces a signed, sized input to two gates that a maker already needs — pause or widen after being hit, lean against a quote-driven move with no trade behind it — and it is the first measurement here that says what to set them to rather than merely that they are needed.
What generalises
Visible constraint violations are not edges. Ladder violations occur in two thirds of samples and persist for a median of five minutes. That looks like a latency opportunity and is not one: the median resting edge is 0.20c and the crossing cost is an order of magnitude larger. The violation is visible precisely because it is not worth removing.
Check what your labels mean before trusting a bootstrap. A calibration effect survived three legitimate statistical controls and died to the observation that a third of the contracts were not the instrument we assumed. No amount of resampling detects a systematically wrong label.
Significance on a mechanically-linked family measures the mechanism. A t-statistic of 9 on legs bound by a sum-to-one constraint is the constraint, not a lead. The control is to re-run on structurally independent books, where the same test returned t = 0.06.
Contradictory findings usually mean an unstated conditioning variable. Continuation and reversal were both real. The resolution was not to pick a winner but to find the split that made both true, and the split turned out to be the most informative single fact about Polymarket.
And the headline one: a fee waiver is a competition magnet. If you are screening a venue and hoping a low-cost segment will let a marginal edge clear, check whether the mispricing survives there before building anything. On this venue it did not, anywhere, in 8,526 measured instances.
What this does not measure
One venue, a three-day measurement window, and a book recorder sampling at one sweep per minute. That cadence is a real limitation: 38.8% of prices moved within 60 seconds and the median quote lifetime sits at the sampling floor, so the tail of the lifetime distribution and any diurnal structure are not resolved. For the arbitrage work this does not matter — a sub-minute violation is untradeable at these costs either way — but it does bound the microstructure conclusions.
There is a coverage bias worth stating plainly. The recorder tracks the reward-eligible universe, and short-dated catalyst markets largely fall outside it: of 1,285 markets resolving inside one 30-hour window, 1,208 had no token in the tape. The catalyst result therefore rests on 79 markets that happened to be covered, and scheduled macro releases were not captured at all. That branch is untested rather than dead.
The fee-free comparison rests on 90 markets, and the trade-versus-quote split on 120. Both are large in observations and modest in markets, which is the axis that matters for clustered standard errors. The direction of each result is well outside its noise floor; the precise magnitudes should be treated as directional.
Finally, this is a screen rather than a complete account. Several branches of the taxonomy remain open and are deliberately not discussed here.