Kalshi AI — Mispricing Audit
Report date 2026-07-28 · Methodology v3.3 (conviction-weighted sizing, entry-band screen, 15% cluster cap, EV-agnostic, resolver-priced) · Horizon 45 days · Capital $2,500 · Picks count toward the public track record and are cleared for auto-copy.
VERDICT: 0 PICKS — 100% CASH
No capital deploys this run. Six candidates survived the mechanical screens and received full diligence — live orderbooks, resolver profiling, and primary-source research (the Ornn GPU index reconstruction, the arena.ai leaderboards fetched directly, Epoch's FrontierMath problem pages). Every one of them failed the v3 gates: either the market's price already sits above my modeled probability (no edge), the fillable book is a rounding error, the entry price falls in a banned band (sub-35¢ tail or coin-flip without HIGH conviction), or the resolution feed is literally unobservable to anyone without a paid data seat. Eight rejects are logged below and machine-tracked to settlement. Holding $2,500 (100%) in cash is the position.
1. How this was researched
Mode: theme-match. "AI" is not a Kalshi category (categories in the mirror: Sports, Crypto, Climate and Weather, Elections, Entertainment, Financials, Commodities, Mentions, Economics, Politics, Science and Technology, Companies, World, Social, Health, Transportation). I selected markets by keyword across all categories: openai, anthropic, claude, "ai", artificial intelligence, gpt, gemini, grok, xai, nvidia, agi, llm, deepseek, chatbot, language model, ai model, superintelligence, llama, mistral, fable, mythos, altman, frontier math, data center, chip export, applied to event titles, market titles, and tickers (an initial pass matching %agi% flooded the set with e-sports teams named "magic" and was tightened). Sports was excluded after inspection.
- Universe pull via the read-only DB mirror (schema fetched first; note: the mirror exposes trading_markets/trading_events/signals/signal_snapshots — no market_snapshots table, so price history came from signal_snapshots plus the live Kalshi API). ~370 matching market rows closing within 45 days, dominated by five GPU compute-price ladder families (~200 strikes) and leaderboard one-of-N events.
- Live state for all 128 plausible markets from the Kalshi public API (the mirror's status lags: several "active" markets are in fact finalized — the Claude Opus release markets resolved YES on Jul 24, GPT-5.6 markets resolved YES, Fable-5-restored resolved YES, FrontierMath solve #2 resolved YES on Jul 27).
- Stage-1 mechanical cuts: vol24 < 1500, spread > 5¢, mention/announcer markets, fully-priced (bid ≥ 95 or ask ≤ 5), finalized, and the 6 AI tickers already live in users' feeds from prior runs (KXGEMINI-GEMI35P-26JUL31, KXGPT-OPEN-26SEP01, KXMATHAI-26JUL31-GEMI, KXTOPMODEL-26JUL31-CLAUF, KXTOPMODEL-26JUL31-CLAUT, KXRTX5090WS-26AUG07 — excluded up front, no re-picks).
- Stage-2 diligence on 6 survivors (2 kept despite borderline volume, flagged): rules text, 14-day price history, live depth-of-book, and three parallel primary-source research threads — (a) the Ornn Compute Price Index and a reconstruction of its July H200/RTX-5090 means from Kalshi's own settled weekly prints, (b) arena.ai text/code leaderboards fetched directly on Jul 28, (c) Epoch AI FrontierMath problem pages, Challenger reports, Gemini 3.5 Pro reporting.
- Stages 2.5–4: resolver profiled and priced per candidate; conviction + entry-band screens applied; result was an empty book.
2. Markets reviewed
| Market / family | Price (YES) | Vol 24h | Disposition |
|---|---|---|---|
| KXH200MS-26JUL-4.500 — H200 July avg > $4.50 (Ornn) | 87 / 88 | 6,309 | Stage-2 candidate → rejected (unverifiable resolver; no edge) |
| KXRTX5090MS-26JUL-0.500 — RTX 5090 July avg > $0.50 | 9 / 10 | 2,532 | Stage-2 candidate → rejected (no edge; 8-contract book) |
| KXTOPMODEL-26JUL31-CLAU5 — opus-5-high Arena #1 Jul 31 | 5 / 7 | 54,851 | Stage-2 candidate → rejected (tail band; resolver ambiguity) |
| KXTECHRANKLISTAICODE-26AUG03-CLAU — Claude #1 Code Arena Aug 3 | 92 / 95 | 1,351 | Stage-2 candidate → rejected (no edge at ask) |
| KXFRONTIER-FRONB-26SEP01 — 3rd FrontierMath open-problem solve | 40 / 41 | 1,210 | Stage-2 candidate → rejected (coin-flip band) |
| KXCHINAAI-26JUL31-ALIB — Alibaba top Chinese AI on Arena | 79 / 84 | 926 | Stage-2 candidate → rejected (liquidity; unverifiable ordering) |
| KXCHAICUTS-26AUG06-T1 — AI #1 Challenger job-cut reason (Jul) | 46 / 51 | 87 | Stage-1 fail, diligenced → logged reject |
| KXGEMINI-GEMI35P-26AUG16 — Gemini 3.5 Pro before Aug 16 | 45 / 50 | 111 | Stage-1 fail, diligenced → logged reject |
| GPU monthly ladders (H200/B200/H100/A100/RTX 5090, ~40 strikes) | 99 / 1 | — | Fully priced at every strike except the two above |
| KXLLM1-26JUL31-A · KXMATHAI-CLAU · KXCODEAI-CLAU | 96–99 | — | Fully priced (bid ≥ 95) |
| KXCLAUDE-OPUS-* · KXGPT-OPENB · KXFABLERESTORE · KXFRONTIER-FRON · KXCLAUDE-CLAU-* | — | — | Already finalized (all YES) — mirror status stale |
| KXCLAUDE-NXTMYTH-26SEP01 (35/41) · KXAIREVIEW (9/14) · KXANTHROPICRESCIND (2/3) · KXFABLEDISABLED-AUG31 (4/5) · KXVERARELEASE (no book) · KXIPO* (0/1) | — | <350 | Stage-1 fails: illiquid, wide, or tail-priced |
| KXFEDMENTION-AI · KXHEARINGMENTION-* · KXTRUMPSAYCOMPANY-* | — | — | Mention/announcer-noise markets — skipped by rule |
| Existing feed picks (6 AI tickers) | — | — | Excluded up front — already published, re-picking would duplicate |
3. Picks — none deployed
0 PICKS This is the methodology working, not failing. The v1/v2 resolved record showed the book bleeding exactly where this week's temptations sit: coin-flips (FrontierMath at 41¢, Challenger at 51¢), tails with "obvious" modeled edge (opus-5-high at 7¢), and leaderboard-snapshot markets whose resolver flipped on a mid-month launch — the single costliest loss in program history was a leaderboard bet like the ones rejected here. The one family with genuinely deep liquidity (GPU monthlies) is being priced by traders reading a live per-hour data feed that is login-gated to everyone else; my best independent reconstruction says the two live strikes are priced fairly to slightly rich, not cheap.
What would have changed the answer: (a) observable Ornn hourly data showing the July H200 mean already ≥ $4.55 — that would have made KXH200MS-4.500 YES a HIGH-conviction favorite; (b) an Epoch status flip on the Book Graphs problem before Sep 1 pushing FRONB toward mechanical; (c) a listed claude-opus-5-max leg in the TOPMODEL event, which would have been the clean way to express the Arena thesis. None of those exist today.
4. Recommended $2,500 portfolio
| Position | Conviction | Entry band | Cost | % of capital | Max payout |
|---|---|---|---|---|---|
| Cash | — | — | $2,500 | 100% | $2,500 |
| Total deployed: $0 · Cash held: $2,500 (100%) · Blended EV: n/a · Dollar edge: $0. EV figures on rejects are recorded for calibration only and did not drive sizing (v3 is EV-agnostic by rule). | |||||
Cluster exposure (cap: 15% = $375 per cluster)
| Cluster | Cost | % of capital | Cap |
|---|---|---|---|
| ornn-gpu-july-mean | $0 | 0% | 15% |
| arena-text-jul31 | $0 | 0% | 15% |
| arena-code-aug03 | $0 | 0% | 15% |
| frontiermath-third-solve | $0 | 0% | 15% |
| challenger-ai-jobcuts · gemini-35-pro-release | $0 | 0% | 15% |
Conviction exposure
| Tier | Cost deployed | Per-pick cap |
|---|---|---|
| HIGH | $0 | $375 (15%) |
| MEDIUM | $0 | $175 (7%) |
| LOW | never deploys (v3.3) | — |
Risk profile & execution notes
- Worst case / best case / most likely: $2,500 / $2,500 / $2,500. An all-cash book cannot lose; the cost is opportunity, priced below.
- Watchlist triggers (for the next run, not standing orders):
- KXFRONTIER-FRONB-26SEP01 — if Epoch flips Ramsey Numbers for Book Graphs to Solved (Dualverse's infinite-family construction has been under author review since Jun 25), YES becomes near-mechanical; the NO book at 59¢ was 147+535 deep today, so entry should still be available on the news minute-scale.
- KXH200MS-26AUG strikes (August monthly, will list ~Aug 1) — if the H200 spike to ~$5 holds, the August low strikes open mispriced the way July's did in early July; this month's lesson is to price them before the crowd's 59¢ repricing, not after.
- Gemini 3.5 Pro at the Aug 12 Made-by-Google event — a launch would move KXGEMINI-GEMI35P-26AUG16/31 and, with a lag, every Arena-snapshot market; the books there are thin enough that news beats them.
- No hedges required — nothing is on.
5. What I rejected and why
All eight entries below are machine-logged in picks.json and shadow-tracked to settlement — if the rejects outperform, the screen itself is on trial. Three cuts came specifically from the Stage-3.5 entry-band screen (one tail, two coin-flips): that discipline is the visible difference between v3 and the book that lost money.
KXH200MS-26JUL-4.500 — H200 July monthly average > $4.50 · REJECTED
Receipts. Kalshi's own settled weekly markets print the Ornn Friday closes (primary data): Jul 3 $4.06 → Jul 10 $4.27 → Jul 17 $4.79 → Jul 24 $4.55, with the Jul-31 daily ladder now centered near $5.0–5.3 after a spike. Trapezoid interpolation of these prints — a method that back-tests to within $0.01 (RTX 5090 June: est. 0.740 vs settled 0.74) and $0.03 (H200 June: est. ~3.80 vs settled 3.77) against the June settlements — puts the July month-to-date mean at ≈ $4.45–4.52 and the final mean at ≈ $4.50 ± 0.05. That straddles the $4.505 cutoff. The market repriced this contract from 29¢ (Jul 26) to 88¢ (Jul 27, +59¢ in ~36h) on live hourly data I cannot see. Buying YES at 88 means paying above my own estimate on faith in someone else's feed; buying NO at 12–13 means betting my interpolation error against informed flow — and is a sub-35¢ tail besides. Both sides fail. The single cleanest way this reject looks wrong: intraweek hourly values ran hot above the Friday closes all month and the mean is already ≥ $4.55 — in which case the 88¢ crowd simply read the meter.
KXRTX5090MS-26JUL-0.500 — RTX 5090 July monthly average > $0.50 · REJECTED
Friday prints ground down $0.52 → $0.51 → $0.50 → $0.49 through July; trapezoid full-month mean ≈ $0.500–0.505 — within half a cent of the cutoff. The market agrees (YES fell 80¢ → 9¢ over the month as the mean converged on the strike). At 91¢ the NO side offers no edge over my 87%, and the book is decisive on its own: 8 contracts available at 91, then a gap to 96 — a $200 position would pay ≥ 5¢ slippage. Fails Stage-3's fill test outright.
KXTOPMODEL-26JUL31-CLAU5 — claude-opus-5-high is Arena #1 on Jul 31, 10:00 ET · REJECTED — TAIL BAND (v3.2)
This is the market that explains the week. Anthropic launched Claude Opus 5 on Jul 24 (TechCrunch, Axios, Fortune); Arena listed the variants on Jul 27; and the no-style-control board fetched directly today reads: #1 claude-opus-5-max 1512±12 (2,386 votes), #2 claude-opus-5-high 1505±8 (6,159 votes), #3 claude-opus-4-6-thinking 1503±4. The winning string — opus-5-max — is not a listed leg, which is why a mutually-exclusive event sums to ~12¢ while company-level "Anthropic top" trades 99¢, and why the old favorite (opus-4-6-thinking) collapsed from ~80¢ to 2¢ in a week. With a 7-point gap, overlapping CIs, and both scores built on tiny vote counts, opus-5-high overtaking its sibling by Friday is far likelier than 7% — my estimate ~18%. It is still a hard reject: sub-35¢ tail (the band went 0-for across the resolved record; v3.2 removed the lottery slot), stacked on a resolver whose UI no longer matches its own rules text. Logged and shadow-tracked — if it hits, the record will say so. This family is also the exact shape of the program's costliest historical loss (right about the world, wrong about a leaderboard snapshot).
KXTECHRANKLISTAICODE-26AUG03-CLAU — Claude #1 on LM Code Arena, Aug 3, 10:00 ET · REJECTED
The board today: Claude holds #1 (opus-5-max, 1725), #3 (opus-5-high, 1670) and #4 (fable-5, 1629); the only live threat is Moonshot's kimi-k3-max at 1682. For NO, Kimi must finish above every Claude — it already sits above two of them, so the thesis reduces to "opus-5-max's preliminary 1725 regresses ~45 points in six days while opus-5-high stays put." Plausible enough that 94% is my honest number — which the ask already pays for. Deep-favorite band, zero edge, thin book: pass. Gemini 3.5 Pro landing and topping the code board by Aug 3 is a rounding-error risk (Opus 5 took 3 days just to list).
KXFRONTIER-FRONB-26SEP01 — a third FrontierMath Open Problem solved before Sep 1 · REJECTED — COIN-FLIP BAND
Genuinely interesting and honestly coin-flip. The concrete catalyst: on the Book Graphs problem, Dualverse AI's infinite-family construction has been under author review since Jun 25 — a positive review plus an Epoch status flip before Sep 1 resolves YES. Against it: two solves in six months is the entire base rate, Epoch is deliberately slow (5–7 weeks solution→flip), and the review has already run a month without news. 45% vs 41¢ is not a v3 trade: the 35–60¢ band requires HIGH conviction and ≥15¢, and this is MEDIUM at best. This is the band that hit ~33% historically. Shadow-tracked.
KXCHINAAI-26JUL31-ALIB — Alibaba top Chinese AI company on Arena (no style control), Jul 31 · REJECTED
Priced 81/17 Alibaba/Moonshot, but the one direct observation I have cuts the other way: on the default text board, Moonshot's kimi-k3-max (1486±10) is the highest-ranked Chinese model in the top 15 and no Qwen appears at all. Resolution uses the no-style-control Score view, where I could not see below the top 10 — so either the crowd knows Qwen leads on that specific view, or a 17¢ Moonshot leg is the mispricing. Unverifiable in the time available, volume under screen, book gaps beyond 33 contracts: reject with the honest label "couldn't check the resolver's actual view." Logged without a probability.
KXCHAICUTS-26AUG06-T1 — "Artificial Intelligence" is the #1 job-cut reason in Challenger's July report · REJECTED — COIN-FLIP BAND + LIQUIDITY
AI has been the #1 reason four straight months — May was a blowout (38,579 cuts, ~40% of all cuts; Fox Business, CNBC), but June narrowed sharply: AI 14,029 vs Market/Economic Conditions 12,470 (Challenger June report). A fifth month is better-than-even, not safe; one seasonal restructuring wave flips a 1,559-cut margin. 58% at 51¢ in the coin-flip band without HIGH conviction fails Stage 3.5 — and the 87-contract daily volume fails Stage 1 before that. Logged for the shadow track.
KXGEMINI-GEMI35P-26AUG16 — Gemini 3.5 Pro (or greater) released before Aug 16 · REJECTED — COIN-FLIP BAND + LIQUIDITY
The NO case is well-sourced: three missed targets (June GA → mid-July → Jul 17), Bloomberg's Jul 16 report of a scrapped-and-rebuilt base model over hallucination rates, and Jul 21's stopgap Flash releases relieving launch pressure (TechCrunch). The YES case is one date: the Aug 12 Made-by-Google event sits four days inside the window and is a natural stage. That tension is a true coin-flip leaning NO — exactly what v3 refuses to fund at 7¢ of edge. The Aug 31 sibling (71/78) prices the "or greater by end of August" door more richly and is wider still. Logged.
6. Sources
- Kalshi public API (live prices, orderbooks, and settled Ornn expiration prints): api.elections.kalshi.com/trade-api/v2 · read-only DB mirror for universe/rules/history
- Arena leaderboards, fetched directly Jul 28: text · text (no style control) · code · leaderboard changelog
- Claude Opus 5 launch (Jul 24): TechCrunch · Axios · Fortune; GPT-5.6 (Jul 9): TechCrunch · OpenAI
- Ornn / OCPI: ornn.com · dashboard.ornnai.com · Bloomberg Terminal listing (PRNewswire) · ICE OCPI futures (BusinessWire) · Tunguz on launch-driven GPU spikes · AIMultiple cross-check · Thunder Compute
- FrontierMath: Epoch open problems · 2-adic Galois (Solved; Fable 5 Jun 9, GPT-5.5 Pro Jun 24) · Book Graphs (Dualverse review since Jun 25) · Epoch on solve #1 (Mar 23) · padicIGP repo
- Gemini 3.5 Pro: TechCrunch Jul 21 · Bloomberg-sourced delay coverage · Made by Google, Aug 12
- Challenger reports: May 2026 PDF · June 2026 · CNBC
- Anthropic/government context (for completeness on rejected/illiquid markets): CNBC Jun 12 · WaPo Jun 30 · Fortune Jul 1 · Mayer Brown on the DoD designation