Kalshi AI — Mispricing Audit

Report date 2026-07-28 · Methodology v3.3 (conviction-weighted sizing, entry-band screen, 15% cluster cap, EV-agnostic, resolver-priced) · Horizon 45 days · Capital $2,500 · Picks count toward the public track record and are cleared for auto-copy.

VERDICT: 0 PICKS — 100% CASH

No capital deploys this run. Six candidates survived the mechanical screens and received full diligence — live orderbooks, resolver profiling, and primary-source research (the Ornn GPU index reconstruction, the arena.ai leaderboards fetched directly, Epoch's FrontierMath problem pages). Every one of them failed the v3 gates: either the market's price already sits above my modeled probability (no edge), the fillable book is a rounding error, the entry price falls in a banned band (sub-35¢ tail or coin-flip without HIGH conviction), or the resolution feed is literally unobservable to anyone without a paid data seat. Eight rejects are logged below and machine-tracked to settlement. Holding $2,500 (100%) in cash is the position.

1. How this was researched

Mode: theme-match. "AI" is not a Kalshi category (categories in the mirror: Sports, Crypto, Climate and Weather, Elections, Entertainment, Financials, Commodities, Mentions, Economics, Politics, Science and Technology, Companies, World, Social, Health, Transportation). I selected markets by keyword across all categories: openai, anthropic, claude, "ai", artificial intelligence, gpt, gemini, grok, xai, nvidia, agi, llm, deepseek, chatbot, language model, ai model, superintelligence, llama, mistral, fable, mythos, altman, frontier math, data center, chip export, applied to event titles, market titles, and tickers (an initial pass matching %agi% flooded the set with e-sports teams named "magic" and was tightened). Sports was excluded after inspection.

  1. Universe pull via the read-only DB mirror (schema fetched first; note: the mirror exposes trading_markets/trading_events/signals/signal_snapshots — no market_snapshots table, so price history came from signal_snapshots plus the live Kalshi API). ~370 matching market rows closing within 45 days, dominated by five GPU compute-price ladder families (~200 strikes) and leaderboard one-of-N events.
  2. Live state for all 128 plausible markets from the Kalshi public API (the mirror's status lags: several "active" markets are in fact finalized — the Claude Opus release markets resolved YES on Jul 24, GPT-5.6 markets resolved YES, Fable-5-restored resolved YES, FrontierMath solve #2 resolved YES on Jul 27).
  3. Stage-1 mechanical cuts: vol24 < 1500, spread > 5¢, mention/announcer markets, fully-priced (bid ≥ 95 or ask ≤ 5), finalized, and the 6 AI tickers already live in users' feeds from prior runs (KXGEMINI-GEMI35P-26JUL31, KXGPT-OPEN-26SEP01, KXMATHAI-26JUL31-GEMI, KXTOPMODEL-26JUL31-CLAUF, KXTOPMODEL-26JUL31-CLAUT, KXRTX5090WS-26AUG07 — excluded up front, no re-picks).
  4. Stage-2 diligence on 6 survivors (2 kept despite borderline volume, flagged): rules text, 14-day price history, live depth-of-book, and three parallel primary-source research threads — (a) the Ornn Compute Price Index and a reconstruction of its July H200/RTX-5090 means from Kalshi's own settled weekly prints, (b) arena.ai text/code leaderboards fetched directly on Jul 28, (c) Epoch AI FrontierMath problem pages, Challenger reports, Gemini 3.5 Pro reporting.
  5. Stages 2.5–4: resolver profiled and priced per candidate; conviction + entry-band screens applied; result was an empty book.

2. Markets reviewed

Market / familyPrice (YES)Vol 24hDisposition
KXH200MS-26JUL-4.500 — H200 July avg > $4.50 (Ornn)87 / 886,309Stage-2 candidate → rejected (unverifiable resolver; no edge)
KXRTX5090MS-26JUL-0.500 — RTX 5090 July avg > $0.509 / 102,532Stage-2 candidate → rejected (no edge; 8-contract book)
KXTOPMODEL-26JUL31-CLAU5 — opus-5-high Arena #1 Jul 315 / 754,851Stage-2 candidate → rejected (tail band; resolver ambiguity)
KXTECHRANKLISTAICODE-26AUG03-CLAU — Claude #1 Code Arena Aug 392 / 951,351Stage-2 candidate → rejected (no edge at ask)
KXFRONTIER-FRONB-26SEP01 — 3rd FrontierMath open-problem solve40 / 411,210Stage-2 candidate → rejected (coin-flip band)
KXCHINAAI-26JUL31-ALIB — Alibaba top Chinese AI on Arena79 / 84926Stage-2 candidate → rejected (liquidity; unverifiable ordering)
KXCHAICUTS-26AUG06-T1 — AI #1 Challenger job-cut reason (Jul)46 / 5187Stage-1 fail, diligenced → logged reject
KXGEMINI-GEMI35P-26AUG16 — Gemini 3.5 Pro before Aug 1645 / 50111Stage-1 fail, diligenced → logged reject
GPU monthly ladders (H200/B200/H100/A100/RTX 5090, ~40 strikes)99 / 1Fully priced at every strike except the two above
KXLLM1-26JUL31-A · KXMATHAI-CLAU · KXCODEAI-CLAU96–99Fully priced (bid ≥ 95)
KXCLAUDE-OPUS-* · KXGPT-OPENB · KXFABLERESTORE · KXFRONTIER-FRON · KXCLAUDE-CLAU-*Already finalized (all YES) — mirror status stale
KXCLAUDE-NXTMYTH-26SEP01 (35/41) · KXAIREVIEW (9/14) · KXANTHROPICRESCIND (2/3) · KXFABLEDISABLED-AUG31 (4/5) · KXVERARELEASE (no book) · KXIPO* (0/1)<350Stage-1 fails: illiquid, wide, or tail-priced
KXFEDMENTION-AI · KXHEARINGMENTION-* · KXTRUMPSAYCOMPANY-*Mention/announcer-noise markets — skipped by rule
Existing feed picks (6 AI tickers)Excluded up front — already published, re-picking would duplicate

3. Picks — none deployed

0 PICKS  This is the methodology working, not failing. The v1/v2 resolved record showed the book bleeding exactly where this week's temptations sit: coin-flips (FrontierMath at 41¢, Challenger at 51¢), tails with "obvious" modeled edge (opus-5-high at 7¢), and leaderboard-snapshot markets whose resolver flipped on a mid-month launch — the single costliest loss in program history was a leaderboard bet like the ones rejected here. The one family with genuinely deep liquidity (GPU monthlies) is being priced by traders reading a live per-hour data feed that is login-gated to everyone else; my best independent reconstruction says the two live strikes are priced fairly to slightly rich, not cheap.

What would have changed the answer: (a) observable Ornn hourly data showing the July H200 mean already ≥ $4.55 — that would have made KXH200MS-4.500 YES a HIGH-conviction favorite; (b) an Epoch status flip on the Book Graphs problem before Sep 1 pushing FRONB toward mechanical; (c) a listed claude-opus-5-max leg in the TOPMODEL event, which would have been the clean way to express the Arena thesis. None of those exist today.

4. Recommended $2,500 portfolio

PositionConvictionEntry bandCost% of capitalMax payout
Cash$2,500100%$2,500
Total deployed: $0 · Cash held: $2,500 (100%) · Blended EV: n/a · Dollar edge: $0. EV figures on rejects are recorded for calibration only and did not drive sizing (v3 is EV-agnostic by rule).

Cluster exposure (cap: 15% = $375 per cluster)

ClusterCost% of capitalCap
ornn-gpu-july-mean$00%15%
arena-text-jul31$00%15%
arena-code-aug03$00%15%
frontiermath-third-solve$00%15%
challenger-ai-jobcuts · gemini-35-pro-release$00%15%

Conviction exposure

TierCost deployedPer-pick cap
HIGH$0$375 (15%)
MEDIUM$0$175 (7%)
LOWnever deploys (v3.3)

Risk profile & execution notes

5. What I rejected and why

All eight entries below are machine-logged in picks.json and shadow-tracked to settlement — if the rejects outperform, the screen itself is on trial. Three cuts came specifically from the Stage-3.5 entry-band screen (one tail, two coin-flips): that discipline is the visible difference between v3 and the book that lost money.

KXH200MS-26JUL-4.500 — H200 July monthly average > $4.50 · REJECTED

Would-have side: YES @ 88¢ (bid 87 / ask 88 · vol24 6,309 · OI 4,755 · ~204 contracts fillable at 88)
My probability: ~68% vs market 88% — my model is below the price
Resolver: Ornn Compute Price Index (OCPI), arithmetic mean of hourly USD values for July, rounded to 2dp — so YES effectively requires a raw mean ≥ $4.505. Priced risks: unobservable feed (dashboard is JS/login-gated, no public API) −10 pts; borderline rounding cutoff −8 pts; last-3-day path dependence −6 pts. Deductions > 15 pts → automatic Stage-3 reject even before the edge math.

Receipts. Kalshi's own settled weekly markets print the Ornn Friday closes (primary data): Jul 3 $4.06 → Jul 10 $4.27 → Jul 17 $4.79 → Jul 24 $4.55, with the Jul-31 daily ladder now centered near $5.0–5.3 after a spike. Trapezoid interpolation of these prints — a method that back-tests to within $0.01 (RTX 5090 June: est. 0.740 vs settled 0.74) and $0.03 (H200 June: est. ~3.80 vs settled 3.77) against the June settlements — puts the July month-to-date mean at ≈ $4.45–4.52 and the final mean at ≈ $4.50 ± 0.05. That straddles the $4.505 cutoff. The market repriced this contract from 29¢ (Jul 26) to 88¢ (Jul 27, +59¢ in ~36h) on live hourly data I cannot see. Buying YES at 88 means paying above my own estimate on faith in someone else's feed; buying NO at 12–13 means betting my interpolation error against informed flow — and is a sub-35¢ tail besides. Both sides fail. The single cleanest way this reject looks wrong: intraweek hourly values ran hot above the Friday closes all month and the mean is already ≥ $4.55 — in which case the 88¢ crowd simply read the meter.

KXRTX5090MS-26JUL-0.500 — RTX 5090 July monthly average > $0.50 · REJECTED

Would-have side: NO @ 91¢ effective (YES 9/10 · vol24 2,532 · OI 6,802)
My probability (NO): ~87% vs 91¢ ask — negative edge
Resolver: same Ornn hourly-mean instrument; YES needs raw mean ≥ $0.505. Priced risks: unobservable feed −10 pts; cutoff within ~$0.005 of my estimate −8 pts.

Friday prints ground down $0.52 → $0.51 → $0.50 → $0.49 through July; trapezoid full-month mean ≈ $0.500–0.505 — within half a cent of the cutoff. The market agrees (YES fell 80¢ → 9¢ over the month as the mean converged on the strike). At 91¢ the NO side offers no edge over my 87%, and the book is decisive on its own: 8 contracts available at 91, then a gap to 96 — a $200 position would pay ≥ 5¢ slippage. Fails Stage-3's fill test outright.

KXTOPMODEL-26JUL31-CLAU5 — claude-opus-5-high is Arena #1 on Jul 31, 10:00 ET · REJECTED — TAIL BAND (v3.2)

Would-have side: YES @ 7¢ (bid 5 / ask 7 · vol24 54,851 · OI 35,173 · 242 fillable at 5–7)
My probability: ~18% vs market 7% — positive modeled edge, rejected anyway
Resolver: arena.ai text leaderboard, Rank (UB), "Remove Style Control" view, snapshot Jul 31 10:00 ET; ties break by Arena Score. Priced risks: UI drift (the toggle the rules name no longer exists — arena.ai now has an "Adjustments" dropdown; on the default view #1 is claude-fable-5, on no-style-control it's claude-opus-5-max) −8 pts; low-vote score volatility −6 pts.

This is the market that explains the week. Anthropic launched Claude Opus 5 on Jul 24 (TechCrunch, Axios, Fortune); Arena listed the variants on Jul 27; and the no-style-control board fetched directly today reads: #1 claude-opus-5-max 1512±12 (2,386 votes), #2 claude-opus-5-high 1505±8 (6,159 votes), #3 claude-opus-4-6-thinking 1503±4. The winning string — opus-5-maxis not a listed leg, which is why a mutually-exclusive event sums to ~12¢ while company-level "Anthropic top" trades 99¢, and why the old favorite (opus-4-6-thinking) collapsed from ~80¢ to 2¢ in a week. With a 7-point gap, overlapping CIs, and both scores built on tiny vote counts, opus-5-high overtaking its sibling by Friday is far likelier than 7% — my estimate ~18%. It is still a hard reject: sub-35¢ tail (the band went 0-for across the resolved record; v3.2 removed the lottery slot), stacked on a resolver whose UI no longer matches its own rules text. Logged and shadow-tracked — if it hits, the record will say so. This family is also the exact shape of the program's costliest historical loss (right about the world, wrong about a leaderboard snapshot).

KXTECHRANKLISTAICODE-26AUG03-CLAU — Claude #1 on LM Code Arena, Aug 3, 10:00 ET · REJECTED

Would-have side: YES @ 95¢ (bid 92 / ask 95, but only 21 contracts at 95, then 301 at 96 · vol24 1,351 — under screen)
My probability: ~94% vs ~95–96¢ effective — edge ≤ 0
Resolver: arena.ai Code leaderboard, brand-level ("Claude", any model), Aug 3 10:00 ET. Priced risks: two low-vote scores on top (opus-5-max 1725 preliminary; kimi-k3-max 1682, listed Jul 16) −4 pts; new-entrant listing before Aug 3 −2 pts.

The board today: Claude holds #1 (opus-5-max, 1725), #3 (opus-5-high, 1670) and #4 (fable-5, 1629); the only live threat is Moonshot's kimi-k3-max at 1682. For NO, Kimi must finish above every Claude — it already sits above two of them, so the thesis reduces to "opus-5-max's preliminary 1725 regresses ~45 points in six days while opus-5-high stays put." Plausible enough that 94% is my honest number — which the ask already pays for. Deep-favorite band, zero edge, thin book: pass. Gemini 3.5 Pro landing and topping the code board by Aug 3 is a rounding-error risk (Opus 5 took 3 days just to list).

KXFRONTIER-FRONB-26SEP01 — a third FrontierMath Open Problem solved before Sep 1 · REJECTED — COIN-FLIP BAND

Would-have side: YES @ 41¢ (bid 40 / ask 41 · vol24 1,210 — under screen · OI 1,025)
My probability: ~45% vs market 41% — ~4¢ edge, far below the ≥15¢ the band demands
Resolver: Epoch AI's FrontierMath Open Problems status page; "solved" means a new distinct problem flips (re-solves of the Ramsey-hypergraphs or 2-adic-Galois problems don't count, per rules). Priced risks: Epoch's confirmation lag is the whole game — solve #2's underlying solutions date to Jun 9 (Claude Fable 5) and Jun 24 (GPT-5.5 Pro), but the status only flipped ~Jul 27 −7 pts; "solved as stated" may require more than the partial construction under review −5 pts.

Genuinely interesting and honestly coin-flip. The concrete catalyst: on the Book Graphs problem, Dualverse AI's infinite-family construction has been under author review since Jun 25 — a positive review plus an Epoch status flip before Sep 1 resolves YES. Against it: two solves in six months is the entire base rate, Epoch is deliberately slow (5–7 weeks solution→flip), and the review has already run a month without news. 45% vs 41¢ is not a v3 trade: the 35–60¢ band requires HIGH conviction and ≥15¢, and this is MEDIUM at best. This is the band that hit ~33% historically. Shadow-tracked.

KXCHINAAI-26JUL31-ALIB — Alibaba top Chinese AI company on Arena (no style control), Jul 31 · REJECTED

Would-have side: YES @ 84¢ (bid 79 / ask 84 · vol24 926 — under screen · 33 contracts at 84, then 85/87)
My probability: not formed — data gap

Priced 81/17 Alibaba/Moonshot, but the one direct observation I have cuts the other way: on the default text board, Moonshot's kimi-k3-max (1486±10) is the highest-ranked Chinese model in the top 15 and no Qwen appears at all. Resolution uses the no-style-control Score view, where I could not see below the top 10 — so either the crowd knows Qwen leads on that specific view, or a 17¢ Moonshot leg is the mispricing. Unverifiable in the time available, volume under screen, book gaps beyond 33 contracts: reject with the honest label "couldn't check the resolver's actual view." Logged without a probability.

KXCHAICUTS-26AUG06-T1 — "Artificial Intelligence" is the #1 job-cut reason in Challenger's July report · REJECTED — COIN-FLIP BAND + LIQUIDITY

Would-have side: YES @ 51¢ (bid 46 / ask 51 · vol24 87 — Stage-1 fail, diligenced anyway)
My probability: ~58% vs market ~51% — ~7¢ edge, below the band's bar
Resolver: Table 4 of the initially-published July 2026 Challenger report (~Aug 6); ties resolve YES; revisions ignored.

AI has been the #1 reason four straight months — May was a blowout (38,579 cuts, ~40% of all cuts; Fox Business, CNBC), but June narrowed sharply: AI 14,029 vs Market/Economic Conditions 12,470 (Challenger June report). A fifth month is better-than-even, not safe; one seasonal restructuring wave flips a 1,559-cut margin. 58% at 51¢ in the coin-flip band without HIGH conviction fails Stage 3.5 — and the 87-contract daily volume fails Stage 1 before that. Logged for the shadow track.

KXGEMINI-GEMI35P-26AUG16 — Gemini 3.5 Pro (or greater) released before Aug 16 · REJECTED — COIN-FLIP BAND + LIQUIDITY

Would-have side: NO @ ~55¢ effective (YES bid 45 / ask 50 · vol24 111 — Stage-1 fail, diligenced anyway)
My probability (NO): ~62% vs ~55¢ — ~7¢ edge, below the band's bar
Resolver: public release (closed beta excluded; high-cost tier acceptable) of a model called Gemini 3.5 Pro or greater — note the "or greater" widens the YES door.

The NO case is well-sourced: three missed targets (June GA → mid-July → Jul 17), Bloomberg's Jul 16 report of a scrapped-and-rebuilt base model over hallucination rates, and Jul 21's stopgap Flash releases relieving launch pressure (TechCrunch). The YES case is one date: the Aug 12 Made-by-Google event sits four days inside the window and is a natural stage. That tension is a true coin-flip leaning NO — exactly what v3 refuses to fund at 7¢ of edge. The Aug 31 sibling (71/78) prices the "or greater by end of August" door more richly and is wider still. Logged.

6. Sources