Kalshi AI — Mispricing Audit
2026-07-21 · v3.3 methodology (conviction-weighted, EV-agnostic sizing; Stage 2.5 resolver pricing) · $2,500 capital · 45-day horizon · picks count toward the public track record and are cleared for auto-copy
1. How this was researched
Mode: theme-match. Kalshi has no AI category, so I selected markets by keyword across all categories, then curated by hand. SQL terms searched over trading_events.title and trading_markets.title: openai, anthropic, claude, "ai", artificial intelligence, chatgpt, gpt, gemini, grok, xai, nvidia, deepseek, llama, llm, superintelligence, sam altman, robot, waymo, leaderboard, arena — with Sports/Crypto/Weather/Mentions categories excluded to kill false hits (player names, etc.). This surfaced ~36 AI-theme events: model-release timing (Claude Opus/Claude 5/Mythos, GPT-6/5.6, Gemini 3.5 Pro), leaderboard snapshots (Best AI, Top Coding/Math AI, Best Chinese AI), NVIDIA GPU compute-price ladders (A100/B200/H100/H200/RTX), AI policy (Anthropic DoD designation, Fable 5 US access, Trump AI-review EO), and corporate events (Anthropic/OpenAI/Waymo IPO announcements).
Pipeline per the v3 flow: live quotes and orderbooks from the Kalshi public API; rules text from the DB mirror (the mirror carries no market_snapshots table, so 14-day history came from Kalshi candlesticks); primary-source news via web search; every surviving candidate got a Stage 2.5 resolver profile with risks priced as explicit point deductions. Six tickers already live in the feed (Gemini 3.5 Pro Jul 24/31 NO, Arena Math Gemini NO, Top-model CLAUF YES / CLAUT NO, RTX 5090 NO) were excluded from the candidate set up front.
World-state anchors established during diligence (from finalized Kalshi markets + primary sources): Claude 5 family released (market resolved YES); Fable 5 + Mythos 5 shipped Jun 9, Sonnet 5 Jun 30, Opus 4.8 May 28; GPT-5.6 went GA Jul 9; Fable 5 US access restored before Jul 24 (resolved YES).
2. Markets reviewed
| Market | Question | Price (YES) | v24 / OI | Outcome of review |
|---|---|---|---|---|
| KXGPT-OPEN-26SEP01 | GPT-6 public before Sep 1 | 25 / 26 | 1,903 / 15,362 | PICK — BUY NO @ 75 |
| KXCODEAI-26JUL31-CLAU | Claude tops Datacurve DeepSWE Jul 31 | 26 / 31 | 2,025 / 28,370 | Reject — fairly priced (+2¢) |
| KXCODEAI-26JUL31-CHAT | ChatGPT tops Datacurve DeepSWE Jul 31 | 62 / 70 | 875 / 14,826 | Reject — fairly priced (+3¢) |
| KXMATHAI-26JUL31-CLAU | Claude #1 Arena Math Jul 31 | 88 / 93 | 1,515 / 13,777 | Reject — negative edge at ask |
| KXCHINAAI-26JUL27-ALIB | Alibaba best Chinese AI (Arena RSC) Jul 27 | 74 / 75 | 1,670 / 1,438 | Reject — resolver view unverifiable |
| KXCLAUDE-OPUS-26JUL31 | Next Claude Opus before Jul 31 | 77 / 81 | 683 / 1,038 | Reject — rumor-priced, no edge |
| KXCLAUDE-OPUS-26JUL24 | Next Claude Opus before Jul 24 | 38 / 79 | 1,962 / 2,397 | Reject — 41¢ spread unfillable |
| KXCLAUDE-OPUS-26AUG14 | Next Claude Opus before Aug 14 | 76 / 89 | 14 / 533 | Reject — illiquid, 13¢ spread |
| KXCLAUDE-NXTMYTH-26SEP01 | Next Mythos-class model before Sep 1 | 37 / 43 | 0 / 3,845 | Reject — low conviction, definitional |
| KXCLAUDE-MYTH-26SEP01 | Mythos public release before Sep 1 | 5 / 6 | 1,257 / 51,632 | Reject — 1¢ edge in >90 band |
| KXB200WS-26JUL24-7.000 | B200 > $7/hr Jul 24 (Ornn) | 57 / 58 | 1,473 / 1,318 | Reject — coin-flip, index unreadable |
| KXCHAICUTS-26AUG06-T1 | AI #1 reason, Challenger July report | 45 / 51 | 0 / 1,129 | Reject — coin-flip band, no volume |
| KXFRONTIER-FRON-26SEP01 | AI solves Frontier Math problem by Sep 1 | 30 / 34 | 204 / 2,109 | Reject — low conviction |
| KXGEMINI-GEMI35P-26AUG31 | Gemini 3.5 Pro before Aug 31 | 71 / 77 | 629 / 1,540 | Reject — no timing edge |
| KXTECHRANKLISTAICODE-26JUL27-KIMI | Kimi #1 LM Code Arena Jul 27 | 94 / 95 | 2,309 / 1,472 | Reject — fully priced |
| KXANTHROPICRESCIND-26JUN-26AUG01 | Anthropic DoD designation rescinded by Aug 1 | 8 / 14 | 40 / 1,946 | Reject — sub-35¢ tail (hard band) |
| KXTOPMODEL-26JUL31-CLAUT | opus-4-6-thinking top-ranked Jul 31 | 90 / 93 | 9,744 / 21,875 | Excluded — ticker already a live feed pick |
| KXLLM1-26JUL31-A | Claude best AI in July | 98 / 99 | 8,930 / 63,259 | Stage-1 skip — fully priced (≥95) |
Stage-1 mechanical cuts not shown: ~150 GPU-ladder strikes with near-zero 24h volume, IPO-announcement deep-NO markets (≤3¢ asks), weekly leaderboard legs at 0–3¢, and the Fable-5-disabled / AI-review-EO tails.
3. The pick
Pick 1 — KXGPT-OPEN-26SEP01 · BUY NO @ 0.75 · MEDIUM CONVICTION
Will OpenAI release GPT-6 before Sep 1, 2026? — closes 2026-09-01T03:59Z (41 days)
Mispricing thesis. The market gives a 25–26% chance that OpenAI takes a generational GPT-6 to public release within 41 days — six weeks after GPT-5.6 went generally available on Jul 9, and with no GPT-6 announcement, model card, or date in existence. The crowd read the 5.6 launch as evidence of an accelerating ramp toward 6 (YES spiked 12¢→39¢ on launch day and held ~40¢ for a week); the release-cadence evidence says the opposite — a fresh flagship GA is what OpenAI ships instead of a generational jump.
Evidence.
- GPT-5.6 entered limited preview Jun 26, 2026 and went GA Jul 9, 2026 (TechCrunch, Jul 9; CNBC, Jul 8; Wikipedia: GPT-5.6).
- As of Jul 11, 2026: "still no public release of ChatGPT 6 / GPT-6" — no announcement, no model card, no date (Fello AI GPT-6 tracker).
- 2026 cadence has been point releases inside the GPT-5 family: 5.4 → 5.5 (Apr 23) → 5.6 (Jul 9), ~2.5 months apart. On-cadence, the next drop lands late September — and would more plausibly be a 5.7 (FindSkill tracker: "5.6 shipped, GPT-6 in Q4?"; tracker consensus puts GPT-6 late-2026/2027 with December modal).
- The resolver already demonstrated its definition on this exact series: GPT-5.6's government-restricted June preview did not resolve the GPT-5.6 market; the Jul 9 public GA did (KXGPT-OPENB-26JUL31 finalized YES). A gated GPT-6 preview before Sep 1 would not count — only public availability does.
Stage 2.5 — resolver profile. Rules: "If OpenAI releases a model called GPT-6 or greater before Sep 1, 2026 → YES"; secondary: "Release must be to the public, outside of a closed beta, though limiting it to a high-cost subscription tier is acceptable." Mechanism: first-occurrence release determination from public announcements; no snapshot exposure, no third-party scoreboard. Definitional boundary is naming ("GPT-6 or greater") plus public availability. Priced risks: surprise generational launch under competitive pressure from Anthropic's Claude-5 sweep of the leaderboards, −2 pts (5.6's preview→GA took only 2 weeks, so an early-August announcement could still make it); naming stunt / "GPT-6-preview" public tier, −1 pt. Evidence-based ~93% → final model_prob 90%.
Tail risk (cleanest single loser). OpenAI announces GPT-6 at a surprise August event and pushes it to public GA within ~2–3 weeks, replicating the 5.6 preview-to-GA speed. That is the whole bear case, and it is real — which is why this is MEDIUM, not HIGH.
Price history & entry context. 14-day candles: 11¢ (Jul 8) → 39¢ on 5.6 launch day (Jul 9, 7.4k vol) → 35–42¢ plateau for a week → faded to 25–26¢ over Jul 18–21 as no announcement materialized. Honest caveat per the 48h screen: the last 3 days moved in my direction (NO 62¢→75¢), so part of the edge is already eaten; what remains is still +15¢ against my number. Book at entry: YES 25 bid / 26 ask, spread 1¢; NO fillable ≤75¢: ~2,246 contracts (160 @ 74¢ + 2,086 @ 75¢) — the 233-lot fills at ≤75¢ with no slippage. v24 1,903; OI 15,362.
4. Recommended $2,500 portfolio
| # | Market | Action | Limit | Contracts | Cost | Conviction / band | Max payout | EV¢ / EV%* |
|---|---|---|---|---|---|---|---|---|
| 1 | KXGPT-OPEN-26SEP01 | BUY NO | 75¢ | 233 | $174.75 | MEDIUM · 60–90¢ favorite | $233.00 | +15¢ / +20% |
| — | Cash reserve | HOLD | — | — | $2,325.25 | — | — | — |
| Totals | $174.75 deployed (7.0%) · $2,325.25 cash (93.0%) | — | $233.00 | blended +20% on deployed | ||||
*EV figures are recorded for calibration only. Per v3, position size was set by conviction tier (MEDIUM → ≤7% of capital) and entry band — not by EV. Dollar edge on deployed capital: ≈ +$34.95 expected (0.90 × $233 − $174.75).
Cluster exposure (cap: 15% of capital per cluster)
| Cluster | Cost | % of capital | Cap check |
|---|---|---|---|
| gpt6-release-timing | $174.75 | 7.0% | OK (<15%) |
Conviction exposure
| Tier | Cost deployed | % of capital |
|---|---|---|
| HIGH | $0 | 0% |
| MEDIUM | $174.75 | 7.0% |
| LOW | $0 (never deploys, v3.3) | 0% |
Risk profile
- Worst case: GPT-6 goes public before Sep 1 → book loses $174.75 (−7.0% of capital). No single event can flip the run's sign beyond that; 93% of capital is never at risk.
- Best case / most likely: no GPT-6 by Sep 1 → +$58.25 (+2.3% on total capital, +33% on deployed cost). My 90% estimate makes this the modal outcome.
- Concentration: a single thesis, deliberately. The alternative was padding with coin-flips and unverifiable leaderboard snapshots — the exact shapes that produced the historical losses.
Execution notes
- Work a resting NO limit at 75¢ (do not cross above it): 160 contracts clear at 74¢, the balance at 75¢; the 233-lot fills without slippage against ~2,246 contracts of visible depth.
- Invalidation triggers — exit or stop adding: (a) any official OpenAI communication naming GPT-6 with a date or a public waitlist; (b) an OpenAI event announcement for August with generational framing; (c) YES re-pricing above 40¢ on volume, which would imply news I haven't seen.
- Watchlist for opportunistic adds from the cash reserve: KXCODEAI-26JUL31-CLAU NO becomes attractive above ~80¢ effective if the next Opus ships after ~Jul 28 (too late for a Datacurve eval before the Jul 31 10:00 ET snapshot); KXMATHAI-26JUL31-CLAU YES becomes attractive at ≤88¢ if Gemini 3.5 Pro is still unreleased by ~Jul 27.
5. What I rejected and why
All fifteen rejects below are machine-logged in picks.json and shadow-tracked to settlement — if they outperform the pick, the screen adds nothing and we want to know. The dominant reject reasons this week were Stage 2.5 resolver risk (leaderboard-snapshot markets whose exact resolver view — LMArena with Remove Style Control, Datacurve DeepSWE, the Ornn GPU index — I could not independently verify or could not price within the 15-pt deduction cap) and the entry-band screen.
Rejected on resolver risk / no edge after deductions
- KXCODEAI-26JUL31-CLAU (would-be NO @ 74¢, my prob 76%) — the most interesting non-pick. Resolver is Datacurve DeepSWE (not LM Arena): current board has OpenAI's gpt-5-6-sol[max] #1 at 72.7%, claude-fable-5[max] #2 at 69.7% (BenchLM DeepSWE snapshot, Jul 17). Claude YES at 26–31¢ looks rich until you price the channel: Kalshi gives the next Opus ~80% by Jul 31, rumors claim big agentic-coding gains, and Datacurve posted kimi-k3 one day after its Jul 16 launch — so a Jul-2x Opus release plausibly gets evaluated before the Jul 31 snapshot. Decomposed: P(release in time ~70%) × P(evaluated in time ~75%) × P(beats 72.7% ~40%) + re-eval risk ≈ 24% Claude-tops. Market prices 26–31%. Edge +2¢. Correctly priced — reject.
- KXCODEAI-26JUL31-CHAT (would-be YES @ 70¢, my prob 73%) — same event, same math from the other side, plus it also loses to a Kimi jump. +3¢ is not an edge; same cluster as above.
- KXMATHAI-26JUL31-CLAU (would-be YES @ 93¢ ask, my prob 87%) — Arena Math #1 is claude-opus-4-6-thinking today, and any Claude counts (brand-level). But this is the exact market family that produced the program's costliest historical loss, and the deductions are live: Gemini 3.5 Pro launch + Arena listing before the Jul 31 snapshot −7, Kimi K3 (5 days of votes, math-strong lineage) drift −4, GPT-5.6 vote accumulation −2. Fair ≈ 87 vs 93 ask = negative edge; the Math-category margin at the resolver's Remove-Style-Control view could not be independently read (mirrors show conflicting, style-controlled views).
- KXCHINAAI-26JUL27-ALIB (would-be YES @ 75¢) — resolver is Arena Text with Remove Style Control, highest Chinese company, snapshot Jul 27 10:00 ET. Qwen 3.7 Max leads Chinese entrants today, but Kimi K3 (launched Jul 16, "largest open-weight model ever") is still accumulating votes and mirrors disagree about the order (Swfte vs DataLearner). Unpriceable within the deduction cap → Stage 2.5 reject.
- KXCLAUDE-OPUS-26JUL31 (would-be YES @ 81¢, my prob 80%) — the market's ~80% is built on leaks, not an announcement: "Honeycomb" strings in Cursor (Jul 9), Vertex AI signals (Jul 14), the Opus 4.7-fast removal scheduled Jul 24 (explainx rumor roundup; tokenmix signal trace). My estimate lands on top of the market's. No edge either side — and the NO side at 19–23¢ is a sub-35¢ tail, a hard reject band.
- KXCLAUDE-MYTH-26SEP01 (would-be NO @ 95¢, my prob 96%) — Mythos 5 shipped Jun 9 to approved organizations only; the market trading at 5–6¢ six weeks later (with early-close enabled) proves the resolver does not count restricted access as a public release. But that leaves ~1¢ of edge in the >90 band with residual definitional risk. Not worth capital.
- KXB200WS-26JUL24-7.000 (57/58¢) — resolver is the Ornn B200 index at 4 PM ET Jul 24; the dashboard is JS-only and no mirror publishes the live index. Adjacent strikes imply spot ≈ $7.00 exactly, in a regime where B200 spot rose 114% in six weeks (Tunguz). A coin-flip on an unreadable index can never be HIGH conviction → band rule rejects it.
Cut by the entry-band / conviction screen
- KXANTHROPICRESCIND-26JUN-26AUG01 (YES @ 14¢) — sub-35¢ tail, hard reject with no exceptions (v3.2). Also 6¢ spread on 40 contracts of daily volume.
- KXCHAICUTS-26AUG06-T1 (45/51¢) — coin-flip band without HIGH conviction (no primary read on Challenger's July category mix), zero 24h volume.
- KXCLAUDE-NXTMYTH-26SEP01 (37/43¢) — coin-flip band + a genuine definitional trap: if the rumored next Opus ships positioned as Mythos-class, NO loses on branding. Low conviction = logged reject (v3.3).
- KXFRONTIER-FRON-26SEP01 (30/34¢) — no defensible basis to out-forecast the market on research-breakthrough timing. Low conviction.
- KXGEMINI-GEMI35P-26AUG31 (71/77¢) — no independent edge on Google's ship date; 6¢ spread; the cluster already carries two live NO picks in the feed.
- KXTECHRANKLISTAICODE-26JUL27-KIMI (94/95¢) — effectively fully priced, with a live risk that a new Opus enters the LM Code Arena board before the Jul 27 snapshot.
Cut on liquidity
- KXCLAUDE-OPUS-26JUL24 — 38/79 quote: a 41¢ spread cannot be crossed within the 3¢ slippage budget.
- KXCLAUDE-OPUS-26AUG14 — 13¢ spread on 14 contracts of daily volume.
6. Sources
- BenchLM — DeepSWE leaderboard mirror (snapshot Jul 17, 2026)
- VentureBeat — DeepSWE benchmark and the Claude Opus loophole finding
- Datacurve DeepSWE (resolution source for KXCODEAI)
- TechCrunch — OpenAI launches GPT-5.6 family (Jul 9, 2026)
- CNBC — OpenAI to publicly release GPT-5.6, ending government limits (Jul 8, 2026)
- Axios — GPT-5.6 Sol/Terra/Luna restricted preview (Jun 26, 2026)
- Wikipedia — GPT-5.6
- Fello AI — "Still No GPT-6" tracker (Jul 11, 2026)
- FindSkill — GPT-6 release-date tracker (Q4 modal)
- LifeArchitect — GPT-6 (2026)
- ScriptByAI — Anthropic Claude release timeline (Fable 5 / Mythos 5 Jun 9; Sonnet 5 Jun 30; Opus 4.8 May 28)
- explainx — Claude Opus 5 release rumors (July 2026)
- tokenmix — tracing the four Opus-5 signals (Honeycomb, Vertex)
- Swfte — LMArena leaderboard mirror (Jul 20, 2026)
- DataLearner — AI model leaderboard mirror (Jul 16, 2026)
- Tunguz — GPU spot prices surge 114% in six weeks
- Ornn compute index (resolution source for GPU ladders; JS-only)