# Kalshi AI — Mispricing Audit — 2026-07-21 Date: 2026-07-21 Source: https://kalshi-audits.pages.dev/2026-07-21-kalshi-ai-mispricing-audit --- # Kalshi AI — Mispricing Audit 2026-07-21 · v3.3 methodology (conviction-weighted, EV-agnostic sizing; Stage 2.5 resolver pricing) · $2,500 capital · 45-day horizon · picks count toward the public track record and are cleared for auto-copy TLDR: The AI category is efficient this week. Of 18 markets diligenced, one pick survives every screen: KXGPT-OPEN-26SEP01 BUY NO @ 75¢ (GPT-6 will not get a public release before Sep 1). $174.75 deployed (7.0%), $2,325.25 held as cash (93%). Fifteen diligenced candidates are logged as shadow-tracked rejects — most cut by resolver-risk pricing (Stage 2.5) or the entry-band screen. ## 1. How this was researched **Mode: theme-match.** Kalshi has no `AI` category, so I selected markets by keyword across all categories, then curated by hand. SQL terms searched over `trading_events.title` and `trading_markets.title`: _openai, anthropic, claude, "ai", artificial intelligence, chatgpt, gpt, gemini, grok, xai, nvidia, deepseek, llama, llm, superintelligence, sam altman, robot, waymo, leaderboard, arena_ — with Sports/Crypto/Weather/Mentions categories excluded to kill false hits (player names, etc.). This surfaced ~36 AI-theme events: model-release timing (Claude Opus/Claude 5/Mythos, GPT-6/5.6, Gemini 3.5 Pro), leaderboard snapshots (Best AI, Top Coding/Math AI, Best Chinese AI), NVIDIA GPU compute-price ladders (A100/B200/H100/H200/RTX), AI policy (Anthropic DoD designation, Fable 5 US access, Trump AI-review EO), and corporate events (Anthropic/OpenAI/Waymo IPO announcements). Pipeline per the v3 flow: live quotes and orderbooks from the Kalshi public API; rules text from the DB mirror (the mirror carries no `market_snapshots` table, so 14-day history came from Kalshi candlesticks); primary-source news via web search; every surviving candidate got a Stage 2.5 _resolver profile_ with risks priced as explicit point deductions. Six tickers already live in the feed (Gemini 3.5 Pro Jul 24/31 NO, Arena Math Gemini NO, Top-model CLAUF YES / CLAUT NO, RTX 5090 NO) were excluded from the candidate set up front. **World-state anchors established during diligence** (from finalized Kalshi markets + primary sources): Claude 5 family released (market resolved YES); Fable 5 + Mythos 5 shipped Jun 9, Sonnet 5 Jun 30, Opus 4.8 May 28; GPT-5.6 went GA Jul 9; Fable 5 US access restored before Jul 24 (resolved YES). ## 2. Markets reviewed | Market | Question | Price (YES) | v24 / OI | Outcome of review | |---|---|---|---|---| | KXGPT-OPEN-26SEP01 | GPT-6 public before Sep 1 | 25 / 26 | 1,903 / 15,362 | PICK — BUY NO @ 75 | | KXCODEAI-26JUL31-CLAU | Claude tops Datacurve DeepSWE Jul 31 | 26 / 31 | 2,025 / 28,370 | Reject — fairly priced (+2¢) | | KXCODEAI-26JUL31-CHAT | ChatGPT tops Datacurve DeepSWE Jul 31 | 62 / 70 | 875 / 14,826 | Reject — fairly priced (+3¢) | | KXMATHAI-26JUL31-CLAU | Claude #1 Arena Math Jul 31 | 88 / 93 | 1,515 / 13,777 | Reject — negative edge at ask | | KXCHINAAI-26JUL27-ALIB | Alibaba best Chinese AI (Arena RSC) Jul 27 | 74 / 75 | 1,670 / 1,438 | Reject — resolver view unverifiable | | KXCLAUDE-OPUS-26JUL31 | Next Claude Opus before Jul 31 | 77 / 81 | 683 / 1,038 | Reject — rumor-priced, no edge | | KXCLAUDE-OPUS-26JUL24 | Next Claude Opus before Jul 24 | 38 / 79 | 1,962 / 2,397 | Reject — 41¢ spread unfillable | | KXCLAUDE-OPUS-26AUG14 | Next Claude Opus before Aug 14 | 76 / 89 | 14 / 533 | Reject — illiquid, 13¢ spread | | KXCLAUDE-NXTMYTH-26SEP01 | Next Mythos-class model before Sep 1 | 37 / 43 | 0 / 3,845 | Reject — low conviction, definitional | | KXCLAUDE-MYTH-26SEP01 | Mythos public release before Sep 1 | 5 / 6 | 1,257 / 51,632 | Reject — 1¢ edge in >90 band | | KXB200WS-26JUL24-7.000 | B200 > $7/hr Jul 24 (Ornn) | 57 / 58 | 1,473 / 1,318 | Reject — coin-flip, index unreadable | | KXCHAICUTS-26AUG06-T1 | AI #1 reason, Challenger July report | 45 / 51 | 0 / 1,129 | Reject — coin-flip band, no volume | | KXFRONTIER-FRON-26SEP01 | AI solves Frontier Math problem by Sep 1 | 30 / 34 | 204 / 2,109 | Reject — low conviction | | KXGEMINI-GEMI35P-26AUG31 | Gemini 3.5 Pro before Aug 31 | 71 / 77 | 629 / 1,540 | Reject — no timing edge | | KXTECHRANKLISTAICODE-26JUL27-KIMI | Kimi #1 LM Code Arena Jul 27 | 94 / 95 | 2,309 / 1,472 | Reject — fully priced | | KXANTHROPICRESCIND-26JUN-26AUG01 | Anthropic DoD designation rescinded by Aug 1 | 8 / 14 | 40 / 1,946 | Reject — sub-35¢ tail (hard band) | | KXTOPMODEL-26JUL31-CLAUT | opus-4-6-thinking top-ranked Jul 31 | 90 / 93 | 9,744 / 21,875 | Excluded — ticker already a live feed pick | | KXLLM1-26JUL31-A | Claude best AI in July | 98 / 99 | 8,930 / 63,259 | Stage-1 skip — fully priced (≥95) | Stage-1 mechanical cuts not shown: ~150 GPU-ladder strikes with near-zero 24h volume, IPO-announcement deep-NO markets (≤3¢ asks), weekly leaderboard legs at 0–3¢, and the Fable-5-disabled / AI-review-EO tails. ## 3. The pick ### Pick 1 — KXGPT-OPEN-26SEP01 · BUY NO @ 0.75 · MEDIUM CONVICTION _Will OpenAI release GPT-6 before Sep 1, 2026?_ — closes 2026-09-01T03:59Z (41 days) Current price (YES)25 bid / 26 ask → NO fillable at 74–75¢ My probability (NO)90%  vs  market-implied 74–75% Entry band75¢ effective — **60–90¢ favorite band** (the deployable band) Edge+15¢/contract, +20% EV (recorded for calibration only — did NOT set size) Size233 contracts × 75¢ = $174.75 (7.0% of capital — MEDIUM tier cap) Max payout$233.00 (profit +$58.25 if NO) Clustergpt6-release-timing **Mispricing thesis.** The market gives a 25–26% chance that OpenAI takes a _generational_ GPT-6 to public release within 41 days — six weeks after GPT-5.6 went generally available on Jul 9, and with no GPT-6 announcement, model card, or date in existence. The crowd read the 5.6 launch as evidence of an accelerating ramp toward 6 (YES spiked 12¢→39¢ on launch day and held ~40¢ for a week); the release-cadence evidence says the opposite — a fresh flagship GA is what OpenAI ships _instead of_ a generational jump. **Evidence.** - GPT-5.6 entered limited preview Jun 26, 2026 and went GA Jul 9, 2026 ([TechCrunch, Jul 9](https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/); [CNBC, Jul 8](https://www.cnbc.com/2026/07/08/openai-expanding-gpt-5point6-ai-model-release-ending-government-limits.html); [Wikipedia: GPT-5.6](https://en.wikipedia.org/wiki/GPT-5.6)). - As of Jul 11, 2026: "still no public release of ChatGPT 6 / GPT-6" — no announcement, no model card, no date ([Fello AI GPT-6 tracker](https://felloai.com/all-we-know-about-chatgpt-6/)). - 2026 cadence has been point releases inside the GPT-5 family: 5.4 → 5.5 (Apr 23) → 5.6 (Jul 9), ~2.5 months apart. On-cadence, the next drop lands late September — and would more plausibly be a 5.7 ([FindSkill tracker](https://findskill.ai/blog/gpt-6-release-date/): "5.6 shipped, GPT-6 in Q4?"; tracker consensus puts GPT-6 late-2026/2027 with December modal). - The resolver already demonstrated its definition on this exact series: GPT-5.6's government-restricted June preview did _not_ resolve the GPT-5.6 market; the Jul 9 public GA did (KXGPT-OPENB-26JUL31 finalized YES). A gated GPT-6 preview before Sep 1 would not count — only public availability does. **Stage 2.5 — resolver profile.** Rules: "If OpenAI releases _a model called GPT-6 or greater_ before Sep 1, 2026 → YES"; secondary: "Release must be to the public, outside of a closed beta, though limiting it to a high-cost subscription tier is acceptable." Mechanism: first-occurrence release determination from public announcements; no snapshot exposure, no third-party scoreboard. Definitional boundary is naming ("GPT-6 or greater") plus public availability. _Priced risks:_ surprise generational launch under competitive pressure from Anthropic's Claude-5 sweep of the leaderboards, −2 pts (5.6's preview→GA took only 2 weeks, so an early-August announcement could still make it); naming stunt / "GPT-6-preview" public tier, −1 pt. Evidence-based ~93% → **final model_prob 90%**. **Tail risk (cleanest single loser).** OpenAI announces GPT-6 at a surprise August event and pushes it to public GA within ~2–3 weeks, replicating the 5.6 preview-to-GA speed. That is the whole bear case, and it is real — which is why this is MEDIUM, not HIGH. **Price history & entry context.** 14-day candles: 11¢ (Jul 8) → 39¢ on 5.6 launch day (Jul 9, 7.4k vol) → 35–42¢ plateau for a week → faded to 25–26¢ over Jul 18–21 as no announcement materialized. Honest caveat per the 48h screen: the last 3 days moved _in my direction_ (NO 62¢→75¢), so part of the edge is already eaten; what remains is still +15¢ against my number. Book at entry: YES 25 bid / 26 ask, spread 1¢; NO fillable ≤75¢: ~2,246 contracts (160 @ 74¢ + 2,086 @ 75¢) — the 233-lot fills at ≤75¢ with no slippage. v24 1,903; OI 15,362. ## 4. Recommended $2,500 portfolio | # | Market | Action | Limit | Contracts | Cost | Conviction / band | Max payout | EV¢ / EV%* | |---|---|---|---|---|---|---|---|---| | 1 | KXGPT-OPEN-26SEP01 | BUY NO | 75¢ | 233 | $174.75 | MEDIUM · 60–90¢ favorite | $233.00 | +15¢ / +20% | | — | Cash reserve | HOLD | — | — | $2,325.25 | — | — | — | | Totals | $174.75 deployed (7.0%) · $2,325.25 cash (93.0%) | — | $233.00 | blended +20% on deployed | | | | | *EV figures are recorded for calibration only. Per v3, position size was set by conviction tier (MEDIUM → ≤7% of capital) and entry band — not by EV. Dollar edge on deployed capital: ≈ +$34.95 expected (0.90 × $233 − $174.75). ### Cluster exposure (cap: 15% of capital per cluster) | Cluster | Cost | % of capital | Cap check | |---|---|---|---| | gpt6-release-timing | $174.75 | 7.0% | OK (<15%) | ### Conviction exposure | Tier | Cost deployed | % of capital | |---|---|---| | HIGH | $0 | 0% | | MEDIUM | $174.75 | 7.0% | | LOW | $0 (never deploys, v3.3) | 0% | ### Risk profile - **Worst case:** GPT-6 goes public before Sep 1 → book loses $174.75 (−7.0% of capital). No single event can flip the run's sign beyond that; 93% of capital is never at risk. - **Best case / most likely:** no GPT-6 by Sep 1 → +$58.25 (+2.3% on total capital, +33% on deployed cost). My 90% estimate makes this the modal outcome. - **Concentration:** a single thesis, deliberately. The alternative was padding with coin-flips and unverifiable leaderboard snapshots — the exact shapes that produced the historical losses. ### Execution notes - Work a resting **NO limit at 75¢** (do not cross above it): 160 contracts clear at 74¢, the balance at 75¢; the 233-lot fills without slippage against ~2,246 contracts of visible depth. - **Invalidation triggers — exit or stop adding:** (a) any official OpenAI communication naming GPT-6 with a date or a public waitlist; (b) an OpenAI event announcement for August with generational framing; (c) YES re-pricing above 40¢ on volume, which would imply news I haven't seen. - **Watchlist for opportunistic adds from the cash reserve:** KXCODEAI-26JUL31-CLAU NO becomes attractive above ~80¢ effective if the next Opus ships after ~Jul 28 (too late for a Datacurve eval before the Jul 31 10:00 ET snapshot); KXMATHAI-26JUL31-CLAU YES becomes attractive at ≤88¢ if Gemini 3.5 Pro is still unreleased by ~Jul 27. ## 5. What I rejected and why All fifteen rejects below are machine-logged in `picks.json` and shadow-tracked to settlement — if they outperform the pick, the screen adds nothing and we want to know. The dominant reject reasons this week were **Stage 2.5 resolver risk** (leaderboard-snapshot markets whose exact resolver view — LMArena with Remove Style Control, Datacurve DeepSWE, the Ornn GPU index — I could not independently verify or could not price within the 15-pt deduction cap) and the **entry-band screen**. ### Rejected on resolver risk / no edge after deductions - **KXCODEAI-26JUL31-CLAU (would-be NO @ 74¢, my prob 76%)** — the most interesting non-pick. Resolver is _Datacurve DeepSWE_ (not LM Arena): current board has OpenAI's gpt-5-6-sol[max] #1 at 72.7%, claude-fable-5[max] #2 at 69.7% ([BenchLM DeepSWE snapshot, Jul 17](https://benchlm.ai/benchmarks/deepSwe)). Claude YES at 26–31¢ looks rich until you price the channel: Kalshi gives the next Opus ~80% by Jul 31, rumors claim big agentic-coding gains, and Datacurve posted kimi-k3 _one day_ after its Jul 16 launch — so a Jul-2x Opus release plausibly gets evaluated before the Jul 31 snapshot. Decomposed: P(release in time ~70%) × P(evaluated in time ~75%) × P(beats 72.7% ~40%) + re-eval risk ≈ 24% Claude-tops. Market prices 26–31%. Edge +2¢. Correctly priced — reject. - **KXCODEAI-26JUL31-CHAT (would-be YES @ 70¢, my prob 73%)** — same event, same math from the other side, plus it also loses to a Kimi jump. +3¢ is not an edge; same cluster as above. - **KXMATHAI-26JUL31-CLAU (would-be YES @ 93¢ ask, my prob 87%)** — Arena Math #1 is claude-opus-4-6-thinking today, and any Claude counts (brand-level). But this is the exact market family that produced the program's costliest historical loss, and the deductions are live: Gemini 3.5 Pro launch + Arena listing before the Jul 31 snapshot −7, Kimi K3 (5 days of votes, math-strong lineage) drift −4, GPT-5.6 vote accumulation −2. Fair ≈ 87 vs 93 ask = negative edge; the Math-category margin at the resolver's Remove-Style-Control view could not be independently read (mirrors show conflicting, style-controlled views). - **KXCHINAAI-26JUL27-ALIB (would-be YES @ 75¢)** — resolver is Arena Text _with Remove Style Control_, highest Chinese company, snapshot Jul 27 10:00 ET. Qwen 3.7 Max leads Chinese entrants today, but Kimi K3 (launched Jul 16, "largest open-weight model ever") is still accumulating votes and mirrors disagree about the order ([Swfte](https://www.swfte.com/lmarena) vs [DataLearner](https://www.datalearner.com/en/leaderboards)). Unpriceable within the deduction cap → Stage 2.5 reject. - **KXCLAUDE-OPUS-26JUL31 (would-be YES @ 81¢, my prob 80%)** — the market's ~80% is built on leaks, not an announcement: "Honeycomb" strings in Cursor (Jul 9), Vertex AI signals (Jul 14), the Opus 4.7-fast removal scheduled Jul 24 ([explainx rumor roundup](https://explainx.ai/blog/claude-opus-5-release-speculation-july-2026); [tokenmix signal trace](https://dev.to/tokenmixai/i-traced-4-claude-opus-5-signals-the-release-date-still-isnt-real-yet-2f2j)). My estimate lands on top of the market's. No edge either side — and the NO side at 19–23¢ is a sub-35¢ tail, a hard reject band. - **KXCLAUDE-MYTH-26SEP01 (would-be NO @ 95¢, my prob 96%)** — Mythos 5 shipped Jun 9 to approved organizations only; the market trading at 5–6¢ six weeks later (with early-close enabled) proves the resolver does not count restricted access as a public release. But that leaves ~1¢ of edge in the >90 band with residual definitional risk. Not worth capital. - **KXB200WS-26JUL24-7.000 (57/58¢)** — resolver is the Ornn B200 index at 4 PM ET Jul 24; the dashboard is JS-only and no mirror publishes the live index. Adjacent strikes imply spot ≈ $7.00 exactly, in a regime where B200 spot rose 114% in six weeks ([Tunguz](https://tomtunguz.com/b200-gpu-pricing-spot-market-model-releases/)). A coin-flip on an unreadable index can never be HIGH conviction → band rule rejects it. ### Cut by the entry-band / conviction screen - **KXANTHROPICRESCIND-26JUN-26AUG01 (YES @ 14¢)** — sub-35¢ tail, hard reject with no exceptions (v3.2). Also 6¢ spread on 40 contracts of daily volume. - **KXCHAICUTS-26AUG06-T1 (45/51¢)** — coin-flip band without HIGH conviction (no primary read on Challenger's July category mix), zero 24h volume. - **KXCLAUDE-NXTMYTH-26SEP01 (37/43¢)** — coin-flip band + a genuine definitional trap: if the rumored next Opus ships positioned as Mythos-class, NO loses on branding. Low conviction = logged reject (v3.3). - **KXFRONTIER-FRON-26SEP01 (30/34¢)** — no defensible basis to out-forecast the market on research-breakthrough timing. Low conviction. - **KXGEMINI-GEMI35P-26AUG31 (71/77¢)** — no independent edge on Google's ship date; 6¢ spread; the cluster already carries two live NO picks in the feed. - **KXTECHRANKLISTAICODE-26JUL27-KIMI (94/95¢)** — effectively fully priced, with a live risk that a new Opus enters the LM Code Arena board before the Jul 27 snapshot. ### Cut on liquidity - **KXCLAUDE-OPUS-26JUL24** — 38/79 quote: a 41¢ spread cannot be crossed within the 3¢ slippage budget. - **KXCLAUDE-OPUS-26AUG14** — 13¢ spread on 14 contracts of daily volume. ## 6. Sources - [BenchLM — DeepSWE leaderboard mirror (snapshot Jul 17, 2026)](https://benchlm.ai/benchmarks/deepSwe) - [VentureBeat — DeepSWE benchmark and the Claude Opus loophole finding](https://venturebeat.com/technology/deepswe-blows-up-the-ai-coding-leaderboard-crowns-gpt-5-5-and-finds-claude-opus-exploiting-a-benchmark-loophole) - [Datacurve DeepSWE (resolution source for KXCODEAI)](https://deepswe.datacurve.ai/) - [TechCrunch — OpenAI launches GPT-5.6 family (Jul 9, 2026)](https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/) - [CNBC — OpenAI to publicly release GPT-5.6, ending government limits (Jul 8, 2026)](https://www.cnbc.com/2026/07/08/openai-expanding-gpt-5point6-ai-model-release-ending-government-limits.html) - [Axios — GPT-5.6 Sol/Terra/Luna restricted preview (Jun 26, 2026)](https://www.axios.com/2026/06/26/openai-gpt-sol-terra-luna-trump) - [Wikipedia — GPT-5.6](https://en.wikipedia.org/wiki/GPT-5.6) - [Fello AI — "Still No GPT-6" tracker (Jul 11, 2026)](https://felloai.com/all-we-know-about-chatgpt-6/) - [FindSkill — GPT-6 release-date tracker (Q4 modal)](https://findskill.ai/blog/gpt-6-release-date/) - [LifeArchitect — GPT-6 (2026)](https://lifearchitect.ai/gpt-6/) - [ScriptByAI — Anthropic Claude release timeline (Fable 5 / Mythos 5 Jun 9; Sonnet 5 Jun 30; Opus 4.8 May 28)](https://www.scriptbyai.com/anthropic-claude-timeline/) - [explainx — Claude Opus 5 release rumors (July 2026)](https://explainx.ai/blog/claude-opus-5-release-speculation-july-2026) - [tokenmix — tracing the four Opus-5 signals (Honeycomb, Vertex)](https://dev.to/tokenmixai/i-traced-4-claude-opus-5-signals-the-release-date-still-isnt-real-yet-2f2j) - [Swfte — LMArena leaderboard mirror (Jul 20, 2026)](https://www.swfte.com/lmarena) - [DataLearner — AI model leaderboard mirror (Jul 16, 2026)](https://www.datalearner.com/en/leaderboards) - [Tunguz — GPU spot prices surge 114% in six weeks](https://tomtunguz.com/b200-gpu-pricing-spot-market-model-releases/) - [Ornn compute index (resolution source for GPU ladders; JS-only)](https://dashboard.ornnai.com/) **Data:** Kalshi public trade API (live quotes, orderbooks, candlesticks) and a read-only Kalshi DB mirror (market rules, event metadata), both queried 2026-07-21. Web sources as linked above; leaderboard mirrors are third-party and may lag or restyle the official resolver views — treated accordingly. **Disclaimer:** All probabilities are subjective estimates. Prediction-market contracts can and do resolve to zero; nothing here is financial advice. Rejected candidates are logged and shadow-tracked to settlement so the selection screen itself is testable.