Kalshi AI — Mispricing Audit

2026-07-21 · v3.3 methodology (conviction-weighted, EV-agnostic sizing; Stage 2.5 resolver pricing) · $2,500 capital · 45-day horizon · picks count toward the public track record and are cleared for auto-copy

TLDR: The AI category is efficient this week. Of 18 markets diligenced, one pick survives every screen: KXGPT-OPEN-26SEP01 BUY NO @ 75¢ (GPT-6 will not get a public release before Sep 1). $174.75 deployed (7.0%), $2,325.25 held as cash (93%). Fifteen diligenced candidates are logged as shadow-tracked rejects — most cut by resolver-risk pricing (Stage 2.5) or the entry-band screen.

1. How this was researched

Mode: theme-match. Kalshi has no AI category, so I selected markets by keyword across all categories, then curated by hand. SQL terms searched over trading_events.title and trading_markets.title: openai, anthropic, claude, "ai", artificial intelligence, chatgpt, gpt, gemini, grok, xai, nvidia, deepseek, llama, llm, superintelligence, sam altman, robot, waymo, leaderboard, arena — with Sports/Crypto/Weather/Mentions categories excluded to kill false hits (player names, etc.). This surfaced ~36 AI-theme events: model-release timing (Claude Opus/Claude 5/Mythos, GPT-6/5.6, Gemini 3.5 Pro), leaderboard snapshots (Best AI, Top Coding/Math AI, Best Chinese AI), NVIDIA GPU compute-price ladders (A100/B200/H100/H200/RTX), AI policy (Anthropic DoD designation, Fable 5 US access, Trump AI-review EO), and corporate events (Anthropic/OpenAI/Waymo IPO announcements).

Pipeline per the v3 flow: live quotes and orderbooks from the Kalshi public API; rules text from the DB mirror (the mirror carries no market_snapshots table, so 14-day history came from Kalshi candlesticks); primary-source news via web search; every surviving candidate got a Stage 2.5 resolver profile with risks priced as explicit point deductions. Six tickers already live in the feed (Gemini 3.5 Pro Jul 24/31 NO, Arena Math Gemini NO, Top-model CLAUF YES / CLAUT NO, RTX 5090 NO) were excluded from the candidate set up front.

World-state anchors established during diligence (from finalized Kalshi markets + primary sources): Claude 5 family released (market resolved YES); Fable 5 + Mythos 5 shipped Jun 9, Sonnet 5 Jun 30, Opus 4.8 May 28; GPT-5.6 went GA Jul 9; Fable 5 US access restored before Jul 24 (resolved YES).

2. Markets reviewed

MarketQuestionPrice (YES)v24 / OIOutcome of review
KXGPT-OPEN-26SEP01GPT-6 public before Sep 125 / 261,903 / 15,362PICK — BUY NO @ 75
KXCODEAI-26JUL31-CLAUClaude tops Datacurve DeepSWE Jul 3126 / 312,025 / 28,370Reject — fairly priced (+2¢)
KXCODEAI-26JUL31-CHATChatGPT tops Datacurve DeepSWE Jul 3162 / 70875 / 14,826Reject — fairly priced (+3¢)
KXMATHAI-26JUL31-CLAUClaude #1 Arena Math Jul 3188 / 931,515 / 13,777Reject — negative edge at ask
KXCHINAAI-26JUL27-ALIBAlibaba best Chinese AI (Arena RSC) Jul 2774 / 751,670 / 1,438Reject — resolver view unverifiable
KXCLAUDE-OPUS-26JUL31Next Claude Opus before Jul 3177 / 81683 / 1,038Reject — rumor-priced, no edge
KXCLAUDE-OPUS-26JUL24Next Claude Opus before Jul 2438 / 791,962 / 2,397Reject — 41¢ spread unfillable
KXCLAUDE-OPUS-26AUG14Next Claude Opus before Aug 1476 / 8914 / 533Reject — illiquid, 13¢ spread
KXCLAUDE-NXTMYTH-26SEP01Next Mythos-class model before Sep 137 / 430 / 3,845Reject — low conviction, definitional
KXCLAUDE-MYTH-26SEP01Mythos public release before Sep 15 / 61,257 / 51,632Reject — 1¢ edge in >90 band
KXB200WS-26JUL24-7.000B200 > $7/hr Jul 24 (Ornn)57 / 581,473 / 1,318Reject — coin-flip, index unreadable
KXCHAICUTS-26AUG06-T1AI #1 reason, Challenger July report45 / 510 / 1,129Reject — coin-flip band, no volume
KXFRONTIER-FRON-26SEP01AI solves Frontier Math problem by Sep 130 / 34204 / 2,109Reject — low conviction
KXGEMINI-GEMI35P-26AUG31Gemini 3.5 Pro before Aug 3171 / 77629 / 1,540Reject — no timing edge
KXTECHRANKLISTAICODE-26JUL27-KIMIKimi #1 LM Code Arena Jul 2794 / 952,309 / 1,472Reject — fully priced
KXANTHROPICRESCIND-26JUN-26AUG01Anthropic DoD designation rescinded by Aug 18 / 1440 / 1,946Reject — sub-35¢ tail (hard band)
KXTOPMODEL-26JUL31-CLAUTopus-4-6-thinking top-ranked Jul 3190 / 939,744 / 21,875Excluded — ticker already a live feed pick
KXLLM1-26JUL31-AClaude best AI in July98 / 998,930 / 63,259Stage-1 skip — fully priced (≥95)

Stage-1 mechanical cuts not shown: ~150 GPU-ladder strikes with near-zero 24h volume, IPO-announcement deep-NO markets (≤3¢ asks), weekly leaderboard legs at 0–3¢, and the Fable-5-disabled / AI-review-EO tails.

3. The pick

Pick 1 — KXGPT-OPEN-26SEP01 · BUY NO @ 0.75 · MEDIUM CONVICTION

Will OpenAI release GPT-6 before Sep 1, 2026? — closes 2026-09-01T03:59Z (41 days)

Current price (YES)
25 bid / 26 ask → NO fillable at 74–75¢
My probability (NO)
90%  vs  market-implied 74–75%
Entry band
75¢ effective — 60–90¢ favorite band (the deployable band)
Edge
+15¢/contract, +20% EV (recorded for calibration only — did NOT set size)
Size
233 contracts × 75¢ = $174.75 (7.0% of capital — MEDIUM tier cap)
Max payout
$233.00 (profit +$58.25 if NO)
Cluster
gpt6-release-timing

Mispricing thesis. The market gives a 25–26% chance that OpenAI takes a generational GPT-6 to public release within 41 days — six weeks after GPT-5.6 went generally available on Jul 9, and with no GPT-6 announcement, model card, or date in existence. The crowd read the 5.6 launch as evidence of an accelerating ramp toward 6 (YES spiked 12¢→39¢ on launch day and held ~40¢ for a week); the release-cadence evidence says the opposite — a fresh flagship GA is what OpenAI ships instead of a generational jump.

Evidence.

Stage 2.5 — resolver profile. Rules: "If OpenAI releases a model called GPT-6 or greater before Sep 1, 2026 → YES"; secondary: "Release must be to the public, outside of a closed beta, though limiting it to a high-cost subscription tier is acceptable." Mechanism: first-occurrence release determination from public announcements; no snapshot exposure, no third-party scoreboard. Definitional boundary is naming ("GPT-6 or greater") plus public availability. Priced risks: surprise generational launch under competitive pressure from Anthropic's Claude-5 sweep of the leaderboards, −2 pts (5.6's preview→GA took only 2 weeks, so an early-August announcement could still make it); naming stunt / "GPT-6-preview" public tier, −1 pt. Evidence-based ~93% → final model_prob 90%.

Tail risk (cleanest single loser). OpenAI announces GPT-6 at a surprise August event and pushes it to public GA within ~2–3 weeks, replicating the 5.6 preview-to-GA speed. That is the whole bear case, and it is real — which is why this is MEDIUM, not HIGH.

Price history & entry context. 14-day candles: 11¢ (Jul 8) → 39¢ on 5.6 launch day (Jul 9, 7.4k vol) → 35–42¢ plateau for a week → faded to 25–26¢ over Jul 18–21 as no announcement materialized. Honest caveat per the 48h screen: the last 3 days moved in my direction (NO 62¢→75¢), so part of the edge is already eaten; what remains is still +15¢ against my number. Book at entry: YES 25 bid / 26 ask, spread 1¢; NO fillable ≤75¢: ~2,246 contracts (160 @ 74¢ + 2,086 @ 75¢) — the 233-lot fills at ≤75¢ with no slippage. v24 1,903; OI 15,362.

4. Recommended $2,500 portfolio

#MarketActionLimitContractsCostConviction / bandMax payoutEV¢ / EV%*
1KXGPT-OPEN-26SEP01BUY NO75¢233$174.75MEDIUM · 60–90¢ favorite$233.00+15¢ / +20%
Cash reserveHOLD$2,325.25
Totals$174.75 deployed (7.0%) · $2,325.25 cash (93.0%)$233.00blended +20% on deployed

*EV figures are recorded for calibration only. Per v3, position size was set by conviction tier (MEDIUM → ≤7% of capital) and entry band — not by EV. Dollar edge on deployed capital: ≈ +$34.95 expected (0.90 × $233 − $174.75).

Cluster exposure (cap: 15% of capital per cluster)

ClusterCost% of capitalCap check
gpt6-release-timing$174.757.0%OK (<15%)

Conviction exposure

TierCost deployed% of capital
HIGH$00%
MEDIUM$174.757.0%
LOW$0 (never deploys, v3.3)0%

Risk profile

Execution notes

5. What I rejected and why

All fifteen rejects below are machine-logged in picks.json and shadow-tracked to settlement — if they outperform the pick, the screen adds nothing and we want to know. The dominant reject reasons this week were Stage 2.5 resolver risk (leaderboard-snapshot markets whose exact resolver view — LMArena with Remove Style Control, Datacurve DeepSWE, the Ornn GPU index — I could not independently verify or could not price within the 15-pt deduction cap) and the entry-band screen.

Rejected on resolver risk / no edge after deductions

Cut by the entry-band / conviction screen

Cut on liquidity

6. Sources