Kalshi AI — Mispricing Audit

2026-07-24 · Theme: AI · Horizon: 45 days · Capital: $2,500 · Methodology v3.3 (conviction-weighted, EV-agnostic sizing; Stage 2.5 resolver pricing; no tails, no low-conviction deploys). Picks count toward the public track record and are cleared for auto-copy.

TL;DR. The AI category is efficiently priced this week. Of 78 AI-themed events and 24 Stage-2 candidates, exactly one contract cleared every screen: KXTECHRANKLISTAICODE-26JUL27-KIMI (Kimi #1 on LM Code Arena at the Jul 27 snapshot) — a deep favorite backed by a verified 44-point Elo lead with 3 days to resolution. Everything else — including the single largest modeled edge found (+27¢ on Challenger AI job cuts) — fails a v3 screen and is machine-logged as a shadow-tracked reject. Deployment: $279 (11.2%); cash: $2,221 (88.8%). Cash is a position.

1 · How this was researched

The fable-5 resolver anomaly (why this report is conservative on leaderboard markets). The live Arena text leaderboard (Jul 21 data) shows claude-fable-5 at #1 (1507±6) over claude-opus-4-6-thinking (1505±4) — yet the market prices fable-5-top-on-Jul-27 at 1–2¢ and opus-4-6-thinking at 91–92¢, and the settled record agrees with the market: the Jul 6, Jul 13 and Jul 20 weeklies all finalized YES for opus-4-6-thinking and NO for fable-5, including Jul 20 — one day before the leaderboard data date on which I can see fable-5 ahead. I could not fully reconcile the resolver's view (Rank-UB mechanics, style-control toggle state, or a snapshot-timing artifact) with the public page. Mid-July the changelog also shows Arena recalculated fable-5's scores to count only post-Jul-1 votes. This is precisely the market family that produced this program's costliest historical loss, so every leaderboard candidate here carries explicit resolver deductions, and one (Chinese AI monthly) was rejected outright on unpriceable resolver mechanics.

2 · Markets reviewed

TickerQuestion (short)Yes bid/ask ¢Vol 24hVerdict
KXTECHRANKLISTAICODE-26JUL27-KIMIKimi #1 on LM Code Arena, Jul 2791 / 921,186PICK — BUY YES
KXCHAICUTS-26AUG06-T1AI = #1 job-cut reason, July Challenger report45 / 470REJECT — band screen (best modeled edge found)
KXCHINAAI-26JUL31-ALIBAlibaba top Chinese AI on Arena text, Jul 3174 / 751,926REJECT — resolver risk
KXCODEAI-26JUL31-CHATChatGPT #1 on Datacurve DeepSWE, Jul 3186 / 904,886REJECT — no edge at fillable ask
KXMATHAI-26JUL31-CLAUClaude #1 on Arena math board, Jul 3192 / 96161REJECT — no edge at ask
KXTOPMODEL-26JUL27-CLAUTclaude-opus-4-6-thinking top model, Jul 2791 / 922,147REJECT — unreconciled resolver view
KXFRONTIER-FRON-26AUG01AI solves a FrontierMath open problem by Aug 17 / 10308REJECT (NO side) — slippage
KXGEMINI-GEMI35P-26AUG16Gemini 3.5 Pro released before Aug 1664 / 66208REJECT (NO side) — coin-flip band
KXB200MS-26JUL-6.500B200 July avg > $6.50/hr (Ornn)40 / 411,106REJECT — resolver data unverifiable
KXH200MS-26JUL-4.500H200 July avg > $4.50/hr (Ornn)40 / 411,428REJECT — resolver data unverifiable
KXRTX5090MS-26JUL-0.500RTX 5090 July avg > $0.50/hr (Ornn)49 / 502,148REJECT — resolver data unverifiable
KXFABLEDISABLED-27-26AUG31Fable 5 US access disabled again before Aug 314 / 90REJECT (NO side) — no edge, illiquid
KXANTHROPICRESCIND-26JUN-26AUG01DoD/WH rescinds Anthropic designation by Aug 15 / 1140REJECT — no edge, wide spread
KXCLAUDE-NXTMYTH-26SEP01Next Mythos-class model before Sep 130 / 3562REJECT — no evidence, illiquid
KXLLM1-26JUL31-AAnthropic best AI in July98 / 99427Stage-1 cut — fully priced
KXIPOANTHROPIC-DATE-26SEP01Anthropic confirms IPO before Sep 13 / 520,270Stage-1 cut — fully priced
KXCLAUDE-MYTH-26SEP01Anthropic releases Mythos before Sep 13 / 42,047Stage-1 cut — fully priced
KXVERARELEASE-VERARUBIN-*NVIDIA Vera Rubin release timing2 / 980Stage-1 cut — no market
KXAIREVIEW-26JUN-SEP01Mandatory federal pre-release AI review by Sep 19 / 15102Stage-1 cut — volume + spread
KXCHINAAI-26JUL27-ALIBAlibaba top Chinese AI (weekly), Jul 2793 / 941,696No edge at 94 — see monthly reject
KXTOPMODEL-26JUL31-*, KXMATHAI-26JUL31-GEMI, KXGEMINI-GEMI35P-26JUL31, KXGPT-OPEN-26SEP01Excluded — already live feed picks
KXFEDMENTION / KXHEARINGMENTION / KXTRUMPSAYCOMPANYExcluded — mention markets

3 · The pick

Pick 1 — KXTECHRANKLISTAICODE-26JUL27-KIMI · BUY YES @ 0.93 limit · HIGH CONVICTION

"Will Kimi be the #1 model on the LM Code Arena Leaderboard on Jul 27, 2026 at 10:00 AM ET?" — current 91¢ bid / 92¢ ask.

My probability: 95%  vs  market implied: ~92% (ask) / 93¢ at fillable size
Conviction tier / entry band: HIGH · deep favorite (>90¢) — thin edge by construction, sized by conviction
Cluster: kimi-k3-code-arena

Resolver: LM Code Arena leaderboard, brand-level ("Kimi" = any Kimi model), snapshot Jul 27, 2026 10:00 AM ET; ties resolve by the publisher's methodology, else shared rank paying $1/n. Priced risks: −3 pts (a brand-new frontier entrant added to Code Arena and out-scoring 1678 with enough votes inside 3 days, or a methodology surprise), −2 pts (an Arena score recalculation hitting kimi-k3 — Arena demonstrably recalculated claude-fable-5's scores mid-July), −0.5 pts (ordinary vote drift). Evidence-based ~99.5 → model_prob 95.

Mispricing thesis. This is a pay-the-cents lock-harvest, not a narrative bet: kimi-k3 leads the named resolver by 44 Elo points (1678 vs claude-fable-5's 1634, gpt-5.6-sol-xhigh 1630 — Jul 21 data), which is multiple confidence intervals; ordinary vote flow cannot close it in 3 days. The identical market settled YES last week (Jul 20), confirming the resolver already sees Kimi #1.

Evidence. Tail risks (cleanest way to lose).
Numbers: 300 contracts @ 93¢ limit = $279.00 cost · max payout $300.00 · edge +2¢ · EV ≈ +2.2% (~$6; recorded for calibration only — did NOT set size) · ~0.5¢/contract fees at this price point
Liquidity / entry context: yes_bid 91 (103 shown), yes_ask 92 (≈17 fillable), ≈400 more at 93 → ≈417 contracts fillable at ≤93¢; volume 24h 1,186; total volume 5,810; open interest 3,154; spread 1¢.
Price history: intraday history unavailable (no snapshots table in this mirror); API last=92, previous close 91 — no adverse 48h move; market opened Jul 20 post-settlement of last week's YES.
Disclosed screen deviation: 24h volume is 1,186 vs the Stage-1 ≥1,500 line. I kept the pick because the screen's purpose (fillability) is directly verified by the book — 417 contracts rest at ≤93¢ against a 300-contract order — and total volume/OI are healthy. Flagged for the record; if you copy-trade mechanically, treat this as the one judgment call in this report.

4 · Recommended $2,500 portfolio

#TickerActionLimitContractsCost% capMax payoutEdge ¢EV %Tier / band
1KXTECHRANKLISTAICODE-26JUL27-KIMIBUY YES93¢300$279.0011.2%$300.00+2+2.2%HIGH · >90¢
Cash reserve$2,221.0088.8%Held — no other candidate cleared the screens
Total300$2,500.00100%$300.00+2+2.2%

EV cents / EV % are recorded for calibration only; under v3 they do not drive contract counts (the resolved record showed stated EV was anti-predictive). Size was set by conviction tier (HIGH cap $375) and visible book depth, then trimmed to 300 to stay inside resting liquidity.

Cluster exposure

ClusterCost% of capitalCap (15%)
kimi-k3-code-arena$279.0011.2%$375 ✓

Conviction exposure

TierCost deployed% of capitalPer-pick cap
HIGH$279.0011.2%$375 (15%)
MEDIUM$00%$175 (7%)
LOW— never deploys (v3.3)

Risk profile

Execution notes

5 · What I rejected and why

Every Stage-2 survivor below is machine-logged in picks.json and shadow-tracked to settlement — if the rejects outperform the pick, the screen adds nothing, and that finding feeds the next methodology revision. Rejects cut specifically by the Stage-3.5 entry-band screen are marked [BAND].

The one that hurt to cut

Resolver-risk rejects (Stage 2.5)

No-edge rejects (market already right)

Liquidity / slippage rejects

Entry-band screen cuts [BAND]

6 · Sources

Data: Kalshi read-only DB mirror + Kalshi public API (quotes/orderbooks as of 2026-07-24 UTC; the mirror carries no intraday snapshot history) + primary web sources above. All probabilities are subjective estimates; prediction-market contracts can and do resolve to zero. Methodology v3.3: conviction-weighted sizing, EV-agnostic; sub-35¢ tails and low-conviction leans never deploy; rejects are machine-logged and shadow-tracked. This report is research, not investment advice.