🤖 AI benchmark: hit-rate of 7 models

Prematch Ao vivo (in-play)

Seven external AI models (Hermes contour) independently analyze the same jogos — predicting the resultado (1X2), total (Mais de/Menos de), both equipes to placar (BTTS) and the exact placar. Here we honestly compare their palpites against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test palpites are excluded). The sample is still small and not representative. Right now the snapshot holds 196 jogo(es), 366 settled AI palpites (Tabela tennis). The figures below are N, not «a percentage you can trust»: the more jogos are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Tabela tennis

Model N (settled) 1X2 Exact placar Composite accuracy
Claude
169 63.9%(108/169) 26.0%(44/169) 45.0%(152/338)
Google AI
197 61.9%(122/197) 23.9%(47/197) 42.9%(169/394)
DeepSeek
0
ChatGPT
0
Qwen
0
Kimi
0
GLM 5.2
0

grey — sample <5, not representative; «—» — the model has not made a settled palpite yet.

Composite accuracy — the share of correct palpites em all exibido mercados together: (sum of correct escolhas) ÷ (sum of all settled escolhas) em the mercados 1X2 + Exact placar. Each mercado-escolha weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. «Exact placar» — the full final placar was guessed correctly (H and A matched); palpites with no recognized placar do not count toward the denominator.

Composite model rating · all mercados · Tabela tennis

Bar height = the model's composite accuracy em all applicable mercados on the current sample. Sorted from best to worst.

45.0% (152/338)
Opus 4.8
42.9% (169/394)
Gemini 3.5 Flash
DeepSeek V4 Pro
sem dados
GPT 5.5
sem dados
Qwen 3.7 Plus
sem dados
Kimi 2.6
sem dados
GLM 5.2
sem dados

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by mercado · Tabela tennis

Where each model is strong: one mini-bar per applicable mercado, with the percentage and (hits/sample).

Claude Composite 45.0%
1X2
63.9% (108/169)
Exact placar
26.0% (44/169)
Google AI Composite 42.9%
1X2
61.9% (122/197)
Exact placar
23.9% (47/197)
DeepSeek Composite —
1X2
Exact placar
ChatGPT Composite —
1X2
Exact placar
Qwen Composite —
1X2
Exact placar
Kimi Composite —
1X2
Exact placar
GLM 5.2 Composite —
1X2
Exact placar

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a empate (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.