Curiosities

Google's Next Gemini Pro Is Almost Certain to Crack 40% on Humanity's Last Exam

The benchmark was designed to humble machines permanently — the fact that its ceiling is now a trading question reveals how fast that premise is aging.

Source: Polymarket market “Next Google Gemini Pro Model: Humanity’s Last Exam Debut?”

Leading outcome 40%+ 94% Likely
24h move ▲ 44.5 pts 40%+
Traded 24h $22K $20K all time
Resolves by 2026-12-31

Something has shifted decisively in how informed observers assess the frontier of machine intelligence. The money staking real claims on the performance of Google's next Gemini Pro model has all but settled one question: whether it will clear 40% on Humanity's Last Exam, a benchmark assembled from the hardest problems human experts could devise across mathematics, science, and the humanities. That outcome is now treated as a near-certainty.

What makes this worth pausing on is not the percentage itself but what the benchmark was supposed to be. Humanity's Last Exam was constructed explicitly as a ceiling — a set of problems chosen because no existing AI system was expected to master them. It was designed to stay hard. The fact that a prediction market now prices a 40% pass rate as essentially settled, and puts the odds of a 50% pass at roughly even, means the market's collective judgment is that the ceiling is being raised faster than its architects anticipated.

The sharp repricing in the last 24 hours — a near-50-point lurch toward the 40% threshold — suggests this was not gradual drift but a specific update. That kind of move typically reflects something concrete: a leak, a benchmark preview, an internal demonstration, or sufficiently credible technical reporting to move people with enough conviction to stake meaningful money quickly. The 50% threshold sits at about 62%, a strong lean but not a certainty, implying the crowd believes the next model will be formidable without being quite sure how formidable. The 55% outcome, by contrast, has collapsed — suggesting the market sees a ceiling somewhere in the low-to-mid fifties, at least for this generation.

The anthropological curiosity here runs deeper than any single model's score. That a liquid market exists at all for AI performance on a test designed to be permanently beyond AI reach is its own signal about where collective anxiety — and collective hope — currently sits. People are not merely curious whether machines are getting smarter; they are curious at what rate, on what timeline, and whether the institutions built to measure that progress can keep up. A benchmark meant to be a fixed star is turning out to be a moving one, and markets are, in their blunt way, pricing the motion.

What the cluster reveals, read whole, is a consensus that Google's next Gemini Pro represents a genuine generational step — not a marginal refinement. The near-certainty at 40% and the strong lean at 50% together describe a model that the money believes will land somewhere in a range that would have seemed implausible even eighteen months ago. The open question is whether 55% or beyond remains out of reach for this iteration, or whether the market has simply not yet received the information that would move that final threshold. If credible benchmarking data surfaces before resolution and shows scores in the mid-fifties, expect that upper band to reprice sharply. If the model lands in the low fifties, the market's current read will prove remarkably well-calibrated — and the more unsettling question will be what comes next.

Where the money stands

40%+ 94% ▲ 44.5
45%+ 92% 0.0
50%+ 62% 0.0
55%+ 24% ▼ 25.5
View the market on Polymarket ← Front Page