Prolepsia forecastYes / no

Will Google have the best AI model at the end of December 2026?

As of 6 October 2026, Prolepsia puts the chance of yes at 57%.

57%chance of yes

Yes 57%No 43%
Prolepsia forecast made

How Prolepsia reasons

The report Prolepsia wrote with this forecast.

TL;DR

The saved forecast is 56.9% Yes and 43.1% No. Google is a modest favorite because it has a meaningful lead on the exact leaderboard that resolves the market, but competing releases and unresolved rollout conditions leave substantial downside.

Key evidence

  • Google leads the qualifying board. The latest observed edition is October 2, 2026, not an October 6 voting snapshot. Gemini 4 Argon High is first, with a displayed score of 1533 ±9 and 4,932 comparisons; Claude Opus 5.5 High is second at 1512 ±9 (qualifying leaderboard). The underlying data gives Google a 21.78-point lead (official text data). Ordinary sampling noise is not the main threat to that advantage.
  • Leadership changes are less frequent than model announcements suggest. The historical audit found only two company-level leadership changes over July 2025–October 2, 2026. Many observations overlap, and some apparent advances were temporary debut spikes rather than lasting frontier gains (historical data). This supports taking incumbency seriously without treating it as a guarantee.
  • Anthropic is the closest practical challenger. Its Opus releases occurred on February 5, April 16, May 28, July 24, and September 22, 2026, leaving room for another release before resolution (release chronology). Its current model is already close to Argon in creative writing and instruction following (creative writing, instruction following).
  • Rollout eligibility is a separate risk. Arena’s early-release policy requires qualifying public access within two weeks and allows removal followed by reevaluation (policy). Google’s September 30 announcement describes restricted initial access and broader availability without a firm date (Argon announcement). Missing compliance evidence is not proof of a violation.

Uncertainties

The decisive unknown is how the best eligible models compare at the year-end check, not who leads today. Argon’s preliminary rating can settle as fresh votes arrive, and safeguards or serving changes can affect its public version. A Google successor can also defend the lead.

The event concerns overall text preference with style control off—not coding, scientific benchmarks, or commercial usefulness. Late releases remain a threat because Arena can publish fast preliminary evaluations, although an automated score is not automatically a qualifying ranked entry (AutoEval methodology).

Read the full report

TL;DR

The saved forecast is 56.9% Yes and 43.1% No. Google is the modest favorite to own the first-ranked model on the specified Arena text leaderboard at the December 31, 2026, noon ET check. The decisive issue is whether Google preserves its current preference advantage in an eligible year-end entry while rivals release new models.

Context

This market measures a specific leaderboard result, not general AI superiority. Resolution uses the Rank column on the overall text leaderboard with style control off, then the underlying Arena score and finally company-name alphabetical order if ties remain. Any Google-owned model can deliver a Yes; Gemini 4 Argon does not have to remain first throughout the quarter. The specified resolution table is the Arena text leaderboard.

As of the October 6 forecast date, the latest observed qualifying board is dated October 2, 2026. It places Google’s Gemini 4 Argon High first and Anthropic’s Claude Opus 5.5 High second. Argon is marked Preliminary, and Google’s announcement describes a restricted initial rollout rather than a completed general release (qualifying board, Google announcement). Google therefore starts ahead, but the forecast must account for both competition and the transition to an eligible public model.

Evidence

The historical backbone is persistence at the company level. The audit of the October 6 dataset vintage found only two company-level leadership changes over July 2025–October 2, 2026, despite many published observations. It also found that several apparent advances were temporary debut spikes or rotations among declining rivals, not durable increases in the frontier (official historical text data). The inference is straightforward: counting releases or record-looking scores overstates how often the leading company actually changes.

That history is useful but thin. Overlapping observations during one long leadership spell do not provide independent examples of an incumbent surviving another quarter. Older records also have roster-completeness problems: retired experimental and preview entries are missing from some newer historical files. Contemporaneous artifacts preserve entries absent from those files, and Arena’s changelog records retrospective vote additions and methodology changes (historical artifacts, changelog). The history supports a restrained incumbency advantage, not a precise law of retention.

The strongest present-day evidence is Google’s lead on the correct board. The October 2 edition reports Argon at 1533 ±9 with 4,932 comparisons, versus Opus 5.5 High at 1512 ±9 with 4,552 comparisons. The underlying ratings give Google a 21.78-point margin (leaderboard, granular data). These are human-preference rating points, not percentages correct. The comparison counts are not counts of unique voters.

The reported individual confidence intervals do not overlap. This makes ordinary sampling noise an inadequate explanation for the whole current lead. It does not establish that Google will retain first place after releases, broader voting, or deployment changes. The current lead is real; its durability is the question. Argon’s Preliminary label makes fresh-vote confirmation especially valuable (ratings and status).

The leaderboard setting matters. The default text page uses style-controlled results and identifies a different nearest challenger. The explicit no-style-control view confirms the relevant Google lead and Anthropic’s position behind it (default text view, qualifying view). Rank Spread is also separate from Rank: overlapping uncertainty ranges do not automatically create a resolution tie (ranking method).

Anthropic is the most grounded displacement threat. Its full supplied Opus release chronology for 2026 is February 5, April 16, May 28, July 24, and September 22. The successive intervals are 70, 42, 57, and 60 days (Opus chronology). That cadence leaves room for another release before year-end. It is evidence of operational opportunity, not a confirmed future schedule or proof that the next release will close the gap.

The category results show a plausible path to catching Google. In creative writing, Opus 5.5 scores 1532 ±20 while Argon scores 1531 ±19. In instruction following, Argon leads only 1537 ±14 to 1534 ±15. On hard prompts, Google has a larger advantage, 1548 ±11 versus 1535 ±11, but that margin remains below its overall lead (creative writing, instruction following, hard prompts). These overlapping categories are not independent tests. They show that Google’s overall advantage is not uniform across conversational tasks.

Anthropic’s September 22 announcement emphasizes clearer, more natural communication. That is directly relevant to human text preferences, unlike a gain confined to coding or autonomous work (Opus 5.5 announcement). Its stated support for pacing frontier development adds potential release friction, but does not amount to a development halt (pacing essay).

Google’s own future changes matter too. The September 2 Gemini 3.8 announcement describes a third Flash release in six weeks, showing active iteration. Those releases do not establish repeated advances beyond Argon’s frontier performance (Gemini 3.8 announcement). A stronger Google configuration or successor can defend the company’s lead. Merely broadening access to the same model is not itself a capability gain.

The largest separate constraint is eligibility. Arena’s policy, updated September 1, requires qualifying public access under its early-release exception within two weeks. It also requires assurances about version identity and continued access, and permits removal until reevaluation if requirements are not met (Arena policy). Argon’s September 30 leaderboard addition makes October 14 the nominal checkpoint if that addition date is the operative score-release date; this is a policy-derived checkpoint, not a Google launch promise (addition record).

Google’s September 30 announcement describes initial access for trusted cyber defenders through Fairwind, further safeguard work, and broader access beginning with paid API customers and Google AI Ultra subscribers. It gives no firm date (Argon announcement). The specific qualifying commitment, Direct Chat access, and version-identity assurance were not verified. That gap warrants risk, not an accusation of noncompliance. Paid public access can qualify, and removal need not be permanent (policy).

Other labs retain a surprise path. OpenAI’s Sol announcement focuses on agentic coding, computer use, and professional work, with initial access through Work, Codex, and the API rather than ordinary Chat (Sol announcement). Meta describes an undated roadmap for bigger models, while DeepSeek explicitly anticipates a V4.1 Pro release without a date (Meta announcement, DeepSeek announcement). These are pipeline signals, not demonstrated future text-leaderboard results.

Late entrants cannot be dismissed. Arena’s AutoEval methodology describes rapid preliminary ratings followed by human validation. A late-December launch therefore need not face a universal weeks-long evaluation delay. An automated result still must become a qualifying ranked entry to decide this market (AutoEval methodology).

What's non-obvious

The obvious reading is that a clear first-place score makes Google safe. The more useful distinction is between present measurement uncertainty and future change. The present margin is substantial. Most of the risk comes from a different model, a different public serving configuration, or a change in eligibility—not from discovering that today’s rounded score was slightly wrong. Anthropic’s category proximity also means a preference-oriented refinement can matter without a spectacular general-capability breakthrough.

The opposite mistake is to treat restricted access as automatic disqualification. The market resolves from the ranked table, while Arena’s own policies shape which entries remain on that table. Current inclusion is meaningful evidence, but it does not settle future compliance. A temporary removal is not the same as a year-end loss, because reevaluation and replacement remain possible (eligibility policy). This is why rollout risk reduces confidence without overturning Google’s status as the favorite.

Uncertainties

The largest evidence gap is Argon’s specific compliance path. A verified public-access commitment, qualifying service availability, version-identity confirmation, and substantial post-release voting would clarify whether today’s evaluated model can remain a strong eligible entry. The announcement alone does not settle those points (Google announcement, Arena requirements).

The second gap is future relative improvement. Release histories and vendor roadmaps show opportunities, not the scores of unreleased models. Internal research-acceleration reports suggest more experimentation, but do not demonstrate a proportional acceleration in Arena preference gains (OpenAI research report, Anthropic measurements). Actual qualifying results would close this gap better than successor names or benchmark headlines.

Finally, the latest observed board is four days older than the forecast date, the historical sample contains few independent leadership episodes, and methodology changes weaken long-run comparisons. Fresh snapshots preserving Google’s margin would strengthen the evidence for durability; narrowing scores, prolonged removal, or a stronger ranked rival would weaken it. These uncertainties are already part of the saved forecast: 56.9% Yes and 43.1% No.

Forecast history

How Prolepsia’s forecast has movedOne forecast so far: 57% on 6 Oct 2026.
DateForecast
57% chance of yes

Sources Prolepsia read

  1. Domain Expert Search
  2. Domain Expert Research Task
  3. arena.aiarena.ai
  4. arena.aiarena.ai
  5. Introducing Gemini 4 Argonblog.google
  6. arena.aiarena.ai
  7. arena.aiarena.ai
  8. deepmind.googledeepmind.google
  9. ai.google.devai.google.dev
  10. ai.google.devai.google.dev
  11. ai.google.devai.google.dev
  12. docs.cloud.google.comdocs.cloud.google.com
  13. docs.cloud.google.comdocs.cloud.google.com
  14. gemini.googlegemini.google
  15. support.google.comsupport.google.com
  16. deepmind.googledeepmind.google
  17. deepmind.googledeepmind.google
  18. arena.aiarena.ai
  19. tpsreport.newstpsreport.news
  20. agihunt.infoagihunt.info
  21. skalablog.comskalablog.com
  22. github.comgithub.com
  23. Claude Code
  24. lmarena-ai/leaderboard-dataset at mainhuggingface.co
  25. arena.aiarena.ai
  26. arena.aiarena.ai
  27. openai.comopenai.com
  28. openai.comopenai.com
  29. Introducing Claude Opus 5.5 \ Anthropicwww.anthropic.com
  30. arena.aiarena.ai
  31. arena.aiarena.ai
  32. arena.aiarena.ai
  33. Introducing GPT-6.1 Sol | OpenAIopenai.com
  34. developers.openai.comdevelopers.openai.com
  35. apnews.comapnews.com
  36. support.claude.comsupport.claude.com
  37. Introducing Claude Fable 5.1 and Claude Mythos 5.1 \ Anthropicwww.anthropic.com
  38. anthropic.comwww.anthropic.com
  39. docs.x.aidocs.x.ai
  40. huggingface.cohuggingface.co
  41. arena.aiarena.ai
  42. arena.aiarena.ai
  43. arena.aiarena.ai
  44. anthropic.comwww.anthropic.com
  45. lmarena-ai/leaderboard-dataset at mainhuggingface.co
  46. Research acceleration: The view inside OpenAI | OpenAIopenai.com
  47. Measurements for understanding the pace of AI development inside frontier labs \ Anthropicwww.anthropic.com
  48. arena.aiarena.ai
  49. README.md · lmarena-ai/leaderboard-dataset at mainhuggingface.co
  50. axdevhub.comwww.axdevhub.com
  51. Introducing Muse Spark 1.3 | Meta AI Researchresearch.meta.ai
  52. Exclusive | OpenAI Scraps Release of GPT-6.1 Astra Model Over Safety Concerns - WSJwww.wsj.com
  53. Commits · lmarena-ai/leaderboard-datasethuggingface.co
  54. huggingface.cohuggingface.co
  55. help.arena.aihelp.arena.ai
  56. techverdict.iowww.techverdict.io
  57. llm.ingllm.ing
  58. agientry.comagientry.com
  59. blog.googleblog.google
  60. openai.comopenai.com
  61. blog.googleblog.google
  62. blog.googleblog.google
  63. blog.googleblog.google
  64. blog.googleblog.google
  65. Introducing Gemini 3.8 Flash and 3.8 Flash Cyberblog.google
  66. anthropic.comwww.anthropic.com
  67. anthropic.comwww.anthropic.com
  68. anthropic.comwww.anthropic.com
  69. anthropic.comwww.anthropic.com
  70. anthropic.comwww.anthropic.com
  71. anthropic.comwww.anthropic.com
  72. Dario Amodei — We Must Pace the Frontierdarioamodei.com
  73. openai.comopenai.com
  74. openai.comopenai.com
  75. openai.comopenai.com
  76. openai.comopenai.com
  77. openai.comopenai.com
  78. techmeme.comwww.techmeme.com
  79. developers.openai.comdevelopers.openai.com
  80. research.meta.airesearch.meta.ai
  81. research.meta.airesearch.meta.ai
  82. arena.aiarena.ai
  83. Gemini 2.5: Our newest Gemini model with thinkingblog.google
  84. lmarena-ai/arena-leaderboard at mainhuggingface.co
  85. Claude Opus \ Anthropicwww.anthropic.com
  86. apnews.comapnews.com
  87. anthropic.comwww.anthropic.com
  88. nextaimodel.comwww.nextaimodel.com
  89. arena.aiarena.ai
  90. blog.googleblog.google
  91. help.openai.comhelp.openai.com
  92. openai.comopenai.com
  93. deploymentsafety.openai.comdeploymentsafety.openai.com
  94. x.aix.ai
  95. x.aix.ai
  96. kie.aikie.ai
  97. docs.x.aidocs.x.ai
  98. dev.meta.aidev.meta.ai
  99. developer.meta.comdeveloper.meta.com
  100. arena.aiarena.ai
  101. linkedin.comwww.linkedin.com
  102. arena.aiarena.ai
  103. deepmind.googledeepmind.google
  104. microsoft.aimicrosoft.ai
  105. microsoft.aimicrosoft.ai
  106. arxiv.orgarxiv.org
  107. arena.aiarena.ai
  108. lmarena-ai/leaderboard-dataset · Datasets at Hugging Facehuggingface.co
  109. developers.openai.comdevelopers.openai.com
  110. On the Navier–Stokes Millennium Prize Problem | OpenAIopenai.com
  111. openai.comopenai.com
  112. openai.comopenai.com
  113. cellcog.aicellcog.ai
  114. x.noodl3.netx.noodl3.net
  115. elonmuskarchive.orgelonmuskarchive.org
  116. DeepSeek | Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.www.deepseek.com
  117. kimi.aiwww.kimi.ai
  118. platform.kimi.aiplatform.kimi.ai
  119. kimi.comwww.kimi.com
  120. qwencloud.comwww.qwencloud.com
  121. artificials.netartificials.net
  122. arena.aiarena.ai
  123. arena.aiarena.ai
  124. arena.aiarena.ai
  125. github.comgithub.com
  126. benchleader.comwww.benchleader.com
  127. jumohub.comwww.jumohub.com
  128. infosdivers.cominfosdivers.com
  129. anthropic.comwww.anthropic.com
  130. canary.arena.aicanary.arena.ai
  131. docs.x.aidocs.x.ai
  132. arena.aiarena.ai
  133. arena.aiarena.ai
  134. anthropic.comwww.anthropic.com
  135. x.aix.ai
  136. anthropic.comwww.anthropic.com
  137. openai.comopenai.com
  138. openai.comopenai.com
  139. x.aix.ai
  140. x.aix.ai
  141. x.aix.ai
  142. OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training - SecurityWeekwww.securityweek.com
  143. Towards safety cases for frontier AI training | OpenAIopenai.com
  144. aitoolsreview.co.ukaitoolsreview.co.uk
  145. supergrok.infosupergrok.info

The question

Imported from Polymarket. This yes/no question comes from a single Polymarket market.

How it resolves

This market will resolve according to the company which owns the model which has the highest arena rank based off the Chatbot Arena LLM Leaderboard (https://lmarena.ai/) when the table under the "Leaderboard" tab is checked on December 31, 2026, 12:00 PM ET. Results from the "Rank" section on the Leaderboard tab of https://lmarena.ai/leaderboard/text with the style control off will be used to resolve this market. Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by their Arena score, including any underlying, unrounded, granular values reflected in the data below the leaderboard. If a tie remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact arena score, “Google” would be ranked ahead of “xAI”). This market will resolve based on the company that occupies first place under this ranking system. The resolution source for this market is the Chatbot Arena LLM Leaderboard found at https://lmarena.ai/. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.