# Will Google have the best AI model at the end of December 2026?

> As of 6 October 2026, Prolepsia puts the chance of yes at 57%.

- Forecaster: Prolepsia (https://prolepsia.com)
- Forecast made: 2026-10-06
- Question type: Yes / no
- Resolves: 2026-12-31
- Status: Open
- Forecasts: 1
- Page: https://prolepsia.com/forecasts/saber-relay-citadel/will-google-have-the-best-ai-model-at-the-end-of-december-2026
- Open in the app: https://app.prolepsia.com/questions/7a12a457-29cf-4f95-b28a-22ffa66610f3

## Forecast

- Yes: 57%
- No: 43%

## How Prolepsia reasons

## TL;DR
The saved forecast is **56.9% Yes and 43.1% No**. Google is the modest favorite to own the first-ranked model on the specified Arena text leaderboard at the December 31, 2026, noon ET check. The decisive issue is whether Google preserves its current preference advantage in an eligible year-end entry while rivals release new models.

## Context
This market measures a specific leaderboard result, not general AI superiority. Resolution uses the Rank column on the overall text leaderboard with style control off, then the underlying Arena score and finally company-name alphabetical order if ties remain. Any Google-owned model can deliver a Yes; Gemini 4 Argon does not have to remain first throughout the quarter. The specified resolution table is the [Arena text leaderboard](https://lmarena.ai/leaderboard/text).

As of the October 6 forecast date, the latest observed qualifying board is dated October 2, 2026. It places Google’s Gemini 4 Argon High first and Anthropic’s Claude Opus 5.5 High second. Argon is marked Preliminary, and Google’s announcement describes a restricted initial rollout rather than a completed general release ([qualifying board](https://arena.ai/leaderboard/text/overall-no-style-control), [Google announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)). Google therefore starts ahead, but the forecast must account for both competition and the transition to an eligible public model.

## Evidence
The historical backbone is persistence at the company level. The audit of the October 6 dataset vintage found only two company-level leadership changes over July 2025–October 2, 2026, despite many published observations. It also found that several apparent advances were temporary debut spikes or rotations among declining rivals, not durable increases in the frontier ([official historical text data](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/tree/main/text)). The inference is straightforward: counting releases or record-looking scores overstates how often the leading company actually changes.

That history is useful but thin. Overlapping observations during one long leadership spell do not provide independent examples of an incumbent surviving another quarter. Older records also have roster-completeness problems: retired experimental and preview entries are missing from some newer historical files. Contemporaneous artifacts preserve entries absent from those files, and Arena’s changelog records retrospective vote additions and methodology changes ([historical artifacts](https://huggingface.co/spaces/lmarena-ai/arena-leaderboard/tree/main), [changelog](https://arena.ai/company/leaderboard-changelog)). The history supports a restrained incumbency advantage, not a precise law of retention.

The strongest present-day evidence is Google’s lead on the correct board. The October 2 edition reports Argon at 1533 ±9 with 4,932 comparisons, versus Opus 5.5 High at 1512 ±9 with 4,552 comparisons. The underlying ratings give Google a 21.78-point margin ([leaderboard](https://arena.ai/leaderboard/text/overall-no-style-control), [granular data](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/tree/main/text)). These are human-preference rating points, not percentages correct. The comparison counts are not counts of unique voters.

The reported individual confidence intervals do not overlap. This makes ordinary sampling noise an inadequate explanation for the whole current lead. It does not establish that Google will retain first place after releases, broader voting, or deployment changes. **The current lead is real; its durability is the question.** Argon’s Preliminary label makes fresh-vote confirmation especially valuable ([ratings and status](https://arena.ai/leaderboard/text/overall-no-style-control)).

The leaderboard setting matters. The default text page uses style-controlled results and identifies a different nearest challenger. The explicit no-style-control view confirms the relevant Google lead and Anthropic’s position behind it ([default text view](https://arena.ai/leaderboard/text), [qualifying view](https://arena.ai/leaderboard/text/overall-no-style-control)). Rank Spread is also separate from Rank: overlapping uncertainty ranges do not automatically create a resolution tie ([ranking method](https://arena.ai/blog/ranking-method)).

Anthropic is the most grounded displacement threat. Its full supplied Opus release chronology for 2026 is February 5, April 16, May 28, July 24, and September 22. The successive intervals are 70, 42, 57, and 60 days ([Opus chronology](https://www.anthropic.com/claude/opus)). That cadence leaves room for another release before year-end. It is evidence of operational opportunity, not a confirmed future schedule or proof that the next release will close the gap.

The category results show a plausible path to catching Google. In creative writing, Opus 5.5 scores 1532 ±20 while Argon scores 1531 ±19. In instruction following, Argon leads only 1537 ±14 to 1534 ±15. On hard prompts, Google has a larger advantage, 1548 ±11 versus 1535 ±11, but that margin remains below its overall lead ([creative writing](https://arena.ai/leaderboard/text/creative-writing-no-style-control), [instruction following](https://arena.ai/leaderboard/text/instruction-following-no-style-control), [hard prompts](https://arena.ai/leaderboard/text/hard-prompts-no-style-control)). These overlapping categories are not independent tests. They show that Google’s overall advantage is not uniform across conversational tasks.

Anthropic’s September 22 announcement emphasizes clearer, more natural communication. That is directly relevant to human text preferences, unlike a gain confined to coding or autonomous work ([Opus 5.5 announcement](https://www.anthropic.com/claude-opus-5-5)). Its stated support for pacing frontier development adds potential release friction, but does not amount to a development halt ([pacing essay](https://darioamodei.com/post/we-must-pace-the-frontier)).

Google’s own future changes matter too. The September 2 Gemini 3.8 announcement describes a third Flash release in six weeks, showing active iteration. Those releases do not establish repeated advances beyond Argon’s frontier performance ([Gemini 3.8 announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber)). A stronger Google configuration or successor can defend the company’s lead. Merely broadening access to the same model is not itself a capability gain.

The largest separate constraint is eligibility. Arena’s policy, updated September 1, requires qualifying public access under its early-release exception within two weeks. It also requires assurances about version identity and continued access, and permits removal until reevaluation if requirements are not met ([Arena policy](https://arena.ai/blog/policy)). Argon’s September 30 leaderboard addition makes October 14 the nominal checkpoint if that addition date is the operative score-release date; this is a policy-derived checkpoint, not a Google launch promise ([addition record](https://arena.ai/company/leaderboard-changelog)).

Google’s September 30 announcement describes initial access for trusted cyber defenders through Fairwind, further safeguard work, and broader access beginning with paid API customers and Google AI Ultra subscribers. It gives no firm date ([Argon announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)). The specific qualifying commitment, Direct Chat access, and version-identity assurance were not verified. That gap warrants risk, not an accusation of noncompliance. Paid public access can qualify, and removal need not be permanent ([policy](https://arena.ai/blog/policy)).

Other labs retain a surprise path. OpenAI’s Sol announcement focuses on agentic coding, computer use, and professional work, with initial access through Work, Codex, and the API rather than ordinary Chat ([Sol announcement](https://openai.com/index/introducing-gpt-6-1-sol/)). Meta describes an undated roadmap for bigger models, while DeepSeek explicitly anticipates a V4.1 Pro release without a date ([Meta announcement](https://research.meta.ai/blog/introducing-muse-spark-1-3), [DeepSeek announcement](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)). These are pipeline signals, not demonstrated future text-leaderboard results.

Late entrants cannot be dismissed. Arena’s AutoEval methodology describes rapid preliminary ratings followed by human validation. A late-December launch therefore need not face a universal weeks-long evaluation delay. An automated result still must become a qualifying ranked entry to decide this market ([AutoEval methodology](https://arena.ai/blog/autoeval-scores)).

## What's non-obvious
The obvious reading is that a clear first-place score makes Google safe. The more useful distinction is between present measurement uncertainty and future change. The present margin is substantial. Most of the risk comes from a different model, a different public serving configuration, or a change in eligibility—not from discovering that today’s rounded score was slightly wrong. Anthropic’s category proximity also means a preference-oriented refinement can matter without a spectacular general-capability breakthrough.

The opposite mistake is to treat restricted access as automatic disqualification. The market resolves from the ranked table, while Arena’s own policies shape which entries remain on that table. Current inclusion is meaningful evidence, but it does not settle future compliance. **A temporary removal is not the same as a year-end loss**, because reevaluation and replacement remain possible ([eligibility policy](https://arena.ai/blog/policy)). This is why rollout risk reduces confidence without overturning Google’s status as the favorite.

## Uncertainties
The largest evidence gap is Argon’s specific compliance path. A verified public-access commitment, qualifying service availability, version-identity confirmation, and substantial post-release voting would clarify whether today’s evaluated model can remain a strong eligible entry. The announcement alone does not settle those points ([Google announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/), [Arena requirements](https://arena.ai/blog/policy)).

The second gap is future relative improvement. Release histories and vendor roadmaps show opportunities, not the scores of unreleased models. Internal research-acceleration reports suggest more experimentation, but do not demonstrate a proportional acceleration in Arena preference gains ([OpenAI research report](https://openai.com/index/research-acceleration-view-inside-openai/), [Anthropic measurements](https://www.anthropic.com/institute/measuring-pace-of-ai-development)). Actual qualifying results would close this gap better than successor names or benchmark headlines.

Finally, the latest observed board is four days older than the forecast date, the historical sample contains few independent leadership episodes, and methodology changes weaken long-run comparisons. Fresh snapshots preserving Google’s margin would strengthen the evidence for durability; narrowing scores, prolonged removal, or a stronger ranked rival would weaken it. These uncertainties are already part of the saved forecast: **56.9% Yes and 43.1% No**.

## Forecast history

| Date | Forecast |
|---|---|
| 2026-10-06 | 57% chance of yes |

## Sources Prolepsia read

- Domain Expert Search
- Domain Expert Research Task
- [arena.ai](https://arena.ai/company/leaderboard-changelog)
- [arena.ai](https://arena.ai/leaderboard/text/overall-no-style-control)
- [Introducing Gemini 4 Argon](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon)
- [arena.ai](https://arena.ai/blog/policy)
- [arena.ai](https://arena.ai/direct)
- [deepmind.google](https://deepmind.google/fairwind-program)
- [ai.google.dev](https://ai.google.dev/gemini-api/docs/models)
- [ai.google.dev](https://ai.google.dev/gemini-api/docs/changelog)
- [ai.google.dev](https://ai.google.dev/gemini-api/docs/pricing)
- [docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/google-models)
- [docs.cloud.google.com](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models)
- [gemini.google](https://gemini.google/release-notes)
- [support.google.com](https://support.google.com/gemini/answer/16275805)
- [deepmind.google](https://deepmind.google/models/model-cards)
- [deepmind.google](https://deepmind.google/models/gemini)
- [arena.ai](https://arena.ai/leaderboard/text/overall)
- [tpsreport.news](https://tpsreport.news/models/gemini-4-argon)
- [agihunt.info](https://agihunt.info/en/daily/2026-10-03)
- [skalablog.com](https://skalablog.com/p/what-is-gemini-4-argon-and-who-can-use-it)
- [github.com](https://github.com/lmarena/lmarena.github.io/blob/main/_pages/model_list.md)
- Claude Code
- [lmarena-ai/leaderboard-dataset at main](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/tree/main/text)
- [arena.ai](https://arena.ai/leaderboard/text/creative-writing-no-style-control)
- [arena.ai](https://arena.ai/leaderboard/text/instruction-following-no-style-control)
- [openai.com](https://openai.com/news/research)
- [openai.com](https://openai.com/products/release-notes)
- [Introducing Claude Opus 5.5 \\ Anthropic](https://www.anthropic.com/claude-opus-5-5)
- [arena.ai](https://arena.ai/leaderboard/text/hard-prompts-no-style-control)
- [arena.ai](https://arena.ai/leaderboard/text/coding-no-style-control)
- [arena.ai](https://arena.ai/leaderboard/text/math-no-style-control)
- [Introducing GPT-6.1 Sol \| OpenAI](https://openai.com/index/introducing-gpt-6-1-sol)
- [developers.openai.com](https://developers.openai.com/api/docs/models/gpt-6.1-sol)
- [apnews.com](https://apnews.com/article/5afb865b2cddc439efdcf31ebdc406a5)
- [support.claude.com](https://support.claude.com/en/articles/12138966-release-notes?version=published)
- [Introducing Claude Fable 5.1 and Claude Mythos 5.1 \\ Anthropic](https://www.anthropic.com/claude-fable-and-mythos-5-1)
- [anthropic.com](https://www.anthropic.com/claude-fable-and-mythos-5-1?_bhlid=e7aaf2df3e24697dbf12bef2c1d5f8b38dc62074)
- [docs.x.ai](https://docs.x.ai/developers/release-notes)
- [huggingface.co](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/viewer/text_style_control/latest)
- [arena.ai](https://arena.ai/leaderboard/text)
- [arena.ai](https://arena.ai/leaderboard/text/creative-writing)
- [arena.ai](https://arena.ai/leaderboard?arena=text)
- [anthropic.com](https://www.anthropic.com/system-cards)
- [lmarena-ai/leaderboard-dataset at main](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/tree/main)
- [Research acceleration: The view inside OpenAI \| OpenAI](https://openai.com/index/research-acceleration-view-inside-openai)
- [Measurements for understanding the pace of AI development inside frontier labs \\ Anthropic](https://www.anthropic.com/institute/measuring-pace-of-ai-development)
- [arena.ai](https://arena.ai/blog/autoeval-scores)
- [README.md · lmarena-ai/leaderboard-dataset at main](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/blob/main/README.md)
- [axdevhub.com](https://www.axdevhub.com/models)
- [Introducing Muse Spark 1.3 \| Meta AI Research](https://research.meta.ai/blog/introducing-muse-spark-1-3)
- [Exclusive \| OpenAI Scraps Release of GPT-6.1 Astra Model Over Safety Concerns - WSJ](https://www.wsj.com/tech/ai/openai-chatgpt-model-release-cancel-safety-5a2f9f42)
- [Commits · lmarena-ai/leaderboard-dataset](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/commits/main)
- [huggingface.co](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset/tree/main/text_style_control)
- [help.arena.ai](https://help.arena.ai/articles/7011479247-how-to-see-ai-rankings-in-arena-leaderboards-2-0-wip)
- [techverdict.io](https://www.techverdict.io/articles/gemini-4-rumors-release-date-2026)
- [llm.ing](https://llm.ing/)
- [agientry.com](https://agientry.com/en/leaderboard)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro)
- [openai.com](https://openai.com/index/gpt-5-6)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber?trk=article-ssr-frontend-pulse_x-social-details_comments-action_comment-text)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash?trk=public_post_comment-text)
- [Introducing Gemini 3.8 Flash and 3.8 Flash Cyber](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber)
- [anthropic.com](https://www.anthropic.com/news/claude-opus-4-6)
- [anthropic.com](https://www.anthropic.com/news/claude-opus-4-7)
- [anthropic.com](https://www.anthropic.com/news/claude-opus-4-8)
- [anthropic.com](https://www.anthropic.com/claude/fable)
- [anthropic.com](https://www.anthropic.com/news/claude-opus-5)
- [anthropic.com](https://www.anthropic.com/claude-sonnet-5-5)
- [Dario Amodei — We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier)
- [openai.com](https://openai.com/index/introducing-gpt-5-5)
- [openai.com](https://openai.com/research/index)
- [openai.com](https://openai.com/index/gpt-6-astra)
- [openai.com](https://openai.com/index/introducing-gpt-6-sol-and-luna)
- [openai.com](https://openai.com/index/devday-2026-recap)
- [techmeme.com](https://www.techmeme.com/260928/p37)
- [developers.openai.com](https://developers.openai.com/api/docs/models)
- [research.meta.ai](https://research.meta.ai/)
- [research.meta.ai](https://research.meta.ai/blog/developing-capable-models-responsibly)
- [arena.ai](https://arena.ai/blog/ranking-method)
- [Gemini 2.5: Our newest Gemini model with thinking](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-model-thinking-updates-march-2025)
- [lmarena-ai/arena-leaderboard at main](https://huggingface.co/spaces/lmarena-ai/arena-leaderboard/tree/main)
- [Claude Opus \\ Anthropic](https://www.anthropic.com/claude/opus)
- [apnews.com](https://apnews.com/article/open-ai-artificial-intelligence-altman-trump-astra-5afb865b2cddc439efdcf31ebdc406a5)
- [anthropic.com](https://www.anthropic.com/news)
- [nextaimodel.com](https://www.nextaimodel.com/)
- [arena.ai](https://arena.ai/leaderboard/text?styleControl=false)
- [blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon?trk=article-ssr-frontend-pulse_little-text-block)
- [help.openai.com](https://help.openai.com/en/articles/6825453-chatgpt-release-notes)
- [openai.com](https://openai.com/news)
- [deploymentsafety.openai.com](https://deploymentsafety.openai.com/gpt-6-1-sol)
- [x.ai](https://x.ai/news)
- [x.ai](https://x.ai/news/grok-4-7)
- [kie.ai](https://kie.ai/blog/what-is-grok-4-8)
- [docs.x.ai](https://docs.x.ai/developers/models)
- [dev.meta.ai](https://dev.meta.ai/models/muse-spark)
- [developer.meta.com](https://developer.meta.com/ai/models/muse-spark)
- [arena.ai](https://arena.ai/leaderboard)
- [linkedin.com](https://www.linkedin.com/company/googledeepmind)
- [arena.ai](https://arena.ai/text/direct)
- [deepmind.google](https://deepmind.google/models/evals-methodology/gemini-4-argon)
- [microsoft.ai](https://microsoft.ai/news/mai-image-2-6-launches-at-no-2-on-arena-ahead-of-google-meta-and-xai)
- [microsoft.ai](https://microsoft.ai/news/introducing-mai-image-2-5)
- [arxiv.org](https://arxiv.org/html/2504.20879v2)
- [arena.ai](https://arena.ai/blog/our-response)
- [lmarena-ai/leaderboard-dataset · Datasets at Hugging Face](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset)
- [developers.openai.com](https://developers.openai.com/api/docs/changelog?mw_entry=2026-09-29-99a06ddef837dc53)
- [On the Navier–Stokes Millennium Prize Problem \| OpenAI](https://openai.com/index/navier-stokes-solution)
- [openai.com](https://openai.com/index/pacing-model-development-cyber-capabilities)
- [openai.com](https://openai.com/index/path-to-astra)
- [cellcog.ai](https://cellcog.ai/blog/openai-bel)
- [x.noodl3.net](https://x.noodl3.net/elonmusk/status/2099458047408013751)
- [elonmuskarchive.org](https://elonmuskarchive.org/posts/2103160462472892536)
- [DeepSeek \| Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.](https://www.deepseek.com/en/news/deepseek-v4-1-flash)
- [kimi.ai](https://www.kimi.ai/blog/kimi-k3)
- [platform.kimi.ai](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
- [kimi.com](https://www.kimi.com/blog/kimi-k3)
- [qwencloud.com](https://www.qwencloud.com/models/qwen3.8-max)
- [artificials.net](https://artificials.net/tests/arena-text)
- [arena.ai](https://arena.ai/blog/factuality-in-arena)
- [arena.ai](https://arena.ai/faq)
- [arena.ai](https://arena.ai/blog/leaderboard-changelog)
- [github.com](https://github.com/lmarena/arena-rank/blob/main/arena_rank/models/bradley_terry.py)
- [benchleader.com](https://www.benchleader.com/benchmarks/lmarena_text)
- [jumohub.com](https://www.jumohub.com/leaderboards/arena)
- [infosdivers.com](https://infosdivers.com/lmarena-ai)
- [anthropic.com](https://www.anthropic.com/news/claude-fable-5-mythos-5)
- [canary.arena.ai](https://canary.arena.ai/leaderboard/agent/pareto)
- [docs.x.ai](https://docs.x.ai/developers/grok-4-7)
- [arena.ai](https://arena.ai/leaderboard/text/overall-no-style-control?trk=article-ssr-frontend-pulse_little-text-block)
- [arena.ai](https://arena.ai/blog/llm-judge-self-preference)
- [anthropic.com](https://www.anthropic.com/news/claude-sonnet-5?trk=public_post_comment-text)
- [x.ai](https://x.ai/news/grok-4-7?_bhlid=924dcd63f9ad5091a8fe335b921dd0b1177615a6)
- [anthropic.com](https://www.anthropic.com/news/claude-opus-5?conversion=ppt-n8n)
- [openai.com](https://openai.com/index/gpt-5-6?trk=article-ssr-frontend-pulse_little-text-block)
- [openai.com](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6?trk=article-ssr-frontend-pulse_little-text-block)
- [x.ai](https://x.ai/news/series-e)
- [x.ai](https://x.ai/news/grok-4-5?_bhlid=263676315a044f4f80c797d4dcde989d1d15b71f&trk=article-ssr-frontend-pulse_little-text-block)
- [x.ai](https://x.ai/news/grok-4-6?trk=public_profile__reactions-text)
- [OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training - SecurityWeek](https://www.securityweek.com/openai-calls-off-gpt-6-1-astra-launch-details-safety-cases-for-frontier-training)
- [Towards safety cases for frontier AI training \| OpenAI](https://openai.com/index/towards-safety-cases-for-frontier-ai-training)
- [aitoolsreview.co.uk](https://aitoolsreview.co.uk/insights/next-claude-model)
- [supergrok.info](https://supergrok.info/blog/grok-5-release-date)

## The question

Imported from Polymarket.
This yes/no question comes from a single Polymarket market.

### How it resolves

This market will resolve according to the company which owns the model which has the highest arena rank based off the Chatbot Arena LLM Leaderboard (https://lmarena.ai/) when the table under the "Leaderboard" tab is checked on December 31, 2026, 12:00 PM ET.

Results from the "Rank" section on the Leaderboard tab of https://lmarena.ai/leaderboard/text with the style control off will be used to resolve this market.

Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by their Arena score, including any underlying, unrounded, granular values reflected in the data below the leaderboard. If a tie remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact arena score, “Google” would be ranked ahead of “xAI”). This market will resolve based on the company that occupies first place under this ranking system.

The resolution source for this market is the Chatbot Arena LLM Leaderboard found at https://lmarena.ai/. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.

---

Prolepsia’s forecasting engine made this forecast: it researched the question, weighed the evidence and wrote the report on this page. A probability is not a promise; this one changes as the evidence does.
