Google tests Gemini 4 Carbon variant matching Anthropic’s Opus 5.5 in coding
Google is internally testing multiple variants of its Gemini 4 model, with the Carbon variant reportedly matching Anthropic’s Opus 5.5 in coding performance. The company is refining its flagship reasoning model, Argon, while preparing updates to its apps ahead of the launch.
Editor, Lazyfounder

30 SEC SUMMARY
- Google is internally testing Gemini 4 variants named Argon, Barium, and Carbon, with Carbon reportedly matching Anthropic’s Opus 5.5 in coding performance.
- Carbon, deployed on Google’s internal coding platform Jetski, is said to outperform Argon in programming tasks.
- Argon is Google’s frontier reasoning model, while Carbon and Barium appear to be iterative updates within the same family.
- Google is preparing for the Gemini 4 launch with updates to its apps, including new modes and adjustable reasoning intensity.
- Recursive self-improvement is cited as a key driver for rapid advancements in AI model performance.
TABLE OF CONTENTS
KEY HIGHLIGHTS
- Google is testing Gemini 4 variants named Argon, Barium, and Carbon, with Carbon deployed internally on its coding platform Jetski.
- Carbon is reportedly outperforming Argon in programming tasks and is compared to Anthropic’s Opus 5.5, while early Argon versions were likened to Opus 5.
- Argon is Google’s flagship reasoning model, with Flash, Omni, and Gemma serving other roles in the Gemini lineup.
- Google’s apps are being updated ahead of the Gemini 4 launch, including new modes like "Automatic" and "Ultra," and adjustable reasoning intensity settings.
- A Google DeepMind employee referenced recursive self-improvement as a factor in the rapid development of AI models.
Gemini 4 variants under internal testing
According to The Decoder, Google is internally testing multiple variants of its upcoming Gemini 4 model, identified as Argon, Barium, and Carbon. These variants are part of an effort to refine performance ahead of a broader release.
Carbon, the latest addition, was recently deployed on Google’s internal coding platform Jetski. Sources suggest it demonstrates improved performance in programming tasks compared to Argon, the flagship reasoning model in the Gemini 4 family.
Carbon’s performance and comparisons
Carbon is reportedly capable of matching the coding performance of Anthropic’s Opus 5.5, according to internal observations cited by The Decoder. This represents a significant step up from early versions of Argon, which some employees compared to Anthropic’s older Opus 5 model.
While Argon is positioned as Google’s frontier reasoning model, Carbon and Barium appear to be incremental updates or checkpoints within the same development cycle, rather than distinct product tiers.
Preparations for Gemini 4 launch
Google is updating its suite of applications in anticipation of the Gemini 4 launch. Changes include a new "Automatic" mode for 3-series models in the Gemini app, as well as adjustable reasoning intensity settings ranging from low to high.
Google AI Studio has also introduced an "Ultra" mode, which promises "advanced skills and tools," though details about its specific capabilities remain unclear.
Logan Kilpatrick, a Google employee, confirmed that the team is focused on optimizing Argon’s performance, suggesting that the model could ship under that name even if Carbon is its internal designation.
Recursive self-improvement in AI development
Vedant Misra, an employee at Google DeepMind, referenced recursive self-improvement as a driver behind the rapid advancements in AI model performance. This approach involves models iteratively refining their own capabilities, potentially accelerating development timelines.
What this means
Lazyfounder analysis — our interpretation, not reported fact.
Google’s iterative testing of Gemini 4 variants highlights a broader trend in AI development: the pursuit of incremental improvements rather than single breakthroughs. By refining models like Carbon within the Argon family, Google is signaling that performance gains—particularly in specialized tasks like coding—are now a key battleground in the AI race.
For founders and operators, this approach underscores the importance of vertical optimization. A model that excels in coding or reasoning may not need to outperform rivals across every benchmark to be valuable—especially if it can be seamlessly integrated into existing workflows, like Google’s Jetski platform.
The focus on recursive self-improvement also suggests that AI development is becoming more dynamic, with models capable of refining themselves faster than human-led tuning. This could shorten the cycle between major releases and raise the bar for competitors like Anthropic and OpenAI.
Key takeaways
- Founders building AI-powered tools should prioritize task-specific performance over general benchmarks, as Google’s focus on coding excellence demonstrates a viable niche strategy.
- Internal testing and iterative updates can yield rapid improvements, so operators should consider phased rollouts or beta programs to refine product-market fit.
- Recursive self-improvement in AI models may accelerate development cycles, creating pressure to stay agile and adapt to faster-paced releases from competitors.
- Adjustable reasoning intensity and new modes in Google’s apps suggest a shift toward customizable AI interactions, which could inspire similar features in other AI-driven products.
FAQ
Why is Google testing multiple Gemini 4 variants?
Google is testing variants like Argon, Barium, and Carbon to refine performance, particularly in coding and reasoning tasks. This approach allows the company to compare and iterate on specific capabilities before a broader release.
How does Carbon compare to Anthropic’s Opus 5.5?
According to internal observations, Carbon is reportedly matching Anthropic’s Opus 5.5 in coding performance. This is a notable improvement over early versions of Argon, which were compared to Anthropic’s older Opus 5 model.
Will Carbon be released as a separate product?
Carbon and Barium are likely incremental updates within the Argon family rather than standalone products. Google may ultimately ship the final model under the Argon name.
What changes is Google making to its apps ahead of Gemini 4’s launch?
Google is introducing new features such as an "Automatic" mode for 3-series models, adjustable reasoning intensity settings, and an "Ultra" mode in Google AI Studio promising advanced skills and tools.
Related on Lazyfounder
Sources
- The Decoder · 2026-10-10
Google's Gemini 4 "Carbon" model is reportedly matching Anthropic's Opus 5.5 coding performance
This story is an original summary drafted with AI by Lazyfounder from the reporting listed above and checked by automated validation. Facts are attributed to their original publishers; sections marked as analysis are Lazyfounder's. Where a source is in another language, facts were machine-translated and quotations are reported, not reproduced. Read the original coverage via the links, and see our AI policy and corrections policy.
About the author
Editor, Lazyfounder
Tarun Mottlia edits LazyFounders, covering Indian startups, funding rounds, AI and product launches. Every story on the site is AI-assisted and checked against its cited sources before publication.
More stories by Tarun MottliaGet the LazyFounder Brief
Startup, funding and AI news in a five-minute read. Join the early-access list.


