Google Leak: Gemini 4 Reportedly Struggles with Some Programming Tasks. Company Denies It

  • Some Google employees reportedly see a discrepancy between Gemini 4's test results and its performance in practice
  • The objections concern certain programming tasks; Google rejects the assessment
  • Gemini 4 Argon also increases the output limit to up to 1 million tokens

Sdílejte:
Jakub Kárník
Jakub Kárník
2. 10. 2026 10:30
Ilustrační náhled: Gemini 4 reportedly 'struggles' despite benchmark scores, some say
Advertisement

A high benchmark score and a reliable programming assistant might not be the same thing. Such a discrepancy is suggested by a leak of internal evaluations of the new Gemini 4, reported by Bloomberg. According to the report, some Google employees encountered programming tasks with which the model did not perform as well as its test results would suggest. Google has pushed back against this assessment.

Ahead of the Competition in Tests, but No Consensus Within Google

During the presentation of Gemini 4 Argon, Google showed results in which the new model surpassed competitors in several key areas, including OpenAI’s GPT-6 Astra. However, leaked evaluations reveal a divided view within the company. One employee speaks of a broad consensus that Gemini belongs at the technological forefront. Others fear it will continue to lag behind Anthropic and OpenAI models.

For those interested in a code-writing assistant, this is a more significant debate than the ranking itself. The resulting score can facilitate initial orientation, but it doesn’t tell you how well the model will handle their specific tasks. However, even differing employee opinions are not enough to conclude that the new Gemini is generally a weaker choice for programming.

Longer Output Also Comes with a Specific Price

Besides the debate about quality, Argon brings a specific change: the limit for generated output increases to up to 1 million tokens. This refers to the space for content that the model creates. Google also highlights the new model’s capabilities in cyber defense.

A higher ceiling can be useful for extensive outputs where their length is an issue. However, it doesn’t by itself address the correctness of the generated code, which is the focus of the current dispute.

In your opinion, which programming tasks does AI most frequently fail at?

Sources: 9to5google.com

About the author

Jakub Kárník

Jakub is known for his endless curiosity and passion for the latest technologies. His love for mobile phones started with an iPhone 3G, but nowadays… More about the author

Jakub Kárník
Sdílejte: