Google Leak: Gemini 4 Reportedly Struggles with Some Programming Tasks. Company Denies It Home News Some Google employees reportedly see a discrepancy between Gemini 4's test results and its performance in practice The objections concern certain programming tasks; Google rejects the assessment Gemini 4 Argon also increases the output limit to up to 1 million tokens Sdílejte: Jakub Kárník Published: 2. 10. 2026 10:30 Advertisement A high benchmark score and a reliable programming assistant might not be the same thing. Such a discrepancy is suggested by a leak of internal evaluations of the new Gemini 4, reported by Bloomberg. According to the report, some Google employees encountered programming tasks with which the model did not perform as well as its test results would suggest. Google has pushed back against this assessment. Ahead of the Competition in Tests, but No Consensus Within Google During the presentation of Gemini 4 Argon, Google showed results in which the new model surpassed competitors in several key areas, including OpenAI’s GPT-6 Astra. However, leaked evaluations reveal a divided view within the company. One employee speaks of a broad consensus that Gemini belongs at the technological forefront. Others fear it will continue to lag behind Anthropic and OpenAI models. For those interested in a code-writing assistant, this is a more significant debate than the ranking itself. The resulting score can facilitate initial orientation, but it doesn’t tell you how well the model will handle their specific tasks. However, even differing employee opinions are not enough to conclude that the new Gemini is generally a weaker choice for programming. Longer Output Also Comes with a Specific Price Besides the debate about quality, Argon brings a specific change: the limit for generated output increases to up to 1 million tokens. This refers to the space for content that the model creates. Google also highlights the new model’s capabilities in cyber defense. A higher ceiling can be useful for extensive outputs where their length is an issue. However, it doesn’t by itself address the correctness of the generated code, which is the focus of the current dispute. In your opinion, which programming tasks does AI most frequently fail at? Sources: 9to5google.com About the author Jakub Kárník Jakub is known for his endless curiosity and passion for the latest technologies. His love for mobile phones started with an iPhone 3G, but nowadays… More about the author Sdílejte: Gemini Google Umělá inteligence You might be interested in That's amazing! The 6-year-old PlayStation 5 gets AI upscaling, you can use it (for now) in these games Jakub Kárník 06:30 Galaxy Buds On leave ears open. Samsung added voice memos and head gestures Jakub Kárník 1. 10. This is what the Pixel 11a should look like. Leaked renders show a familiar design Adam Kurfürst 1. 10. Xiaomi's washer-dryer features a heat pump for the first time. It reportedly saves almost three-quarters of electricity during drying Adam Kurfürst 30. 9. Pixel Watch to get free blood pressure and insulin resistance trends. Older models will also receive them Jakub Kárník 30. 9. Some Pixels are experiencing contactless payment failures. According to users, a restart helps Jakub Kárník 30. 9.