[Explanation] Computer Use is included as standard in Gemini 3.5 Flash

IT

Google has equipped its flagship AI model, “Gemini 3.5 Flash,” with a “Computer Use” feature that recognizes screen content and directly controls PCs and mobile devices as a standard tool. This integration is expected to significantly lower the barriers to “agent-based automation,” where AI autonomously performs complex tasks.

Key Points of the Announcement and the Significance of Native Integration

On June 24, 2026, Google announced that its flagship lightweight AI model, “Gemini 3.5 Flash,” now comes standard with the “Computer Use” feature, which automatically recognizes computer screens and performs operations. The most notable aspect of this announcement is that advanced features previously offered only in the dedicated model for specific uses, “Gemini 2.5 Computer Use,” have been natively integrated into the mainstream Flash model that is widely used daily.

This allows developers to build complex agents that include screen manipulation within a single model, without the hassle or overhead of calling “separate models dedicated to computer operation.” Google claims that this integration has achieved its best-ever performance in agent-based screen operation tasks. Standard equipment on flagship models marks a structural turning point, as AI evolves from merely a “text generation tool” to an “agent that autonomously uses tools (such as PCs and apps).”

An autonomous mechanism that “sees, thinks, and manipulates” the screen.

The “Computer Use” feature built into Gemini 3.5 Flash enables AI to perform processes similar to those used by humans operating smartphones or PCs. Specifically, when you give the AI a target and a screenshot of the screen, it analyzes the image and suggests next actions such as “Click this coordinate” or “Enter specific characters here.” By repeating the cycle where the developer’s program executes the proposed operation and sends a new screen back to the AI after execution, the task autonomously progresses until it is completed.

Please refer to the diagram below.

Figure 1

The overwhelming versatility of this feature lies in its ability to operate even on older systems (legacy systems) without dedicated integration interfaces or in complex web browser workflows as long as you see the screen. To put it simply, regardless of the refrigerator’s manufacturer or model number, it offers flexible service like a butler who can check the contents and take out the milk. Since the AI itself does not directly run web browsers but instead executes AI proposals through code, it is designed to be easily integrated into existing development environments.

Accelerating agent development and applying it in practice

Improved Developer Convenience and Agent Performance

Gemini 3.5 Flash already featured powerful built-in tools such as grounding with Google Search and Google Maps, as well as code execution. With the addition of “Computer Use” as a standard tool alongside these, developers can seamlessly assemble a series of advanced actions—”research,” “think,” and “manipulate the screen”—within a single model. Because there is no disconnection between tools, it offers the advantage of significantly reducing the complexity of enterprise automation pipeline design.

In terms of performance, despite being a lightweight and high-speed Flash model, it performs well in benchmarks such as “OSWorld-Verified,” which measures screen operation capabilities, reaching a practical level. Additionally, the Flash model’s unique speed—four times faster per second output tokens per second compared to competitors’ cutting-edge models—provides a significant advantage in screen operations that require fast agent loops. We now have an environment where building low-cost, fast, and multifunctional AI agents can be achieved more easily than ever before.

Specific use cases expected in business settings

Google cites the automation of long-term, complex business tasks—such as continuous software testing and knowledge work across multiple specialized applications—as key use cases for “Computer Use.” For example, in web application development sites, AI autonomously and repeatedly performs user interface (UI) tests, dramatically reducing the workload of quality control. Additionally, in administrative tasks, it is possible to replace routine tasks that humans used to perform for hours—such as moving back and forth between multiple browser tabs, Excel, and business systems to input and integrate data—with agents.

Please refer to the diagram below.

Figure 2

Real-world applications have already begun, with reports of cases such as banks extracting information from complex documents exceeding 100 pages to speed up account opening procedures, and fintech companies autonomously managing and executing tax filing workflows that span several weeks. By simply giving instructions in natural language, it is becoming increasingly realistic to use it as a personal AI agent to handle complex tasks in daily digital life, such as booking appointments on websites or handling complex invoices.

Safety Measures and Future Social Impact

Multi-layered safeguard against malicious attacks

AI’s ability to freely manipulate screens comes with new security risks. In particular, “prompt injection” attacks, where AI reads and executes malicious instructions hidden within web pages under operation, pose a serious threat. In response, Google conducts targeted adversarial training on Gemini 3.5 Flash to increase its resistance to generating harmful content and unauthorized instructions. For enterprise users, it offers two powerful safeguard options based on a “Defense-in-Depth” approach.

The first is a system that always requires explicit human confirmation before performing highly confidential operations or irrevocable actions such as deleting information. The second is a defensive feature that automatically stops the task the moment indirect prompt injection is detected. Google recommends combining these features with execution in secure sandbox environments and strict access controls to ensure reliable practical use in enterprise environments.

Labor Market Transformation and Prospects for Agent AI

The spread of “operable AI” is expected to have a significant impact on social structures. According to survey data from the first half of 2026, 14% of all workers have already lost their jobs mainly due to AI, with routine tasks among the “white-collar middle class” especially facing the wave of automation. Standardizing PC operations with Gemini 3.5 Flash could further accelerate the replacement of knowledge labor by automating even the “information bridging between applications” that previously required human intervention.

The next point of interest is the trends of the flagship model, the Gemini 3.5 Pro, scheduled for release in June 2026. The Pro model is expected to achieve inference accuracy surpassing Flash and feature even longer context windows, enabling agent tasks that require more complex and advanced decision-making. The shift to an “agent economy,” where AI masters tools and works 24/7, 365 days a year, will fundamentally redefine how we work and what we define in professions.

[#Gemini #AIエージェント #Google #ComputerUse #自動化 #生成AI #労働市場 #DX]

コメント

Copied title and URL