The Agent That Watches Your Screen: Google’s Gemini Now Reads Every App — Including Your Bank

Aug 13, 2026 | google ai

In a nutshell

everything on the web starts with the domain

Last night Google launched the Pixel 11, and buried under the camera specs and the folding screens was the announcement that actually matters. Google's AI can now see everything on your phone's display and act on it for you — booking, buying, calling, ordering, across apps that were never built to work with it. It is genuinely useful, and it quietly opens a door that will be very hard to close.

What Google Actually Announced

At Made by Google in New York on 12 August, Google unveiled the Pixel 11 line — four phones starting at $899, a new Pixel Watch 5, and an AirTag rival called Pixel Tag — all running a 2-nanometre Tensor G6 chip. But the company's own framing was unambiguous: the hardware is the vehicle, and Gemini Intelligence is the point. Google's product chief called it a totally new way to interact with a computer — one built not merely to answer questions but to get things done.

What "getting things done" means here is a genuine escalation. Gemini Intelligence on the Pixel 11 is agentic: it can order groceries, book rides, manage reservations, and phone businesses to complete tasks on your behalf, now across more than forty apps. The AI runs partly in Google's cloud on its Gemini models and partly on the device itself on smaller Gemma models, and Google leans hard on the privacy language of local processing. But the headline capability is not local, and it is not small.

The Mechanism, and Its Price: Screen-Parsing

Here is the part that deserves far more scrutiny than the camera bar. The way Gemini acts across apps that have no built-in AI integration is a technique called screen-parsing. Gemini reads the interface of whatever app is on your screen — visually, the way a human would — and then taps, types, and navigates it directly. No developer has to cooperate; if you can see it, Gemini can operate it.

Sit with what that requires. For this to work, Gemini must be able to see and interact with every app visible on your display — including your banking app, your private messages, and your shopping accounts. Google has added a sensible guardrail: the assistant asks for a confirmation tap before it executes an action. But the confirmation covers the action, not the seeing. To act across your apps, the system must first be able to read across your apps, and that is a new category of access to the most sensitive surfaces of your digital life. The convenience is real. So is the fact that you are granting an AItier-level view of everything you do on your phone.

Meta's Mirror Image

The timing could not be sharper, because just two days ago the opposite philosophy walked on stage. Meta released Muse Glimmer, a capable agent designed to run entirely on your own hardware, air-gapped, sending nothing to anyone's servers. Google's Gemini Intelligence is the mirror image: powerful, deeply integrated, partly cloud-based, and woven through a vertically controlled stack that Google owns end to end — the Tensor chip, the Android operating system, the Gemini and Gemma models, and the Pixel hardware. One vision puts the agent on hardware you control and answers to you. The other puts a more capable agent inside an ecosystem the company controls, that sees everything and reports a great deal back. Neither is simply good or bad, but they are genuinely different bets about who holds the reins, and this week Europe got to watch both placed within forty-eight hours of each other.

Claim and Counter-Claim

The optimistic case is strong and should not be dismissed as naive. This is the most useful phone AI yet built — an assistant that actually completes real tasks instead of demoing them, that works across apps developers never wired up, and that runs many operations locally on a chip explicitly designed to keep data on the device. For a lot of people, an agent that books the reservation and orders the groceries is a genuine gift of time, and the confirmation tap is a real, if partial, safeguard.

The skeptical case is just as grounded. An AI with standing permission to read every screen you open is the most intimate surveillance surface ever shipped to a consumer phone, and a confirmation tap governs what it does, not what it sees. Screen-parsing is precisely the capability that security researchers spent this summer warning about — an agent that operates human interfaces directly is also an agent that a malicious screen, a spoofed app, or a prompt injection can potentially hijack. And it deepens a dependency most users never chose: one company now mediates your chip, your operating system, your assistant, and your actions, while Apple prepares to rebuild Siri on the very same Gemini models, concentrating extraordinary influence in a single stack. The honest synthesis: the capability is a real advance and a real exposure at once, and which one dominates depends on engineering, enforcement, and trust that cannot be verified from a keynote.

The European Perspective

For Europe, this launch lands squarely on three of its most active nerves at once. On data protection, an assistant that must read your banking app and your private messages to function is a GDPR question of the first order, and it is worth watching closely whether the full agentic feature set even ships in the EU at launch or arrives late, as Apple Intelligence and Meta's AI did before it, precisely because European rules force harder answers.

On the AI Act, an agent taking autonomous multi-step actions in the real world is exactly the kind of system the transparency and risk rules were written to govern, and the confirmation-tap model will be tested against them. And on competition, a single company controlling chip, OS, model, and hardware — while a rival giant leases its intelligence from that same company — is a concentration the Digital Markets Act exists to scrutinise.

The deeper European point is the one that ties this week together. Two days ago Meta offered a vision of AI you own and run yourself; last night Google offered a vision of AI that owns the whole stack and watches the whole screen. Europe cannot build either at frontier scale. But it can decide which vision its rules encourage, and which they constrain. That choice — between the agent that answers to you and the agent that answers to the platform — may matter more for European autonomy than any gigafactory. The tools are being handed to us fully formed. The only sovereignty left is deciding, clearly and in time, on whose terms we accept them.

We are not first. We are right.