Tools

Google DeepMind demos Gemini-powered pointer that interprets cursor context

Google DeepMind demonstrated an experimental Gemini-powered pointer that lets users point at on-screen elements (images, maps, menus, PDFs, webpages) and issue natural-language commands.

Google DeepMind unveiled an experimental pointer that uses the Gemini language model to turn the mouse cursor into an AI control surface. The idea is simple: you point the cursor at an image, map, menu, PDF or webpage element and then issue a natural-language command. The system reads the object indicated by the cursor, the surrounding screen context, and the instruction together, allowing it to resolve deictic references like “this,” “that,” “here,” and “there.”

How this could affect real-world use

The commercial implications are straightforward. If the technology becomes reliable, users could interact with AI directly in-place instead of opening a separate chatbot window. That would make any webpage, document, image, shopping cart, travel search or enterprise dashboard an AI-enabled surface.

With integration into Google products such as Chrome, Workspace and Android — and potential future hardware — Gemini could shift from being an app to becoming an operating habit: the model would mediate between what the user points at and the user’s intent across many on-screen objects.

Why it matters

The key takeaway is that the mouse gains a kind of ‘‘brain’’: the layer between user intent and on-screen objects could be controlled by a language model. Were this approach to work reliably, the next interface competition would center not simply on chat versus search, but on who controls that intermediary layer and how AI can access every element on the screen.

The demonstration is experimental; questions remain about practical deployment and reliability, as well as privacy, security and usability considerations.