Following OpenAI DevDay and an accompanying podcast, OpenAI product leads outlined a new stack for "Computer Use" that combines personal assistants (Dots), a lower‑cost, faster model variant (GPT‑6.1 Sol), and new APIs for agents and low‑latency decisioning. Ari Weinstein (Product & Engineering, Computer Use, OpenAI) and Nikunj Handa (Product, API, OpenAI) walked through the features, engineering tradeoffs and intended developer experience.
What are Dots and the personal cloud machines?
Ari Weinstein explained that each Dot personal assistant now has access to its own Linux virtual machine in the cloud. This goes beyond a cloud browser: a Dot can run full desktop applications and use a web browser. The intent is to let an agent operate software in the same way a human would, so users can delegate tasks safely and broadly.
Weinstein gave a concrete example: configuring a complex meal‑prep order that used to take him two hours was completed by a Dot in 15 minutes using GPT‑6.1 Sol.
GPT‑6.1 Sol: cost and speed
Weinstein said GPT‑6.1 Sol is particularly good for Computer Use because of cost and speed advantages: compared with Astra the model runs at about one‑fifth the cost overall, and for Computer Use specifically about one‑seventh the cost, figures mentioned during the keynote.
How Computer Use changed recently
Weinstein emphasized significant progress in recent months: models not only can start tasks reliably but are much better at debugging, retrying and introspecting failures. The field has moved forward by combining techniques — screenshots, accessibility trees, direct DOM access, Playwright and generated JavaScript — which together improve speed and robustness.
He highlighted that agents often generate JavaScript that is executed to perform multiple steps at once rather than issuing single GUI actions sequentially, which yields major speedups.
App shots, accessibility and richer context
The "app shots" feature is more than a screenshot: it captures raw text and an accessibility representation from the application, giving the model a richer, token‑efficient context. Weinstein noted that accessibility technologies originally designed for users with assistive needs also make language models far better at using software.
Closing the development loop: build, test, QA
Ari pointed out that Computer Use allows agents to not only write code but also test the software they build, completing much of the development lifecycle and reducing manual QA work for humans.
Agents API, trust, permissions and safety
Bringing Computer Use into the Agents API lets third‑party developers build on the same Computer Use harness OpenAI uses. Weinstein warned that trust and safety remain critical: apps should request user consent for consequential actions (payments, sensitive operations) and limit agent access to only required sites and applications.
Nikunj Handa: API updates — async tool calls, mid‑turn steering, WebSockets
Nikunj Handa explained several API innovations tied to GPT‑6 and the new model family:
- Async tool/function calling: tool invocations can run asynchronously so the model does not have to pause reasoning while waiting for external operations to complete. The model can continue producing thoughts and check back later.
- Mid‑turn steering: developers can inject messages while the model is still reasoning, enabling dynamic instruction changes during a single turn.
- WebSockets: bidirectional communication reduces overhead for real‑time interactions and tool calls.
Handa said these features are effective when the model and harness are trained/engineered together and that WebSockets play an important role for tool‑driven, interactive apps.
UltraFast and the inference stack
Handa described UltraFast as an effort to push inference latency down. The inference team made many optimizations to improve efficiency and cut costs (for example, large price reductions for Luna were attributed to inference improvements). The focus has shifted toward squeezing out the lowest possible latencies on frontier models.
Decisions API: fast decision models inspired by Jev
Handa acknowledged Jev as an inspiration and said the Decisions API was developed quickly in response to internal demand for very fast classification and decisioning. The initial Decisions API is implemented on top of existing Luna weights, with inference optimizations, structured outputs and parallel/batched evaluation to achieve low time‑to‑first‑decision. OpenAI intends to iterate on model and calibration behavior based on usage.
Decisions API is aimed at very fast classification workloads, benefits from Luna's vision capabilities, and pairs well with GPT Live for snappy, real‑time agent experiences.
Caching, pre‑warming and long‑lived agent threads
Handa described improvements to caching in the Responses API: a general cache‑hit guarantee within 30 minutes is in place, and a preview offering included a 12‑hour caching guarantee for some customers. A pre‑warming feature lets developers pay an upfront cache‑write fee to prepare responses for a forthcoming time window.
He urged developers to be cache‑aware (use prompt and cache diagnostics) to lower costs and improve responsiveness, especially for persistent personal agents.
Context compaction and /compact
To manage very large context lengths, OpenAI provides compaction (context compression) solutions. Agents API includes compaction in the harness. Responses API supports server‑side compaction (automatic when a token threshold is reached) and a /compact endpoint for explicit developer control. New compaction techniques are being tested and parts are already available in the Codex harness.
Where Computer Use is headed
Weinstein said Computer Use is already faster than the average human for many tasks, and the next frontier is to become ‘‘literally superhuman’’ — matching or exceeding expert software users. Reaching that goal requires simultaneous advances in models, inference, harness engineering and representation.
Summary and developer call to action
OpenAI packaged several capabilities at DevDay — Dots with personal cloud Linux machines, GPT‑6.1 Sol with stated cost advantages, Agents API integration for Computer Use, and a low‑latency Decisions API built on Luna weights. Key engineering themes include async tool calling, mid‑turn steering, WebSockets, UltraFast inference, app shots combining accessibility/DOM data and generated JavaScript, caching/pre‑warming and compaction for long‑running agents.
OpenAI invited developers to try the new Agents and Decisions APIs and to provide feedback on performance, caching and higher‑level platform primitives as the company evolves what it calls an "AI cloud." Participants on the podcast included Ari Weinstein and Nikunj Handa; the conversation was hosted by Vibhu and Swyx.



