Model launches

Hark launches Handoff agent for autonomous web tasks, claims top benchmark score and lower pricing

Hark unveiled Handoff, a "computer use agent" that autonomously navigates the open web to complete tasks like ordering food or booking flights.

Hark launches Handoff agent for autonomous web tasks, claims top benchmark score and lower pricing

Hark, the secretive startup founded earlier this year by serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a "computer use agent" (CUA) the company says can autonomously navigate the open web to complete end-to-end tasks such as ordering food on DoorDash, booking flights on United and Delta, or messaging job candidates on LinkedIn.

Sign-ups opened today at hark.com, and Hark plans to make the software available later this month as part of its initial platform release.

Performance claims and pricing

Hark says Handoff achieved a 97.7 score on the Online-Mind2Web (OM2W) benchmark, a third-party test with a human-evaluated leaderboard. The company cited comparisons of 97.7 for Handoff versus 92.8 for OpenAI's GPT 5.4, 84.1 for Anthropic's Claude Opus 4.8, and 69 for Google's Gemini 2.5 Pro. Hark also published token pricing that it says is less than one-tenth that of competing frontier models: $0.18 per million input tokens and $2.37 per million output tokens, compared with example prices of $5 and $30 for GPT 5.5. Hark reported per-turn model latency of 0.8 seconds.

For every request, Handoff launches a dedicated virtual machine with its own browser, file system, and terminal. Users can connect existing accounts so the agent can log in and act using saved addresses, payment methods, and history.

Benchmark caveats and omitted comparisons

The benchmark comparisons Hark shared exclude the most recent frontier models. Hark compared Handoff to GPT 5.5, GPT 5.4, Opus 4.8 and Gemini 2.5 Pro — but did not include current leaders such as OpenAI's GPT-5.6 and Anthropic's Opus 5, nor established open-source contenders like DeepSeek V4, Kimi K3, and Qwen3.8-Max. Those newer models have not published OM2W results to the benchmark's public leaderboard, so Hark's "top-ever" claim cannot be checked against the strongest available systems.

This omission matters because the newest frontier models have shown the largest gains in computer use tasks: for example, on the related OSWorld 2.0 benchmark, Anthropic's Opus 5 scores roughly 70.6% versus 55.7% for Opus 4.8, the latter being the model Hark used for comparison.

The latency figures Hark cites also come with caveats: the 6.8-second and 6-second per-turn numbers for GPT 5.5 and Opus 4.8 were measured by Hark in Hark's own harness with the competing models set to their highest (and slowest) reasoning levels. No independent latency measurements are available for comparison.

Even inside Hark's chosen comparison set there are exceptions: on WebTailBench v2, one of the three benchmarks in Hark's table, GPT 5.5 scores 72.3 while Handoff scores 68.6. Two of the three benchmarks (WebTailBench and an unnamed internal evaluation) were run inside Hark's harness, with pass rates computed by Hark's internal LLM judge — conditions Hark controls.

Nevertheless, Hark's pricing advantage is clearer: Anthropic's Opus 5 list price remains $5 per million input and $25 per million output, so Handoff's roughly tenfold cost savings would persist against current frontier pricing if its performance comparisons also hold.

Training pipeline and data questions

Hark's research preview describes a training pipeline of supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm, according to materials shared with VentureBeat before today's announcement. The company acknowledges it has only performed post-training so far and that pre-training is planned for later this year.

That indicates Handoff is built on top of a base model Hark did not pre-train. Hark has not specified which base model it uses or the mix of proprietary and open data in its training set.

Another significant question for enterprise customers is who can access the dedicated virtual machines and the files created on them. A Hark spokesperson said "security and privacy is a primary focus, but this is a technical preview," adding the company will share more information when the product reaches market at the end of the summer.

Founding, funding and related companies

Hark is Adcock's fourth company. He previously co-founded the talent marketplace Vettery (sold in 2018 for roughly $100 million), the air-taxi maker Archer Aviation, and the humanoid robotics company Figure AI.

Hark raised a $700 million Series A in May 2026 at a $6 billion valuation, led by Parkway Venture Capital with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest. Adcock seeded Hark with $100 million of his own funds and remains founder and CEO of both Figure and Hark, a company spokesperson confirmed.

The spokesperson said Hark models "are being trained on the Figure robots," but that Adcock has no plans to combine the two companies.

Context on Adcock's public claims

Adcock's promotional approach has attracted skeptics. In April 2025, Fortune correspondent Jason Del Rey reported that Figure's widely publicized BMW partnership was more limited than Adcock's public descriptions of a robot "fleet" performing "end-to-end operations"; BMW spokesperson Steve Wilson said a single Figure robot was practicing picking up parts during non-production hours. By June 2026, BMW said the Figure 02 robot had supported production of more than 30,000 BMW X3 vehicles over a ten-month period and that the next-generation Figure 03 robot was being deployed at the plant for a parts-sequencing logistics role.

Adcock called the Fortune story "mischaracterizations and downright lies" on the social network X and threatened a defamation suit. Two months later TechCrunch reported that Adcock skipped a promised live demo at a tech conference and avoided questions about the BMW deal onstage.

Those controversies do not prove Handoff's numbers wrong. The agent may indeed perform very well, and if its pricing holds it would undercut major labs' pricing.

Demo video

Hark published an approximately four-minute announcement video on YouTube and social media. In the clip, Adcock sits in a sparsely furnished warehouse space and speaks a request aloud to Handoff (for example, ordering roses and cherry blossoms). The video shows Handoff navigating a florist's website to place an order while Adcock narrates that, unlike a typical chatbot, Handoff "is always working, it's looping," and that he uses it for "all of my recruiting efforts end to end." Hark's announcement blog post includes additional demos shown in real time and at 5x speed.