Two independent reports — from Anthropic and Sunday Robotics — indicate that scaling up large, general-purpose AI models can significantly improve the capabilities and generalization of physical robots.
Anthropic: Opus line makes rapid gains
Anthropic tested how its Opus models affect a quadruped robot’s ability to perform a set of tasks. In August 2025, Claude Opus 4.1 was unable to solve the tasks autonomously; humans using the model were roughly twice as effective as those without it, but completing the full task set still required 181 minutes.
By May 2026, Opus 4.7 completed all but one task autonomously in 9 minutes and 35 seconds. The single remaining failure involved repositioning a ball to its starting spot after contact — a task that humans had also found difficult. Anthropic states that with more time and additional scaffolding, current generations of Claude would likely be able to solve the remaining task as well.
Anthropic emphasizes these gains were not the result of a targeted robotics development program but emerged from general model scaling: stronger base models can produce downstream improvements in robot performance as a natural side effect.
Sunday Robotics: big pretrained model plus small high-quality data
Sunday Robotics argues that the path to robot generalization is a strong pretrained base model augmented by small amounts of high-quality in‑house data. Their new model, ACT‑2, follows the company’s recipe: scale pretraining, then perform rapid post‑training iterations using minimal internal data.
Sunday reports that as the pretrained base model gets stronger, the reliability improvements obtained from a small set of in‑house iterations transfer more effectively to unseen real‑home environments. Remaining deployment gaps stem from edge cases and failures that appear only after repeated real-world policy execution; their post‑training loop specifically targets these gaps.
In tests, Sunday’s robots achieved a 99.1% success rate, performing 778 successful folds across nine garment types. Simple garments such as shorts and T‑shirts were easiest; more complex items like blouses were harder but still showed success rates above 90%. Sunday says it will deploy its Memo system to families through a Beta Program in the fall.
Sunday also notes that ACT‑2 often surprises the team: the same base model is learning a broader set of household skills, including vacuuming, toy organization, fastening zippers, turning pants inside out, and coffee preparation.
Why this matters
Adoption of non‑industrial robots has been limited by brittleness and poor generalization. Industrial robots succeed in tightly scripted environments where generalization is unnecessary; home robots need much greater flexibility. The Anthropic and Sunday findings suggest that increasing the intelligence of general‑purpose base models, combined with targeted in‑house tuning, can narrow the generalization gap and enable broader, more reliable robot behavior.
Both teams caution that challenges remain — especially rare failure modes and edge cases that surface only after repeated real‑world use — but the reported results indicate that model scaling plus focused fine‑tuning can deliver tangible improvements for real‑world robotics.
Sources
Anthropic: Project Fetch: Phase Two (Anthropic blog)
Sunday Robotics: ACT‑2 Preview: Generalizing Reliability
Videos: Sunday YouTube channel
(Synthesis based on the teams’ public blog posts and demonstrations.)



