Researchers at the Okinawa Institute of Science and Technology (OIST) have developed a new training approach for robots that rewards curiosity and leads to faster acquisition of language-based tasks compared with conventional methods. The team published their findings in the journal Science Advances.
Method: PV-RNN combined with reinforcement learning
Instead of relying on large language model-style next-word prediction, the researchers used a brain-inspired architecture called PV-RNN (Predictive-coding Variational Recurrent Neural Network). PV-RNN is designed to help an agent maintain a stable internal model of the world while correctly understanding its environment.
To encourage discovery, the team combined PV-RNN with reinforcement learning. Robots received external rewards for completing language tasks and internal rewards for satisfying curiosity by exploring unfamiliar situations that required updating their internal models. Balancing stability and exploration in this way allowed the robots to discover more efficient solutions for complex tasks.
Faster learning and spontaneous, playful behaviour
In the experiments, curiosity-driven systems required roughly half the time to learn language-based tasks compared with robots trained by traditional methods. During training, the curious robots displayed unexpected, playful behaviour: even after mastering assigned tasks they continued to experiment, toppling objects despite receiving no direct instruction to do so. The researchers found that this spontaneous exploration accelerated language learning and likened it to play-driven knowledge acquisition seen in young children.
Diverse linguistic exposure is necessary
The team also found that curiosity alone was insufficient; it needed to be paired with a rich and varied linguistic environment for learning to be effective. Robots exposed to only 48 language combinations achieved about 25 percent generalization, while those trained with 180 combinations reached around 85 percent generalization. This contrast demonstrates that greater variety in linguistic exposure substantially improves the ability to understand unseen instructions.
U-shaped learning curve and handling exceptions
The researchers observed that the curiosity-driven robots exhibited a U-shaped learning curve similar to patterns documented in child language acquisition. To test this, they deliberately reversed the meanings of two language tasks. Initially, robots learned the exceptions correctly, but as they generalized broader linguistic rules their performance temporarily degraded before recovering. The emergence of this behaviour surprised the team because the AI model did not include an explicit mechanism for handling exceptions or overgeneralization.
Implications and conclusions
OIST's results suggest that rewarding curiosity and continuously updating internal models, combined with reinforcement learning, can produce faster, more adaptable and human-like behaviours in robots. The study highlights that effective language learning requires not only intrinsic motivation but also diverse linguistic input. The researchers believe such approaches could contribute to more adaptive artificial intelligence systems in the future.
Publication and coverage
The work was reported in Science Advances and has been summarized in technology outlets such as Interesting Engineering and HVG.



