Tools

AI-generated text

Interactive AI Avatars and the Corporate Future of Digital Twins

Synthesia has created interactive, deterministic AI avatars that can answer a limited set of questions and be used for training, PR, and enterprise interaction.

Interactive AI Avatars and the Corporate Future of Digital Twins

When Alexandru Voica, Head of Corporate Affairs at the video-generation startup Synthesia, sent me a link this summer to the newest member of their PR team, I was surprised: it was an interactive virtual avatar of him trained to answer common press questions about Synthesia’s activities and how the technology works. The day before, I had been on a panel where PR people asked whether I minded AI-generated pitch text. Voica’s avatar felt like a step beyond that — for me it was an advanced use case of AI in PR.

About Synthesia and its products

Synthesia, originally based in the U.K. and now with a presence in New York, is among the fast-growing digital avatar startups alongside D-ID, HeyGen, and Colossyan. Earlier this year the company reached a $4 billion valuation, and last year it reported having crossed $100 million in annual recurring revenue (ARR).

The company enables enterprises to build interactive training videos using AI avatars. It recently launched a product called Roleplay Sessions that lets employees practice scenarios — for example sales pitches — with an interactive avatar that responds and scores their answers.

How my avatar was made

In September I visited Synthesia’s new New York office. When they asked if I wanted my own AI avatar, I immediately said yes. The production took place in a small studio inside the office: they took multiple photos of me and recorded a two-minute sample of my voice. I consented to the creation of these avatars and my digital twin was produced.

They made two kinds of avatars for me: a personal avatar that simply reads whatever script I give it (available with and without glasses), and two interactive avatars (also with and without glasses) capable of listening and responding.

For the interactive avatar we selected one of my articles — a piece about why venture-backed startups commit more fraud than non-VC-backed startups — and trained the avatar to answer questions only about that story. That avatar is deterministic: it will respond only with answers related to the trained article.

Technical stack in brief

Synthesia’s interactive avatar combines several models: a voice-to-text model converts speech to text; an agentic language model interprets the text and can choose actions; a text-to-voice model synthesizes audio responses; and a video model (built by Synthesia) animates the avatar during speech. Synthesia uses its own video and voice models but allows customers to pick alternatives from other labs such as Cartesia, ElevenLabs, Google, or OpenAI. Enterprises can also choose their preferred cloud for hosting or pay Synthesia to host the avatars.

Synthesia’s product lineup can be described as three types:

  • A video-creation and distribution platform with classic avatars where a user types a script and the avatar reads it aloud.
  • An agentic platform called Sessions for interactive roleplays or surveys with avatars.
  • An API platform that lets customers combine Synthesia’s video and voice models with other services to build interactive avatars or other products.

First impressions and reactions

It took a few days to produce my avatars. With the personal avatar I first typed a generic script to test my AI voice; I had it talk about fall arriving in New York. The synthesized voice was fairly accurate and did not pick up a hoarseness I had in the recording.

Non-technical friends found the personal avatar interesting and creepy. The interactive avatar, trained only on the venture-fraud story, would redirect any off-topic or personal questions back to the article. My friends thought the interactive voice was less similar to mine than the personal avatar, though still close enough to feel unsettling. My mother called it “amazing” and tried to stump it with questions only she or my father would know; the model did not answer those and always steered them back to the story.

Implications for journalism and corporate life

This experience made me think about what journalism’s future could look like. Would audiences accept news presented by avatars? One investor I asked immediately rejected the idea. There is notable pushback today against low-quality AI output that has spread across social media and other news-sharing platforms. Others were less definitive — questions remain whether avatars could augment or even replace journalists, and whether CEOs would prefer speaking to an AI avatar of a journalist rather than the person.

For me, journalism is about connecting with people, reporting, and researching, and — above all — trust. Trust seems difficult to outsource entirely to AI.

Outside journalism, cloning oneself might be appealing in corporate settings: you needn’t catch up on work after a holiday if a version of you can always respond to routine queries. How avatar use evolves across corporate America remains to be seen.

Closing thoughts

After the novelty wore off, I studied my avatar and found myself waiting — perhaps for it to blink, to say something new, to smile, or to indicate that it "knows". Deterministic avatars will never do that, but a nondeterministic, chatbot-powered twin could provoke different psychological effects. I told one investor that my generation, Gen Z, might not easily adapt to digital twins; they feel like science fiction brought to life. Personally, avatars feel less jarring than humanoid robots — at least with an avatar you can log off if things get strange.

For now, Synthesia has produced a personal avatar for me that summarizes top stories on our site — my first digital representative.