Model launches

Leaked Gemini Omni shows improved on-screen text stability in video but mixed visual quality

A leaked build of Google’s Gemini Omni, surfaced via the Gemini app, reveals a native video model capable of generation and chat-driven editing.

A leaked instance of Google’s Gemini Omni, surfaced through the Gemini app, reveals a native video model designed for generation and chat-based editing. Among the demos, the standout was not flashy anime sequences or glossy visuals but a pragmatic example: a professor writing formulas on a chalkboard where the written text remained readable and stable throughout the clip.

Practical implications

Text consistency has been one of the most persistent failure modes in AI video generation—characters, numbers and symbols often warp or become illegible across frames. The leaked Gemini Omni demos suggest the model targets this issue directly: in the shown example, the chalkboard formulas remained coherent and legible, which could represent a genuine improvement over current video models.

The leak also showcased real-time editing capabilities such as object replacement and watermark removal triggered by a simple textual prompt. These features point toward a shift in video editing from traditional timeline-based tools toward conversational, prompt-driven controls.

Limitations visible in the leak

The demos are not uniformly strong. The anime-style samples and general image quality sometimes appear rough or poor. The leak does not claim that Google has solved all problems in video generation; visual fidelity and certain types of generative outputs still show notable weaknesses.

Why this matters

The leaked material does not assert that Google has fully solved video generation. However, achieving readable, stable on-screen text addresses a particularly challenging bottleneck and could have outsized practical impact compared with the relatively crude visual demos. If Gemini Omni reliably renders stable text in motion, that would help applications in education, technical tutorials, and editing workflows—even as overall graphic quality continues to improve.

Key takeaway

The leak suggests Google may have made meaningful progress on a difficult technical constraint—stable, legible text inside generated video—while broader image quality and style consistency remain works in progress.