1 Sep 2026
Citation Bureau
Vol. I
No. 300

AI improvements in video models come primarily from language model gains and data/pipeline bug fixes, not from novel video architecture advances.

The case

According to a rumor, Noam Shazeer dramatically improved a Google DeepMind training run by looking at the codebase and finding bugs because he knew where to look.

“There's a rumor that right after Noam Shazeer joined GDM, which he's now left, they had a new really good training run, and the reason why is that Noam Shazeer just looked at their code base and found a bunch of bugs, because he just knew where to look.”
Ryan Greenblatt · 11 Aug 2026

Improvements in video model quality come primarily from language model gains, not from video model architecture itself.

“I have a pretty big claim the visual intelligence are actually mostly coming from language like these video models especially from now since the diffusion model technology is more mature like every time you see there's some improvement on these models I would say mostly the gain comes from language model not coming from the v the video model itself like the video description models themselves.”
Ethan He · 1 Jun 2026

The pushback

Text and pictures will remain the dominant data modalities for AI for at least 10 years, surpassing world models or video.

“Nothing's going to beat text and pictures in 10 years, right?”
Mark Cuban · 21 Jul 2026

Topics

AI ModelsLLMsVideo Generation

Citation Bureau · compiled from attributed public discussion. Last updated 2026-08-11.