OpenAI · Official · Blog Article
Video generation models as world simulators
Summary
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes. Our largest model, Sora, is capable of generating a minute of high fidelity video. Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to OpenAI News.
Hype check
2/5Worth a look
Rated low: routine. Worth knowing, not worth rearranging your day for.
Prior coverage our earlier items on the same thing
- New embedding models and API updatesOpenAI News
- Weak-to-strong generalizationOpenAI News
- New models and developer products announced at DevDayOpenAI News
- OpenAI Red Teaming NetworkOpenAI News
- OpenAI partners with Scale to provide support for enterprises fine-tuning modelsOpenAI News
- Function calling and other API updatesOpenAI News
Questions people ask
- Where can I read the full story?
- On OpenAI News. The "Read this on OpenAI News" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for ChatGPT?
- We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and…
More from OpenAI News 1017 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to OpenAI News.


