DELE-w0.5 重新定义了世界模型在机器人操作中的作用。
视频生成
共 94 条相关资讯 · 来自历史归档
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." Instead of running code, the system generates the user interface frame…

You can now create decent video faster than you watch it. This is the start of... something. We’re not sure what.
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, langu…
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native join…
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and…
We look at Gemini Omni 1.1 Flash, Google's production update to its native multimodal video generation and editing model. We break down what changed: scene extension now reads up t…
AI 点评 · 原生多模态视频编辑升级,场景延长与4K增强直击创作痛点,实用价值显著。
Official implementation of "LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation."
国产算力跑通视频生成规模化应用
AI 点评 · 国产算力突破视频生成瓶颈,规模化落地价值凸显。
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cu…
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they ei…

Google's Gemini Omni 1.1 Flash video model now analyzes up to ten seconds of existing footage instead of just the last second for more consistent scene extensions. Scenes can be ex…
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible tr…
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: object permanence, the…
Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing meth…
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should ado…
少数派的近期动态新一季少数派会员启航,更新权益,更多惊喜,还有实体纪念卡。点击了解能让AI助手通过自然语言指令直接与您的Quote/0摘录墨水屏交互的DotSkill已上线。点击了解你可能错过的文章快 ... 查看全文
Alibaba's video generation model Wan3.0 creates clips up to 30 seconds long from text, PDFs, and PowerPoint files. A 30-second 1080p clip costs $6. Alibaba's quarterly profit dropp…
AI 点评 · 多模态输入直出30秒视频,定价透明,国产大模型商业化提速信号明确。
8月24日,阿里巴巴视频生成大模型Wan3.0正式上线。
With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engin…
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended…

Current world models like Sora or Genie only simulate physics and ignore what people think, want, or feel. The new "Mental World Modeling" framework adds mental variables like beli…
AI 点评 · 忽视人类信念的世界模型,预测行为必失准,新框架补足心智维度。
AI 点评 · 视频生成跨界设计工具,直击Adobe与Canva的护城河,行业格局或将生变。
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a cohe…
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may…
Official RMD implementation for cross-resolution diffusion distillation and accelerated video generation.
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic ground…

AI production companies like Promise are setting up shop around Hollywood's historic studios, using real-time backgrounds and other AI tools to cut film costs. Netflix already uses…
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on ch…
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min…
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, howev…
Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control…
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos pr…
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces ar…
LTX-2.5 brings frontier video generation to local NVIDIA hardware: 6.8-second clips, native multishot, day-one ComfyUI, open weights. The post The Video Production Stack Now Fits o…
AI 点评 · 本地开源视频生成首次媲美前沿水平,多镜头与ComfyUI支持将大幅降低创作门槛。
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video genera…
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60…
8月6日,阿里巴巴视频生成大模型Wan 3.0开启公测
Black Forest Labs 正式推出 FLUX 3 视频生成模型,研究人员发现 iCloud Private Relay 存在 IP 泄露风险等。 查看全文

Black Forest Labs has launched FLUX 3 Video, which generates Full HD clips up to 20 seconds long with native audio and lip-synced dialogue in more than 14 languages. It can also re…
10秒1080P,成本只要5毛钱
Open agentic prompt-expansion harness for image and video generation, bridging polished demos, public APIs, and deployable workflows.
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inher…
狂刷3小时!
AI 点评 · AI巨头沉迷短视频,揭示人性与算法共谋的荒诞现实。
🎬 Curated MiniMax H3 video generation prompts — cinematic, ads, anime, UGC, product videos, and more. Includes playable examples and creator attribution.
文|王毓婵 兰杰 编辑|乔芊 36氪独家获悉,曾爱玲入职哔哩哔哩(下称“B站”),担任AI视频生成业务负责人,向CEO陈睿汇报。 36氪就此事向B站方面求证,对方暂无回应。 曾爱玲 B站此前已公开表示,AI投入主要聚焦视频理解、视频推荐和辅助视频创作等方向。曾爱玲入职后,或将参与相关业务。 曾爱玲的个人主页显示,她曾在腾讯混元&AI Lab团队和国际数字经济…
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly f…

7 月 28 日,OpenAI CEO 萨姆 · 奥尔特曼表示,自己曾因研究 TikTok 的产品机制而逐渐沉迷,甚至在一个周六下午连续刷了约 3 个小时,最终不得不删除这款应用。 奥尔特曼在《Relentless》播客节目中称,OpenAI 此前筹备一款类似 TikTok、以 AI 生成视频为主要内容的社交应用 Sora。为了了解短视频平台的运行方式及用户…
AI 点评 · 科技领袖也难逃算法沉迷,凸显AI时代注意力争夺的普遍性与产品设计的力量。
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variation…
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score…
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this…
The AI lab Midjourney continues to expand its purview beyond image and video generation.
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a developer prioritizes quality, speed or cost…
AI 点评 · 解决生成式媒体模型选择难题,自动平衡质量、速度与成本,实用价值高。
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movem…
AI 点评 · 用图结构控制多对象交互,精准生成动态视频,突破文本与运动控制的局限。
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SA…
Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and imme…
Kandinsky WM 1.0 — a family of models for Physical AI. Image-to-video generation for autonomous driving, robotics & general physics.
A curated list of reinforcement learning, preference optimization, and reward-driven post-training and alignment methods for video generation.
Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiote…
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-wor…
AI 点评 · 训练数据对文生视频影响被低估,该研究首次系统控制变量,揭示数据质量比规模更关键。
CLI & async Python library for free AI chat, image & video generation.
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition r…
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on…
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisio…
文 | 周鑫雨 编辑 | 张雨忻 智能涌现从多个独立信源处获悉, 截至 2026 年 7 月, 智谱的 ARR(年度经常性收入)已经达到 10 亿美元 。 截至发稿前,针对上述信息,智谱未回复。 过去一年,AI Coding 和视频生成模型已经成为全球造血能力最强的 AI 赛道。 海外,Anthropic 的 Claude Code 仅发布半年,ARR 就飙…
AI 点评 · 智谱ARR半年飙升15倍达10亿美元,印证AI商业化进入爆发期,行业格局加速重塑。
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the o…
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, bu…
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a dat…
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-dr…
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning,…
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can li…
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consis…
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in c…
AI 点评 · 视频生成模型突破任务局限,迈向通用视觉,预示AI基础模型新范式。
视频生成的下一站,或是机器人大脑
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct f…
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low…
2026年,AI视频生成赛道已迈入全面爆发的成熟竞速期,Seedance 2.0 的出圈更是让AI视频变成一场“全民狂欢”。 短短两年间,AI视频从最初几秒的碎片化模糊画面,到如今分钟级长视频的连贯叙事、真实物理世界的精准还原,AI视频工具的迭代速度远超预期,AI视频工具完成了从“能用”到“好用”再到“专业”的三级跳,让创意落地的门槛降至新低——专业团队能用…
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visu…
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these m…
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, gener…

IT之家 7 月 3 日消息,据 AI 普瑞斯消息,字节豆包视频生成模型 Seedance 2.5 预计 7 月 6 日上线体验中心,将在一周后开放 API 。 据IT之家此前报道, 字节豆包视频生成模型 Seedance 2.5 发布于 6 月 23 日,该模型目前处于全球企业内测阶段。 据介绍,Seedance 2.5 在单段生成长度、多素材参考、视频编…
AI 点评 · 字节视频模型快速从内测走向开放,行业落地节奏加快,值得关注其能力上限。
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models exce…
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions…
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ)…
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ)…
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world models is the lack of me…
Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between t…
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or…
Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. T…