通义千问 Qwen
共 29 条相关资讯 · 来自历史归档
细粒度标签+ 20 种方言

IT之家 7 月 17 日消息,今日,阿里通义实验室发布 Wan-Streamer v0.2 模型,让 AI 像真人一样边听、边看、边回应。 IT之家从官方介绍获悉,它是面向实时双工交互的端到端全模态理解与生成模型,把“听、看、说、演”统一进单个 Transformer 中。 极致低延迟: 端到端响应延迟 550ms (200ms 模型延迟 + 350ms…
AI 点评 · 端到端延迟仅550ms,让AI视频通话逼近真人实时交互体验,技术突破极具商用价值。
阿里发布实时语音模型 Qwen-Audio-3.0-Realtime、腾龙发布 12-20mm F2.8 镜头等。 查看全文
AI 点评 · 苹果AI入华合规,阿里实时语音模型发布,硬件与AI双线突破值得关注。
The deal, which was rumored to be in the works last year, marks an important step for Apple's AI ambitions in a key market.

IT之家 7 月 15 日消息,@PrismML 官方账号今天(7 月 15 日)发布博文,宣布推出 Bonsai 27B 模型,基于 Qwen 3.6 27B 模型微调, 在保留 90% 智能水平的情况下,可以在 12GB 内存的 iPhone 上原生运行。 Qwen 3.6 27B 模型进一步提升本地 AI 的能力,重点涵盖多步推理、结构化工具使用、长上…
AI 点评 · 端侧大模型突破内存瓶颈,iPhone变身AI终端,本地智能体验迎来质变。
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks:…
AI 点评 · Qwen模型线性注意力优化,或成高效大模型推理新突破口。

IT之家 7 月 10 日消息,科技媒体 The Information 昨日(7 月 9 日)发布博文,报道称苹果公司正接洽 PrismML 初创公司, 评估在 iPhone 上直接运行更大规模 AI 模型的可行性。 IT之家查询公开资料,PrismML 是加州理工衍生出来的 AI 初创公司,核心突破为原生 1-bit 模型压缩技术,模型体积可压缩至全精度…
AI 点评 · 苹果布局端侧大模型,小公司技术或成关键突破。
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowledge. We identify th…
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instructio…
Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen…
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling…
Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search, target tracking, a…
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether t…
边走、边看、边思考
We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual tra…
甩开视觉内卷
Few-step distillation has become an effective strategy for accelerating advanced visual generative models, yet prior work has largely focused on distillation objectives. In this work, we revisit few-s…
下一代CUA训练范式
AI 点评 · CUA训练范式直击Agent工具选择瓶颈,复旦与通义合作开辟新路径。
下一代CUA训练范式
AI 点评 · 解决Agent工具选择难题,复旦与通义提出全新训练思路,推动智能体实用化。
We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while preserving baseline task performance. On a Qwen 2.5 1…
AI 点评 · 揭示LoRA微调模型的安全漏洞,提出首个令牌级后门检测方法,对防范AI供应链攻击意义重大。
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks,…
AI 点评 · 统一视觉语言与动作建模,突破单一任务限制,推动机器人跨场景泛化能力跃升。