AAI Search
← 返回首页
话题

多模态

353 条相关资讯 · 来自历史归档

模型发布/更新NEW
昨天
匿名 AI 模型再现:又一 Alpha 上线 OpenCode,线索指向智谱

IT之家 9 月 5 日消息, 在 Ox Alpha 确认为 GLM-5.3-Flash 原生多模态模型后 ,又一款名为 Omen Alpha 的模型昨日(9 月 4 日)上线 OpenCode Go 订阅计划 ,种种线索表明同样来自智谱公司。 费用方面,该模型目前仅对 OpenCode Go 订阅用户开放,月费 10 美元 (IT之家注:现汇率约合 67.…

AI 点评 · 智谱连推匿名模型,或预示多模态矩阵加速扩张,订阅模式成新战场。

产品发布/更新
8/31 21:19
sutiankang/aster

Learn and build modern AI in one native PyTorch stack: LLMs, VLM/VLA, diffusion & flow matching, world models, and agents—with LoRA, RL post-training, distillat…

技巧与观点
8/27 07:52
Qwen3.8-Flash-Next

Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty bi…

论文研究
8/27 04:00
UI-Venus-2 Technical Report

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due t…

模型发布/更新
8/26 13:44
not much happened today

**Z.ai** launched **GLM-5.3-Flash**, a natively multimodal model with a **1M-token context window**, **320B total parameters / 18B active parameters**, under the **MIT License**. I…

产品发布/更新
8/25 10:46
Tencent/WeMM-Embedding

WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.

论文研究
8/15 04:00
MOSS-VL Technical Report

We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the languag…

行业动态
8/14 23:13
阿里开源 Qwen3.8-27B 模型,编程、办公场景表现超越 Qwen3.7-Plus

IT之家 8 月 14 日消息,今天(14 日)晚间,阿里“千问大模型”公众号宣布,Qwen3.8-27B 模型正式开源,所有开发者、科研机构和企业均可自由下载、部署和使用。 IT之家附体验地址: Hugging Face 魔搭社区 根据介绍,27B(270 亿参数)是 全球 AI 社区呼声最高 的模型尺寸,Qwen3.8-27B 是原生多模态稠密(Dens…

AI 点评 · 开源27B尺寸兼顾性能与部署成本,编程办公双场景越级表现,或成中小开发者首选。

产品发布/更新
8/14 02:57
ysr666/dsh-vision-router

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG t…

产品发布/更新
8/10 22:32
cobusgreyling/Muse-Glimmer

Introducing Muse Glimmer: open-weight 30B agentic multimodal model that runs on your device (Meta). Interactive local agent lab + guide. Apache 2.0 · on-device…

技巧与观点
8/10 13:44
not much happened today

**Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on loc…

产品发布/更新
8/5 14:51
KuaaMU/mcp-vision-bridge

MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude…

模型发布/更新一手源
8/4 20:00
Introducing Shieldstral.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

模型发布/更新
8/4 14:32
Kimi K3与DeepSeek V4之间,隔着原生多模态的时间差

文 | 李炤锋 编辑 | 张雨忻 “长链任务如果只通过代��层面的反馈,误差可能会不断累积,最终效果会非常差。”谈及原生多模态的意义,一位多模态研究员表示,“视觉是一种更准确的反馈,也更贴近用户意图。” 过去一年,Coding与Agent能力不断改写大模型的排名,也成为AI最快兑现商业价值的场景之一。与此同时,随着Agent开始接管更多长链任务,越来越多的通…

技巧与观点
8/4 13:44
not much happened today

**Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning…

模型发布/更新
8/3 13:44
Qwen 3.8 Max

**Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing…

行业动态
7/31 23:23
国内唯一做多模态长记忆的公司,融资数千万,押注主动智能|涌现新项目

文|王欣逸 编辑|张雨忻 一句话介绍 国内唯一做多模态长记忆的公司——丘脑智能,推出原生多模态记忆基座,押注AI从通用走向个性化,最终走向主动智能。 主动智能,指的是AI能在足够了解用户的基础上,在合适的时间、以恰当的方式主动跟用户交互。要实现主动智能,Memory是必须要跨过的门槛。 融资情况 近日,丘脑智能已完成数千万元种子轮融资,投资方包括深圳一线基金…

AI 点评 · 多模态长记忆是主动智能关键门槛,资本押注稀缺赛道,看点十足。

论文研究
7/29 04:00
Metis: Memory Foundation Model

Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However…

论文研究
7/28 04:00
Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the ar…

论文研究一手源
7/28 01:46
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted datab…

AI 点评 · 评测视觉语言模型理解结构化ER图的能力,填补了AI辅助数据库设计的评估空白。

论文研究
7/27 04:00
Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states…

产品发布/更新
7/24 18:16
亚马逊升级 Alexa+ AI 助手,可完成购物、餐厅预订、叫车等任务

IT之家 7 月 24 日消息,亚马逊官方今日宣布升级 Alexa+ AI 助手, Alexa+ 支持自然对话、信息记忆和任务执行 ,可帮助用户管理日程、总结邮件、生成播客、控制智能家居,并完成购物、餐厅预订、叫车等日常任务。 Alexa+ 已从单纯的语音助手升级为一个具备多模态能力(视觉识别、生成式 AI 创作)、超强记忆力,并打通了众多第三方现实服务接口…

AI 点评 · AI助手从对话升级到行动,打通现实服务,智能助手终于能真正“办事”了。

行业动态
7/24 08:00
8点1氪丨段永平称10年内大概率不会卖泡泡玛特;中国数学家王虹、邓煜获得菲尔兹奖;宜家回应甩卖8处物业:不代表退出中国市场

今日热点导览 混元多模态理解负责人胡瀚离职创业,原团队或将聚焦世界模型 极氪回应“海外锁车”事件 客服回应滔搏暴力打折甩卖耐克库存:没有收到降价通知 哈兰德和亚马尔2.2亿欧元身价破纪录 张雪峰女儿再接手三家公司股份 TOP 3 大新闻 段永平:10年内大概率不会卖泡泡玛特 7月23���,段永平在社交媒体平台雪球上发表了他对近期投资操作的最新想法。雪球上有…

AI 点评 · 段永平罕见长线看好,泡泡玛特投资逻辑获大佬背书,市场风向标意义显著。

论文研究
7/24 01:59
3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reason…

AI 点评 · 将隐式与显式几何融合进视觉语言模型,突破2D局限,实现精细3D空间推理。

论文研究
7/24 01:35
MIRROR: Learning from the Other View for Multi-Modal Reasoning

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diag…

AI 点评 · 多模态推理新突破,利用跨视角学习提升视觉语言模型几何问题解决能力。

行业动态
7/23 16:07
独家|混元多模态理解负责人胡瀚离职创业,原团队或将聚焦世界模型

文 | 周鑫雨 编辑 | 张雨忻 《智能涌现》独家获悉,近期,腾讯混元多模态理解负责人胡瀚提出了离职。 此前,他曾担任微软亚洲研究院视觉计算组首席研究员。2025 年初加入腾讯后,负责视觉大模型的研究。在后续的调整中,他加入大语言模型部旗下的“Frontier”前沿技术研究组,负责多模态理解的相关研究,汇报给姚顺雨。 据了解,胡瀚还曾承担世界模型的研发工作。…

AI 点评 · 顶级多模态人才离职创业,折射出世界模型赛道竞争白热化。

论文研究
7/22 04:00
ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) polic…

AI 点评 · 将自然语言描述与移动追踪结合,突破传统视觉追踪限制,提升具身智能的实用性与交互性。

技巧与观点
7/21 22:22
Nativ: Run AI models locally on your Mac

Nativ: Run AI models locally on your Mac Prince Canuma is the developer behind the excellent MLX-VLM Python library for running vision-LLMs using MLX on a Mac. I'm really excited a…

AI 点评 · 本地运行AI模型,保护隐私且无需联网,苹果用户的新利器。

产品发布/更新
7/21 15:32
Today-Hbw/rag-platform

Enterprise multi-source RAG platform with multimodal embedding, hybrid vector/BM25 retrieval, MCP & RBAC support

行业动态
7/17 18:45
2026最受投资人关注人工智能/具身智能企业50揭晓

人工智能正在进入一个新的产业周期。 过去一年,大模型能力持续演进,生成式AI、多模态交互、智能体等技术方向快速推进;而具身智能也从早期的技术探索阶段,逐渐步入产业验证的深水区,机器人开始成为人工智能与现实世界的重要载体。 市场率先给出了回应。据36氪研究院测算,中国具身智能市场规模已从2018年的2133亿元增长至2025年的9150亿元,2026年有望突破…

行业动态
7/17 15:22
腾讯发布具身 VLM 基座模型 Hy-Embodied-VLM-1.0,A3B 规模整体性能接近上一代 A32B 模型

IT之家 7 月 17 日消息,腾讯 Robotics X 实验室、福田实验室联合腾讯混元打造的第二代具身 VLM 基座模型 Hy-Embodied-VLM-1.0 昨日正式发布。 官方表示,在覆盖 37 个评测任务的具身能力评测体系中,Hy-Embodied-VLM-1.0 在物理状态理解、动作 — 变化推理、时序与自适应推理三大维度分别取得 68.6、6…

AI 点评 · 参数规模锐减十倍,性能逼近上一代大模型,具身智能走向高效实用化。

行业动态
7/17 10:39
36氪首发 | 港科大博士创业做机器人全身触觉系统,红杉、瓴智、智元共同押注

作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,全身多模态融合触觉解决方案公司模感科技(MoSense)近日完成数千万元天使轮融资,投资方包括红杉中国、高瓴创投及智元机器人。本轮融资资金将主要用于加速研发、团队扩充、算力投入及量产测试体系建设。 模感科技成立于2026年5月,总部注册于上海,在深圳前海设有研发中心,聚焦机器人全身多模态触觉感知系统研发。公司正式…

AI 点评 · 红杉、高瓴、智元联手押注,机器人触觉赛道技术壁垒高、应用前景广。

论文研究
7/17 01:38
Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design…

AI 点评 · 多模态医疗AI的可靠性设计突破,用胃肠镜案例揭示可解释性关键。

行业动态
7/13 16:34
字节探索自动驾驶,Seed世界模型团队负责|36氪独家

36氪从多位产业人士处获悉,字节跳动正探索进入自动驾驶领域。这一项目目前由Seed旗下周畅的世界模型团队负责。据了解,Seed旗下不仅有周畅的多模态模型、世界模型等团队,还有大语言模型方向。 而自动驾驶与世界模型的技术路线有交叠之处。 另有消息人士告诉36氪,业务方向上,字节有意布局的自动驾驶场景有无人物流,这一业务隶属于字节旗下的火山引擎汽车行业线。 部分…

行业动态
7/13 10:39
对话Om AI赵天成:多年坚守,押注物理AI原生的「流式」未来

一个从未见过监控画面的多模态模型,却比在监控数据上练了多年的小模型“老将”更懂监控。这不是科幻电影,这是2023年Om AI联汇的一场“无心插柳”,也是CEO兼首席科学家赵天成博士更加坚信“多模态训练方式能为物理开放世界带来泛化性”的关键节点。彼时,AI行业正在追求以大语言模型为核心的生成式AI。 三年后,这个多模态模型演变成了VLX——全球首个面向物理AI…

AI 点评 · 押注物理AI原生流式架构,多模态泛化性突破传统小模型局限。

产品发布/更新
7/12 23:12
Meta 发布多模态推理模型 Muse Spark 1.1,强化 AI 智能体任务能力

IT之家 7 月 12 日消息,Meta 于 7 月 9 日正式发布适用于 AI 智能体的多模态推理模型 Muse Spark 1.1 版本,重点提升了模型在智能体任务中的规划、协同与执行能力,并增强了工具调用、代码开发、应用操作能力。 Meta 表示,Muse Spark 1.1 强化了多智能体协作机制,由主智能体负责收集信息、制定计划,再将任务拆分并分配…

AI 点评 · 多智能体协作机制是AI落地的关键突破,Meta这次强化了任务拆解与分工能力。

论文研究
7/11 00:42
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often in…

AI 点评 · 用强化学习微调动作块,提升机器人接触操作的鲁棒性和泛化能力。

产品发布/更新
7/7 16:38
caseclose/cma-harness

Cognitive-structured Multimodal Agent (CMA-Harness): a memory-centric agent for long-horizon multimodal understanding, generation, and editing — externalizing v…

行业动态
7/7 09:05
用AI“复刻”人类细胞、预判药效,「华源智因」获千万级人民币种子轮融资|36氪首发

文|胡香赟 编辑|海若镜 36氪获悉,AI虚拟细胞(AIVC)企业华源智因近期已完成千万级人民币种子轮融资。本轮融资由水木创投领投,募集资金将主要用于多模态测序底层技术迭代,进一步拓展与头部三甲医院的合作,以及团队扩充等。此外,华源智因团队已计划启动新一轮融资。 华源智因创始团队由资深医药产业从业者、计算生物学研发人员组成,并邀请到深圳国家基因库等单位专家组…

产品发布/更新
7/6 10:52
EPFL-VILAB/Modus

MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities [ICML 2026]

产品发布/更新
7/5 13:38
zhiweio/EagleRAG

Search knowledge by what documents mean and how they look — not one or the other.

产品发布/更新
7/4 16:20
JT-Sun/UAVReason

🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation

论文研究
7/2 04:00
Gemma 4 Technical Report

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite feat…

论文研究
6/30 04:00
Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigat…

论文研究
6/29 04:00
Orca: The World is in Your Mind

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multimodal readout interf…

产品发布/更新
6/27 14:12
odebo/CC-Vision

Give non-multimodal Claude Code main models the ability to see pasted screenshots — a ~200-line UserPromptSubmit hook.

产品发布/更新
6/23 15:11
vancyland/DataClaw0

DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).

产品发布/更新
6/17 13:11
kaistmm/SeeandSniff

[ECCV 2026] Official Pytorch implementation for See & Sniff: Learning Visuo-Olfactory Representations

产品发布/更新
6/13 14:50
ratschlab/DeepSpotM

Multimodal foundation model predicting transcriptome-wide virtual spatial transcriptomics from histology.

模型发布/更新一手源
1/26 19:08
Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL!

QWEN CHAT GITHUB HUGGING FACE MODELSCOPE DISCORD We release Qwen2.5-VL, the new flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL. To tr…