
IT之家 7 月 21 日消息, 荣耀 Robot Phone 已正式开启预约 ,该机搭载第五代骁龙 8 至尊版芯片,其核心创新在于机身顶部集成了一套行业最小的四自由度(4DoF)钛合金机械云台系统。云台系统配备微型电机,体积比主流方案缩小 70%。 IT之家注意到,荣耀全球首席营销官关海涛今日晒出了自己的荣耀 Robot Phone 真机,并分享了新机的拍…
共 301 条相关资讯 · 来自历史归档

IT之家 7 月 21 日消息, 荣耀 Robot Phone 已正式开启预约 ,该机搭载第五代骁龙 8 至尊版芯片,其核心创新在于机身顶部集成了一套行业最小的四自由度(4DoF)钛合金机械云台系统。云台系统配备微型电机,体积比主流方案缩小 70%。 IT之家注意到,荣耀全球首席营销官关海涛今日晒出了自己的荣耀 Robot Phone 真机,并分享了新机的拍…
Gritt is coming out of stealth with $34 million and plans to automate the hardest tasks on construction sites.
Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global…
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts…

IT之家 7 月 20 日消息,科技记者 Joanna Stern 上个月撰文称,在自己的新书《I Am Not a Robot》正式出版后的几天内,她就在 Apple Books 上发现了 10 本冒充该书的盗版版本。虽然她成功让这些作品下架,但很快又有新的盗版书重新出现。 《纽约时报》记者 Kashmir Hill 也发现,亚马逊上出现了一本关于她本人的…
From open models to real-time simulation, AI and graphics breakthroughs are transforming media, content creation and robotics.
7月17日-20日,一起在WAIC2026现场,看见人工智能真正进入产业深处。 过去一年,围绕AI行业的讨论正在变得更具体。大模型能力仍在持续迭代,但外界关注的重点,已经不再只停留在模型参数、模型发布和单点能力展示上。随着智能体、具身智能、空间智能、AI基础设施等方向不断演进,行业开始更频繁地追问:AI如何进入真实流程,如何完成复杂任务,又如何在产业场景中形…
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports…
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains cha…
7月18日,在2026世界人工智能大会(WAIC)上,腾讯面向具身智能与智能体领域带来多项产品技术的升级发布。在具身智能领域,腾讯正式升级发布具身智能全栈方案,贯穿云底座、模型层、平台层与应用层,全面助力机器人本体及系统开发商提质提效;在智能体领域,基于个人与企业提效需求,推出差异化的全矩阵解决方案。其中,面向企业用户的腾讯云企业级智能体开发平台ADP4.0…
Agility is opening a new training center for its Digit robots in Fremont, California.
AI 点评 · 选择特斯拉大本营设立培训中心,Agility Robotics此举意在抢夺物流机器人市场先机。
人工智能正在进入一个新的产业周期。 过去一年,大模型能力持续演进,生成式AI、多模态交互、智能体等技术方向快速推进;而具身智能也从早期的技术探索阶段,逐渐步入产业验证的深水区,机器人开始成为人工智能与现实世界的重要载体。 市场率先给出了回应。据36氪研究院测算,中国具身智能市场规模已从2018年的2133亿元增长至2025年的9150亿元,2026年有望突破…
36氪获悉,7月17日,2026世界人工智能大会(WAIC)召开。网易旗下专注工程机械领域的具身智能品牌网易灵动,将极具科技感的“智能座舱”搬入展区,座舱内工作人员跨越千里,可远程操控真实电厂、港口的生产一线设备。同时,网易灵动还集中展现了具身智能规模化落地高危场景的最新成果,以及“人机混编”“黑灯工地”“一人多机”等自主研发技术。
AI 点评 · 人机混编技术突破高危场景,远程操控颠覆传统作业模式,具身智能商业化落地加速。

The CEO of Foundation Future Industries, which counts the president’s son as its chief strategy adviser, tells WIRED it’s exploring some “kinetic things.”
AI 点评 · 埃里克·特朗普背书的人形机器人公司涉足军事,政治与科技的交汇点值得警惕。

IT之家 7 月 17 日消息,据央视财经报道,在 2026 世界人工智能大会上,银河通用机器人创始人兼首席技术官王鹤预测,通过海量综合数据训练,具身智能基础模型能在未经专门训练的任务上直接达到 70%-80% 成功率,媲美当年数字模型的对话水平;而强大的预训练模型加上高效后训练范式,正是敲开这一里程碑的关键条件。 王鹤判断,按当前数据积累速度和模型收敛趋势…
AI 点评 · 具身智能将从实验室走向产业,2028年或成关键拐点,值得关注技术突破与商业落地。

IT之家 7 月 17 日消息,腾讯 Robotics X 实验室、福田实验室联合腾讯混元打造的第二代具身 VLM 基座模型 Hy-Embodied-VLM-1.0 昨日正式发布。 官方表示,在覆盖 37 个评测任务的具身能力评测体系中,Hy-Embodied-VLM-1.0 在物理状态理解、动作 — 变化推理、时序与自适应推理三大维度分别取得 68.6、6…
AI 点评 · 参数规模锐减十倍,性能逼近上一代大模型,具身智能走向高效实用化。
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,具身智能世界模型公司日冕开物(北京日冕机器人有限公司)近期完成连续两轮种子轮融资,融资合计金额达数亿元人民币,由鼎峰科创、远图未来、百度风投、沃衍资本、武岳峰科创、万林国际共同参与投资。同时,新一轮融资也在同步交割中。 此前融资资金主要用于自研世界模型 LaMPA 的研发迭代、强化学习体系建设,以及数据闭环和产品…
AI 点评 · 前蔚来华为核心团队跨界创业,三个月获数亿融资,凸显具身智能赛道热度。

Hyundai aims to deploy 25,000 Atlas robots starting with US factories in 2028.
AI 点评 · 人形机器人威胁就业引发罢工,折射出AI替代焦虑与技术落地的现实冲突。
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct a…
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whe…
AI 点评 · 探索大模型安全新维度,揭示文本安全与物理行动的鸿沟,为具身智能风险防控开辟新思路。

IT之家 7 月 16 日消息,荣耀 AI 首席科学家、首席 AI 官黄非昨日官宣入驻微博,并在今日分享了一段 Robot Phone 新机的 AI 功能演示视频。 黄非表示:“荣耀 Robot Phone,‘活’了。刚和团队对完最后几组 Demo,只能说, 这次我们把对未来科技的想象,真的照进了现实 。” IT之家注意到,黄非在视频中,对荣耀 Robot…
为破解具身智能行业发展瓶颈构建了新一代“进化底座”
AI 点评 · 具身智能领域迎来关键突破,从模型训练到真机部署实现全链路能力升级。
Robotaxi第一股文远知行孵化
AI 点评 · 具身智能赛道迎来首个基础设施供应商,复制英伟达与宁德时代的成长逻辑。

General-purpose robots and autonomous machines are moving from research labs to real-world mass-market deployment, creating demand for compact, power-efficient AI supercomputers ca…
AI 点评 · 英伟达推Jetson Thor,瞄准机器人量产与边缘AI,算力能效比成关键看点。
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world…
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visu…
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero…
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen en…
This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-A…
Home to leading manufacturers, robotics pioneers, infrastructure builders and iconic gaming companies, of course, Japan is one of the world’s centers of AI — building across the fu…
FSD、Robotaxi共用一套模型
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action ge…
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint lan…
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotat…
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial services ambitions, its increasingly complicated relationship with Waymo, its new AV Labs…
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surroundi…

“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.
AI 点评 · AI代理自动生成3D训练场景,大幅降低机器人数据采集成本,加速家务机器人落地。
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track the…
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consis…
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe th…
为了帮你看清具身数据行业,我们总结了以下十个行业现状
AI 点评 · 数据成AI新燃料,百亿资本涌入具身智能赛道,揭示产业变现路径。
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this int…
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification,…
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often in…
AI 点评 · 用强化学习微调动作块,提升机器人接触操作的鲁棒性和泛化能力。
AI 点评 · 蚂蚁自研具身模型,从零预训练实现动作原生,突破机器人智能边界。
开源第四弹:LingBot-VA 2.0

Amid live coding sessions and Silicon Valley optimism, the UN’s AI for Good summit wrestled with an urgent question: Can global governance catch up before the technology races beyo…
AI 点评 · 联合国AI峰会展示前沿科技,核心挑战在于全球治理能否跟上技术飞速发展。

Preclinical trial is testing the feasibility of humanoid robots in surgery.
AI 点评 · 远程操控人形机器人完成活体手术,标志着医疗自动化从概念迈向临床的关键一步。

The soft, oddly intimate home-chore robot has been given some very tactile hands.
AI 点评 · 人形机器人触觉灵巧手取得突破,家庭服务场景落地加速。

MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water.
AI 点评 · 自组装水上机器人集群,展现未来动态建筑与自主协作的潜力。
视频生成的下一站,或是机器人大脑
36氪获悉,聚焦工业制造领域的具身智能机器人企业 「昇视唯盛」(3Srobotics)正式宣布完成数亿元B轮融资,本轮融资由上海半导体产投、金桥基金领投,零一创投、新鼎资本、中关鼎华及老股东微光创投持续加码跟投。 资金将主要用于焊接具身智能大脑与自主小脑运动控制系统的迭代升级,扩大自有智造工厂产能,扩充研发与市场服务团队,加速产品全国多行业规模化落地。 昇视…
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer prec…
General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to build smarter robots with minimal real-world…
AI 点评 · 用游戏数据训练机器人,可能加速物理AI突破,降低开发门槛。
具身测评界的珠峰来了:RoboDojo
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Exist…
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visu…

Top robotics researchers and founders explain how robot autonomy is evolving.

Top robotics researchers and founders explain how robot autonomy is evolving.
“共识与非共识”的深度思辨

Open source AI has shown how quickly developers can innovate when models, data and tools are shared. Robotics has the same opportunity, but advancements in physical AI development…
来自蚂蚁灵波
AI 点评 · 开源机器人学习框架迭代,强化想象与评估闭环,加速AI物理世界应用落地。
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lac…
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy,…
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware…
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth,…
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these m…
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narro…
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmabil…
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Lang…
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current obs…
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-f…
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing design…
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensio…
一段视频,生成无限训练场景
AI 点评 · 用真实视频自动生成多样训练数据,突破仿真向现实迁移的成本瓶颈。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,具身智能公司 「光象科技」 宣布完成累计数亿元天使轮融资。 最新一轮由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、SeeFund、亿宸资本、上市公司行云科技等头部财投与产投深度参与,老股东零一创投、L2F光源创业者基金持续加注。 本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人在工…
AI 点评 · 清华系创业团队获数亿元融资,聚焦具身智能落地汽车产业,技术底蕴与产业场景结合是关键看点。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,通用餐饮具身机器人公司「影智XBOT」连续完成数亿元两轮融资——其中,A轮的2亿元融资由香港简坤资本GPTX出资,B轮融资为3-5 亿元人民币,由多支政府基金、美元基金和产业投资方共同参与出资。 这是目前餐饮垂直机器人领域规模最大的一笔融资之一。 在此之前,「影智XBOT」还完成了一轮天使融资,出资人阵容豪华——…
AI 点评 · 小米前高管创业项目获数亿融资,林斌、黎万强加持,餐饮机器人赛道热度可见一斑。
近日,光象科技宣布完成累计数亿元天使轮融资,最新一轮融资由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、See Fund、亿宸资本、上市公司行云科技等头部财投与头部产投深度参与,老股东零一创投、L2F光源创业者基金等持续加注。据悉,本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人产品的商业化交付。(每日经济新闻)
AI 点评 · 资本密集押注物理原生模型,具身智能从实验室走向量产的关键信号。
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing re…
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on loc…
今日热点导览 苹果拟于今明两年推出至少五款新iPhone AI版支付宝开放公测,上线72项办事技能 特朗普回应“利用职位牟利”:股市在涨,大家都在赚钱 LV起诉茉莉奶白,茉莉奶白被判赔1030万元 6月赴日航班取消1488个 安克创新港股上市首日破发 TOP3大新闻 证监会同意宇树科技科创板IPO注册 7月2日,证监会网站显示,同意宇树科技首次公开发行股票注…
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems s…
具身智能数采迎来了新范式
Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To reduce policy-call frequency while preserving temporal coherence, most generative pol…
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, an…
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted…
We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and the physical hardwa…
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world ro…
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembl…
全新的持续学习范式
四大互联网分别领投
Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts to a different but similar robot (e.g., from Pand…
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling,…
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary prefe…

When Jaiveer Singh talks about robots, he doesn’t begin with spectacle. He begins with infrastructure: the boards inside machines, the software that lets developers see through a r…
大公司: 保时捷中国回应取消多门店销售授权 近日,有媒体报道称,安徽芜湖、山东济宁、江苏淮安及广西南宁部分保时捷中心,将于6月30日终止经销业务,这意味着上述门店将失去保时捷官方独立销售授权,此事引起广泛关注。6月30日,保时捷中国向记者表示,山东济宁、江苏淮安、广西南宁兴宁保时捷中心将于6月30日终止经销业务,安徽芜湖保时捷中心将于7月31日终止销售业务,…
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,具身智能公司纽娲机器人近日完成5000万元天使轮融资,由蓝湖资本领投,不同资本、共青城朴一投资跟投。两个月前,纽娲机器人曾完成由Plug and Play中国基金领投的种子轮融资。 纽娲机器人(下称“纽娲”)成立于2026年2月,半年不到的时间已先后获得多家财务、产业基金投资。创始人杨睿刚博士,长期从事3D视觉、…

South Korea targets physical AI lead and commercial humanoid robots by 2028.
AI 点评 · 韩国万亿投资布局AI芯片与人形机器人,彰显其抢占下一代科技制高点的战略雄心。
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introd…
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector…
Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively beyond their intended functions, the robot faces the challeng…
刚融了2亿美元,冲到了具身榜单第一
The startup, Proception, is taking a unique approach to collecting training data to tackle one of the hardest problems in robotics: hands.
大湾区首个200亿具身智能独角兽诞生!“最像特斯拉”智平方吸金50亿,全矩阵顶级资本重仓
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.
We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the r…
6月26日,北京人形机器人创新中心慧思开物平台的双大脑模型天鹕(Pelican-VL)和我悟(WoW)同步完成北京市网信办最新一批生成式人工智能服务备案。北京人形将正式启动慧思开物全系列模型Token服务,计划分阶段面向产业客户、科研机构、开发者全面开放API调用能力。(界面新闻)
AI 点评 · 具身智能迈入备案监管时代,标志人形机器人商业化落地加速。

IT之家 6 月 27 日消息,据央视新闻 6 月 25 日报道,市场监管总局正会同相关部门,加快智能体等前沿技术领域标准制定速度,动态完善适配产业发展的人工智能国家标准矩阵。 报道称,目前正在抓紧制定的国家标准,除智能体外,还有 具身智能、世界模型、本体模型 等前沿技术标准,算力基础设施、高质量数据集、仿真测试平台、深度学习编译器、开源模型平台等底座类标准…
AI 点评 · 政策加速标准制定,将推动智能体与具身智能产业规范化发展,抢占技术制高点。

To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.
文 | 海若镜 访谈 | 海若镜 巴芮 6月盛夏,在清华无锡研究院智能产业创新中心,我们见到了张亚勤院士。他匆匆赶来,一进门,就建议让室内温度降得更低些。 访谈中,聊起当下具身智能、AI投资创业热潮,张亚勤也觉得应该降降温,“更冷静些,不要急躁”。 五年前,张亚勤创建了清华大学智能产业研究院(以下简称AIR),聚集了多位有AI产业经验的知名教授。据介绍,如今…
We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it one of the most pro…
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a v…
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce ph…
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pretrained VLMs whose capacity far exceeds what is needed for sho…
作者|黄楠 编辑|袁斯来 6月24日,通用具身智能企业RoboScience机器科学通用具身大模型发布,首次完整披露自研Visics大模型的技术架构VLOA(Vision-Language-Object-Action),并展示了模型在家具拼装、灵巧抓取、动态流水线等多项真实场景的应用。 大语言模型有标准的文本Token,自动驾驶有统一的视觉或点云表征,这些基…
Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because they are typically conditioned only on current obs…
Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy jointly. This design inherits strong visual and li…

In this post, we demonstrate the architecture and approach Loka used to solve a common frustration: robotic, slow voice assistants that cause customers to hang up, damaging brand r…
AI 点评 · Loka用亚马逊新模型打造自然流畅语音代理,解决机器人语音痛点,技术方案值得借鉴。
Agility Robotics, the humanoid robotics startup that spun out of Oregon State University in 2015, expects to generate $620 million in proceeds.
AI 点评 · 人形机器人商业化加速,SPAC上市模式凸显资本对具身智能赛道的重注。
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection…
有智青年挑战赛暨全国AI+场景应用大赛决赛在WAVES新浪潮大会期间举行,汇聚多支青年团队围绕AI与数字经济前沿场景展开角逐,展现青年创业者的技术探索与落地能力。 2026年,AI全面进入“行动者”时代。当大模型、智能体、具身智能从实验室的技术概念,全面走进千行百业的产业落地场景,AI工具的平民化与普及化,依托开源生态、轻量化开发工具与普惠算力,不断压低创新…
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).
英伟达:不造机器人,但要帮具身企业造好机器人

Researchers combined an efficient algorithm with dedicated hardware to rapidly generate 3D maps for navigation using minimal memory and power.

US autoworkers union warns of robot automation as dark factory future looms.
AI 点评 · 机器人替代工人趋势加速,汽车巨头裁员后引入自动化,工会警示未来工厂隐患。
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal understanding, requ…
当人形机器人在春晚上扭秧歌,在马拉松赛场上健步如飞,一个尴尬的现实浮出水面:机器人会动、很灵活,但不能干活——前者考验运动控制能力,后者还需要机器人真正理解物理环境,记住物体位置、推理空间关系。 感知瓶颈是今天几乎所有具身智能从业者都无法回避的现实。成立不到三个月的「映界科技」(MirrorSpace)正试图补齐具身智能产业链中的这个短板。映界科技由三位平均…
We introduce ShotcreteDepth, a bi-modal dataset from the construction domain that captures both an active shotcreting process and general construction environments. The dataset comprises stereo RGB im…
Long-horizon tasks are common in real-world robotic deployments, yet failure detection for such tasks remains underexplored. Detecting failures in long-horizon robotic tasks is particularly challengin…
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration da…
2026年还没过半,具身智能的融资就快要超过去年全年了,超过一半的钱都流入了机器人的“大脑”。
AI 点评 · 资本加速涌入具身智能大脑赛道,百亿估值成新门槛,技术竞争聚焦模型能力。
AI 点评 · 配送机器人遭抵制暴露技术落地与人居环境的矛盾,折射公众对自动化浪潮的真实态度。
Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execution efficiency. While existing efforts predomina…
Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangl…
The startup trains embodied AI and world models using Medal’s dataset of 2 billion videos per year from 10 million monthly active users.
AI 点评 · 用2亿小时人类游戏视频训练具身AI,突破数据瓶颈,估值翻倍至20亿美金。
Quick framing, since the post is long: I did robotic manipulation research at OpenAI from 2017–2020, and the tabletop setup back then cost roughly 10x this one and took a team to r…
大公司: 滴滴自动驾驶参加伦敦MOVE 2026大会 6月17至18日,MOVE 2026大会在英国伦敦召开,滴滴自动驾驶在会上分享了来自中国的自动驾驶落地实践。在AI技术方面,滴滴自动驾驶已实现L4级全栈核心技术的自主可控;硬件方面,与广汽埃安联合打造的新一代Robotaxi车型R2已于今年1月交付,正持续在广州和北京等地开展道路测试;在场景应用上,自去年…
文 | 孙小雯 访谈 / 编辑 | 海若镜 「暗涌Waves」独家获悉,侵入式脑机接口公司「芯生视界」近日完成近亿元人民币种子轮融资。本轮融资由经纬创投领投,星连资本、燕缘创投、水木创投跟投。 当下,侵入式脑机接口已经在治疗瘫痪、脑控外设等医疗场景落地,验证长期植入的安全、有效。与此同时,AI Agent和具身智能技术加速进化,也放大了市场对脑机接口的期待:…
作者|黄楠 编辑|袁斯来 硬氪获悉,具身智能企业穹彻智能(Noematrix)近日完成新一轮数亿元融资,本轮融资由无锡数据集团领投,投资方包括上海交通大学AI未来基金(创业基金)、上海创之智科技有限公司(上海创智学院全资子公司)、一村资本等。 这也是公司近半年来完成的又一轮融资。此前穹彻智能已获得多家机构加持,包括Prosperity7 Ventures、红…
AI 点评 · 机器人选择AI大脑,决定了未来人机交互的信任与安全边界。
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretr…
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intellig…
World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies…
Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over ti…
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deployi…
Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search, target tracking, a…

Nvidia's self-improvement program for robots enlists teams of AI coding agents.
AI 点评 · 英伟达用AI编程智能体教机器人装显卡,展现了自进化系统的潜力,是机器人自主学习的突破。
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain a…
If physical AI is going to match the accomplishments of LLMs, there's a data problem that needs to be solved.
砸2亿只为“喂”数据?星海图在WDC上扔出三个信号弹。
The next humanoid robot might not have a head. It might not have legs. It might even sit on a wheeled base and fold down like a deck chair. But, as Genesis AI puts it, "humanoid ro…
半年三连发:从开源到端侧再到训练场
作者 | 邱晓芬 编辑 | 袁斯来 过去半年,国内具身智能赛道经历了一场静悄悄的重心转移:聚光灯从硬件本体的“自由度竞赛”,逐渐移向决定机器人智能上限的深水区。 只是,当行业反复讨论“机器人能否通过暴力堆数据复刻大语言模型 ScalingLaw”时,上海创智学院副教授、智元机器人首席科学家罗剑岚,给出了一个并不随大流的判断:具身智能不能简单照搬大语言模型的发…

A new spatial memory system for robots efficiently captures details about the objects they see while exploring their environment.
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acq…
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future token…
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-ce…
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether t…
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain a…
Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a gen…
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this end, recent methods leverage foundation models. But they typ…
文|周鑫雨 编辑|张雨忻 梳理近半年的成果,大晓机器人董事长、商汤科技联合创始人王晓刚,滔滔不绝聊了10多分钟。 成立于2025年7月,大晓机器人(ACE ROBOTICS)是具身领域姗姗来迟的入局者。但一年来,这位新玩家成了赛道的“卷王”: 在模型侧,大晓新发布的具身大脑——世界模型“开悟(Kairos)3.0”,在4项全球具身智能基准测试中取得SOTA;…
边走、边看、边思考
World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic…
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-t…
Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchical interface for con…
过去两年,具身智能行业最热的关键词是人形机器人。但未来几年,更重要的关键词或许会变成另一件事:智能生产力。
作者 | 邱晓芬 编辑 | 袁斯来 过去几个月,“世界模型”(World Model)从学术黑话迅速膨胀成AI和机器人行业里的关键词。 行业的目光转向背后是切实的焦虑。 一方面,经过了过去两年的野蛮生长,具身智能暴露了当前AI在物理世界中的短板——机器人能识别物体,却不懂“推杯子会掉”;能听懂指令,却无法预判“拧瓶盖需要多大的力”。世界模型正是试图补上这个短…
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,海洋具身智能公司「世航智能」完成A轮融资,融资金额超过10亿元,这也是目前全球海洋机器人领域规模最大的单轮融资。本轮融资由两家芯片公司「摩尔线程」和「昆仑芯」的产业投资方上河动量基金、新加坡国有投资平台Vertex Growth 、上市公司大洋电机等资方出资。 另外,金沙江创投也在本轮追加投资,这已是其创始人朱啸…
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up tho…
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and…
We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual tra…
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale ego…
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners…
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Actio…
Robotic systems perceive the world through multiple input modalities -- including visual camera streams and natural language instructions -- and must select appropriate actions based on these signals.…

Full autonomy is rare, but Ukraine is installing AI modules on drones and robots.
AI 点评 · 乌克兰首次实战验证全自主AI无人机,标志战场自主攻击武器进入新阶段。
LLM agents are increasingly built not as single model calls, but as scaffolded systems that combine reasoning, memory, reflection, action execution, and learning. While such scaffolds often improve pe…
AI 点评 · 具身智能或成中国AI弯道超车关键,大模型竞争远未结束。
Unlike humanoid robots designed around a fixed form — think Boston Dynamics — Theker's machines are built to be reconfigured.
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and super…
Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focu…
The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world inter…
World models that capture how actions induce physical change enable scalable robot learning without reliance on embodiment-specific action labels. Pixel-space video models provide broad visual priors…
Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. We present Neural External Torque Estimation (NEXT), a data-driven method…
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe th…
Physiological awareness is important for service, social, and assistive robots that interact with humans in everyday environments. Remote photoplethysmography (rPPG) enables non-contact heart-rate (HR…
Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance. However, current HiL…
We propose Ambient Diffusion Policy, a simple and principled method for imitation learning from suboptimal data in robotics. High-quality, task-specific robot data is expensive and time-consuming to c…
作者|黄楠 编辑|袁斯来 硬氪获悉,具身智能世界模型公司「千诀科技」日前完成数亿元A轮融资,本轮由京铭资本领投,山东新动能、山东财金资本、元禾厚望、芯能创投、南创投、英诺天使基金、尚势资本、仁爱集团、玄素投资等机构共同投资,投资方阵容汇集了国家队、产业方、市场化基金及家族办公室。Maple Pledge枫承资本长期出任私募股权融资顾问。 资金将重点用于自研世…
I'm desperate for a personal AI assistant, but do I really want to become the kind of person who can't function without the friendly robot voice in my phone?
AI 点评 · 探讨人类对AI助手的真实需求,揭示依赖与自主的微妙平衡,引人深思。

In this post, we show how to train robot policies for the Unitree H1 humanoid with NVIDIA Isaac Lab on Amazon SageMaker AI across two compute options: Amazon SageMaker HyperPod and…
AI 点评 · 结合NVIDIA仿真与AWS云服务,为机器人强化学习提供可扩展的云端训练方案。
10天跻身具身独角兽
Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simulated and real robot control. While they are known…
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointin…
轻视它就会错过整个时代

IT之家 6 月 8 日消息,理想汽车今日宣布,Livis Day 理想汽车软件与人工智能发布会将于 6 月 15 日 16:30 举行,将探讨“具身智能到底是什么? 理想在软件和人工智能领域会驶向哪里?”。 IT之家注意到,全新一代理想 L9 Livis 车型已经于上个月上市,定价为 50.98 万元,智能化成为该车主要宣传点之一。该车搭载的自研马赫 M1…
AI 点评 · 车企首度公开探讨具身智能,预示智能驾驶向通用机器人技术延伸的新趋势。

NVIDIA and LG Group are building an AI factory to accelerate LG Group’s next wave of AI-driven businesses, spanning robotics, autonomous driving, data center technologies and GPU c…
AI 点评 · 英伟达与LG共建AI工厂,加速机器人、自动驾驶等实体AI落地,展现行业巨头强强联合的新范式。

NVIDIA and Doosan Group are expanding their collaboration to advance new opportunities across physical AI, robotics and AI factory infrastructure, spanning Doosan Robotics, Doosan…
AI 点评 · 英伟达与韩国巨头联手,加速物理AI与机器人技术落地,AI工厂基建成新战场。
World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing wo…
Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely on low-dimensional st…
Recent progress in robot manipulation has been largely driven by learning from large-scale demonstrations. For humanoid robot loco-manipulation tasks, however, existing data sources force an unsatisfy…
World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy to encode task-relevant temporal structure in its…
We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated parts is essential to enhance robotic manipulatio…
Robotic simulators are a cornerstone of modern research in aerial robotics, serving both as a vehicle for the development of new control algorithms and as the data source for training reinforcement le…
具身智能的房地产开发商来了!

Home to cutting-edge sovereign AI infrastructure and robotics innovators, as well as one of the world’s most passionate gaming communities, South Korea is one of the world’s center…
AI 点评 · NVIDIA与韩国强强联手,推动主权AI与机器人创新,展现亚洲技术新格局。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪独家获悉,具身智能企业「原力灵机」近期完成新一轮融资,资方主要为数家大模型公司,包括 智谱、阶跃星辰、商汤科技。 此外, 华勤、上汽恒旭 等产业投资方持续加注 。 「原力灵机」 是一家通用具身大模型公司,2025年3月由旷视科技联合创始人兼CTO 唐文斌 创立,团队核心创始成员为旷视科技原班人马。 有意思的是,此次融资…

Robot demonstrations can distort public perceptions of robotic capabilities.
Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scenarios. To bridge this gap, we present AnchorWorld…
Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels. How…
Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning, executing, monitoring, and recovering from fai…
For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typica…
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VL…
The California startup released the fourth-generation of its home assistance robot, Stretch.
Amazon has announced a new version of its fully autonomous warehouse robot, Proteus, that will interact using language instead of code. The expanded capabilities come as part of a…
Robotaxi双雄同日官宣
作者|黄楠 编辑|袁斯来 硬氪获悉,戴盟机器人近日完成亿元A轮融资,由汇川技术旗下产业基金汇川产投与中国电信联合投资。资金将用于进一步打造超大规模含物理交互信息数据集,加速物理世界模型研发、并驱动真实物理场景下的数据飞轮与商业闭环。 戴盟机器人于2023年正式运营,核心团队长期聚焦机器人灵巧操作与物理交互智能领域。联合创始人兼首席科学家王煜教授曾任港科大机器…
Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the structural mismatch…
Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follows: how well do these…
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subg…
Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA) models, and expect broader generalisation. In th…
Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We attribute this gap…
What makes a robot gripper useful isn’t that it can pick up one object — it’s that it can pick up the next one, and the one after that, with a tool it’s never held before. What mak…
At CVPR, NVIDIA is unveiling new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems. The core…
在潜在空间中完成信息传导
Nav2 planners & controllers battle head-to-head. ROS 2 · ONNX · browser game.
36氪获悉,具身大脑公司“星源智”宣布完成新一轮融资,至今已累计融资10亿元人民币。本轮投资方包括松禾资本、创东方、华控基金等机构,中车资本、北工投资、国君创新投、江西金控等国资以及产业方埃泰克、恒兴集团、奇安投资,同时,老股东元生创投连续三轮追加投资;北京智源研究院持续加持。本轮融资将重点投入三大方向:下一代具身大脑与世界模型的核心技术研发、产品规模化量产…
AI 点评 · 多路资本押注具身智能赛道,彰显“星源智”技术实力与商业化前景。
作者|黄楠 编辑|袁斯来 硬氪获悉,绳驱AI机器人公司星尘智能(Astribot)近日完成B轮系列融资,三个月内三轮累计融资额超10亿元,投资方包括梁溪科创产业二期母基金(博华资本管理)、扬州龙投芯粒、中博聚力、中科创达、科德教育、某头部上市企业及国科投资等老股东持续追投。 目前星尘智能估值已突破百亿元,这也是深圳诞生的又一家具身智能百亿独角兽。此前,公司投…
AI 点评 · 融资超10亿、估值破百亿,彰显绳驱AI机器人的资本热度和行业赛道潜力。
作者丨欧雪 编辑丨袁斯来 硬氪获悉,高危作业领域具身智能公司杭州旷行科技有限公司(以下简称“旷行科技”)近期已完成数千万元市场化第一轮融资(Pre-A轮),由财通资本和商汤国香投资。本轮资金将主要用于算法研发、产品矩阵完善及市场拓展。 旷行科技2025年成立于杭州,是一家专注为高危工业工程领域(资源矿山、能源电力、油气化工、交通城建)提供“机器人+AI大脑”…
AI 点评 · 浙大教授团队获财通商汤投资,高危场景具身机器人商业化提速。
Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scal…
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a…
Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks. However, its practical use still depends heavily on human designed reward fun…
As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not. We introduce…
In robotics systems, vast amounts of visual data are easily captured at high resolution using low-cost, low-power hardware. Yet, limited bandwidth and on-device compute resources prevent full utilizat…
Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matters is not what the buildings look like; it is wher…
自下而上,从底层本能出发
字节跳动多模态负责人周畅管理范围再次扩大,原由李航负责的SeedRobotics团队已向周畅汇报月余,李航现以顾问身份负责学术合作方向。字节也正在招聘具身智能技术负责人,负责机器人业务整体规划,职级定位为L8,对标阿里P10-P11,将向周畅汇报。该岗位候选人主要来自头部具身智能创业公司技术负责人。(晚点 LatePost)
AI 点评 · 架构调整显示字节加速整合资源,具身智能成战略核心,技术负责人招聘透露行业人才争夺升级。
软银正初步洽谈,拟参与支持德国工业机器人初创企业Agile Robots约8亿美元的新一轮融资。知情人士透露,软银集团有意投资超过3亿美元,目前谈判仍处于早期阶段,最终金额和条款可能会发生变化。(新浪财经)
AI 点评 · 软银重注工业机器人赛道,显示对AI自动化前景的强烈信心。
具身智能的落地,正在从实验室走向最真实、最繁忙的物理世界。 而元节智能(AtomBite.AI)选择了一个看起来并不性感、但足够真实的场景:餐饮后厨。 36氪获悉,具身智能公司元节智能近日完成千万级种子轮融资,由英诺科创基金领投,水木清华校友种子基金、知名投资人个人跟投。资金将主要用于餐饮场景具身世界模型研发及核心产品落地。 元节核心团队在成立公司前有过较长…
AI 点评 · 从真实餐饮场景切入具身智能,前美团技术负责人带队,有望实现行业首个世界模型落地。
从仿真到多形态真机,具身模型如何落地复杂家庭柔性物体操作?
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks ei…
In robotics systems, vast amounts of visual data are easily captured at high resolution using low-cost, low-power hardware. Yet, limited bandwidth and on-device compute resources prevent full utilizat…
While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robots are expected to choose actions that prioritize…
Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals, competency, and willingness to cooperate. Safety f…
AI 点评 · 提出可信推理下的宽容安全机制,为交互机器人提供可验证的信念空间神经安全过滤器,兼顾安全与效率。

Lawsuit seeks $12,000 from startup that allegedly damaged home in robot tests.
OpenAI CEO山姆·奥特曼在社交平台发布OpenAI Robotics招聘信息,称公司正在寻找杰出的全栈硬件、运营、系统及机器学习工程师,共同编程并制造对社会真正有用的机器人。奥特曼表示,人工智能应当能够在现实世界中帮助人类。短期内,OpenAI专注于研发能够协助技术工人建设未来基础设施的机器人;长远来看,公司设想未来的每个人都能拥有一个可以完成各种需…
AI 点评 · OpenAI重启机器人研发,结合大模型能力或推动实用型AI硬件落地。

IT之家 6 月 1 日消息,OpenAI CEO 萨姆 · 奥尔特曼(Sam Altman)今日在 X 平台发布 OpenAI Robotics 招聘信息,称公司正在招聘优秀的全栈硬件、运营、系统及机器学习工程师, 研发和制造出对人类社会有用的机器人 。 IT之家获悉,奥尔特曼表示,人工智能应当能够在现实世界中帮助人类。 短期内,OpenAI 专注于利用机…
AI 点评 · OpenAI从软件进军实体机器人,标志AI从虚拟走向物理世界的关键一步。
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction. Yet robot fine-tuni…
AI 点评 · 评估VLA模型语义理解能力的关键基准,揭示机器人动作预测的深层缺陷。
Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstructured real-world environments. Yet, building an af…
Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household service robotics, a…
17800小时的真机数据
AI 点评 · 开创性开源世界模型,刷新具身智能训练规模,推动机器人学习范式变革。
17800小时的真机数据
AI 点评 · 刷新了具身智能数据规模,开源模型让更多研究者能探索真实机器人交互。
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present τ_0-World Model (τ_0-WM), a unif…
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,乘物机器人(深圳)有限公司(以下简称“乘物机器人”)近日完成天使轮融资,由中国台湾工业自动化与智能机器人解决方案领域龙头企业和椿科技战略投资,华君资本担任独家财务顾问。 乘物机器人成立于2025年,总部位于深圳,专注工业具身智能技术研发与产品解决方案,具备从软硬件研发、数据采集、模型训练、场景部署与维护的一体化技…
AI 点评 · 服务富士康半年营收超两千万,展现工业具身智能落地潜力,值得关注商业化路径。
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,乘物机器人(深圳)有限公司(以下简称“乘物机器人”)近日完成天使轮融资,由中国台湾工业自动化与智能机器人解决方案领域龙头企业和椿科技战略投资,华君资本担任独家财务顾问。 乘物机器人成立于2025年,总部位于深圳,专注工业具身智能技术研发与产品解决方案,具备从软硬件研发、数据采集、模型训练、场景部署与维护的一体化技…
AI 点评 · 服务富士康半年营收超两千万,印证工业具身智能从技术到商业化的快速落地。
Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essential. However, existi…
AI 点评 · 评测视觉语言模型在物理场景下的抗压能力,填补机器人感知鲁棒性空白。

The latest twist in paying humans to wear head cameras for robot training data.
AI 点评 · 用真实家务数据训练机器人,开创数据采集新模式,隐私换技术值得深思。

The latest twist in paying humans to wear head cameras for robot training data.
36氪获悉,宇树科技发文称,5月31日,宇树科技具身智能体验馆亚洲首店将正式登陆上海,门店汇聚G1人形机器人、R1人形机器人、Go2机器狗全系列C端产品。
5月29日,理想汽车基座模型部门完成新一轮组织调整:新增3个具身智能相关二级部门,分别为具身工程、具身交互、具身行为;同时自动驾驶变成独立二级部门。与上述部门平行的,还有数据引擎、仿真与测评等部门。调整后,自动驾驶、具身工程、具身行为3个部门直接由基座模型负责人詹锟管理,向理想汽车CTO谢炎汇报。(晚点Auto)
AI training startup Shift wants to clean your home for free. The catch - because, despite what its website says, there's always a catch - is that it will record cleaners as they sc…
AI 点评 · 用免费清洁换取机器人训练数据,揭示真实场景数据对AI落地的核心价值。
AI training startup Shift wants to clean your home for free. The catch - because, despite what its website says, there's always a catch - is that it will record cleaners as they sc…
AI 点评 · 免费家政换取机器人训练数据,揭示数据采集新商业模式与隐私权衡。
从按帧学动作,到按「事件」理解世界
AI 点评 · 将AI从动作模仿提升到理解事件因果链,是具身智能迈向通用化的重要突破。
AI 点评 · 初创公司用Airbnb测试机器人却将其损毁,暴露AI产品真实场景测试的伦理与法律风险。
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliabil…
AI 点评 · 提出轨迹捉迷藏方法,主动发现VLA模型运行时的隐藏故障信号,提升机器人可靠性。
Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over fu…
AI 点评 · 用压力测试场景驱动视频世界模型,提升机器人策略评估的鲁棒性和改进效果。
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog…
AI 点评 · 三模态动态引导的机器人感知新范式,突破静态视觉预训练局限,开启灵巧操作新可能。
The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emph…
AI 点评 · 机器人创造性解决问题的新基准,揭示现实环境中的适应力短板,推动AI迈向更高阶智能。
Robotics is entering a new phase: moving from controlled demos and scripted automation toward generalizable, reliable embodied autonomy in the real world. At the International Conf…
AI 点评 · 英伟达推动机器人从仿真走向真实世界,标志着通用具身智能迈向关键落地阶段。
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks,…
AI 点评 · 统一视觉语言与动作建模,突破单一任务限制,推动机器人跨场景泛化能力跃升。
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog…
AI 点评 · 三模态动态引导感知,突破静态视觉局限,让机器人操控更精准。
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configuration…
AI 点评 · 统一框架实现多模态条件生成,突破固定模态限制,显著提升运动生成灵活性与泛化能力。
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recog…
AI 点评 · 开创性提出三模态动态引导,让机器人感知从静态识别转向动态交互,突破传统视觉编码局限。
Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive de…
Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly te…
Media compression standards have reached a plateau in terms of the rate-distortion-complexity trade-off, limiting the ability to offload expensive AI perception to the cloud in applications like robot…

Hugging Face debuts $2,500 bipedal robot project for builders and researchers.
AI 点评 · 开源低成本人形机器人项目,大幅降低研发门槛,推动双足行走技术普及。

Hugging Face debuts $2,500 bipedal robot project for builders and researchers.
AI 点评 · 开源低成本的人形机器人腿,将大幅降低双足机器人研究门槛,激发创新实验。
Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions. Robotics foundation models, vision-languag…
AI 点评 · 聚焦物理AI安全盲区,系统梳理运行时动作授权机制,为自主系统风险防控提供关键学术支撑。

A recap of the 2026 I/O Dialogues, where leaders discuss the future of AI, quantum computing, robotics and creativity.
AI 点评 · 科技巨头齐聚探讨AI、量子计算与机器人未来,前沿观点碰撞,指明行业发展方向。

A recap of the 2026 I/O Dialogues, where leaders discuss the future of AI, quantum computing, robotics and creativity.
AI 点评 · 行业领袖共议AI、量子计算与机器人未来,揭示技术融合新趋势。
Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Although these numbers ap…

Figure AI's 24/7 livestream showcases human soft spot for humanoid robots.
AI 点评 · Figure AI直播展示人形机器人分拣包裹,引发全网围观,折射人类对类人智能的天然好奇。

Figure AI's 24/7 livestream showcases human soft spot for humanoid robots.

"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of…
AI 点评 · 挑战语言中心主义,揭示具身认知对通用智能的核心价值,重塑AI发展路径。