
IT之家 9 月 6 日消息,特斯拉 AI 负责人 Ashok Elluswamy 本周表示,特斯拉 Robotaxi 自动驾驶无人出租车距离“实现 24 小时全天候运营”已经不远。针对一名希望在深夜使用 Cybercab 出行的用户,他在 X 平台回复称,待“v15 计划中的下一项技术”完成整合后,这项能力将在“下个月左右”上线。 不过,Elluswamy…
AI 点评 · 无人驾驶商业化提速,技术落地与运营监管仍待检验。
共 328 条相关资讯 · 来自历史归档

IT之家 9 月 6 日消息,特斯拉 AI 负责人 Ashok Elluswamy 本周表示,特斯拉 Robotaxi 自动驾驶无人出租车距离“实现 24 小时全天候运营”已经不远。针对一名希望在深夜使用 Cybercab 出行的用户,他在 X 平台回复称,待“v15 计划中的下一项技术”完成整合后,这项能力将在“下个月左右”上线。 不过,Elluswamy…
AI 点评 · 无人驾驶商业化提速,技术落地与运营监管仍待检验。

IT之家 9 月 5 日消息,科沃斯(ECOVACS)在 IFA 2026 期间宣布公司已成为全球第一大家用机器人品牌,其产品已经进入全球 180 个市场、覆盖超过 3800 万户家庭。 科沃斯表示,公司早在 2006 年正式启用如今的品牌名称,并于 2009 年推出首款家用扫地机器人。经过多年发展,公司不断拓展各类应用场景,在石头科技、iRobot、Swi…
AI 点评 · 海外营收过半是里程碑,印证中国服务机器人全球化战略的成效。
The round is being raised just months after the robot data startup exited from stealth.
AI 点评 · 机器人数据赛道爆发力惊人,三个月估值翻至12亿美元,资本抢筹信号明确。
DELE-w0.5 重新定义了世界模型在机器人操作中的作用。
On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two.
AI 点评 · 虚实融合成AI新战场,Nvidia与机器人同台,看点在于前沿技术如何落地现实场景。
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions. This paper presents…

Christopher was sick of being ghosted by AI recruiters. So he unleashed ChatGPT on his robot interviewer.
9月2日,前字节跳动强化学习专家、前腾讯Robotics X智能体中⼼负责⼈孙鹏博⼠正式加入星尘智能
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics…
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks. To address this bo…
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predi…
Hey HN, I’m Antonio from Nori Robotics ( https://norirobotics.com ). We build a $1,688 bimanual mobile robot in San Francisco for robotics developers and researchers. I started wor…
参数冻结,能力暴涨
近期,滴滴自动驾驶新一代 Robotaxi R2正式开启无人载客测试服务
灵犀智涌用一台由Demo级本体组装而成的机器人,成为工业场景赛除行业头部企业外唯一获奖的机器人公司。
The U.S. is shutting out more foreign-made drones and robots. China’s scale means the global competition may simply move elsewhere.
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) alread…
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce…
Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real e…
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limit…

The company is testing robots on tasks that can performed by technicians.
AI 点评 · 机器人进数据中心干活,Meta试水自动化运维,效率与成本拐点将至。
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and…

Plus: Hackers target over 100 US water systems, ICE puts in an order for robot dogs, and you’ll never guess what “MrChildPorn” was arrested for.
AI 点评 · 网络安全威胁升级至关键基础设施,AI巨头预警凸显技术双刃剑效应。

Anthropic's Model Hardware Standard (MHS) gives AI agents a unified interface to physical devices like robotic arms and lab instruments. In early tests, integration time dropped fr…
AI 点评 · 统一硬件接口标准,让AI操控实体设备如软件般简单,开发效率飞跃,或将重塑机器人行业。
Pollen Robotics, the Bordeaux robotics team at Hugging Face, opened pre-orders for Microduck — a 25 cm bipedal robot where every movement is a neural policy trained in MuJoCo and e…
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being…

The company is testing robots that can swap cables, reset servers, and take on other tasks performed by technicians, fueling concerns among some workers that their jobs could be at…

IT之家 8 月 28 日消息,Anthropic 昨日(8 月 27 日)发布博文,宣布以研究预览形式,发布模型硬件标准(Model Hardware Standard,MHS), 路透社、CNBC 等媒体解读认为这标志着该公司首次公开进军具身智能与物理 AI(Physical AI)领域。 Physical AI(物理人工智能 / 实体 AI)是指能够感…
AI 点评 · 首个AI公司主导硬件标准,具身智能赛道迎来新变量。
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich s…
Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 in…
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is cost…

Beijing’s endlessly delightful Robot Games featured tons of impressive stunts. But the most mind-blowing tricks challenged the humanoid’s brain, not its brawn.
AI 点评 · 人形机器人竞速吸睛,但镊子精细操作更显AI大脑突破,看点在于智能与灵巧的融合。
Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. Whet…
Hi HN, we're Sam and Alex, founders of Risklytics ( https://risklytics.ai ). We're both on leave from Harvard, and we run an insurance brokerage for companies building robots, dron…
Gates is mostly in the Responsible AI camp, but there are a few ideas in here we hadn't heard before.
Robot bodies are waiting for their AI brains to catch up.
具身智能迈向GPT时刻
球技丝滑、现场爆满
The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation.

Record-breaking robot races are less substantial than household chore challenges.
AI 点评 · 人形机器人竞技热度飙升,但炫技背后实用家务能力才是落地关键。
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task ca…
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a…
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to re…
近日,范式正式举办 PhanthyMotus 生态社区共建计划发布会,宣布其首个通用具身Agent底座从“开源”迈入“多方共建”新阶段。
AI 点评 · 巨头抱团共建具身智能底座,生态整合或成行业分水岭。

Humanoid robots are having a moment in China. The popular machines are part of the country’s strategy to bring artificial intelligence into daily life. Embedding the technology int…
机器人大脑新SOTA
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for effi…
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models genera…
General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valu…
AI 点评 · 顶级风投加注,估值60亿,通用AI进军机器人赛道,资本风向标值得紧盯。
Generalist AI has released GEN-1.5, a robot foundation model that learns a new physical task from a single demonstration. Drop 3–12 seconds of sensorimotor data into its 30-second…
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reas…
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was dem…
香港的大学里冒出了一批很特别的人
AI 点评 · 香港教授创业潮,产学研融合新样本,折射高校创新生态。
多多支持像王兴兴这样优秀的具身机器人创业者
AI 点评 · 天使轮眼光再押注具身智能,揭示早期投资风向与产业孵化逻辑。
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiat…
端侧部署解决了具身大脑能否装进身体的问题。那么,同一个「大脑」,如何快速适配工业、商用和家庭三类机器人呢?
AI 点评 · 从仿真到真机,优必选展示具身智能落地全链路,行业价值关键在量产验证。
直击WRC
AI 点评 · 陆行具身系统突破轮、足、车多形态融合,或成机器人移动技术新标杆。
让具身智能技术真正落地千行万业。
AI 点评 · 具身智能落地路径首次被系统拆解,产业协同模式或成行业新标杆。
机器人换了身体,智能还能留下多少?

Waymo built its own chip for its robotaxis, cutting its reliance on Nvidia. The article Waymo builds its own chip for its robotaxis, cutting its reliance on Nvidia appeared first o…
明略科技(2718.HK)与海康机器人联合参展2026WRC,聚焦商业服务领域展示具身智能落地进展。
比会做一个动作更难的,是把一整件事连续做完。
We present PhysCaP, a Physics-Informed Code-as-Policy agent for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive…

Robotics startup Generalist AI has unveiled GEN-1.5, an AI model that teaches robots new tasks from a single demonstration. The article GEN-1.5: Generalist AI teaches robots new ta…

Unitree Robotics rose 460 percent in its Shanghai IPO, hitting a valuation of around $50 billion. But an FT report shows much of the demand for its robots comes from state-backed t…
一个打通链路的优秀的通用技术底座
Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets. W…
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models wit…

During a recent visit to Generalist AI, I watched a robotic arm improvise and use a banana as a tool.
AI 点评 · 实时学习能力是AI从工具到伙伴的关键跨越,现场即兴创新更显真章。

IT之家 8 月 19 日消息,小米官方今日宣布, 小米新一代人形机器人亮相 2026 世界机器人博览会 ,登陆「具身花园」展台。 IT之家注意到,小米官方刚刚放出小米新一代人形机器人与观众实时互动的视频。据介绍,从递花、握手到碰拳、比心,这些看似简单的动作背后, 由大模型驱动机器人自主理解、判断并完成动作 。 根据规划,8 月 20 日起,展台将对普通观众…
超维动力KAI全栈具身智能硬核登场
构建全栈物理AI基础设施
SDK for robotics teams to verify the quality of their data used for AI model training.
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by…

“We're not necessarily building in a dogmatic fashion towards full autonomy.”
AI 点评 · 机器人工厂切入钢铁制造,前SpaceX工程师带来降本增效新思路。
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action sp…
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: err…
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-…
The greatest invention in pet tech in recent years is the litter robot. A machine that scoops your kitties' poop so you don't have to - what else could a cat owner possibly want? H…
8月17日,具身智能初创公司共生知行发布双足人形机器人驾驶卡丁车Demo

When Xander first met Moxie, she taught him that when he was anxious, he could calm down by exhaling through his lips so that he buzzed like a bee. They practiced breathing like dr…
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain lar…
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make…
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual…
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains…
World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that trains robot controllers entirely in virtual environments. From a single real-world…
AI 点评 · 用模拟生成海量训练变体,突破真实数据瓶颈,让机器人训练更高效灵活。

The Unitree G1 has found online fame as a relatively affordable robot that can charm a crowd. But can it ever hold down a real job?
AI 点评 · 四足机器人G1凭亲民价格走红网络,但能否胜任实际工作成关键看点。
Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using t…
As heterogeneous robotic systems deploy across diverse urban zones, maintaining safety amid complex human-robot interactions remains a critical challenge. We present a unified framework that bridges s…
An autonomous robot efficiently exploring an unknown environment, such as looking for water sources on Mars, faces two simultaneous demands: building an accurate information map while quickly finding…
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process a…
AI 点评 · 一站式打通机器人开发全链路,降低门槛,值得开发者关注。
重新定义物理AI数据基础设施

Dyna Robotics has released Dyna-2, a world-action model pre-trained on more than one million hours of egocentric human video. The technical report establishes three results: a scal…
影石发布 Insta360 X6、SpaceXAI 推出 Grok Bot 等。 查看全文
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos pr…
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effec…
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followe…
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observ…

IT之家 8 月 12 日消息,在今晚的荣耀 Robot Phone 全球新品发布会上, 全球首款机器人手机 —— 荣耀 Robot Phone 正式发布 , 定价 9999 元起 。 天马微电子宣布, 为本款机型独家供应屏幕,深度定制天马天工屏高端 OLED 显示方案 。天马表示,这是天工屏首次应用于具备具身交互能力的 AI 终端,适配机器人手机的创作、随…
具身模型一小时狂拣1816件异形包裹
从柔性本体走向跨本体基础智能,让任务与世界知识延续
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessi…
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics ar…
资本正从具身本体集体涌向触觉
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, h…
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework th…
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harnes…
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-t…
Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics. While latent world models offer…
Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, sm…
AI 点评 · 14MB级端侧智能体,开启手机手表家居机器人的本地AI新纪元。
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility a…
“我们90天就成了独角兽”
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in tech…
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexp…
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor bench…
采集、仿真、训练和评测,闭环了
AI 点评 · 数据闭环打通,国产具身智能冲刺百万小时里程碑。
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewi…
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted t…
这套中国原创理论,为具身智能开出了新路。
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, l…
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existin…
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more…
On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two.
NVIDIA released Alpamayo 2 Super, a 34B vision-language-action model for autonomous driving, under OpenMDW-1.1 — a permissive license covering fine-tuning, derivatives and commerci…
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods re…
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, howe…
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the re…
An analysis of the last seven years of Tesla earnings calls shows just how little attention Musk pays to Tesla's car business.
%20China-Free%20Robot.jpg)
Ati Robotics assembles its robots in India and uses just a few Chinese parts—a strategy that could pay off as the Trump administration cracks down on Chinese humanoids.
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that…
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Humanoid robots usually elicit more cringe…
AI 点评 · 美国AI保护主义蔓延至机器人领域,或重塑全球产业格局。
AI 点评 · 家庭场景海量数据开源,为机器人“读现实”提供关键教材,具身智能落地提速。
World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future alone does not justify changing the action that a bimanual policy would execute. We…
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has sho…
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as c…
Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gau…

IT之家 8 月 1 日消息,7 月 31 日,国际数据公司(IDC)数据显示,2025 年中国工业具身智能机器人市场规模约为 57.4 亿元 ,其中以工业机器人为载体的具身智能应用市场规模约 36.2 亿元,成为当前产业商业化落地的主要方向。随着产业竞争从机器人本体性能逐步转向模型、数据、工程化和场景落地能力的综合竞争,工业具身智能机器人正在成为智能制造发…
AI 点评 · 工业具身智能首破57亿,赛道从硬件比拼转向数据与场景落地。
具身智能的τ(bushi)
AI 点评 · 具身智能商业化落地,机器人保洁定价对标人工,性价比成关键看点。
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing…

Google Deepmind's Gemini Robotics 2 is its most advanced vision-language-action model yet, built to control everything from tabletop robots to full-body humanoids. Gemini Robotics…
AI 点评 · 多形态机器人统一操控,视觉语言行动模型再突破,通用机器人时代加速到来。
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Clo…
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator…

IT之家 7 月 31 日消息,谷歌 DeepMind 昨日(7 月 30 日)发布博文,宣布推出 Gemini Robotics ER 2 模型, 是其面向机器人的最强具身推理(embodied reasoning)模型。 定位方面,该模型可视为机器人的“高级大脑”,协调处理与人类交流、理解物理世界并规划多步骤任务等。该模型后续会向更低层级的视觉-语言-动…
AI 点评 · 突破性实现多机器人协同与连续视频理解,标志着具身智能向通用化迈出关键一步。

Director of CSAIL and MIT professor honored for her contributions to robotics, artificial intelligence, and autonomous systems.
AI 点评 · 机器人AI领军人物获巴伐利亚最高科技奖,凸显国际对AI前沿研究的高度认可。

Researchers fear AI is moving too fast, while Mark Zuckerberg is worried about who owns it. Plus: Inside Black Forest Labs’ push into robotics.
AI 点评 · AI巨头竞速引发安全与所有权双重焦虑,凸显行业失控风险。
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator…
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangl…

Gemini Robotics 2 includes three models, but only one is publicly available right now.
AI 点评 · 谷歌发布三款模型,仅一款开放,聚焦灵巧操作与安全,引领机器人AI新突破。
Google DeepMind has released Gemini Robotics 2, the intelligence layer for its next generation of robots. The release ships three models: a vision-language-action model for whole b…
AI 点评 · 三个实体AI模型同时发布,标志全身控制、灵巧操作与多机器人协作迈入新阶段。
Google DeepMind says the latest version of its Gemini Robotics AI model can "control entire humanoid robots." While the previous model focused on controlling a humanoid robot's upp…
AI 点评 · 谷歌DeepMind新模型实现机器人全身控制,标志着具身智能从局部操作迈向整体协调。

The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collabora…

The FCC is blocking imports of new Chinese humanoid robots and robot dogs. But the rule's broad definition also sweeps in Roombas, robotic lawn mowers, and delivery bots. The artic…
A German startup sent a camera-wearing chef to my apartment. In exchange for a free lunch, I let them record every chop and stir to train future humanoids.
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interact…

Government ban on foreign-made robots may hinder instead of help US robotics.
AI 点评 · 禁令冲击美国机器人产业,揭示供应链依赖与本土创新困境。
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as…
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a…
Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduc…
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract…
作者|黄楠 编辑|袁斯来 硬氪获悉,柔性触觉感知企业「尧乐科技」近日完成Pre-A+新一轮融资,本轮融资由鼎和高达领投,上市公司常熟汽饰、祖龙娱乐跟投,云道资本担任长期独家财务顾问。资金将主要用于柔性织物传感技术研发、产品性能迭代升级,加快数据手套产能拓展和批量交付,并完善数据采集、标定与接口能力,加速面向具身智能、世界模型与智能座舱等场景落地。 尧乐科技以…
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capt…
In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the black-box nature of deep neural networks. This lack…
AI 点评 · 用物理知识蒸馏强化学习策略,兼顾性能与可解释性,推动安全关键领域应用。

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe no…
Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics AI at Meta
The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
一次技术与场景的深度耦合
Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics AI at Meta
机器人学会「想清楚之后再行动」!

IT之家 7 月 27 日消息,荣耀手机官方今日宣布,全球首款机器人手机 —— 荣耀 Robot Phone 定档 8 月 12 日发布 ,由阿莱联合研发。 据IT之家此前报道, 荣耀 Robot Phone 已开启预约 ,搭载第五代骁龙 8 至尊版芯片。该机的核心创新在于机身顶部集成了一套行业最小的四自由度(4DoF)钛合金机械云台系统。云台系统配备微型电…
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states…
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture t…
Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model t…
项目数据模型均开源
AI 点评 · 触觉数据补齐机器人感知短板,开源模式加速具身智能落地。
触觉:具身智能的下一个 Scaling Law?
AI 点评 · 触觉标准化突破,让机器人“手感”可复制,补齐具身智能关键短板。
这是机器人领域最好的时代
不做人形,抓住真实用户需求
7月24日,据报道,日本软银集团正考虑收购瑞士AI机器人初创公司Gravis Robotics AG,旨在进一步加码支撑机器人技术的具身智能与人工智能领域。软银计划对这家初创公司建立大量股权头寸,并将其归入其正在筹建的新AI及机器人实体公司“Roze”旗下。交易可能分阶段推进,包括购买现有股份以及随着时间推移注入新资金,最终对Gravis的估值可能超过5亿美…
AI 点评 · 软银加码具身智能,高估值收购预示AI机器人赛道加速整合。
AI 点评 · 具身智能赛道爆发,数据服务成新风口,商业模式能否跑通是核心看点。

A HUGE win for BFL!
AI 点评 · 填补具身智能数据标准空白,中国力量首次主导国际协作。
具身智能正在从「参数竞赛」进入「架构竞赛」。
AI 点评 · 参数效率革命证明,小模型也能靠架构创新碾压大模型。
一台机器人的「多任务实战」
AI 点评 · 工业AI落地深水区突破,具身智能从概念走向多任务实战。
The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to b…
AI 点评 · 大规模真实桌面数据集助力机器人操作泛化,为通用操控策略提供关键训练资源。
Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.
AI 点评 · Uber创始人再获巨额融资,a16z领投彰显工业AI前景。

Union previously warned automaker that any robot deployment must be negotiated.
AI 点评 · 现代与工人谈判中提及人形机器人,凸显劳资博弈与技术替代的敏感边界。
Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilize…
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distri…
Kandinsky WM 1.0 — a family of models for Physical AI. Image-to-video generation for autonomous driving, robotics & general physics.

Before a healthcare robot can be useful in the real world, it has to learn how the physical world pushes back. Anatomy varies. Instruments bend, press, slip and interact with tissu…
AI 点评 · 开源首个GPU加速医疗物理仿真,降低机器人精准医疗开发门槛。
按照规划,日冕和远图将首先在服务器制造场景验证超级工站能力,随后向更多生产环节扩展。2027年完成百台级部署,未来实现万台级具身智能产品部署。
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requir…
AI 点评 · 用自然语言指令引导机器人抓取复杂场景中的物体,让多形态机器人理解语义和空间关系。
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) polic…
AI 点评 · 将自然语言描述与移动追踪结合,突破传统视觉追踪限制,提升具身智能的实用性与交互性。
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth senso…
AI 点评 · 导航系统规模化部署的关键突破,在于降低传感器依赖并提升跨形态泛化能力。
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current b…
Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasible trajectories, all onboard and in real time. Conventional a…

IT之家 7 月 21 日消息, 荣耀 Robot Phone 已正式开启预约 ,该机搭载第五代骁龙 8 至尊版芯片,其核心创新在于机身顶部集成了一套行业最小的四自由度(4DoF)钛合金机械云台系统。云台系统配备微型电机,体积比主流方案缩小 70%。 IT之家注意到,荣耀全球首席营销官关海涛今日晒出了自己的荣耀 Robot Phone 真机,并分享了新机的拍…
Gritt is coming out of stealth with $34 million and plans to automate the hardest tasks on construction sites.
AI 点评 · 开源机器人操控数据记录系统,降低研发门槛,加速行业标准化进程。
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to commun…
Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global…
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts…

IT之家 7 月 20 日消息,科技记者 Joanna Stern 上个月撰文称,在自己的新书《I Am Not a Robot》正式出版后的几天内,她就在 Apple Books 上发现了 10 本冒充该书的盗版版本。虽然她成功让这些作品下架,但很快又有新的盗版书重新出现。 《纽约时报》记者 Kashmir Hill 也发现,亚马逊上出现了一本关于她本人的…
From open models to real-time simulation, AI and graphics breakthroughs are transforming media, content creation and robotics.
7月17日-20日,一起在WAIC2026现场,看见人工智能真正进入产业深处。 过去一年,围绕AI行业的讨论正在变得更具体。大模型能力仍在持续迭代,但外界关注的重点,已经不再只停留在模型参数、模型发布和单点能力展示上。随着智能体、具身智能、空间智能、AI基础设施等方向不断演进,行业开始更频繁地追问:AI如何进入真实流程,如何完成复杂任务,又如何在产业场景中形…
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports…
Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous operations in real-world scenarios. Under low illumination, reduced contrast, sens…
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains cha…
7月18日,在2026世界人工智能大会(WAIC)上,腾讯面向具身智能与智能体领域带来多项产品技术的升级发布。在具身智能领域,腾讯正式升级发布具身智能全栈方案,贯穿云底座、模型层、平台层与应用层,全面助力机器人本体及系统开发商提质提效;在智能体领域,基于个人与企业提效需求,推出差异化的全矩阵解决方案。其中,面向企业用户的腾讯云企业级智能体开发平台ADP4.0…
Agility is opening a new training center for its Digit robots in Fremont, California.
AI 点评 · 选择特斯拉大本营设立培训中心,Agility Robotics此举意在抢夺物流机器人市场先机。
人工智能正在进入一个新的产业周期。 过去一年,大模型能力持续演进,生成式AI、多模态交互、智能体等技术方向快速推进;而具身智能也从早期的技术探索阶段,逐渐步入产业验证的深水区,机器人开始成为人工智能与现实世界的重要载体。 市场率先给出了回应。据36氪研究院测算,中国具身智能市场规模已从2018年的2133亿元增长至2025年的9150亿元,2026年有望突破…
36氪获悉,7月17日,2026世界人工智能大会(WAIC)召开。网易旗下专注工程机械领域的具身智能品牌网易灵动,将极具科技感的“智能座舱”搬入展区,座舱内工作人员跨越千里,可远程操控真实电厂、港口的生产一线设备。同时,网易灵动还集中展现了具身智能规模化落地高危场景的最新成果,以及“人机混编”“黑灯工地”“一人多机”等自主研发技术。
AI 点评 · 人机混编技术突破高危场景,远程操控颠覆传统作业模式,具身智能商业化落地加速。

The CEO of Foundation Future Industries, which counts the president’s son as its chief strategy adviser, tells WIRED it’s exploring some “kinetic things.”
AI 点评 · 埃里克·特朗普背书的人形机器人公司涉足军事,政治与科技的交汇点值得警惕。

IT之家 7 月 17 日消息,据央视财经报道,在 2026 世界人工智能大会上,银河通用机器人创始人兼首席技术官王鹤预测,通过海量综合数据训练,具身智能基础模型能在未经专门训练的任务上直接达到 70%-80% 成功率,媲美当年数字模型的对话水平;而强大的预训练模型加上高效后训练范式,正是敲开这一里程碑的关键条件。 王鹤判断,按当前数据积累速度和模型收敛趋势…
AI 点评 · 具身智能将从实验室走向产业,2028年或成关键拐点,值得关注技术突破与商业落地。

IT之家 7 月 17 日消息,腾讯 Robotics X 实验室、福田实验室联合腾讯混元打造的第二代具身 VLM 基座模型 Hy-Embodied-VLM-1.0 昨日正式发布。 官方表示,在覆盖 37 个评测任务的具身能力评测体系中,Hy-Embodied-VLM-1.0 在物理状态理解、动作 — 变化推理、时序与自适应推理三大维度分别取得 68.6、6…
AI 点评 · 参数规模锐减十倍,性能逼近上一代大模型,具身智能走向高效实用化。
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,具身智能世界模型公司日冕开物(北京日冕机器人有限公司)近期完成连续两轮种子轮融资,融资合计金额达数亿元人民币,由鼎峰科创、远图未来、百度风投、沃衍资本、武岳峰科创、万林国际共同参与投资。同时,新一轮融资也在同步交割中。 此前融资资金主要用于自研世界模型 LaMPA 的研发迭代、强化学习体系建设,以及数据闭环和产品…
AI 点评 · 前蔚来华为核心团队跨界创业,三个月获数亿融资,凸显具身智能赛道热度。

Hyundai aims to deploy 25,000 Atlas robots starting with US factories in 2028.
AI 点评 · 人形机器人威胁就业引发罢工,折射出AI替代焦虑与技术落地的现实冲突。
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct a…
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whe…
AI 点评 · 探索大模型安全新维度,揭示文本安全与物理行动的鸿沟,为具身智能风险防控开辟新思路。

IT之家 7 月 16 日消息,荣耀 AI 首席科学家、首席 AI 官黄非昨日官宣入驻微博,并在今日分享了一段 Robot Phone 新机的 AI 功能演示视频。 黄非表示:“荣耀 Robot Phone,‘活’了。刚和团队对完最后几组 Demo,只能说, 这次我们把对未来科技的想象,真的照进了现实 。” IT之家注意到,黄非在视频中,对荣耀 Robot…
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
为破解具身智能行业发展瓶颈构建了新一代“进化底座”
AI 点评 · 具身智能领域迎来关键突破,从模型训练到真机部署实现全链路能力升级。
Robotaxi第一股文远知行孵化
AI 点评 · 具身智能赛道迎来首个基础设施供应商,复制英伟达与宁德时代的成长逻辑。

General-purpose robots and autonomous machines are moving from research labs to real-world mass-market deployment, creating demand for compact, power-efficient AI supercomputers ca…
AI 点评 · 英伟达推Jetson Thor,瞄准机器人量产与边缘AI,算力能效比成关键看点。
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world…
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visu…
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero…
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen en…
This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-A…
Home to leading manufacturers, robotics pioneers, infrastructure builders and iconic gaming companies, of course, Japan is one of the world’s centers of AI — building across the fu…
FSD、Robotaxi共用一套模型
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action ge…
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint lan…
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotat…
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning p…
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial services ambitions, its increasingly complicated relationship with Waymo, its new AV Labs…
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surroundi…

“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track the…
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consis…
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe th…
为了帮你看清具身数据行业,我们总结了以下十个行业现状
AI 点评 · 数据成AI新燃料,百亿资本涌入具身智能赛道,揭示产业变现路径。
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this int…
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification,…
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often in…
AI 点评 · 用强化学习微调动作块,提升机器人接触操作的鲁棒性和泛化能力。
AI 点评 · 蚂蚁自研具身模型,从零预训练实现动作原生,突破机器人智能边界。
开源第四弹:LingBot-VA 2.0

Amid live coding sessions and Silicon Valley optimism, the UN’s AI for Good summit wrestled with an urgent question: Can global governance catch up before the technology races beyo…
AI 点评 · 联合国AI峰会展示前沿科技,核心挑战在于全球治理能否跟上技术飞速发展。

Preclinical trial is testing the feasibility of humanoid robots in surgery.
AI 点评 · 远程操控人形机器人完成活体手术,标志着医疗自动化从概念迈向临床的关键一步。

The soft, oddly intimate home-chore robot has been given some very tactile hands.
AI 点评 · 人形机器人触觉灵巧手取得突破,家庭服务场景落地加速。

MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water.
How Claude Performs on Robotics Tasks Anthropic
视频生成的下一站,或是机器人大脑
36氪获悉,聚焦工业制造领域的具身智能机器人企业 「昇视唯盛」(3Srobotics)正式宣布完成数亿元B轮融资,本轮融资由上海半导体产投、金桥基金领投,零一创投、新鼎资本、中关鼎华及老股东微光创投持续加码跟投。 资金将主要用于焊接具身智能大脑与自主小脑运动控制系统的迭代升级,扩大自有智造工厂产能,扩充研发与市场服务团队,加速产品全国多行业规模化落地。 昇视…
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer prec…
General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to build smarter robots with minimal real-world…
AI 点评 · 用游戏数据训练机器人,可能加速物理AI突破,降低开发门槛。
具身测评界的珠峰来了:RoboDojo
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Exist…
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visu…

Top robotics researchers and founders explain how robot autonomy is evolving.

Top robotics researchers and founders explain how robot autonomy is evolving.
“共识与非共识”的深度思辨

Open source AI has shown how quickly developers can innovate when models, data and tools are shared. Robotics has the same opportunity, but advancements in physical AI development…
来自蚂蚁灵波
AI 点评 · 开源机器人学习框架迭代,强化想象与评估闭环,加速AI物理世界应用落地。
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lac…
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy,…
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware…
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth,…
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these m…
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narro…
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmabil…
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Lang…
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current obs…
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-f…
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing design…
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensio…
一段视频,生成无限训练场景
AI 点评 · 用真实视频自动生成多样训练数据,突破仿真向现实迁移的成本瓶颈。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,具身智能公司 「光象科技」 宣布完成累计数亿元天使轮融资。 最新一轮由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、SeeFund、亿宸资本、上市公司行云科技等头部财投与产投深度参与,老股东零一创投、L2F光源创业者基金持续加注。 本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人在工…
AI 点评 · 清华系创业团队获数亿元融资,聚焦具身智能落地汽车产业,技术底蕴与产业场景结合是关键看点。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,通用餐饮具身机器人公司「影智XBOT」连续完成数亿元两轮融资——其中,A轮的2亿元融资由香港简坤资本GPTX出资,B轮融资为3-5 亿元人民币,由多支政府基金、美元基金和产业投资方共同参与出资。 这是目前餐饮垂直机器人领域规模最大的一笔融资之一。 在此之前,「影智XBOT」还完成了一轮天使融资,出资人阵容豪华——…
AI 点评 · 小米前高管创业项目获数亿融资,林斌、黎万强加持,餐饮机器人赛道热度可见一斑。
近日,光象科技宣布完成累计数亿元天使轮融资,最新一轮融资由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、See Fund、亿宸资本、上市公司行云科技等头部财投与头部产投深度参与,老股东零一创投、L2F光源创业者基金等持续加注。据悉,本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人产品的商业化交付。(每日经济新闻)
AI 点评 · 资本密集押注物理原生模型,具身智能从实验室走向量产的关键信号。
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing re…
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on loc…
今日热点导览 苹果拟于今明两年推出至少五款新iPhone AI版支付宝开放公测,上线72项办事技能 特朗普回应“利用职位牟利”:股市在涨,大家都在赚钱 LV起诉茉莉奶白,茉莉奶白被判赔1030万元 6月赴日航班取消1488个 安克创新港股上市首日破发 TOP3大新闻 证监会同意宇树科技科创板IPO注册 7月2日,证监会网站显示,同意宇树科技首次公开发行股票注…
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems s…
具身智能数采迎来了新范式
Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To reduce policy-call frequency while preserving temporal coherence, most generative pol…
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, an…
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted…
We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and the physical hardwa…
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world ro…
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembl…
全新的持续学习范式
四大互联网分别领投
Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts to a different but similar robot (e.g., from Pand…
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling,…
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary prefe…
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introd…
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector…

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe no…
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection…
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration da…
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain a…

Robots have powered manufacturing for decades, yet they stayed single-purpose and thrived only in perfect settings. Previous attempts at intelligent machines overpromised and under…