The round is being raised just months after the robot data startup exited from stealth.
AI 点评 · 机器人数据赛道爆发力惊人,三个月估值翻至12亿美元,资本抢筹信号明确。
共 420 条相关资讯 · 来自历史归档
The round is being raised just months after the robot data startup exited from stealth.
AI 点评 · 机器人数据赛道爆发力惊人,三个月估值翻至12亿美元,资本抢筹信号明确。
Nscale, which recently struck a $45 billion deal with Anthropic, is in talks to raise additional funds in anticipation of an upcoming IPO.
AI 点评 · AI算力资本竞速白热化,Nscale融资或成IPO前估值风向标。
AI 点评 · Transformer 生态剧变,关键人物去向牵动开源推理风向。

Public-market scrutiny will intensify pressure on the Claude maker’s unusual attempt to balance profit and purpose.
AI 点评 · 万亿估值上市背后,特殊治理结构如何平衡盈利与使命成最大看点。
The round came together after the data center developer reportedly secured a $13 billion contract with Jane Street.
The high-profile startup's annual revenue run rate stands at over $100 million.
AI 点评 · 高估值与高增长并存,资本豪赌AI推理赛道,市场风向标意义显著。
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learnin…
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests…

OpenAI recently estimated its Cursor partnership would make more than $1 billion in revenue a year, WIRED has learned. It still walked away after Elon Musk’s SpaceX acquired the AI…

Nvidia plans to acquire Hugging Face for about $12.9 billion, securing the central platform for open AI models. More than 18 million developers and 200,000 companies use the hub. C…

The long-rumored deal will give the chip giant access to—and help it promote—a huge repository of open-source AI models and data sets.
https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face...
The acquisition also leaves Sequoia-backed Serval as the de facto startup leader in AI IT service automation, industry watchers believe.
AI 点评 · 五亿美元并购凸显AI运维自动化赛道竞争白热化。
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for…
Wonderful said it will use its $550 million Series C funding to develop products faster, expand its FDE teams, and meet demand for its products.
This is Adobe's second acquisition out of India after Rephrase.ai in 2023
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands o…
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present Ver…
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remai…
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation metho…
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study loop…

We're excited to share that AWS has been recognized as a Leader in The Forrester Wave: AI Infrastructure Solutions, Q4 2025. In this evaluation of 13 providers, AWS received the hi…
AI 点评 · 云巨头获权威认可,AI基建竞争白热化,选型风向标值得关注。

Andrew Bailey warns G20 finance ministers about inflated AI valuations, growing leverage across markets, and cyber risks from frontier AI models. Cross-investments between AI compa…
AI 点评 · AI泡沫叠加杠杆攀升,金融稳定警报拉响,监管风向生变。
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of…
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress…
Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation…
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tabl…
Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the wrong purchase, can…
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, a…
As backlash grows over Flock's AI surveillance cameras, Texas Governor Greg Abbott has frozen state spending on them. The move came just ahead of the publication of a Texas Tribune…
AI 点评 · AI监控争议升温,州长冻结拨款凸显公共安全与隐私权平衡难题。
Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many obs…

OpenAI is cutting off the AI coding tool Cursor after SpaceX acquired the company, citing Elon Musk's track record of breaking contracts. Cursor co-founder Michael Truell downplays…
AI 点评 · 马斯克收购引发连锁反应,OpenAI切断合作凸显商业信任危机。
There's a lot of capital pouring into the business of giving models away.
AI 点评 · 开源权重模型成资本新宠,揭示AI行业商业模式剧变,值得关注其并购热潮背后的战略布局。
Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale and multilingual analysis. We introduce two corpus-level metrics based on self-supervised (SSL) speec…

Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from see…
AI 点评 · 双重盲测防作弊,机密空间护公平,AI评测可信度迎来革命性升级。
Our decision to wind down our contract providing OpenAI models to Cursor following its acquisition by SpaceX.
Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constrain…
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whe…

Nvidia is nabbing critical infrastructure for open models as interest grows.
AI 点评 · 英伟达揽下开源模型枢纽,AI生态话语权再落一子。
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yie…
Piloting the world's first double-blind AI evaluations

Anthropic has struck a deal worth around 45 billion dollars with British cloud startup Nscale, Bloomberg reports. The article Anthropic locks in 45-billion-dollar compute deal with…
Nvidia has reportedly agreed to buy Hugging Face, the popular open source AI hub, for $12.9 billion in a move that would let Nvidia both protect its chip empire and jump back into…
近日,AI基础设施公司基元律动(TokenRhythm)宣布完成新一轮融资,融资额达数千万美元。
The startup is only a year old but it has already generated a massive amount of hype (and money) while also spurring privacy concerns.
AI 点评 · 成立一年估值25亿美元,资本狂热背后隐私争议成最大看点。

Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whet…
AI 点评 · 评测与框架解耦,兼容OpenTelemetry,让多框架智能体横向对比成为可能。
The startup came out of stealth with $6 million in seed funding and a plan to use LLMs and cybersecurity know-how to make AI queries coherent.
Arga has raised $10 million in a seed funding round that was led by General Catalyst, with participation from Box Group, Emergence, Gradient and SV Angel.

Anthropic wants to sell investors on a theoretical market opportunity worth more than $30 trillion ahead of its planned IPO. The article Anthropic sees a market opportunity of more…
The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation.
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should ado…
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use…
The company's new fundraising total now stands at $232 million.
AI 点评 · 新融资提振信心,稳固开源图像生成赛道领先地位。
What we look for in a wellbeing evaluation www-cdn.anthropic.com
Lica co-founders are going to work on Gamma's new research team.
AI 点评 · 设计人才流向AI产品,预示Gamma正加速布局生成式设计的底层创新。
Funding better evaluations of AI’s impact on wellbeing Anthropic
AI 点评 · AI福祉影响评估获资金支持,伦理监管短板有望补强。
字节、汇川等已入股未来不远机器人最新一轮融资
The president bought when the stock was in the mid-$150 range. SpaceX finished trading on Monday back at its IPO price of $135.
AI 点评 · 总统精准抄底SpaceX,折射政商联动新常态,航天股波动成政策风向标。
GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against…
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested…
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the origina…
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternati…
General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valu…
AI 点评 · 顶级风投加注,估值60亿,通用AI进军机器人赛道,资本风向标值得紧盯。
Hugging Face has reportedly been fielding acquisition offers that would value the company at around $13B. But with the founders' feeling of responsibility to community, doubts aris…
Evaluator-neutral AI evaluation evidence, baseline regression gates, and JSON, JUnit, and SARIF reports for CI.

Nvidia is negotiating an investment in Perplexity at a valuation above $30 billion, more than 50 percent higher than its last funding round, The Information reports. Perplexity's a…
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks…

IT之家 8 月 23 日消息,据外媒 Interesting Engineering 今天报道,美国初创公司 Starcloud 已完成 2.5 亿美元 (IT之家注:现汇率约合 16.86 亿元人民币) A 轮扩展融资,估值达到 23 亿美元 (现汇率约合 155.09 亿元人民币) 。这笔新资金将用于研发配备高性能 GPU 的卫星,为 AI 提供算力。…
AI 点评 · 英伟达加持下,太空算力布局提速,或成AI基建新赛道风向标。
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime co…
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of train…
AI 点评 · 资本加码验证AI材料研发闭环,产业化提速成关键看点。
The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy regulations. While new regulation requirements vary, many include a documentation arti…

Nvidia is paying $6 billion for software that builds AI models from the startup Poolside, and it wants to bring on 109 employees. The article Nvidia is acquiring Poolside's "Model…
Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce scenarios. While recent LLM- and VLM-based simulators have made encouraging progress,…

Unitree Robotics rose 460 percent in its Shanghai IPO, hitting a valuation of around $50 billion. But an FT report shows much of the demand for its robots comes from state-backed t…
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto analysis, and a no-key Replay demo.
SpaceX was reportedly in talks to buy AI coding startup Cognition. SpaceX has already acquired Cursor as it races to catch up to rivals like OpenAI and Anthropic in enterprise AI.
AI 点评 · AI编程赛道并购热升级,SpaceX布局意图明显,折射科技巨头对智能编码工具的争夺白热化。

In a letter to investors, Stripe declares January 1 the "beginning of the singularity" and uses that as a reason to stay private. Not that it needs one: revenue grew 41 percent in…
AI 点评 · 支付巨头以奇点论推迟IPO,折射AI时代资本叙事新逻辑,看点十足。
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code req…
With a looming IPO, intense competition from Anthropic, and Chinese and open-weight rivals nipping at its heels, OpenAI has plenty of reasons to move fast. Instead, it hit the brak…
AI 点评 · 上市压力与开源夹击下,OpenAI急刹车反而凸显战略定力,看它如何在竞速与稳健间找平衡。
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protoco…
Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology…
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited unde…
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study…
Jane Street has installed Etched's first shipped AI cluster system, and was so impressed, it led another massive round, the startup says.
AI 点评 · 机构背书叠加实测惊艳,估值月翻倍印证AI芯片赛道资本热度。
回头看,A社就这样完成了反超
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study…
Groq raised $350 million at a $3.5 billion valuation as the former AI chipmaker pivots to a neocloud business and expands its Nvidia-powered data center footprint.
AI 点评 · 芯片商转型云服务,估值翻倍印证AI算力需求新风口。
Nvidia's investment in SoftBank's data center developer will guarantee its chips power an OpenAI data center.
AI 点评 · 英伟达押注软银数据中心,锁定OpenAI算力订单,跨界资本绑定凸显AI基建竞争白热化。
The funds will allow Wispr to increase its footprint as it ventures into new areas, such as meetings, with its newly released note-taker tool.

According to Bloomberg, Stripe is acquiring the AI startup OpenRouter for more than $7 billion, up from a latest valuation of $1.3 billion. OpenRouter offers access to over 400 AI…
OpenAI joins PORTS-Pike project, expanding community investment and supporting thousands of Southern Ohio jobs
OpenRouter's CEO recently described the startup as Stripe for AI.
AI 点评 · 支付巨头收购AI网关,打通模型调用与商业变现,或重塑AI应用付费生态。
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollo…
将超越马斯克的SpaceX,成为史上上市估值最高的IPO。
AI coding startup Cursor is now officially a part of SpaceX.
AI 点评 · 交易背后暗藏AI编程与航天巨头的战略协同,或重塑行业竞争格局。
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represent…
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execu…
Network-level maintenance planning requires repeated evaluations of equilibrium traffic flows under road capacity reductions. While equilibrium traffic assignment models are well established, their re…
The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail…

Congrats to the team!

IT之家 8 月 14 日消息,彭博社今天(8 月 14 日)发布博文,报道称在收 IPO(首次公开募股)上市前,OpenAI 的 发展势头良好,2026 有望实现年化收入超过 400 亿美元 (IT之家注:现汇率约合 2,702.76 亿元人民币) ,比 2025 年收入(约 200 亿美元 (现汇率约合 1,351.38 亿元人民币) )翻一番。 据彭博…
AI 点评 · 营收翻番印证AI商业化提速,IPO前估值想象空间再度打开。
AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned.
AI 点评 · AI巨头融资热潮折射资本狂热,190亿估值与超额认购凸显行业泡沫风险与算力成本压力。
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process a…
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, howev…
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis…
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions c…
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely measure whether ski…
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences b…
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to charact…

Rapid revenue growth fuels hope Claude maker's IPO is the biggest listing in history
重新定义物理AI数据基础设施
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to charact…
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences b…
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however…
Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation.
AI 点评 · AI编程赛道估值飙升,印证资本对自动化开发工具商业前景的狂热押注。
Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Altimeter Capital.
AI 点评 · AI企业服务融资火爆,资本重注押宝OpenAI生态,企业落地应用成新战场。
Build a custom LLM post-training pipeline using AllenAI’s Open Instruct framework. This comprehensive guide walks through Supervised Fine-Tuning (SFT), Direct Preference Optimizati…
AI 点评 · 端侧模型头部玩家冲击IPO,AI商业化路径获资本验证。
This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.
AI 点评 · AI建站独角兽估值破百亿美元,高增长验证低代码赛道爆发力。
Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center of multiple legal cases.
Blacksmith says revenue has grown more than tenfold over the past year.
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually imple…

OpenAI wrapped up a $7 billion stock buyback, letting current and former employees sell shares at the company's $852 billion valuation. The move is meant to ease pressure on employ…
AI 点评 · 高估值下员工套现,反映AI巨头内部流动性需求,折射行业资本热度与人才留任压力。
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet…

Anthropic is preparing an IPO for September or October, according to the Wall Street Journal, potentially the largest ever. During investor meetings, the company, valued at $965 bi…
真就把GPU当理财产品了~
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, whi…
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assis…
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear h…
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requiremen…
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corrido…
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination ov…
AI 点评 · 对比式多语基准揭示数据污染与本地化短板,为翻译模型鲁棒性评估提供新视角。
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs…

OpenAI acquired NextSlide, the startup that turned prompts, notes, documents, and research into editable presentations. The article OpenAI acquires NextSlide to bring AI-generated…
“我们90天就成了独角兽”
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a re…
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination ov…
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's c…
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corrido…
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific…
A verified 2026 comparison of LLM observability platforms covering tracing depth, evaluation capability, production monitoring, and pricing. The post Top LLM Observability and Eval…
Regression testing and configuration sweeps for RAG pipelines, with the statistics to know whether a change actually helped.
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-ou…
Founder Ahmed Beshry describes the startup’s product as one “that could turn prompts, notes, documents, or research into a polished, editable presentation.”
AI 点评 · OpenAI并购NextSlide,或为ChatGPT注入演示生成能力,AI办公赛道竞争再升级。
Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible epochs to converge on optimization of a goal.
AI 点评 · 用进化算法批量培育AI代理,跨代优化目标,为自主智能体进化提供新范式。
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this…
FDE and AI engineering guide for production systems: value, architecture, evals, security, deployment, and operations.
The Applied AI Field Guide: fieldwork, value engineering, and operations for AI that works beyond the demo.

AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo ch…
AI 点评 · 芯片直写模型权重,速度极致但灵活性受限,AMD此举或为定制化AI推理开辟新路。
Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulatio…
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
AI 点评 · AI安全新战场,OpenAI提前为智能体实战防御定标,值得跟进。
此后模型、算力、数据中心都在合作范围内
AI 点评 · 杭州双子星联手,AI与机器人产业协同将重塑生态,战略意义远超投资本身。
https://ir.amd.com/news-events/press-releases/detail/1296/am... https://chatjimmy.ai/
AI 点评 · 芯片级模型固化,AMD硬刚英伟达推理赛道,收购背后是架构革命信号。
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppre…
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed…
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias de…
White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired…
The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents.
AI 点评 · 创始团队回归掌舵AI产品,标志电商营销赛道竞争升级,战略价值显著。
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of rel…
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Ex…
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predom…
AI 点评 · 华为基因主导智元,谷歌科学家退出,或重塑机器人行业人才格局。
Once, I had some questions about why SpaceX, Elon Musk's healthiest company, acquired xAI, his sickliest one. Now I have some questions about why we're calling the whole thing Spac…
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
AI 点评 · 第三方评测暴露AI安全短板,新防护措施或成行业标准。
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend its…
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend del…
Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data…
【导语】当前,电子级玻纤布价格较2025年低点翻倍、FR-4覆铜板涨幅超270%,AI封装载板部分交期拉长至6个月以上,有机基板供给危机持续加深。在此背景下, 国内玻璃基板领军企业巽霖科技宣布完成近2亿元B轮融资——这已是该公司半年内完成的第三轮次融资。本轮由英诺基金、千乘资本、海目星、光莆股份、同鑫资本等新股东投资,金雨茂物、海通开元、北岸产投等老股东持续…
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,精密减速器企业陶世智能科技有限公司(以下简称“陶世”)近日完成超亿元融资。本轮融资由国创集团,海川聚义,杭州众燊,新智资本参与,德太资本担任财务顾问。融资资金将主要用于产品研发、产线扩建以及机器人领域市场拓展。 陶世成立于2016年,总部位于深圳,主要研发和生产正交90度微型减速器及机器人关节模组。公司核心产品为…
从全球经验来看,很多科技领军企业的崛起,都少不了创业投资的支持。今年上半年,高技术产业投资同比增长4.6%,这一数据背后,是创业投资的活跃度在持续回升。业内人士普遍认为,AI技术的发展,是这一轮创投市场回暖的核心驱动力,资金正加快涌向AI、量子科技、算力等领域。数据显示,2026年上半年中国AI创业公司融资总额突破3000亿元,已超过2025年全年。而资金的…
AI 点评 · 创投风向标转向硬科技,3000亿真金白银验证AI赛道热度。
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entrie…
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We…
Design Arena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs.
AI 点评 · 人类反馈数据成AI军备竞赛新筹码,5百万用户规模验证了品味训练的商业价值。
Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates "a CEO in the United States," the prompt leaves demographic…
文 | 赵京娜 访谈 编辑 | 海若镜 36氪获悉,近日奇点逃逸完成千万级种子轮融资,由星连资本与水木创投联合领投,奇绩创坛跟投。其正在研发AI原生团队协作操作系统Nexus,让人、Agent、任务、知识和工具基于同一份组织状态持续协作,并让系统从每一次协作中有证据地变强。 奇点逃逸创始人兼CEO薛传奕,本科、博士阶段均在清华大学就读,研究方向覆盖强化学习与…
今日热点导览 马斯克关注了DeepSeek的X账号 祥鹏航空回应航班误发过期方便面 OpenAI或将IPO推迟到明年 SpaceX首份财报即将发布 小米多款手机正式涨价 每月10万美元,特朗普“真实社交”售卖“优先访问权” TOP3大新闻 蔡崇信宣布离婚,不涉及出售阿里股份 8月1日,阿里巴巴集团董事会主席、美国职业篮球队布鲁克林篮网主要所有者蔡崇信与妻子吴…
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervisi…
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, whi…
8月2日消息,中信证券研报表示,A股本轮调整更多是拥挤交易的修正,而非韩国式的去杠杆冲击,这体现在:1)整体杠杆情况比较安全,7月上涨股票数量接近全A一半,明显超过6月;2)相比全球历史上典型的去杠杆行情,当前的融资回落幅度并不算大;3)ETF市场呈现持续的资金流入,科技类ETF的流入提供了充足流动性支持。当然,局部的流动性压力依然存在,尤其是部分非核心AI…
AI 点评 · 券商罕见喊话明确转多,杠杆数据支撑逻辑,给市场注入强心剂。
对前沿实验室估值悲观
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse tr…

The European Commission wants to build up to seven AI gigafactories across Europe, backed by around 30 billion euros in public and private funding. For context, the major U.S. tech…
AI 点评 · 欧盟斥巨资建AI工厂,但与美国巨头投入差距悬殊,凸显产业竞争力隐忧。
文|王欣逸 编辑|张雨忻 一句话介绍 国内唯一做多模态长记忆的公司——丘脑智能,推出原生多模态记忆基座,押注AI从通用走向个性化,最终走向主动智能。 主动智能,指的是AI能在足够了解用户的基础上,在合适的时间、以恰当的方式主动跟用户交互。要实现主动智能,Memory是必须要跨过的门槛。 融资情况 近日,丘脑智能已完成数千万元种子轮融资,投资方包括深圳一线基金…
AI 点评 · 多模态长记忆是主动智能关键门槛,资本押注稀缺赛道,看点十足。
作者|黄楠 编辑|袁斯来 硬氪获悉,Physical AI平台公司「昆腾动力(Quantum Dynamics)」近日完成超亿元种子轮融资,本轮由云启资本、商汤科技联合投资。资金将主要用于Physical AI核心技术研发、人才梯队建设及全球化市场拓展,加速其面向物理世界的智能系统从底层模型到场景化落地的全链路构建。多维资本参与项目孵化与团队组建。 昆腾动力…

In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.
7月31日消息,据报道,数据中心开发商Nexus Data Centers正就为其位于美国得州、服务Anthropic的AI数据中心项目筹集约150亿美元资金进行深入谈判,谷歌将为该项目提供融资担保并供应芯片。谷歌已同意在Anthropic无法履行租赁和电力支付义务时,为其相关承诺提供数十亿美元担保。该担保覆盖Anthropic签署的4份数据中心租赁协议,以…
AI 点评 · 巨头资本深度捆绑,谷歌为Anthropic提供担保并供应芯片,预示AI军备竞赛进入基础设施重资产博弈

MIT students and postdocs discussed science funding and research with policymakers in Washington during the MIT Science Policy Initiative’s annual Congressional Visit Days.
AI 点评 · 学生与政策制定者直接对话,推动科研经费议题,展现青年影响决策的实践路径。
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual ha…
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, th…
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is…
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, deve…
Do you remember Friend? The Friend that launched an AI pendant, spent $1.8 million of its $2.5 million in funding to acquire friend.com, and plastered the NYC subway with ads promo…
大公司: 证监会同意沈鼓集团沪市主板IPO注册 36氪获悉,证监会同意沈鼓集团股份有限公司首次公开发行股票并在沪市主板上市的注册申请。 三星电子:已同全球五大数据中心客户签订合同 据报道,三星电子7月30日预测,明年存储芯片供不应求现象将进一步加剧,超大规模云服务商扩大对人工智能(AI)基础设施的投资,带动服务器、固态硬盘、高带宽内存(HBM)需求进一步增长…
Investigating three real-world incidents in our cybersecurity evaluations Anthropic
AI 点评 · 聚焦AI系统真实安全漏洞,提供实用防御经验,对提升行业安全标准有重要参考价值。
今日热点导览 C长鑫成交额达400亿元 国家烟草专卖局约谈爱奇迹(深圳)技术有限公司 外交部回应美国实施先进机器人进口限制 宝马拟在德国裁员数千人,通过自愿离职计划削减成本 美联储宣布维持利率不变 TOP3大新闻 Anthropic首席执行官等多位AI大牛签署联名信,呼吁控制人工智能发展步伐 超过1100名来自领先人工智能公司的高管及员工联名签署了一封公开信…
When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little tidbit about how its investments in the tw…
AI 点评 · 微软披露对Anthropic投资账面盈利32亿,揭示AI投资回报分化趋势。
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answ…
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional co…
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment…
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is…
Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence learning, nonlinear recurrent updates often require repeated circuit evaluations and…

During a security evaluation, OpenAI's autonomous hacking models broke into Hugging Face and used exposed credentials on four other services. Hugging Face reconstructed about 17,60…
大公司: 美的集团:受欧洲持续极端高温影响当地空调需求爆发,空调芜湖、广州双基地一月内新增欧洲订单共20万台 36氪获悉,美的集团在互动平台表示,受欧洲持续极端高温影响,当地空调需求爆发,美的空调芜湖、广州双基地一月内新增欧洲订单共20万台。PortaSplit移动分体空调自6月新增接单超16万台。7月,法国3万台移动空调紧急订单中2万台已完成发运。 新东方…
作者|黄楠 编辑|袁斯来 硬氪获悉,柔性触觉感知企业「尧乐科技」近日完成Pre-A+新一轮融资,本轮融资由鼎和高达领投,上市公司常熟汽饰、祖龙娱乐跟投,云道资本担任长期独家财务顾问。资金将主要用于柔性织物传感技术研发、产品性能迭代升级,加快数据手套产能拓展和批量交付,并完善数据采集、标定与接口能力,加速面向具身智能、世界模型与智能座舱等场景落地。 尧乐科技以…
The deal is Cyera's third acquisition this year.
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on nar…
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained an…
大公司: YouTube与NBC环球达成合作协议,将为美用户提供流媒体捆绑服务 7月27日,美国流媒体平台YouTube与媒体巨头NBC环球(NBCUniversal)宣布深化战略合作。根据双方达成的协议,自2027年初起,美国本土的YouTube Premium订阅用户将无需支付额外费用,直接获取NBC环球旗下流媒体平台Peacock Premium的会员…
作者 | 黄楠 编辑 | 袁斯来 硬氪获悉,人形机器人基础模型公司德塔智能(Delta Intelligence)近日完成近5亿元天使++轮融资,本轮投资方包括多家上市公司产业方和头部财务投资机构。资金将主要用于人形机器人基础模型持续迭代,加快自研数采设备的量产与数据闭环建设,扩充核心研发团队,以推动技术在真实工业场景中的工程化落地与验证。 这也是公司成立半…

Anthropic and OpenAI employees are expected to give generously after their companies go public. “It’s going to be a wild ride,” says one nonprofit leader.
文|邓咏仪 编辑|张雨忻 吴秉哲每天至少复盘一次。 作为北大计算机博士,他同时也是一个高频的投资者。每天盘后,他会回顾当天的判断——哪些预案执行了,哪些被盘面的新信息打乱了,哪些潜意识的决策事后被验证是对的。 这个习惯持续了很多年。但有一个问题是,大部分复盘都没被系统性地沉淀下来。 “你的认知是一种资产,但现在所有东西都停留在人脑里。”吴秉哲说。他和Kand…
Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales.
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, exis…
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenan…
AI 点评 · 混合RAG架构结合评估框架,解决科研设施多源数据检索难题,提升运维知识利用率。
AI 点评 · AI行业高管动态频繁,融资与离职背后揭示竞争白热化,日本拆机认输更显中国AI硬件实力。
今年 WAIC 前夕,月之暗面发布了 Kimi K3,发布即破圈。但资本市场的注意力,更多地落在了另一件事上。 过去半年,这家明星大模型公司估值翻了 6 倍,目标 300 亿美元,同步推进赴港 IPO。而在 2025 年底的跨年夜全员信中,创始人杨植麟写下了一句话:2026 年聚焦 Agent,不以绝对用户数量为目标。 估值半年翻 6 倍的同时,他们也主动放…
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,智能体育硬件公司「一思智能」(AceiiLab)开启批量交付,此前完成超千万元天使轮融资,由零以资本、变量资本、海益资本投资。资金将主要用于产品研发迭代及市场拓展。 一思智能成立于2024年12月,从AI网球机器人切入,尝试构建覆盖硬件、软件、数据、AI教练与运动服务的智能训练生态。公司创始人刘礼谦拥有十余年机器…
这轮融资计划至少募资100亿元
36氪获悉,近日,企业级AI Agent基础设施专属服务商“词元无限”宣布完成天使++轮融资。本轮融资由临芯投资领投,华控基金跟投,这也是词元无限在一个月内完成的第二笔融资,累计融资额已达数亿元人民币。资金将主要用于加速打造其企业级AI Agent基础设施平台,深化与清华大学、北京航空航天大学等高校的联合研究,并持续构建面向Agent应用范式的下一代基础设施…
AI 点评 · 资本密集加注企业级AI Agent赛道,一个月内两轮融资,凸显市场对基础设施层创新的迫切需求。
近日,智能烹饪机器人企业“智谷天厨”官宣完成新一轮近亿元战略融资。本轮融资由招商局创投领投,跃迁资本跟投,云启资本、啟赋资本等老股东超额追投。融资完成后,公司将持续推进物理 AI 研发,并加速中式及泛亚洲餐饮的全球化落地。

IT之家 7 月 27 日消息,据《华尔街日报》昨日报道,英伟达正在与人工智能企业 OpenAI 谈判,拟为其提供约 2500 亿美元(IT之家注:现汇率约合 1.69 万亿元人民币)的融资担保。 据报道,英伟达提供的担保资金,将帮助 OpenAI 租用软银旗下能源公司在美国俄亥俄州南部的 10GW 级数据中心项目。如果算上芯片等成本,该数据中心的投资总额将…
Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches l…
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed…
A production-oriented LLM engineering platform with OpenAI-compatible serving, streaming, observability, evaluation, reproducible experiments, and deterministic…
A production-oriented LLM engineering platform with OpenAI-compatible serving, streaming, observability, evaluation, reproducible experiments, and deterministic…
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of out…
The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief that how AI assistants interact with users i…
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large lang…
The AI lab Midjourney continues to expand its purview beyond image and video generation.
7月24日,据报道,日本软银集团正考虑收购瑞士AI机器人初创公司Gravis Robotics AG,旨在进一步加码支撑机器人技术的具身智能与人工智能领域。软银计划对这家初创公司建立大量股权头寸,并将其归入其正在筹建的新AI及机器人实体公司“Roze”旗下。交易可能分阶段推进,包括购买现有股份以及随着时间推移注入新资金,最终对Gravis的估值可能超过5亿美…
AI 点评 · 软银加码具身智能,高估值收购预示AI机器人赛道加速整合。
AI 点评 · 具身智能赛道爆发,数据服务成新风口,商业模式能否跑通是核心看点。
据报道,债券投资者正寻求在最新一笔120亿美元Meta支持的数据中心交易中获得显著更高的收益率,与九个月前达成的条款相比,市场已将人工智能融资的更高风险计入价格。据知情人士透露,位于德克萨斯州埃尔帕索、容量近1吉瓦的数据中心项目,正计划通过贝莱德旗下特殊目的公司发行债券,初步讨论中收益率超过7%。上述人士补充称,价格讨论仍处于早期阶段,最早可能于下周一正式启…
AI 点评 · AI数据中心融资成本飙升,折射市场对技术泡沫风险的重新定价。

IT之家 7 月 24 日消息,据《华尔街日报》今日援引知情人士消息, 海外支付巨头 Stripe 正与全球最大 AI 聚合平台 OpenRouter 洽谈收购事宜 。 知情人士透露,虽然本次交易的具体条款尚未确定, 但 OpenRouter 的估值可能接近 100 亿美元(IT之家注:现汇率约合 678.3 亿元人民币) 。相比之下,OpenRouter…
AI 点评 · 支付巨头跨界收购AI平台,折射出AI服务与金融基础设施融合的产业趋势。

Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few…
AI 点评 · 工业级AI Agent评估框架落地,错误率从八分之一降至五十分之一,检测时效从小时级缩至分钟级。
Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs required, it says.
AI 点评 · 天才辍学生团队造出颠覆性AI芯片,无需GPU,估值飙至103亿美元。

Initial research projects advance national priorities across natural resources, manufacturing, nuclear physics, and more.

Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And…
AI 点评 · AI加速新药研发,降低失败风险,有望颠覆传统制药周期与成本。
ServiceNow's investment gives BusinessNext a strategic partner to expand its AI-powered banking software globally.
AI 点评 · ServiceNow押注银行AI软件,金融科技布局再下一城,行业整合信号明显。
硬氪获悉,水下AI自然探索科技公司Deeplore近日完成数千万元种子轮融资,由五源资本、顺为资本联合投资。资金将核心用于研发团队扩建与水下AI技术深耕,加速首款AI潜水面镜落地迭代,搭建水下自然探索智能平台的技术底座。 一群在消费电子行业征战十余年的大疆老兵,把目光从天空转向了深海。 Deeplore 2025年诞生于深圳,核心团队大多出自大疆,覆盖产品定…
AI 点评 · 大疆老兵跨界水下AI,首创光学系统,获顶级资本押注,技术落地潜力巨大。
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1121 comments)
AI 点评 · AI巨头间的安全博弈折射出模型评估的致命漏洞,行业标准亟待建立。
图源/企业 本文约 2100 字,建议阅读 5 分钟 作者丨欧雪 编辑丨袁斯来 硬氪获悉,机器人灵巧操作方案提供商拾玥科技(LumiBot)近日完成数千万元天使轮融资。本轮由某上市公司旗下基金领投。资金将主要用于团队扩张、产品迭代与小批量产线建设,以及灵巧操作数据采集与模型研发。 拾玥科技成立于2025年初,专注于为工业制造、精密装配等高精度场景提供灵巧操作…
图源/企业 本文约 1400 字,建议阅读 3 分钟 作者丨欧雪 编辑丨袁斯来 硬氪获悉,工业AI设计研发解决方案供应商「设序科技」于近日正式完成B轮超亿元融资,累计获超3亿元融资,投资方包括深产投、合鼎共及老股东涌铧投资等。融资将用于市场开拓(含出海)及核心模型技术研发。 设序科技成立于2020年,旗下产品“则形AI”是基于自研工业世界模型与自然语言大模型…
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its cor…
AI 点评 · 首个抗污染多领域编码智能体评测基准,为评估AI编程能力提供更可靠标准。
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an eva…
Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasing creators and algorithmic feeds, the st…
AI 点评 · 反算法反广告的私密社交模式获资本认可,或重塑社交产品逻辑。
OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and access to Codex.
AI 点评 · OpenAI在佐治亚州建AI基建,聚焦社区投资与绿色能源,体现科技落地的社会责任。
Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
硬氪获悉,智能船艇公司同舟智航近日完成数千万种子轮融资,由英诺天使和吴中金控联合领投,江苏金桥基金跟投。融资将用于补充核心研发与量产团队人才,同步启动海外品牌建设与渠道开拓。 同舟智航2025年6月成立,核心理念是用人工智能赋能船舶。创始人张呈伦拥有10余年汽车自动驾驶行业经验,曾任华为商用车解决方案部负责人,并创立过规模超百人的矿山自动驾驶公司。 团队自动…
Anthropic and OpenAI's aggressive 2026 acquisition sprees set the stage for a weekend rumor.

IT之家 7 月 22 日消息,彭博社今日报道称,国内大模型独角兽月之暗面计划于 8 月启动上市前最后一轮融资谈判,目标投前估值 500 亿美元(IT之家注:现汇率约合 3389.04 亿元人民币)。 该公司预计将在未来几天完成现一轮融资,本轮融资估值约投前 315 亿美元(现汇率约合 2135.1 亿元人民币)。融资完成后,月之暗面将立即启动新一轮融资洽谈…
AI 点评 · 估值飙升至500亿美元,揭示AI独角兽的资本狂热,上市进程加速值得关注。
国内大模型独角兽月之暗面计划于8月启动上市前最后一轮融资谈判,目标估值为投前500亿美元。据知情人士透露,公司预计将在未来几天完成Kimi K3发布前启动的一轮融资,该轮融资估值约投前315亿美元。融资完成后,月之暗面将立即启动新一轮融资洽谈,这将成为公司赴港上市前的最后一次私募股权融资。公司最快可能于6个月内登陆香港资本市场。 (财联社)
AI 点评 · 估值飙升至500亿美元,凸显大模型赛道资本热度,上市前最后一轮融资或成行业风向标。
https://www.axios.com/2026/07/21/openai-says-hugging-face-br... See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments) https…
AI 点评 · 两大AI巨头联手回应安全事件,凸显模型评估环节的漏洞风险与行业协作必要性。
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain…
AI 点评 · 打破传统检索局限,以评分标准导向提升文档集质量,为AI生成奠定更优基础。
Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently req…

The conversation about AI often centers on algorithms, computing power, or huge investments in new semiconductor fabrication plants and hyperscale data centers. But beneath each of…
AI 点评 · 材料科学突破将催生更高效AI硬件,为算力瓶颈提供根本性解决方案。
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
https://archive.ph/20260720174223/https://asia.nikkei.com/bu...
本文约3100字,建议阅读7分钟 作者 | 彭孝秋 编者按: 《上市前夜》栏目聚焦企业冲刺资本市场的关键时刻。每一份招股书里,都藏着一家企业上市前的野心、周期与隐忧。这是第五期——芯天下。 7月9日,深圳存储芯片设计公司「芯天下」再次递表港交所。事实上,芯天下早在今年1月9日就递交过招股书,但6个月后失效。此次IPO由广发证券、中信证券联席保荐。 芯天下是谁…
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation co…
AI 点评 · 用词级时间戳激活可控逐字识别,解决ASR转录风格不一致导致的解码不稳定问题。
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather…
AI 点评 · 多模态幽默AI研究难点与突破,揭示文化理解与意图识别的新挑战。
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framewor…
大公司: 摩根士丹利成为华尔街AI债务交易头号银行 摩根士丹利已崛起为华尔街构建人工智能(AI)繁荣背后融资架构的主导力量,设计了新的债务和股权融资模式,为数据中心的大规模建设输送数百亿美元资金。据业内高管称,自去年以来,该行已成为主导顾问,牵头了规模最大、最具创新性的AI基础设施融资交易。 三星电子晶圆代工正评估将美国泰勒厂的初期生产规模扩大至原计划的2倍…
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCu…
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profil…
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting bec…
36氪获悉,月之暗面已通知投资者调整公司架构并筹备赴港IPO,有望最快在6个月内完成上市。7月16日,月之暗面发布全球参数规模最大的开源模型Kimi K3,在Code Arena上超过Claude Fable 5和GPT-5.6 So。
3500亿估值曝光
Databricks has remade its image into an AI company and has published research on the cost savings of open-weight AI models for coding.
AI 点评 · 估值188亿美元,从数据平台成功转型AI公司,开源模型成本优势研究或重塑行业格局。
Apple filed a trade secrets lawsuit against OpenAI last Friday, and it’s not messing around. The complaint alleges a pattern of misconduct reaching all the way up to OpenAI’s chief…
AI 点评 · 苹果诉OpenAI窃密,或动摇其IPO前景,揭示AI巨头间法律风险激增。
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evalu…
Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,全身多模态融合触觉解决方案公司模感科技(MoSense)近日完成数千万元天使轮融资,投资方包括红杉中国、高瓴创投及智元机器人。本轮融资资金将主要用于加速研发、团队扩充、算力投入及量产测试体系建设。 模感科技成立于2026年5月,总部注册于上海,在深圳前海设有研发中心,聚焦机器人全身多模态触觉感知系统研发。公司正式…
AI 点评 · 红杉、高瓴、智元联手押注,机器人触觉赛道技术壁垒高、应用前景广。
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,具身智能世界模型公司日冕开物(北京日冕机器人有限公司)近期完成连续两轮种子轮融资,融资合计金额达数亿元人民币,由鼎峰科创、远图未来、百度风投、沃衍资本、武岳峰科创、万林国际共同参与投资。同时,新一轮融资也在同步交割中。 此前融资资金主要用于自研世界模型 LaMPA 的研发迭代、强化学习体系建设,以及数据闭环和产品…
AI 点评 · 前蔚来华为核心团队跨界创业,三个月获数亿融资,凸显具身智能赛道热度。
AI 点评 · RISC-V架构在AI算力需求下获资本押注,智能体时代催生新硬件机遇。
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completio…
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other…
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completio…
AI 点评 · 成本感知评估揭示安全AI实用门槛,突破传统成功率局限,更贴近真实攻防场景。
The growing use of Bitcoin as a decentralized digital asset and investment tool has sparked strong interest in understanding its market behavior. This study presents a new approach to analyze Bitcoin…
AI 点评 · 区块链数据解锁市场情绪,为投资决策提供全新量化视角。

Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that…
AI 点评 · 企业盲目放权AI代理却缺乏可信评估,暴露出部署与安全之间的严重脱节。
Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed the development of ChatGPT, Andrew Dai exp…
AI 点评 · 顶尖AI人才创业获3亿美元估值,人才价值与技术积累成资本关注核心。
文|胡香赟 编辑|海若镜 互联网大厂做医疗AI,越来越偏爱“减重”场景。 近日,蚂蚁集团宣布投资薄荷健康,持股比例超28%,成为薄荷健康最大外部股东。薄荷健康成立于2008年,这家互联网医疗公司通过为用户提供饮食记录、科学减重方案等,聚拢了超2亿用户,收录160万条食物数据条目。 2021年,薄荷健康完成D轮融资,估值为20亿人民币。本次投资,据接近交易人士…
作者|黄楠 编辑|袁斯来 在首款网球发球机卖爆后,庞伯特做了一款多合一AI教练机器人 。 这家公司是在一级市场很受关注的标的,2025年半年内完成三轮融资累计数亿元;旗下产品也在全球拥有了30余万用户,设备累计发球总量超20亿次。 但当下硬件赛道已经不容小而美的公司慢慢成长,网球赛道涌入了越来越多的对手,国内硬件创业团队快速入局,海外老牌专业设备厂商也在加速…
今日热点导览 澳洲将设AI办公室,限制数据中心能源消耗 巴菲特宣布8年内清仓伯克希尔股票 前高通自动驾驶工程副总裁将加入英伟达自动驾驶团队 长鑫科技:上市后仍将保持无实际控制人的控制权架构 谷歌参与建设美国最大太阳能项目,总规模达2.5GW TOP 3大新闻 阿里、百度分别为苹果智能提供AI相关能力支撑 36氪获悉,网信办公布7款端侧生成式AI服务备案,苹果…
据知情人士透露,Anthropic正安排与投资者举行会面,为潜在的大规模IPO做准备。消息人士称,负责此次IPO的承销银行将在未来几周安排投资者与这家Claude聊天机器人开发商会面。此前有报道称,Anthropic最早可能于10月启动IPO。(新浪财经)
AI 点评 · AI独角兽加速上市,Claude开发商估值飙升,或成今年最受瞩目的科技IPO。
AI 点评 · 用Agent手机概念展示商业落地能力,为IPO估值注入强心剂。
The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made before and after SpaceX went public.
AI 点评 · 马斯克承诺遭市场质疑,SpaceX股价回落至发行价,折射行业估值理性回归。
Livestream shopping platform Whatnot has acquired AI startup Shaped, a machine learning company focused on real-time recommendations and search. The deal will bolster Whatnot’s per…
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the metho…

IT之家 7 月 15 日消息,通用人形机器人公司逐际动力 LimX Dynamics 宣布完成 Pre-IPO 轮融资, 融资金额近 2 亿美元 (IT之家注:现汇率约合 13.56 亿元人民币)。过去半年累计完成融资 4 亿美元。 逐际动力创立于 2022 年,总部位于中国深圳,是一家 AI 驱动的人形机器人公司,主要打造全尺寸通用人形机器人,并衍生了包…

OpenAI employees have donated more than $215,000 to a political effort opposing Leading the Future, a group backed by the company’s president, Greg Brockman.
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,上海追知工程科技有限公司(以下简称“追知工科”)近日完成数千万元种子轮融资,由L2F光源创业者基金、尚融资本、一村资本联合投资。本轮融资将主要用于核心产品研发、团队建设及市场拓展。 追知工科成立于2024年2月,是一家 聚焦垂域工业智能体 的科技企业,同时也是上海交通大学成果转化企业、上海人工智能研究院战略孵化企…
距首轮估值上涨37%
The funding discussions point to investor interest in applying AI to make breakthroughs in life sciences.
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tight…
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-dr…
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning,…
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the scor…
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed prog…

IT之家 7 月 15 日消息,据彭博社昨日援引知情人士消息,中国人工智能公司 DeepSeek(深度求索)已开始筹备 IPO(首次公开募股),最快可能在今年内提交上市申请。 据知情人士透露,DeepSeek 正在规划中国内地上市,以便最早于 2027 年完成上市。 公司目前正在与会计师事务所 、 投行顾问展开合作 , 计划在今年 12 月底前完成财务报告…
凌晨三点,一家刚成立不久的AI创业公司,可能已经在同时服务旧金山的客户、采购首尔的技术服务,并与拉各斯的合作伙伴签下合同。这家公司甚至还没有招到第一名全职财务人员,业务却已经跨越多个市场、币种和监管辖区。 AI正在让这样的创业路径成为可能。过去需要市场、运营、客服等一整套全球化团队才能完成的工作,现在借助智能体就能承担相当一部分。新一代初创企业不必再按照“先…
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
作者 | 张子怡 编辑 | 袁斯来 硬氪获悉,桌面级CNC(数控机床)企业「奇塑科技」于近期完成近亿元人民币天使轮融资。本轮融资由商汤国香、首形科技、全新世、奇绩创坛共同投资。本轮资金将主要用于桌面级CNC软硬件产品的研发、量产以及后续的市场营销。 在消费级3D打印机和激光雕刻机已完成大众市场教育,并诞生了诸如拓竹、xTool等多个头部企业后,能够实现“从软…
触觉手套(图源/企业) 本文约 2400 字,建议阅读 6 分钟 作者丨欧雪 编辑丨袁斯来 硬氪获悉,空间智能公司「大衍科技」近期完成数千万元天使轮融资,由松禾清湛领投,浙江省省金控与广州番禺创新基金等国资背景机构参与。资金将主要用于触觉大模型研发、机器人数据产线建设及团队扩张。 「大衍科技」2025年5月成立于浙江桐乡,是一家专注于4D世界重建、世界模型及…
With the cash, the company aims to expand its world model offering and reach customers across geographies.
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
IT之家 7 月 14 日消息,国产 DRAM 存储芯片龙头长鑫科技的科创板 IPO 进程已进入最后冲刺阶段。 长鑫科技将于 2026 年 7 月 16 日正式开启网上申购及网下申购。此次 IPO,长鑫科技初始公开发行 66.88 亿股,占发行后总股本约 10%,拟募资 295 亿元。 这一募资规模不仅使其成为 2026 年以来 A 股市场规模最大的 IPO…
AI 点评 · 国产存储芯片龙头上市,全员持股造富效应显著,折射中国半导体产业资本化里程碑。
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluatio…
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance o…
Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, bu…
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 appli…
36氪获悉,近日,AI潮玩品牌珞博智能(Robopoet)正式宣布完成亿元级Pre-A轮融资。本轮融资由华映资本与广和通联合领投,涂鸦智能跟投,同时老股东红杉中国、金沙江创投持续加码跟投。
AI 点评 · AI潮玩赛道受资本热捧,亿元融资显示市场对软硬件结合的新消费场景信心十足。
36氪获悉,7月13日,全球高效能AI Token生产服务商“趋境科技”(Approaching.AI)宣布完成A轮融资。半年内,公司累计融资金额超过10亿元。本轮融资由河南投资集团汇融基金领投,真知资本、尚势资本、星连资本、上海国方创新、弘晖基金、华控基金、杭州福成等老股东持续超额跟投。资金将主要用于扩大高品质AI Token产能储备、升级自研高效能AI…
AI 点评 · 半年融资10亿,资本重注AI Token赛道,折射出基础层服务商的稀缺价值。
作者|黄楠 编辑|袁斯来 硬氪获悉,威联机器人科技(深圳)有限公司(以下简称“MOVA LINCO”)近日完成数千万元天使融资。融资资金将主要用于AI算法底层技术研发、完善产品量产体系,以及全球化渠道布局和家庭AI生态的持续建设。 MOVA LINCO核心团队来自头部网络通信与智能硬件企业,在路由器、NAS存储及AI算力设备等方面拥有深厚的技术积累与产品经验…
AI 点评 · 切入AI家庭硬件赛道,首款产品瞄准海外市场,技术团队背景扎实,融资加速产品落地。
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods…
根据发行安排,7月16日(下周四),长鑫科技IPO将迎来网上申购,作为国产存储芯片巨头,5月27日IPO过会,6月12日注册生效,长鑫科技此次科创板上市,拟融资金额高达295亿元,将成为今年A股市场规模最大的IPO,也是科创板史上的第二大IPO,仅次于中芯国际的532亿元。长鑫科技网上发行申购上限为167.2万股,顶格申购需配沪市市值1672万元。Wind数…
AI 点评 · 295亿募资额冲击科创板,国产存储龙头上市将重塑芯片行业格局,值得关注。

Researchers cobbled together funding and time to show how quantum computing could aid in the development of drugs to help underserved populations and combat rare diseases.
AI 点评 · 科学家用业余时间结合AI与量子计算,为罕见病研发新药,展现技术公益潜力。
图片来源:视觉中国 作者/杨继云 报道/投资界PEdaily 深圳超级IPO来了。 日前,深圳云豹智能股份有限公司(简称“云豹智能”)创业板IPO申请正式获深交所受理,成为又一家选用创业板第四套上市标准申报上市的企业,冲刺“国产DPU第一股”。 回想6年前,斯坦福博士萧启阳在深圳创办云豹智能,专注于DPU赛道。这个团队启航两个月后,千亿DPU赛道开始爆发。至…
AI 点评 · 腾讯押注国产DPU第一股,体现其布局前沿芯片的战略眼光,赛道爆发前夜入场看点十足。
为了帮你看清具身数据行业,我们总结了以下十个行业现状
AI 点评 · 数据成AI新燃料,百亿资本涌入具身智能赛道,揭示产业变现路径。
7月11日,有消息称腾讯正在洽谈成为通用AI Agent公司Manus的最大股东,据该消息,由腾讯牵头的中方资本组团以约20亿美元估值从Meta手中回购Manus的全部股权。记者向腾讯方面求证,截至发稿腾讯方面暂无回应。另有知情人士向记者透露,此次交易后,腾讯仍将保持少数股东地位,但不会控股。(南方都市报)
AI 点评 · 腾讯罕见以少数股东身份入局AI Agent,或意在布局生态而非控制,战略意图值得玩味。
今日热点导览 “全球首款智能体手机”已备货8万至10万台?知情人士:假的 百亿私募数量达142家,再次刷新历史纪录 三星李在镕拟于7月底赴美会晤英伟达黄仁勋 德国大众拟大裁员,最高或裁减12万个岗位 OpenAI高管层再现变动,首席运营官因病离职 TOP3大新闻 长鑫科技,承销团阵容公布 长鑫科技IPO进入发行倒计时,这家“国产存储第一股”背后的承销团阵容也…
AI 点评 · 国产存储芯片龙头IPO加速,承销团阵容披露凸显市场关注热度。
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactio…
AI 点评 · 十年视觉语言模型进化揭示:精度提升中,视觉认知错误揭示了人类与AI感知的深层差异。
The AI chip boom just produced its biggest Wall Street moment yet. Now SK Hynix and Samsung are being asked to build U.S. factories.
AI 点评 · 韩国存储巨头创美股最大外资IPO,折射AI芯片需求爆发,美国正施压其本土建厂。
We present a physics-constrained machine learning framework for accelerating the direct numerical simulation (DNS) of turbulent reacting flows. The model replaces the direct evaluation of detailed che…
AI 点评 · 结合物理约束与残差数据增强,显著加速湍流燃烧模拟,突破传统计算瓶颈。

IT之家 7 月 10 日消息,韩国存储芯片巨头 SK 海力士今日正式在纳斯达克挂牌交易其美国存托凭证(ADR)。 此次发行定价为每份 149 美元,共计发行 1.779 亿份 ADR,每份 ADR 对应十分之一股 SK 海力士韩国普通股。 SK 海力士 ADR 于 7 月 10 日以“when-issued”(预发行)方式开始交易(预计 7 月 13 日转…
AI 点评 · 体现AI芯片需求火爆,存储巨头海外上市创纪录,首日大涨验证市场对算力产业链的强劲信心。
韩国芯片巨头SK海力士将于本周五登陆科技股云集的纳斯达克交易所,此举将为其打通一条强力全新融资渠道,方便全球投资者布局高速增长的人工智能基础设施赛道。这家全球第二大存储芯片厂商计划通过在纳斯达克发行美国存托凭证(ADR)募资约40万亿韩元(折合265亿美元)。本次发行共计1.779亿份美国存托凭证,发行价定为每份149美元,1份ADR对应其在韩国首尔交易所上…
大公司: SK海力士赴美上市,华尔街投行合计佣金有望达1.4亿美元 参与SK海力士上市项目的投行团队将斩获上亿美元丰厚佣金。这家市值突破万亿美元的韩国芯片巨头即将登陆美股,本次发行有望跻身史上规模最大IPO行列。两名熟悉该交易的人士透露,高盛、花旗等承销本次纳斯达克股票发行的投行,总佣金池规模或将突破1.4亿美元;佣金由两部分构成,一是募资总额0.5%的基础…
文 | 胡香赟 编辑 | 海若镜 36氪获悉,前传奇生物创始人/首席科学家范晓虎创办的生物科技公司深圳湾岛细胞,近日完成1.4亿人民币A轮融资,本轮融资由松禾资本领投、东方富海跟投。 过去一年,AI像一束强光照亮着资本市场。光芒之外,很多赛道在阴影中被重新定价,创新药板块正是其中之一。潮水褪去,人群对新药好药的需求仍在,创业者们也正在更严苛的估值坐标系里,渡…
AI推理芯片初创公司Positron正与投资者洽谈分两阶段融资约7.5亿美元,首轮预计估值35亿美元,第二轮估值约50亿美元。若顺利完成,这家仅成立三年的公司估值将在五个月内增长两倍以上。(新浪财经)
AI 点评 · AI推理芯片赛道持续升温,三年初创估值飙至50亿美元,凸显资本对算力基础设施的狂热追逐。
We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of pop…
Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these…

北京时间 7 月 9 日,据路透社报道,SpaceX 和其他 AI 公司的首次公开招股 (IPO) 创造的财富,推动了私人航空需求的新一轮增长。 造富效应刺激私人飞机需求 航空律师阿曼达 · 阿普尔盖特 (Amanda Applegate) 上个月取消了年度假期,因为 AI 创业公司和 SpaceX 创造的大量财富,引发了一波科技投资者购买私人飞机的热潮,使…
特斯拉FSD(Full Self-Driving)每迭代出一个新版本,自动驾驶公司Momenta的创始人曹旭东都会飞去美国亲自体验。 “他还不光是去坐车,还会找一切能接触到特斯拉的人聊。”接触到曹旭东的人士告诉36氪,马斯克是曹旭东最常提到的人。 在内部,许多员工对曹旭东的一个共同感受,就是“马斯克式CEO”。 创业十年,Momenta赴港股IPO,上市首日…
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments an…

SpaceXAI continues to move faster than any other frontier lab on earth.
36氪获悉,聚焦工业制造领域的具身智能机器人企业 「昇视唯盛」(3Srobotics)正式宣布完成数亿元B轮融资,本轮融资由上海半导体产投、金桥基金领投,零一创投、新鼎资本、中关鼎华及老股东微光创投持续加码跟投。 资金将主要用于焊接具身智能大脑与自主小脑运动控制系统的迭代升级,扩大自有智造工厂产能,扩充研发与市场服务团队,加速产品全国多行业规模化落地。 昇视…
图源/企业 本文约 2200 字,建议阅读 6 分钟 作者丨欧雪 编辑丨袁斯来 硬氪获悉,物理AI企业深度智控(DeepCtrls)近期完成数亿元人民币B轮融资。本轮融资由晶科能源战略投资,国投创新、招银国际联合领投,红杉中国、源码资本、光远资本、招商局创投等老股东持续跟投。 本轮资金将主要用于核心物理AI算法与产品线的持续研发、标准化产品DeepBot的规…

“IT早报”时间,大家好,现在是 2026 年 7 月 9 日星期四,今天的重要科技资讯有: 1、消息称苹果正为在华销售设备测试长鑫 DRAM 内存 苹果的这些举措旨在强化其供应链韧性,提升对存储半导体后市进一步变化的抵御能力。>> 查看详情 2、曝手机厂 Pocket 云台相机可能会上 3nm 骁龙 8 系芯片,给大疆和影石一点小小的机圈震撼 根据此前多方…
AI 点评 · 苹果测试国产DRAM,手机厂商跨界相机芯片,供应链博弈与跨界竞争双看点。
The $300 million round is expected to be led by Menlo Ventures, Sifted reported.
AI 点评 · AI初创估值飙升,Lovable融资显示资本持续押注生成式AI赛道,市场热度不减。

IT之家 7 月 9 日消息,SemiAnalysis 昨日发布的最新报告显示,Anthropic 今年第三季度的利润预计将超过 10 亿美元(IT之家注:现汇率约合 68.02 亿元人民币),并且公司已于今年 6 月 1 日秘密提交了首次公开募股(IPO)申请。若成功上市,Anthropic 将成为迄今规模最大的 AI 实验室 IPO。 该分析指出,Ant…
AI 点评 · 商业价值与上市节奏双重突破,标志AI行业从技术竞赛转向资本变现新阶段。
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute t…
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
AI chip maker SambaNova has raised at an $11 billion valuation months after Intel was rumored to be trying to buy it for about $1.6 billion.
图源/企业 作者丨欧雪 编辑丨袁斯来 硬氪获悉,上海量感智能科技有限公司(以下简称“量感智能”)近日已完成数千万元天使轮融资,由孚腾资本(上海国投旗下)领投,六禾创投跟投。 量感智能公司于2023年9月成立,由上海交通大学量子感知研究所成员孵化。公司聚焦量子传感方向,已推出的产品覆盖惯性导航、气体监测等市场。公司同步深耕量子 AI 技术研发与应用,推动相关算…
文|胡香赟 编辑|海若镜 36 氪获悉,德睿智药近期已完成 5200 万美元B轮融资,投资方包括头部人民币和美元基金,凯乘资本为独家财务顾问。募集资金将用于AI制药引擎Molecule Arts Platform(MAP)升级迭代,完善其多智能体(Multi-Agent)协同体系与临床数据闭环(Clinical Data-in-the-Loop),以及推进自…
Norm Ai周二表示,已在由Khosla Ventures领投的C轮融资中筹集了1.2亿美元,这家法律人工智能初创企业的估值达到12亿美元。(新浪财经)
AI 点评 · 法律AI赛道爆发,Norm Ai高估值融资印证行业从辅助工具向核心决策支持转型。
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed sol…
Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human lan…
Multi-hop retrieval-augmented generation (RAG) acquires evidence sequentially, with each new document potentially revealing missing facts, bridge entities, query defects, or sufficient support for ans…
The company just raised $7 million in seed funding, and is launching its app for iPhone and Android on Tuesday.
大公司: 网传“7月1日新规:新车智驾芯片自主化率不低于70%”系谣言 36氪获悉,近日,有网民发布题为《7月1日新规正式落地!新车智驾芯片自主化率不低于70%》的文章。文中称“7月1日,工信部联合市监总局出台的《新能源汽车车规级AI芯片技术规范》正式强制执行”,并称“国内生产、销售的所有新能源乘用车,智驾芯片全链路关键技术自主化率必须达到70%以上,不达标…
文 | 阿至 太空算力赛道又有新玩家入场。 36氪获悉,太空算力卫星系统解决方案提供商星际原点航天科技(上海)有限公司(以下简称“星际原点”) 已于近期连续完成种子轮、天使轮两轮融资,累计规模为数千万元。 种子轮投资方为九合创投。天使轮由老股东九合创投、梅花创投和上海国投旗下上海科创集团策源基金联合领投,上海天使会以联合投资形式参与。一苇资本担任独家财务顾问…
文|胡香赟 编辑|海若镜 36氪获悉,AI虚拟细胞(AIVC)企业华源智因近期已完成千万级人民币种子轮融资。本轮融资由水木创投领投,募集资金将主要用于多模态测序底层技术迭代,进一步拓展与头部三甲医院的合作,以及团队扩充等。此外,华源智因团队已计划启动新一轮融资。 华源智因创始团队由资深医药产业从业者、计算生物学研发人员组成,并邀请到深圳国家基因库等单位专家组…
SK Hynix is experiencing a boom credited to AI. It will ride that to a multibillion-dollar U.S. IPO, expected to take place on Friday.
AI 点评 · 美国投资者即将分享AI存储芯片巨头SK海力士的IPO红利,体现AI产业链资本化加速。
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with En…
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use thes…
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narro…
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enfo…
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave natural…
36氪获悉,生物基皮革公司「贻如科技」完成超亿元A轮融资,本轮融资由鄂尔多斯集团与和达金服共同领投,巢生资本跟投,易凯资本担任独家财务顾问。本轮资金将主要用于推动公司商业化开拓与产能建设,并依托AI技术加速生物基皮革的技术革新与产品迭代。 贻如科技成立于2021年,以合成生物学为底层技术,创造以生物基皮革为代表的新一代创新生物基材料 ——通过微生物发酵直接生…
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is…
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,具身智能公司 「光象科技」 宣布完成累计数亿元天使轮融资。 最新一轮由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、SeeFund、亿宸资本、上市公司行云科技等头部财投与产投深度参与,老股东零一创投、L2F光源创业者基金持续加注。 本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人在工…
AI 点评 · 清华系创业团队获数亿元融资,聚焦具身智能落地汽车产业,技术底蕴与产业场景结合是关键看点。
作者 | 邱晓芬 编辑 | 袁斯来 硬氪获悉,通用餐饮具身机器人公司「影智XBOT」连续完成数亿元两轮融资——其中,A轮的2亿元融资由香港简坤资本GPTX出资,B轮融资为3-5 亿元人民币,由多支政府基金、美元基金和产业投资方共同参与出资。 这是目前餐饮垂直机器人领域规模最大的一笔融资之一。 在此之前,「影智XBOT」还完成了一轮天使融资,出资人阵容豪华——…
AI 点评 · 小米前高管创业项目获数亿融资,林斌、黎万强加持,餐饮机器人赛道热度可见一斑。
Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, ind…
Mistral AI, which offers some open source AI models, has raised significant funding since its creation in 2023, with the ambition to “put frontier AI in the hands of everyone.”
AI 点评 · Mistral AI以开源策略和巨额融资挑战OpenAI,展现欧洲AI新势力崛起。
近日,光象科技宣布完成累计数亿元天使轮融资,最新一轮融资由珠海科技产业集团、兴证资本、松禾资本、顺禧基金、慕华科创、See Fund、亿宸资本、上市公司行云科技等头部财投与头部产投深度参与,老股东零一创投、L2F光源创业者基金等持续加注。据悉,本轮资金将重点投入物理原生基座模型的研发迭代,并推进具身智能机器人产品的商业化交付。(每日经济新闻)
AI 点评 · 资本密集押注物理原生模型,具身智能从实验室走向量产的关键信号。
AI 点评 · 物理原生模型获资本重注,预示AI向真实世界感知与交互迈出关键一步。
图源/企业 作者丨欧雪 编辑丨袁斯来 硬氪获悉,硅基光电子集成芯片研发商光引科技近期已完成1亿元Pre-A轮融资,投资方包括光子强链基金、善达投资、长飞基金、洛阳英才、中科创星、西安财金。资金将主要用于上海新实验室建设、人才招募及量产推进。 光引科技2021年成立于徐州,核心团队孵化自英国剑桥大学研发团队。创始人程祺翔为剑桥大学博士、剑桥大学副教授,在光子集…
AI 点评 · 剑桥副教授创业获欧莱雅华为合作,技术实力与市场认可双重背书。
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing re…
本文约2700字,建议阅读6分钟 作者 | 彭孝秋 编者按: AI大爆发之际,越来越多公司走向资本市场。每一份招股书翻动的声音里,都藏着一家公司想说与未曾明说的全部。 鉴于此,硬氪特推出「秋声」专栏。秋声取自欧阳修《秋声赋》,借“听秋声”之意,产业冷暖,辨公司成色,记录企业冲刺IPO途中那些被写下与被隐藏的真实。这是我们第七期,硅基流动。 Q2的最后一天,硅…
今日热点导览 苹果拟于今明两年推出至少五款新iPhone AI版支付宝开放公测,上线72项办事技能 特朗普回应“利用职位牟利”:股市在涨,大家都在赚钱 LV起诉茉莉奶白,茉莉奶白被判赔1030万元 6月赴日航班取消1488个 安克创新港股上市首日破发 TOP3大新闻 证监会同意宇树科技科创板IPO注册 7月2日,证监会网站显示,同意宇树科技首次公开发行股票注…
据报道,与Meta Platforms和甲骨文等公司签订AI算力供应合同的数据中心初创公司Crusoe,正洽谈筹集约30亿美元资金,这可能会使该公司的估值提高两倍。知情人士称,Crusoe仍在积极洽谈此轮融资,最终估值尚未确定。投资者预计估值将达到300亿美元左右,这其中包括新投资。该公司此前在去年10月的估值为约100亿美元。(新浪财经)
AI 点评 · 数据中心融资激增印证AI算力需求爆发,Crusoe估值翻倍凸显行业狂热与巨头争相押注。
Just for kicks, I took a look at Jersey Mike's IPO documents. Surely a sandwich shop would have no need to mention AI. But lo and behold.
AI 点评 · AI炒作过度蔓延,连三明治店IPO都蹭热度,暴露行业泡沫风险。
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech…

In this post, we share best practices for reliable multi-turn RL training. We cover how to build a training environment you can trust, set up an external evaluation, design a rewar…
AI 点评 · 强化学习多轮训练实操指南,亚马逊云服务实战经验值得借鉴。
大公司: 奥飞娱乐:上半年净利润同比预增305.30%—413.38% 36氪获悉,奥飞娱乐发布2026年半年度业绩预告,预计2026年上半年归属于上市公司股东的净利润为1.5亿元—1.9亿元,比上年同期增长305.30%—413.38%。 雀巢称消费者因价格压力避开中等包装产品 雀巢表示,在美国消费者专注于产品可负担性、购买者纷纷转向大包装或经济型包装而避…
本文约3300字,建议阅读7分钟 作者 | 彭孝秋 编者按: AI大爆发之际,越来越多公司走向资本市场。每一份招股书翻动的声音里,都藏着一家公司想说与未曾明说的全部。 鉴于此,硬氪特推出「秋声」专栏。秋声取自欧阳修《秋声赋》,借“听秋声”之意,产业冷暖,辨公司成色,记录企业冲刺IPO途中那些被写下与被隐藏的真实。这是我们第六期,古瑞瓦特。 上个月,古瑞瓦特第…
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,深圳可立点科技有限公司(以下简称“可立点科技”)近日完成战略融资,由力合科创领投,江苏中科智能科学技术应用研究院旗下平台跟投。本轮融资将主要用于产品研发迭代、核心团队建设及商业化落地推进。 可立点科技总部位于深圳, 是一家聚焦“AI+机器人”养老场景的科技公司,围绕银发群体布局家庭陪伴与院内康复两大产品线。 目前…

IT之家 7 月 2 日消息,华尔街日报昨日(7 月 1 日)发布博文,报道称马斯克旗下 SpaceX 近期面向其投资者,展示了一款 AI 原型机,外观类似手机, 设计时尚、厚度要比苹果 iPhone 更薄。 SpaceX 目前正在推进大型首次公开募股(IPO)上市,最近向部分投资者及其他利益相关方展示了这款原型机。 IT之家查询原文报道,以及社交平台相关爆…
AI 点评 · 跨界造机引关注,AI硬件比iPhone更薄,马斯克生态布局新信号。
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended soft…
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to co…
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with on-policy…
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on genera…
We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitting between a policy server and the physical hardwa…
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world ro…
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While ma…
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-g…
文|周鑫雨 编辑|张雨忻 《长安的荔枝》,是 97 年清华博导李一鸣很喜欢的故事。 故事里,为了将“一日色变”的鲜荔枝从岭南运到长安,小吏李善德必须解决保鲜、驿站、路线、补给等一系列环环相扣的难题——没有这套完整系统,鲜荔枝寸步难行。 这个设定在唐朝的故事,在李一鸣眼中,却与当下的“世界模型”赛道,形成了巧妙的互文: Physical AI(物理AI)的场景…
硬氪获悉,硅基光子企业「灵动芯光」正式宣布完成数千万元天使++轮融资,本轮融资由磐霖资本领投、同方投资与深天使合作的子基金汇泽天诚跟投。 此次募集的资金将全部聚焦于芯片间光互联核心技术研发与产品落地,重点推进 SmartComb多波长密波光源的产品化及SmartPHY光I/O小芯粒的研发工作。 灵动芯光成立于2022年,是一家致力于硅基光子集成芯片设计与应用…
硬氪获悉,AI抗衰药物研发公司「无尽方舟」已完成数千万元种子轮融资,Monolith领投,九合创投跟投。此次融资将主要用于加速核心药物的工程化生产与猫狗临床验证,并同时推进H2P跨物种平台及干湿结合药物研发平台两大AI系统的建设。 无尽方舟成立于2026年5月,切入的是一个在长寿产业中足够前沿、但验证路径更短的方向:先在犬、猫等宠物身上建立衰老干预药效模型,…
今年夏天,作为一家基金的投资人,你难保不会接到老板的狠话:2026,必须下注一家核聚变企业! 可惜他们手里的TS,大概率递不出去。 一个2025年刚成立的核聚变企业,起初估值5亿,关完一轮涨到30亿,几个月之后可以再涨两三倍。 公司创始人是业界大佬,普通投资人根本见不到本人,最多和其他机构一起拼桌见见CFO。这位CFO私下告诉硬氪,TA入行时觉得3年能做到1…
硬氪获悉,算力基础设施企业华弘数科近日完成数千万Pre-A轮融资,领投方为吴中金控,此次融资将主要用来发展AI应用和算力落地的生态建设、工厂产能的进一步扩大和基地的建设与推进。 华弘数科成立于2021年,总部位于北京,后迁至苏州,专注全液冷边端侧超级计算机研发,是跨界全液冷散热与IT算力的一体化平台。 其创始人罗华曾任同有科技(300302)董事兼副总经理,…
Wayve’s offering is part of a growing trend of AI startups using employee tenders as a strategic tool to attract and retain talent.
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-g…
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduc…
Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks catastrophic diagnostic…
为 A股投资者打造的全球产业链资讯看板 · 12 大赛道一一对应 A股板块(半导体/AI/机器人/新能源车…),覆盖 100+ 权威源,用你自己的大模型每日提炼中文「今日要点」+ 翻译 · 全程本地、零 API key · Local AI news dashboard tracking the global indu…
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from holistic semantic mat…
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382…
Discrete-diffusion protein language model with ESM-2 conditioning: implementation and honest evaluation.
Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such evaluations leave open whether an artificial agent can acquire, stabilize, and use ne…
Evidence-grounded evaluation for AI agents — verifies each claim against the agent's real tool outputs (constrained, evidence-grounded model judgment, not holis…
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmfu…
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability.…

The complex factors that determine the single evaluation number so many focus on. Plus, how this changes in the future.

Meta’s shocking purchase of 49% of Scale AI at a ~$30B valuation shows that money is of no concern for the $100B annual cashflow ad machine. Despite seemingly unlimited resources,…