Hey HN! This is Divit from Almanac (YC S26). We built CodeAlmanac, a wiki for your coding agents that updates as you talk to them. It is open-source, local, and free. Here’s a demo…
AI 编程
共 287 条相关资讯 · 来自历史归档

A new type of malware can worm deep into AI coding systems to steal data and logins—and can flip a “death switch” to destroy files and keep out real users.
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…

Augment Code's Vinay Perneti talks models, harnesses, and context.
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code…
Databricks has remade its image into an AI company and has published research on the cost savings of open-weight AI models for coding.
AI 点评 · 估值188亿美元,从数据平台成功转型AI公司,开源模型成本优势研究或重塑行业格局。
36氪获悉,7月17日,在2026世界人工智能大会(WAIC 2026)上,网易智企携全新升级的一站式企业AI应用服务亮相,集中展示AI Agent编排、AI Coding、AI客服、AI私域助理、AI智能数据与AI Agent安全等企业级AI能力,围绕安全治理、组织协作与业务增长三大场景,呈现企业级AI应用实践。
AI 点评 · 展示企业级AI全栈能力,聚焦安全治理与业务增长三大场景,为行业提供可落地的AI应用标杆。
模型、Harness Engineering与产品的持续进化
AI 点评 · AI编码市场格局剧变,阿里Qoder以绝对优势领先,凸显大模型工程化落地的关键转折点。
文 | 周鑫雨 编辑 | 张雨忻 智能涌现从多个独立信源处获悉, 截至 2026 年 7 月, 智谱的 ARR(年度经常性收入)已经达到 10 亿美元 。 截至发稿前,针对上述信息,智谱未回复。 过去一年,AI Coding 和视频生成模型已经成为全球造血能力最强的 AI 赛道。 海外,Anthropic 的 Claude Code 仅发布半年,ARR 就飙…
AI 点评 · 智谱ARR半年飙升15倍达10亿美元,印证AI商业化进入爆发期,行业格局加速重塑。

IT之家 7 月 17 日消息,Google(谷歌)当地时间 16 日宣布,其 AI 笔记助理 NotebookLM 更名为 Gemini Notebook,以实现人工智能产品的品牌统一。 谷歌表示,NotebookLM 仍然是一款独立的产品,但 现在将在包括 Gemini 应用和 Google 搜索的整个 Google 生态系统中发挥更大的作用 。 Gem…
AI 点评 · 品牌统一战略升级,原生代码支持强化AI笔记工具实用性,值得开发者关注。

Creator says he will "very loudly ignore" those arguing for a ban on AI tools.
AI 点评 · Linus强硬表态,AI编码争议升级,Linux社区自由与创新的核心原则面临考验。
AI 点评 · 多智能体协作编程,突破单Agent局限,是AI自主开发的关键一步。
AI 点评 · AI原生开发全流程革新,企业人机协同研发范式从理论走向实践。
Tool: Mermaid to Unicode box art (grok-mermaid) While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs , a "se…
AI 点评 · 将Mermaid图表转为Unicode字符画,适合终端环境,实用且有趣。

IT之家 7 月 16 日消息,马斯克旗下 SpaceXAI 公司昨日(7 月 15 日)宣布开源 Grok Build, 并将源代码发布至 GitHub 平台。 在官方博文中,SpaceXAI 表示: 开源发布源代码,是构建强大、可靠框架的最直接方法。用户可以阅读源代码,了解其从上下文构建到工具调用分发的完整工作原理。 开源也让框架更容易探索和扩展:如果用…
AI 点评 · 开源代码降低使用门槛,推动编程AI智能体技术快速迭代,值得开发者关注。
OpenAI, which is in the middle of a legal battle with Apple over hardware trade theft allegations, just released a light-up keyboard designed to be paired with its agentic coding a…
AI 点评 · 硬件纠纷未平却推高价键盘,OpenAI跨界硬件野心与Codex生态绑定值得关注。
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While p…
The startup has reached a $120 million annualized revenue run rate and more than 200,000 paying customers.
SpaceXAI's Grok Build AI coding tool was spotted uploading users' entire codebases to Google Cloud before it was reported, and the company turned it off. The Register reports that…
AI 点评 · 隐私漏洞暴露AI编程工具安全隐患,用户数据保护机制亟待加强。

MIT students designed, built, and tested a jet engine with AI copilots, assessing AI’s usefulness in developing high-performance aerospace systems.
AI 点评 · AI在尖端工程中从辅助走向协作,验证了设计复杂系统的潜力。
一手实测这就奉上
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the acti…
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functio…
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods…
Hello HN, I don't post on here much, but wanted to get some eyes on a new project I'm just launching. I think we definitely need one more AI code agent.. I'm a long-term C++ dev, a…

IT之家 7 月 12 日消息,Meta 于 7 月 9 日正式发布适用于 AI 智能体的多模态推理模型 Muse Spark 1.1 版本,重点提升了模型在智能体任务中的规划、协同与执行能力,并增强了工具调用、代码开发、应用操作能力。 Meta 表示,Muse Spark 1.1 强化了多智能体协作机制,由主智能体负责收集信息、制定计划,再将任务拆分并分配…
AI 点评 · 多智能体协作机制是AI落地的关键突破,Meta这次强化了任务拆解与分工能力。
AI 点评 · 3D可视化编码过程,直观追踪AI代理如何理解代码库,革新调试与协作方式。
换更大的模型就等于更聪明? 【导读】 换更大的模型就等于更聪明?这可能是Claude Code用户最深的误会。很多人为此一路换到最贵的Fable,近日,Anthropic,亲手澄清了这个误区。 你有没有过这种时刻:Claude Code写代码写砸了,第一反应,就是赶紧换个更强的模型。 但这一招,很多时候并不管用,甚至是在白花钱。 近日,Anthropic官方…
AI 点评 · 揭示模型性能瓶颈不在参数规模,而在于使用方式,颠覆用户认知。

IT之家 7 月 11 日消息,月之暗面官方昨晚宣布, K2.7 Code 高速版结束 Beta,成为常驻可选模式 。 订阅 Allegretto 及以上会员计划的用户,在 Kimi Code CLI 或其他工具中使用 Coding Plan 进行编程时,无需任何申请,可以直接调用 K2.7 Code HighSpeed。 官方表示,K2.7 Code Hi…
AI 点评 · 编程效率再升级,K2.7 Code高速版免费开放,开发者可无缝调用。

Amid live coding sessions and Silicon Valley optimism, the UN’s AI for Good summit wrestled with an urgent question: Can global governance catch up before the technology races beyo…
AI 点评 · 联合国AI峰会展示前沿科技,核心挑战在于全球治理能否跟上技术飞速发展。
OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.
AI 点评 · OpenAI与微软合作稳固,GPT-5.6成Copilot 365核心,预示AI办公生态新阶段。

“IT早报”时间,大家好,现在是 2026 年 7 月 10 日星期五,今天的重要科技资讯有: 1、OpenAI 最强 AI 模型:GPT-5.6 系列正式上线,纳德拉称微软 Copilot 同步接入 OpenAI 公司 7 月 10 日发布公告,宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及…
AI 点评 · OpenAI模型重大升级,微软深度整合,预示AI竞争进入新阶段。

IT之家 7 月 10 日消息,OpenAI 公司今天(7 月 10 日)发布公告, 宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及 API 中上线 GPT-5.6 系列模型。 在模型方面,IT之家援引博文介绍,OpenAI 本次共发布 3 档模型: 旗舰版 Sol(太阳):每 100 万 T…
AI 点评 · 微软同步接入,意味着AI竞争格局突变,企业级应用迎来新拐点。
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations — the kind of automation that enterprises are increasingly…
AI 点评 · Meta携Spark 1.1切入企业级AI编码自动化,专注大型代码迁移和修复,展现巨头竞争新方向。
After reentering the AI race with its first in-house Muse Spark model in April, Meta is now opening up the doors to developers with a new model that can plug into AI coding softwar…
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
AI 点评 · 微软将GPT-5.6深度集成办公套件,标志AI办公进入新阶段,用户体验将显著提升。
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments an…

IT之家 7 月 9 日消息,SpaceXAI 今日正式发布了其 Grok 4.5 模型,这是该公司首个专门针对编程和智能体任务训练的模型。 据介绍,该模型由 SpaceXAI 与 Cursor 联合完成训练,在提供前沿智能水平的同时,兼具领先的速度与成本效率。马斯克将其称为“Opus 级模型”。 Grok 4.5 面向真实工程场景设计,擅长处理大型代码库以…
AI 点评 · 编程智能体成本砍半效率翻倍,马斯克联手Cursor的定价策略才真值得行业关注。
AI 点评 · AI时代软件工程师的门槛被重新定义,挑战传统开发思维。
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces ro…
AI 点评 · AI自动修复漏洞,提升开发效率与安全性的关键一步。
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revisi…
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use thes…
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-f…
AI 点评 · AI冲击初级程序员岗位,暴露行业技能结构转型危机。
AI 点评 · 孤岛编码实验揭示AI自主编程的进化路径。

IT之家 7 月 4 日消息,科技媒体 AppleInsider 昨日(7 月 3 日)发布博文,报道称在 iOS 27 开发者测试版中, 发现相关文字描述,指向具备摄像头功能的苹果 AirPods 耳机产品。 程序员 Sam Henri Gold 昨日在 X 平台发布推文,指出在 iOS 27 开发者测试版中,发现了代号为 B790 的智能眼镜。IT之家附…
AI 点评 · 苹果在耳机上集成摄像头的尝试,或为空间计算与AI视觉交互开辟新入口。
AI 点评 · Agent正突破编程边界,重塑各行各业工作流程,预示AI自主执行任务的未来。
AI 点评 · AI编程工具成瘾性暴露工程师效率与代码质量间的深层矛盾。
Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would lo…
AI 点评 · 首个开源LLM编程代理框架,简化代码生成与迭代流程,开发者可快速上手实验。
AI 点评 · 用短链约束AI编程,突破游戏创作瓶颈,展现人机协作新思路。
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt…
AI 点评 · 用微虚拟机隔离智能体,提升了云计算安全性与灵活性。
Coding agents don't have long-term memory. But you do have months of full-fidelity agent transcripts stored on your machine. A simple solution that goes a long way: ingest those tr…
今日热点导览 OpenAI据悉迎来重大技术突破:系统优化使模型推理成本减半 苹果首次上架iPhone16e官翻机,仅有黑色与白色可选 国内航线燃油附加费7月5日起大幅下调 巴菲特“断供”盖茨基金会 董明珠喊话股东:家电不换成格力,凭什么要分红 詹姆斯发文告别湖人 TOP3大新闻 二手豪华燃油车价格大跳水:宾利近27万、保时捷15万 7月1日,在青岛市的陆海汽…

“IT早报”时间,大家好,现在是 2026 年 7 月 2 日星期四,今天的重要科技资讯有: 1、美光 CEO 发声“甩锅”:内存供应失衡别只怪我们,客户 2023 年曾压价到三分之一导致 AI 爆发前没能扩产 美光 CEO 梅赫罗特拉表示,当前内存芯片供应失衡不应全怪芯片商,部分客户曾将价格压至三分之一,导致行业在 AI 需求爆发前投资不足。他警告短缺可能…
AI 点评 · 内存价格博弈、小红书自我复制、AI合规争议,多条行业动态折射科技产业链关键变局。
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to co…
AI 点评 · AI Coding颠覆传统开发模式,预示大厂工程师角色将根本性重塑。
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current R…
OpenSquilla 上线后数周内 GitHub star 增至数千量级
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized b…
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iterat…
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
Wix-owned vibe-coding platform Base44 has started rolling out its own AI model — with hopes that it will eventually outperform frontier models.
OpenAI is releasing some sort of device related to its AI-powered coding tool, Codex, on July 15th. In a video posted to X on Monday, OpenAI shows a square-shaped device with sever…
AI 点评 · 聚焦AI编程工具Codex的硬件化,预示OpenAI正从软件向实体设备延伸,可能开启AI开发新范式。
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete re…
AI 点评 · 开源模型自我进化,开启智能编程新范式,降低AI开发门槛。
Cursor has launched a new mobile app for remote oversight over coding agents.
AI 点评 · 移动端管理编码代理,打破办公空间限制,提升开发灵活性。
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding This is an interesting new open weights (MIT licensed) model, the first model release from DeepReinforce. [...] with variants i…
AI 点评 · 首个自搭建代码智能体模型开源,MIT许可降低门槛,或重塑AI编程工具生态。
Open-source local policy, recovery, and audit layer for explicit Codex execution.
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals…
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete re…
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions
AI 点评 · 本地化开源模型替代付费服务,降低AI编码成本与依赖。
AI 点评 · 微软将AI从辅助工具升级为自主执行体,标志着企业级智能体进入常驻操作新阶段。
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo…
AI 点评 · 让编程助手自动匹配最优模型,大幅提升代码生成效率与成本控制。
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo…
Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates them the way it has always evaluated components, one agent at a time, on isolated ben…
OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stack.
Coding为王
精准识别设计系统
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding. H…
Program verifiers play a central role in training coding agents, including selecting trajectories for supervised fine-tuning (SFT) and providing rewards for reinforcement learning (RL). Standard execu…
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the r…
Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet accurate whole-matrix forecasting remains challenging when prediction must be performe…
AI 点评 · 打通AI编程工具与桌面环境,并行Agent工作流将大幅提升开发效率。

IT之家 6 月 25 日消息,据 The information 今日消息,知情人士透露,谷歌正对其不久前成立、 主攻 AI 编程工具的专项攻坚小组进行重组 ,试图缩小与 Anthropic 之间的技术差距。 知情人士表示,这支成立仅数月的攻坚小组将调整谷歌 AI 模型的训练思路, 既要提升模型代码能力,也要强化其生成演示文稿等其他场景的能力 。 本次调整…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以 ‘今年盛夏’ 为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
Figma has revealed some new design and coding product updates at its annual Config conference that aim to help creatives "push their ideas further" and automate tedious tasks with…
AI 点评 · Figma引入AI动效与着色器工具,降低创意门槛,让设计师快速实现高级视觉特效。
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
目前A社约65%的产品代码已经由Claude Tag参与完成
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从‘小众赛道’ 变成‘主流共识’,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以‘今年盛夏’为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
Open-source, self-hostable alternative to Claude Tag — a Slack-style workspace where your team and its AI agents (Claude Code, Codex, GitHub Copilot, and more)…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…

IT之家 6 月 24 日消息,科技媒体 AppleInsider 昨日(6 月 23 日)发布博文,报道称基于 tvOS 27 Beta 首个开发者测试版本挖掘的代码, 苹果公司正酝酿为 Apple TV 和 HomePod 引入 AI 功能。 在 2026 年全球开发者大会(WWDC)的主题演讲中,Apple TV 与 HomePod 两款产品的篇幅不多…
AI 点评 · 苹果AI布局拓展至客厅与智能音箱,预示全场景智能生态加速成形。
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabil…
AI 点评 · 隔离执行环境保障AI智能体安全,降低代码运行风险,推动企业级应用部署。
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have emerged, yet are sti…

This post shows you how to build a conversational protein research assistant that combines three capabilities: Natural language query parsing to extract structured search parameter…
AI 点评 · 用自然语言查询蛋白质数据库,AI助手让科研效率飞跃。
AI 点评 · Angular官方技能让AI生成更规范代码,开发者效率与质量双提升。

IT之家 6 月 23 日消息,据外媒 Cybernews 今天报道,开源多媒体框架 FFmpeg 最近被曝出严重安全漏洞,黑客只需要利用恶意视频文件,就能导致系统崩溃并失去控制,进而远程执行恶意代码。 据报道,这组漏洞目前已被编号为 CVE-2026-8461,评分 8.8/10。黑客可利用 MagicYUV 中的“堆缓冲区越界写入(heap out-of…
编程比肩Opus 4.7
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction t…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
36氪获悉,中信建投研报称,国内模型持续迭代,GLM-5.2、Kimi K2.7 Code强化1M上下文、长程Agent、Agentic Coding和真实工程交付能力,推动国产模型从通用问答转向开发者工具和企业级工作流。Kimi补强国际化运营能力,DeepSeek融资强化头部模型产业化预期,微信AI灰度测试则显示AI入口正从独立App走向超级应用生态,有望…
AI 点评 · 国产模型转向企业级应用,算力需求确定性增强,产业链景气度有望持续。

IT之家 6 月 22 日消息,据 Windows Latest 报道,尽管公众强烈反对,且微软此前在强制预装 Microsoft 365 Copilot 一事上看似流露过些许歉意,但如今这家企业又故态复萌。 根据用户设备上安装的 Office 版本不同,即便从未主动使用,Microsoft 365 Copilot 相关组件也正悄悄重新侵入用户的办公套件。…
AI 点评 · 强制预装AI工具,微软无视用户意愿,反映巨头对用户自主权的侵蚀。
What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and other ``agentic" tools that promise to drive up…
Agent 和人一样离不开闭环。 查看全文
AI 点评 · 以Vibe Coding实现全流程AI开发,验证了Agent闭环工作的可行性,门槛极低值得关注。
Local coding-agent orchestrator — DAG of auto-approved, git-worktree-isolated sub-sessions across LLM providers (Claude/Kimi/Grok/DeepSeek/local). AGPL-3.0.
AI 点评 · Max对比GLM编码能力,展现国产大模型在自主编程任务上最新突破。
UmaDev: A coding agent that works like a real dev team, commanding the Claude Code / Codex / OpenCode you already use.
UmaDev: A coding agent that works like a real dev team, commanding the Claude Code / Codex / OpenCode you already use.
Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.
Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser…
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser…
LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong fixes)…
文|肖漫 编辑|李勤 6月18日,蔚来同时向两代平台车型(包含8款NT2.0平台车型、4款NT2.5平台车型,以及6款NT3.0车型)推送了最新版的世界模型,这意味着,蔚来现在能让同一套复杂的智驾代码,现在能跑在不同代际的芯片上。 软件迭代节奏被硬件绑架曾是一个困扰行业的难题。很多车企无法在不同版本、配置的车型上迭代同一款软件,这带来的结果是,很长时间内只有…
已接入华为鸿蒙生态
Chinese AI lab Z.ai released GLM-5.2 to their coding plan subscribers on June 13th, and then yesterday (June 16th) released the full open weights under an MIT license. Similar in s…
Current AI-driven game development has made substantial progress in asset generation, gameplay design, and web-based game coding, yet project-level code engineering on professional game engines remain…

Nvidia's self-improvement program for robots enlists teams of AI coding agents.
AI 点评 · 英伟达用AI编程智能体教机器人装显卡,展现了自进化系统的潜力,是机器人自主学习的突破。
Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present…

Separately, neither could compete. Now they hope they can.
I can 100% attest to the fact that Qwen3.6-27B is a very capable local model for coding tasks. Over the last month and a half I've been using it almost daily, either on my M2 Ultra…

SearchLeak exploit shows why the industry's approach to LLM security fails over and over.
https://xcancel.com/ID_AA_Carmack/status/2064095424420487226

IT之家 6 月 16 日消息,苹果公司昨日(6 月 15 日)更新支持文档,解释称在 macOS 26.4 系统中, 若用户不常用终端(Terminal),且命令来自网站、聊天智能体、消息或邮件应用,系统可能阻止粘贴。 科技媒体 9to5Mac 指出,用户此前终端里粘贴命令后,系统会先给出安全警告,提示内容可能含有恶意代码。很多人只知道有这个拦截机制,但不…
AI 点评 · 安全机制升级针对非高频操作,防范恶意代码粘贴执行,体现系统防护精细化。
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game gene…
Has anyone here fully swapped Claude/GPT for a local model as their main coding tool, not just for side experiments? If so, please share your setup and performance (e.g tok/s)
6月14日,HarmonyOS创新赛·极客赛道决赛在HDC华为开发者大会(HDC 2026)落下帷幕。 此前,20支入围队伍从众多参赛作品中脱颖而出,在36小时极限编码中抢先使用HarmonyOS 7,围绕AI、3D空间化、用户体验等方向展开比拼。随着最终获奖名单揭晓,比赛的结果已经落定,但比奖项本身更值得被看见的,是 这些作品背后一批年轻开发者正在关注什么…
Large Language Models (LLMs) have significantly advanced the automation of software engineering tasks. One prominent example is code generation, where an LLM produces code in a specified programming l…
AI 点评 · 开源AI编码工具让家庭开发者低成本入门,打破技术垄断。

IT之家 6 月 12 日消息,前天,字节跳动旗下 AI 应用豆包已针对付费订阅开放灰度测试,除此之外,当时还有另一个“任务模式”也开启测试。 直到今天,豆包“任务模式”已经大范围上线。IT之家实测发现,大部分用户打开豆包首页就能看到官方提示。 用户打开豆包 App 后,界面顶部的模式切换选项也已经变更为“快速、专家、任务”,标志着这款 AI 助手的产品功能…
AI 点评 · AI助手从被动问答向主动执行进化,功能集成度提升是关键看点。
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Yet most agents consume repositories almost entirely as text, which differs from how…
Large Language Model (LLM) coding agents have achieved strong results on software engineering tasks, yet repository exploration remains a major bottleneck: locating relevant code consumes substantial…
AI 点评 · 企业级AI正突破辅助编码,转向重塑研发流程与组织架构。
Recursive language models (RLMs) showed that recursion over model calls is an effective strategy for long-context reasoning, and production coding agents have begun to write code that spawns subagents…
AI 点评 · 建立反馈闭环与评测基准,是让AI编码智能体持续进化的核心驱动力。

Agent-EvalKit is an open-source toolkit (Apache 2.0) that makes this evaluation infrastructure available by integrating with AI coding assistants, including Claude Code, Kiro CLI,…
AI 点评 · 开源工具填补AI智能体评估基础设施空白,让开发者能系统化测试性能。
Hi HN, I'm Djoumé. I've been a developer for over 20 years, and like a lot of you I've been coding almost exclusively through an agent in the past few months. It's been amazing to…
通用GPU就能实现

IT之家 6 月 11 日消息,据央视新闻今日报道,国际汽车开放系统架构组织昨天(10 日)在上海宣布,中国自主研发的智能驾驶操作系统,正式成为全球智能驾驶系统公共代码库的核心基线,这也是中国汽车基础软件技术 首次进入全球行业标准 。 IT之家从报道获悉,中国打造的智能驾驶操作系统,就像是智能汽车的“软件地基”,它可以统筹调配车辆芯片算力,保障各类智能驾驶功…
AI 点评 · 中国标准首次跻身全球汽车软件核心基线,标志技术话语权实现关键突破。

IT之家 6 月 11 日消息,Anthropic 昨日推出了旗下首款 Mythos 级人工智能模型 Claude Fable,这款模型已然在微软内部引发担忧。据 The Verge 报道,消息人士透露,由于 Anthropic 出台了新的数据留存规定,微软已限制员工使用 Claude Fable 5。 微软虽迅速为 GitHub Copilot 和 Fou…
AI 点评 · 数据新规引发巨头博弈,微软限制员工使用竞品AI,凸显企业数据主权与AI工具合规性的敏感边界。
Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated in the next. We stud…
Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcripts included.
AI coding agent startup Niteshift has raised a $7 million seed round from a who's who of angels. It's betting companies will want power over, not lock-in with model makers.
毕马威与微软6月9日宣布扩展全球战略合作关系,聚焦企业级AI智能体的规模化部署。根据协议,微软365 Copilot将向毕马威全球逾27.6万名专业人员全面推广;与此同时,毕马威将采用微软Agent 365平台,对其全球组织内及客户端的AI智能体实施统一管理、监控与安全治理。(界面)
AI 点评 · 毕马威27万员工全面接入Copilot,并首创Agent治理框架,预示企业级AI从工具应用迈入系统化
一个更强的模型上桌了
AI 点评 · 性能飞跃式提升,刷新AI编程效率上限,标志行业进入新量级。
TIL: Setting a custom price for a model in AgentsView I've been really enjoying AgentsView by Wes McKinney as a tool for exploring my token usage across different coding agents run…
AI 点评 · 自定义AI模型定价功能,让用户按需付费,提升灵活性与成本控制。
Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile, Grammar-Constrained Decoding (GCD) has been wide…
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the…
Most of Apple's current AI ideas are roughly the same as everyone else's AI ideas. A chatbot you can ask questions; quick ways to create or summarize text; bizarre, borderline cree…
AI 点评 · 用工程闭环打通AI编码到验收,展现企业级落地的真实路径与关键挑战。
Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani…

IT之家 6 月 9 日消息,微软已封锁其托管在 GitHub 平台上的数十个开源项目访问权限。此前有黑客疑似入侵了这些项目,并在代码中植入窃取密码的恶意程序,微软目前正就此展开调查。 此次受影响的项目大多与微软云服务 Azure,以及开发者用于借助人工智能开发应用程序的其他工具相关,例如 Claude Code、Gemini 的命令行界面和 VS Code…
AI 点评 · 开源生态信任危机,黑客直击代码供应链,微软Azure成攻击跳板。

IT之家 6 月 9 日消息,据《连线》杂志报道,就在有人在 Meta 智能眼镜配套应用中发现一段疑似人脸识别算法的休眠代码仅一天后,Meta 便推送更新移除了这段代码。该媒体在核查一款负责智能眼镜部分核心功能的 Meta 人工智能应用代码时,首次发现了这段可疑代码。Meta 内部将其命名为“姓名标签”。简单来说,这款用于通过蓝牙将智能眼镜与手机配对的必备应…
AI 点评 · 隐私与技术的碰撞,Meta快速响应暴露了AI眼镜潜在的隐私风险。
Hi HN! We’re Jimmy and Ray. Jimmy is a Thiel Fellow with a Ph. D. from MIT who has worked on programming tools for 15 years; Ray became VP of Sales at a $2B company when he was 19…
AI 点评 · 聚焦AI编程工具质量,MIT博士与销售奇才联手,技术实力与商业经验兼备。
Multi-agent code generation offers a promising paradigm for autonomous software development by simulating the human software engineering lifecycle. However, system reliability remains hindered by LLM…
AI 点评 · 提出自适应语义熵指标,精准衡量代码质量,突破多智能体编程可靠性瓶颈。
Advanced scientific simulators expose specialized input languages that turn simulation goals into executable configurations, but learning them can cost domain scientists hours to days. We study simula…
AI 点评 · 用代码生成适配器降低科研模拟门槛,让科学家专注研究而非编程。

Amazon Bedrock AgentCore Runtime gives each agent session its own isolated microVM with a persistent workspace, secure tool access through Gateway, and built-in observability—so yo…
AI 点评 · 亚马逊推出隔离微虚拟机托管编码代理,兼顾安全与可观测性,或重塑云上AI开发流程。
🔥🔥🔥 Turn AI-written code into real apps. Nubase is an open-source, AI-native backend platform for AI Coding, agentic applications, and modern product teams:…
Open-source TypeScript terminal coding agent for DeepSeek-V4 — builds on DeepSeek's strong price-performance and ultra-cheap cache pricing, engineering byte-sta…
轻视它就会错过整个时代
Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and co…

IT之家 6 月 7 日消息,华为云现已针对 Agentic AI 时代发布全新云入口“智果园”,新产品支持云码道 CodeArts 代码智能体、华为云 OfficeAce 办公智能体和 WorkAgent 文档智能体。 据介绍,智果园拥有开发、办公等多种关键行业的智能体,可通过智果 AgentArts 平台打造更加实用的智能体,并通过 Skills、AI…
AI 点评 · 华为云整合多模型,降低企业AI应用门槛,凸显生态开放与行业落地潜力。
AI 点评 · 用AI将开发从自由编码转向目标驱动,提升效率与质量,是工程管理新范式。
AI 点评 · 面向Python开发者,Rust语言入门指南填补了性能与安全需求的空白。
Code based orchestration for any coding agent.
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experim…
AI 点评 · 聚焦前沿模型在科研全流程的自主能力,评测框架填补了现有基准空白。
IT之家 6 月 5 日消息,当地时间周四,外媒 404 Media 获得的内部消息显示,谷歌员工正在 猛烈嘲讽 AI 工具 ,其中也包括谷歌自家的 AI 编程工具 Jetski。员工抱怨这些工具 不够可靠,反而让工作更难做 。 今年 4 月,谷歌 CEO 桑达尔 · 皮查伊曾表示,公司当前 有 75% 的新代码由 AI 生成 。但谷歌内部流传的“反 AI”…
AI 点评 · 高层口号与一线体验割裂,暴露AI落地中的真实痛点。
AI workflow automation plugin for intelligent code generation with Claude/Codex
LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks. Existin…
Developers increasingly use AI tools such as ChatGPT, Copilot, and Claude in everyday software workflows, but prior studies often evaluate LLM outputs in isolation rather than examining how developers…
Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tasks as a holistic, binary prediction problem (e.g.,…
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing deceptive performance.…
文|周鑫雨 编辑|张雨忻 杨轩 《智能涌现》从多个信源处独家获悉,2026 年,字节 AI 有四个重要的命题: 加大对世界模型训练的投入,年底前,模型 性能达到现阶段世界模型全球 SOTA(最佳)Google Genie 3 的水平。 视频模型继续保持领先地位, 探索“动态生成”等新方向。 进一步打好 Coding 的地基, 做好 Coding 的 Dogf…
Self-hosted dev sandboxes with preview URLs. One command. No Kubernetes, perfect for coding agents and Saas factories
Self-hosted dev sandboxes with preview URLs. One command. No Kubernetes, perfect for coding agents and Saas factories
IT之家 6 月 3 日消息,科技媒体 404 Media 昨日(6 月 2 日)发布博文,报道称谷歌已联系安卓应用开发者, 希望付费获取私有代码库访问权,用于改进 Gemini、Antigravity 2.0 等开发者工具。 邮件强调,开发者仍保留 100% 知识产权,授权方式为非独占授权,因此项目归属不变,也可继续在其他平台变现。原文未披露具体付款金额、…
AI 点评 · 谷歌付费获取真实代码训练AI,展现提升编程工具准确性的关键路径。
Repo: https://github.com/getpaseo/paseo Homepage: https://paseo.sh/ Discord: https://discord.gg/jz8T2uahpH
AI 点评 · GitHub推出手机端AI编程助手,让代码协作突破桌面限制。
AI 点评 · AI降低编程门槛,让普通人也能用技术改变命运。
AI 点评 · 搜索领域新范式,将搜索转化为代码生成,有望突破现有检索瓶颈。
Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordinate signatures, modu…
AI 点评 · 评估大模型调用新API的能力,填补实用知识缺口,推动智能体从记忆转向推理。
AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficiency for non-English text and structural entropy in…

Some report burning through their whole monthly "AI credit" allotment in a single day.

IT之家 6 月 1 日消息,为加强自主智能体的智能能力,英伟达今日发布了面向全天候运行智能体的全新开源模型与数据集,相关成果由英伟达 Nemotron 联盟联合打造。 据官方介绍,英伟达 Nemotron 3 Ultra 是一款拥有 5500 亿参数的混合专家模型,可为代码开发、科研及企业业务流程中的长效智能体提供顶尖智能能力。相较于同级别主流开源前沿模型…
AI 点评 · 参数规模与推理速度双突破,为智能体部署树立新标杆。
MiniMax M3 今日正式发布。 MiniMax M3 在编程和智能体等专业任务上达到了前沿的能力。它使用了全新注意力架构 MSA (MiniMax Sparse Attention),最高支持 1M 超长上下文。它也是一个原生多模态模型,支持图片和视频的输入,并能操作电脑桌面。 在衡量 Coding 能力的 SWE-Bench Pro 上,MiniMa…
DBeaver 是一个免费开源的通用数据库工具,适用于开发人员和数据库管理员。DBeaver 26.1 已发布,具体更新内容包括: SQL Editor: 新增对 double-curly SQL 参数的支持[#40914] AI 助手:新增对 GitHub Copilot 中新的 OpenAI Codex 模型的支持 Metadata Editor: 修复…
Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences. Existing benchmarks, however, primarily assess whether models refuse un…
The golden age of Microsoft's Github Copilot appears to be at an end.
AI 点评 · 新计费模式引发开发者强烈不满,反映AI工具商业化与用户体验的深层矛盾。
The golden age of Microsoft's GitHub Copilot appears to be at an end.
AI 点评 · 开发者不满新计费模式,GitHub Copilot或面临信任危机。
Molecode presents molecules as code and enables LLMs to operate and reason on chemistry directly.
AI-native coding orchestration platform: unified multi-model agent runtime with stateful sessions, tool governance, and traceable delivery.
Cognition makes Devin, the first and arguably most successful AI coding agent. But famed coder Wu says it isn't designed to supplant human programmers.
AI 点评 · AI编程工具定位辅助而非替代,揭示人机协作新方向。
Cognition makes Devin, the first and arguably most successful AI coding agent. But famed coder Wu says it isn't designed to supplant human programmers.
AI 点评 · AI编程工具定位为人机协作而非替代,创始人观点打破行业焦虑。

Undisclosed addition in jqwik instructed AI coding agents to delete app output.
AI 点评 · 开发者用恶意代码反制AI编码工具,暴露了人机协作中的安全漏洞与信任危机。

Undisclosed addition in jqwik instructed AI coding agents to delete app output.
AI 点评 · 开发者用提示注入反制低代码乱象,揭示AI安全与人类创意间的冲突新战场。
Microsoft is launching a revamped version of Microsoft 365 Copilot, offering a cleaner design that the company claims loads twice as fast. As part of this update, Copilot will prov…
AI 点评 · 提速两倍加界面简化,微软365 Copilot的体验升级将直接影响数亿办公用户效率。
Microsoft is launching a revamped version of Microsoft 365 Copilot, offering a cleaner design that the company claims loads twice as fast. As part of this update, Copilot will prov…
AI 点评 · 微软AI助手提速两倍,界面更清爽,办公效率提升值得期待。
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of R…
Are AI agents tools, co-authors, or researchers? We present a quantified case study ($N=1$): a physicist supervising an AI coding agent (Claude Code, Sonnet and Opus models) over 12 work days and 57 s…
AI 点评 · 物理学家监督AI编码的实证研究,揭示人机协作在科学软件开发中的新边界。
AI 点评 · AI全栈开发效率革命,单人企业级应用落地门槛骤降。
封面来源 | 官方供图 5月27日,快手(证券代码:1024.HK)公布了2026年一季度财报。 财报显示,快手在2026年一季度实现营业收入337亿元,同比增长3.4%,符合市场预期;利润方面,同期实现经营利润36亿元,同比减少15.6%;同期经调整净利润录得34亿元,对应的经调整净利率为10%。 运营层面,2026年一季度,快手的用户增长表现不俗,尤其是…
AI 点评 · 快手用AI视频生成开辟新路,主业稳健下技术变现潜力值得关注。
“87%用户不懂代码,却用AI跑通了一人公司”
AI 点评 · AI降低编程门槛,让非专业者也能创造商业价值,预示行业人才结构颠覆。
Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive de…
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating au…
82 ready-to-deploy Microsoft Copilot Chat agents — paste the instruction block into Copilot Studio and you're live. Writing, HR, PM, IT ops, Sales, Finance, Eng…
Your AI forgets. This remembers. Spec-driven coding harness for vibecoders, product owners, CEOs and real builders — self-improving context memory, 12 agents, 3…
AI 点评 · 用结构化记忆解决AI遗忘痛点,12个智能体协同,适合追求效率的开发者。
Your AI forgets. This remembers. Spec-driven coding harness for vibecoders, product owners, CEOs and real builders — self-improving context memory, 15 agents, 3…
AI 点评 · 用12个智能体构建自进化记忆系统,专为追求高效编码的实干者设计,重新定义AI协作体验。
Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundt…
Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflows.
AI 点评 · 开源协作与前沿模型结合,展现AI编程工具跨环境协调的新可能。
Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflows.
AI 点评 · Warp结合GPT-5.5与开源,探索跨环境编程新范式,值得关注。
Microsoft Copilot Cowork Exfiltrates Files The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this ca…
AI 点评 · 微软AI助手暴露数据安全漏洞,警示企业需重视智能体系统防护。
Microsoft Copilot Cowork Exfiltrates Files The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this ca…
AI 点评 · 揭示AI安全短板:Copilot被利用外泄文件,警示企业需警惕智能助手的数据防护漏洞。
Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation, postmortem). Ships with a 5-scenario regression suite + L…
LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain largely centered on static, isolated, and short-h…
ADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cogn…
AI 点评 · 用树状思维结合剪枝策略,让AI编码代理更接近人类认知模式,提升复杂任务处理效率。
ADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cogn…
AI 点评 · 用树状思维加剪枝策略,让编码代理模拟多动症思考,提升复杂问题解决效率。
Local-first persistent memory for AI coding agents (Claude Code, Cursor, Codex) via MCP. 94.5% LoCoMo recall@10, 70ms p50, multilingual, zero API keys.
AI 点评 · 为AI编程助手提供本地持久记忆,高召回低延迟,无需API密钥即可实现多语言支持。
Local-first persistent memory for AI coding agents (Claude Code, Cursor, Codex) via MCP. 94.5% LoCoMo recall@10, 70ms p50, multilingual, zero API keys.
AI 点评 · 本地优先持久记忆方案,大幅提升AI编码代理效率,无需API密钥,性能指标出色。
AI 点评 · 大模型后端代码生成能力脆弱,揭示智能体在复杂约束下的稳定性短板。
🧭 Architecture-first system design: 26 bilingual tutorials, 25 architecture templates, and 6 end-to-end cases covering distributed systems, AI-native systems,…
OpenAI is named a leader in the 2026 Gartner Magic Quadrant for Enterprise AI Coding Agents, with Codex recognized for innovation and enterprise-scale deployment.
AI 点评 · Gartner权威认证,OpenAI在AI编程代理领域的技术领先性获行业标杆认可。
OpenAI is named a leader in the 2026 Gartner Magic Quadrant for Enterprise AI Coding Agents, with Codex recognized for innovation and enterprise-scale deployment.
AI 点评 · Gartner权威认证,OpenAI编码智能体在创新与规模化部署上领先行业。
Hey HN, We're Gus and Carlos from Runtime ( https://runtm.com ). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other ag…
AI 点评 · 让非工程师也能安全使用AI编码代理,大幅降低团队协作门槛。
Hey HN, We're Gus and Carlos from Runtime ( https://runtm.com ). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other ag…

The vibes were strong at Code with Claude, Anthropic’s two-day event for software developers in London that kicked off on May 19, the same day as Google’s I/O in Palo Alto. (A coin…
AI 点评 · AI编程工具已从辅助进化为主角,开发者生态正被悄然重塑。

The vibes were strong at Code with Claude, Anthropic’s two-day event for software developers in London that kicked off on May 19, the same day as Google’s I/O in Palo Alto. (A coin…
AI 点评 · AI编码能力展示,揭示人机协作开发模式的必然趋势。
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…
AI 点评 · 开源AI部署平台填补市场空白,让编码代理拥有类似Heroku的自动化运维能力,降低开发门槛。
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…
AI 点评 · 开源首个面向AI编码代理的Heroku式平台,填补了代理部署与调试的空白,值得开发者关注。
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…
OpenAI and Dell partner to bring Codex to hybrid and on-premise environments, helping enterprises deploy AI coding agents securely across data and workflows.
AI 点评 · OpenAI与戴尔联手,让企业本地部署AI编程助手,兼顾数据安全与效率提升。
一个写接口文档的AI Agent。支持使用Vibe coding 的方式,编写接口文档,同时自带友好的文档查看工具与接口Mock工具
Minimalistic coding agent written in Rust, optimized for memory footprint and performance
AI 点评 · 零开销AI代理框架,Rust实现兼顾极致性能与低内存消耗,开发者效率新标杆。
Minimal coding agent written in Rust, optimized for memory footprint and performance
AI 点评 · 用Rust打造极简编码代理,专注内存优化与性能,为轻量化AI工具开辟新路径。
A curated list of tools, libraries, MCP servers, and frameworks that power AI coding agents.
AI 点评 · 盘点AI编程代理全生态工具链,开发者和研究者必备的实用资源导航。
A curated list of tools, libraries, MCP servers, and frameworks that power AI coding agents.
AI 点评 · 资源聚合清单,帮你快速找到提升AI编程效率的利器。
OpenSeek - 广度求索: open-source TUI coding agent with multi-provider routing, MCP, LSP, and Plan/Agent/YOLO modes.
AI 点评 · 开源TUI编程智能体,集成多模型路由与MCP协议,创新工作模式值得关注。
OpenSeek - 广度求索: open-source TUI coding agent with multi-provider routing, MCP, LSP, and Plan/Agent/YOLO modes.
Explore how AlphaEvolve's Gemini-powered algorithms are driving impact across business, infrastructure, and science.
Explore how AlphaEvolve's Gemini-powered algorithms are driving impact across business, infrastructure, and science.
AI 点评 · Gemini驱动代码智能体跨领域落地,标志AI从实验室走向产业规模化应用。
Shared memory + orchestration for your coding agents — one MCP server, persistent vector memory, agent registry
Hey HN! I built SimplePDF Copilot: an AI assistant that can interact with the PDF editor. It fills fields, answers questions, focuses on a specific field, adds fields, deletes page…
Hey HN! I built SimplePDF Copilot: an AI assistant that can interact with the PDF editor. It fills fields, answers questions, focuses on a specific field, adds fields, deletes page…
AI 点评 · 客户端AI工具调用实现PDF表单填写,兼顾隐私与效率,突破传统自动化局限。
Give your coding agent the power to write and run agent evals.

How coding agents use tools, memory, and repo context to make LLMs work better in practice
AI 点评 · 拆解编码代理三大核心模块,为LLM落地提供实用框架。

The artificial intelligence coding revolution comes with a catch: it's expensive. Claude Code , Anthropic's terminal-based AI agent that can write, debug, and deploy code autonomou…
AI 点评 · Claude Code收费昂贵,Goose免费替代,AI编程工具价格战打响。

Anthropic released Cowork on Monday, a new AI agent capability that extends the power of its wildly successful Claude Code tool to non-technical users — and according to company in…
AI 点评 · 面向非技术人员的AI代理工具,降低编程门槛,拓展办公自动化应用场景。

Nous Research , the open-source artificial intelligence startup backed by crypto venture firm Paradigm , released a new competitive programming model on Monday that it says matches…
AI 点评 · 开源编程模型NousCoder-14B在Claude Code发布时亮相,挑战闭源巨头,展现社区创新

KV caches are one of the most critical techniques for efficient inference in LLMs in production.
AI 点评 · 从零实现KV缓存,揭示大模型高效推理的核心机制,是深入理解LLM性能优化的必学技术。

Why build LLMs from scratch? It's probably the best and most efficient way to learn how LLMs really work. Plus, many readers have told me they had a lot of fun doing it.
AI 点评 · 从零构建代码大模型,深度理解LLM原理,兼具学习效率与趣味性。