We look at Project HydraFusion, GitHub's research preview that treats workflow selection as an optimization problem rather than a model picker. We break down the three execution pa…
AI 点评 · 多模型动态编排颠覆单一模型选择,按任务实时构建工作流,开启编程助手新范式。
共 311 条相关资讯 · 来自历史归档
We look at Project HydraFusion, GitHub's research preview that treats workflow selection as an optimization problem rather than a model picker. We break down the three execution pa…
AI 点评 · 多模型动态编排颠覆单一模型选择,按任务实时构建工作流,开启编程助手新范式。
TIL: Using Blender with coding agents on macOS I've been having fun with Blender in ChatGPT Codex on my Mac recently. Getting it to work with coding agents is really easy: install…
AI 点评 · 跨工具融合新玩法,展示AI代理操作3D软件的潜力,降低创作门槛。
AI 点评 · 鸿蒙原生AI开发框架落地,展示国产系统端侧智能新路径。
Microsoft's Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new…

new SOTA computer use and coding, 2.5x pricier per token, but WAY cheaper per task, less monitorable. overall, a very successful launch of OpenAI’s new frontier model class.

IT之家 9 月 4 日消息,智谱 AI 昨日宣布,旗下 GLM Coding Plan 正式推出“Flash × ZCode”夜间畅用活动。 即日起至 9 月 20 日,每晚 23:00 至次日 9:00,所有付费套餐用户无需手动开启,系统将自动生效优惠。活动期间,GLM Coding Plan 套餐中的 GLM-5.3-Flash 模型额度消耗规则如下:…
AI 点评 · 夜间免费时段精准覆盖开发者高频场景,降低试用门槛,或催化AI编程工具竞争新策略。

OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the "AGI era." Astra tops benchmarks in math, coding, and cybersecur…
AI 点评 · 首个敢自封开启AGI时代的大模型,性能碾压基准测试,或成行业新分水岭。
Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's ma…
AI 点评 · 模型瘦身两成还提速,端云协同成AI落地新范式,值得开发者关注。
For its new Muse Spark model, intended for operating coding and other agents, Meta is offering an explicit discount averaging out to about 95% for users who "contribute" to the dev…
AI 点评 · 以补贴换数据,Meta为训练智能体模型开辟新路。

OpenAI leaders think the company’s next generation model, which excels at computer use and coding, may mark a major milestone in AI development.
AI 点评 · 模型能力跃升与AGI前景绑定,预示AI应用范式将迎来质变拐点。
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests…

IT之家 9 月 3 日消息,西班牙 AI 公司 Multiverse Computing 最新推出 Quasar 438B 模型,根据 Artificial Analysis 得分, 该模型 Intelligence Index v4.1.1 得分为 43,在参与比较的欧洲模型中排名最高。 IT之家注:该指数综合 9 项评测,覆盖智能体、代码、科学推理、通…
AI 点评 · 欧洲开源大模型格局生变,西班牙新贵以1M长上下文挑战美国主导。

Google's Gemini 3.8 Flash, the third Flash model in six weeks, matches Claude Opus 5 on some agentic coding benchmarks at lower cost. But its "working harder" reasoning burns about…
代码之后,3D建模也开始「一句话交付」。
Today is Claude Fable (and Mythos) 5.1 day . Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announceme…
AI 点评 · 代码与知识工作新标杆,动画生成能力惊艳,值得开发者实测。

Anthropic launches Claude Fable 5.1 and Mythos 5.1, its most capable AI models yet. Fable 5.1 doubles its predecessor's score on Terminal-Bench-Science and improves agentic coding…
AI 点评 · 成本降45%还升级编程科研,Claude新模型性价比拉满,开发者效率或迎质变。
Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-e…
Faithfully translating research papers into repository-level implementations remains challenging because papers often describe methods at a high level, leave implementation assumptions implicit, and r…
The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often…
This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention…
Embedding-based code retrieval is a core component of coding agents and retrieval-augmented code generation, where retrieving correct code matters more than retrieving lexically similar code. Existing…
AI 点评 · AI编程提效受制于交付环节,架构实践揭示关键瓶颈。
AI 点评 · AI编码再进一步,打通研发全链路才是企业落地的关键。
Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context…
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, a…

AI coding assistants like Claude Code and Codex have no sense of time, according to a new study. Both systematically overestimate how long tasks will take. Codex is off by as much…
AI 点评 · 时间感知缺失暴露AI代理致命短板,任务规划误差成规模化应用关键瓶颈。
Stop your coding agent from stalling real work on self-invented bookkeeping - receipts, hashes, locks, certification rituals. Ship first, then verify. Skill for…

OpenAI is cutting off the AI coding tool Cursor after SpaceX acquired the company, citing Elon Musk's track record of breaking contracts. Cursor co-founder Michael Truell downplays…
AI 点评 · 马斯克收购引发连锁反应,OpenAI切断合作凸显商业信任危机。
Coding正在变成Al世界的数字执行力
AI 点评 · 编程门槛骤降,AI正将创意直接转化为数字执行力。
AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly extended through plugin marketplaces, yet the structure, maintenance, and co-evolution…
AI 点评 · 多智能体协作审查代码,规模化落地经验值得借鉴。
腾讯混元刚开源的 Hy4 preview,明显奔着干活去了。 参数更大,跑得也不慢:总参数量 770B,单次推理激活 49B,支持 1M 上下文。而且它把自身训练和优化都包圆了,端到端吞吐较基线提升 31.8%。 代码和办事能力大涨:在 Arena 的代码测试里冲到第 5,开源模型里排第 3。内部盲测 200 多个工程任务,分数压过了 GLM 和 Kimi。…

IT之家 8 月 28 日消息,科技媒体 Linuxiac 昨日(8 月 27 日)发布博文,报道称开源操作系统项目 Asahi Linux 官方宣布, 针对苹果 M3 系列 Mac 电脑的原生 Linux 支持已步入最后冲刺阶段,核心驱动代码已被正式并入 Linux 7.2 内核主线周期。 该媒体指出苹果 M3 系列芯片采用了高度定制化的电源管理、显示控制…
AI 点评 · 底层驱动突破,M3跑Linux不再是梦,开源生态再进一步。
Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attack…
AI 点评 · 自动模式成防注入攻击新防线,Anthropic押注AI编码安全,行业风向标值得紧盯。
Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work,…
Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality…
Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Age…
World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes ra…
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model…
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, termed visual repair co…
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they…
AI 点评 · AI依赖将瓦解编程专长,警示技术反噬与人才断层风险。
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks…
Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger…
Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborati…
As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention, with closed-source LLMs (e.g., Mythos) delivering…
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they…
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and au…
Hi HN- I'm Pablo, the founder of Proliferate (YC S25)! Proliferate ( https://github.com/proliferate-ai/proliferate ) is an open-source, self-hostable AI IDE that lets you work and…
**Ox Alpha** emerged as a mystery model with strong coding and agentic performance, likely a **Zhipu/GLM-family** model such as **GLM-5.3 Vision**. Analysts suggest its gains come…
Hello everyone. I've been working on this experimental editor called Huzzah. I've been working almost exclusively with coding agents since January of this year, and over the past f…
AI 点评 · AI编程工具新范式,直击代码代理协作痛点,值得开发者实测。
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except ins…
Slack is introducing dedicated channels where teams can vibe-code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includ…
SpaceX was reportedly in talks to buy AI coding startup Cognition. SpaceX has already acquired Cursor as it races to catch up to rivals like OpenAI and Anthropic in enterprise AI.
AI 点评 · AI编程赛道并购热升级,SpaceX布局意图明显,折射科技巨头对智能编码工具的争夺白热化。
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scien…
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code…
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into a…

IT之家 8 月 18 日消息,据科技媒体 Windows Central 今天报道,微软暂时移除了 Win11 系统的 Copilot 任务栏搜索框,同时将更多精力放在统一版 Copilot 应用上。 IT之家从原报道获悉,微软曾于去年更新 Windows 11 的任务栏,允许用户将原来的普通搜索框替换成 Copilot 搜索框,让用户直接从任务栏与 Co…
AI 点评 · 统一体验优先,微软战略聚焦或预示AI助手入口新方向。

Secret parameter allowed hackers to steal passwords when a target clicked on a link.
除了首页时间流和侧栏的精选展位,少数派Matrix社区还有很多优秀内容因条件所限无法得到有效曝光,因此我们决定重启Matrix周报,并在此基础上添加更多社区内容、作者投稿新玩意呈现给大家。上周社区速递 ... 查看全文
代码能跑≠游戏能玩
Agent 越来越会写代码了,却还是常常像第一次来到这个项目。
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution envi…
AI 点评 · AI重构研发流程,Rootly颠覆传统代码评审,预示智能体将重塑开发范式。
AI 点评 · AI修复工具反成攻击面,供应链安全新漏洞值得警惕。
DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, a…
AI 点评 · 代码生成告别玄学,回归可验证的工程逻辑。
AI 点评 · AI Coding正从工具走向金融科技研发全流程,闭环落地是行业关键转折。
AI coding startup Cursor is now officially a part of SpaceX.
AI 点评 · 交易背后暗藏AI编程与航天巨头的战略协同,或重塑行业竞争格局。
AI 点评 · AI角色转变:从写代码到带团队,重新定义人机协作新范式。
DeepSeek Harness 零代码桌面端|一键启动,支持 Windows 与 macOS;内置插件发现、热点插件推送、一键安装与管理、AI 智能推荐和视觉增强。
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditio…

Zhipu AI has released GLM-5.3, a model that, according to its own benchmarks, is the most powerful open-weights coding model, with a 50 percent improvement over its predecessor thr…
顺手拿下最强开源安全模型
Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments…
基于 Hybrid RAG 与 LangGraph 的本地代码仓库理解工具,支持中英文提问、语义检索、证据引用和多问题拆分
**Z.ai launched GLM-5.3**, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger b…
Microsoft Copilot will no longer show its emotive yellow blob, Mico, when you use the chatbot's voice mode. In a support page, Microsoft says it's going to move Mico to its Learn L…
AI 点评 · 微软砍掉情感化形象,聚焦功能优先,AI助手交互设计风向标值得关注。
Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipu…

Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the…
AI 点评 · 三周迭代降价五成,性价比与编码能力双重突破,AI大模型竞争进入快车道。
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures i…
Google has released Gemini 3.7 Flash, a refinement of Gemini 3.6 Flash with algorithmic improvements to its reasoning core. It handles text, images, audio, and video across a 1M-to…
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and…
Microsoft is simplifying Copilot by combining its consumer and business apps, and dropping AI-generated podcasts, Group Chats, Deep Research, and its Mico character.

IT之家 8 月 13 日消息,据外媒 The Verge 今天(13 日)晚间报道,微软终于开始整合面向个人和企业用户的两套 Copilot AI 助手,目标是打造一个统一的“超级应用”。 为超级应用“铺路”的第一步,就是从 Copilot 和 Microsoft 365 Copilot 两款应用开始,个人账户和工作账户都会 迁移到新版统一应用 。 新应用…
Microsoft is finally beginning to combine its consumer and commercial Copilot AI assistants into a single "super app" interface, starting with the Copilot and Microsoft 365 Copilot…
Hi HN! We’re Adi and Alex, founders of Bullet, a faster coding agent. Bullet started in a senior year dorm. We were fresh out of working at AppLovin and Citadel, and naturally thou…
SpaceXAI released Grok 4.6 on August 12, 2026 — a post-training upgrade over Grok 4.5, not a larger base model. It ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligenc…
**Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory…
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures i…
Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation.
AI 点评 · AI编程赛道估值飙升,印证资本对自动化开发工具商业前景的狂热押注。
AI 点评 · AI原生语言颠覆编程范式,开发者角色或将重塑。

IT之家 8 月 12 日消息,北京时间今天(12 日)晚间,Grok 4.6 正式发布。新模型在 Grok 4.5 基础上进一步强化长时间运行的智能体任务,以及复杂的交互和视觉工作。 按照发布信息,Grok 4.6 能够持续处理包含大量步骤的复杂任务,包括 资料研究、信息分析、大型代码库处理 ,以及将产品构想转化为完整应用或工作成果。 基准测试方面,Gro…

Microsoft has released MAI Code 1.1 Flash, a code model for GitHub Copilot that's said to be 25 percent more token-efficient at a quarter of the cost of its predecessor. In benchma…
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, manag…
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated…
随着HarmonyOSNEXT的日渐成熟,越来越多的独立开发者和小团队开始将目光投向这片新的生态。借助日益强大的AI辅助编程工具,写出能跑通的代码、画出漂亮的UI界面,门槛已经变得前所未有地低。但当你 ... 查看全文

IT之家 8 月 11 日消息,外媒 9to5Mac 在挖掘苹果 iOS 27 开发者预览版 Beta 5 代码信息时发现苹果正在开发一项名为 Apple Reference Image 的新功能,用于认证 iPhone 拍摄照片的来源真实性。 随着生成式 AI 图像技术快速发展,AI 生成或经过深度修改的图片越来越容易伪装成真实照片,过去需要专业摄影后期人…
AI 点评 · 用系统级认证对抗AI造假,手机照片可信度有了新标准。
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-t…
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a re…

IT之家 8 月 9 日消息,甲骨文(IT之家注:Oracle)现已通知 OpenJDK 开发者,要求项目组不能再提交 AI 生成的代码。 甲骨文对此表示:“ OpenJDK 社区贡献内容不得包含由大语言模型 、 扩散模型或深度学习系统生成的内容 。此处的‘内容’包括但不限于代码、文本、PR、电子邮件沟通、维基和 Bug 报告。” 甲骨文解释称, 本次新规出…
AI 点评 · AI写代码虽热,但开源项目信任与质量更重,此规为业界立下新标杆。
600 行 TypeScript 写成的超级迷你版 pi,让你轻松从 0 写出属于你的 pi-agent
特别声明本文的项目构思、结构设计及相关素材整理均由人工完成。在产品开发与调试过程中,使用GPT-5.6Sol模型作为辅助工具,参与方案讨论、代码编写与问题排查。文章内容基于实际开发过程中的经验与记录, ... 查看全文
AI 点评 · 用AI辅助搭建桌面看板,展示F1数据,是极客与车迷的实用结合。
Tencent Cloud has open-sourced TencentDB Agent Memory v2.0, a team-level memory hub that turns conversations, documents and code into four governed, reusable assets — Chat Memory,…
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evol…
AI 点评 · 开源替代方案冲击成本壁垒,AI编程支出骤降揭示行业拐点。
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled cont…
**OpenAI** escalates its upcoming **Astra** model to "critical" cyber status due to significant advancements in agentic coding and cybersecurity, pausing some activities to strengt…
Microsoft has open sourced code-testing-generator, a polyglot unit-test agent shipping in the MIT-licensed dotnet/skills repository. It reads a repository before writing anything —…
Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generati…
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification fro…
In this tutorial, we explore adaptive experimentation using Meta’s Ax with the modern Client API. We work through a complete workflow where we tune a RandomForest model on a synthe…
AI 点评 · Ax平台实操指南,降低自适应实验门槛,值得算法工程师收藏。
Taking vibe-coding a step further, Naïve claims its infra can automate most of the work in setting up and running a business.
AI 点评 · 从“氛围编程”到公司运营,全流程自动化或颠覆创业服务赛道。

As engineering teams adopt coding agents like Codex, leaders need visibility into adoption, consumption, and reliability. This post shows how to route Codex OpenTelemetry metrics t…
Learn how to build 24/7 automated AI agents & chatbots for your website, SaaS, or Shopify store with zero coding required.

Cloudflare built an AI agent workspace for its employees. Now it’s open source.

Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, a…

Meta released Muse Spark 1.2 along with its own coding agent, Muse Code, which is designed to pick up exactly where it left off after a crash. The cheapest tier runs just 20 cents…
Prime Intellect has open-sourced Prime Agent, a coding and research harness built on two abstractions: the Recursive Language Model, which turns sub-agent calls into functions insi…
Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software.
AI 点评 · 开源AI编码新赛道,Meta入局挑战GitHub Copilot,复杂任务处理成核心看点。

Meta Superintelligence Labs has released Muse Code, a terminal coding agent in beta, powered by the new Muse Spark 1.2 model. Muse Code plans changes, writes code, and validates re…
The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing trade-off, revealing persistent challenges with code quality…
Aixle Flow — orchestrate coding agents through durable, inspectable workflows.
AI 点评 · AI辅助优化传统工具,展现外行跨界创新潜力,效率提升惊人。
Hi HN, this is Shailendra and Karan here. We are building a fast and safe way for coding agents to debug issues live in production. When prod breaks, it lets Cursor, Claude, and ot…
MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude…
2026 年编程导航 AI 编程实战新项目,基于 Taro + React + FastAPI + DeepSeek 的 AI 闯关学习小程序,支持一句话 / 文本主题 AI 出题、闯关答题与即时讲解、联网搜索增强、RAG 私有知识库出题、AI 生图配图、微信登录与通关复盘报告。覆盖 LangChain / LangG…

CopilotKit has published the Channels SDK, an MIT licensed library that runs an existing AG-UI agent inside Slack and Microsoft Teams. Version 0.5.0 ships five platform adapters an…
FuXi is a fast, self-contained AI coding agent that lives in your terminal — edit code, run commands, and drive tools, with cost-aware routing across LLM provid…
文 | 李炤锋 编辑 | 张雨忻 “长链任务如果只通过代��层面的反馈,误差可能会不断累积,最终效果会非常差。”谈及原生多模态的意义,一位多模态研究员表示,“视觉是一种更准确的反馈,也更贴近用户意图。” 过去一年,Coding与Agent能力不断改写大模型的排名,也成为AI最快兑现商业价值的场景之一。与此同时,随着Agent开始接管更多长链任务,越来越多的通…

Qwen is so back!
AWS now allows vibe-coding tool Superblocks to be embedded into the private clouds of AWS customers. It's another step toward decoupling apps from models.
AI 点评 · 云巨头开放私有化部署,标志AI应用与模型解耦加速,生态控制权争夺战升级。
Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet…
AI 点评 · 终端AI助手体验升级,交互更直观,配置门槛降低,开发者效率有望显著提升。
AI 点评 · 命令行与AI结合,开启高效编程新体验。
Hi HN, we’re Bence and Ryan, founders of Hoplite ( https://hoplite.sh ). Hoplite lets you deploy coding agents in the cloud, with a suite of tools that makes it incredibly easy to…
AI 点评 · 云上部署编程代理门槛大降,直击AI开发团队协作痛点,看点在于工具链整合的实操价值。

Kaggle’s AI Agents Intensive with Google brought learners together in a no-cost course to build and deploy the next frontier of AI.

Your next drive-thru order might be taken by a bot. And you might not even notice.
**Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing…
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typicall…
The graph based agentic IDE

A field report from OpenAI and academic partners shows coding agents can modernize neglected research software, with speedups of up to 60x. But the systems are "eloquent, convincin…
AI 点评 · AI编码智能提速60倍,但科学判断仍是人类专属,AI辅助科研的边界值得深思。

A security researcher has demonstrated a worm-like attack on Microsoft Copilot for Word: invisible prompt injections hidden in documents spread automatically into new files every t…
AI 点评 · 生成式AI的隐蔽攻击链曝光,文档蠕虫可劫持Copilot,企业防护需紧急升级。
A low-token visual evidence compiler for text-only coding agents. Convert images into compact Visual Evidence Packets (VEP) for DeepSeek, Codex, Claude Code, an…
LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small…
尽显AGI黄金时代的奢华
AI 点评 · 编程大师对AI代码态度两极分化,揭示行业核心争议。
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, deve…
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue…
a quiet day lets us cover how AI is permeating financial services as the next big vertical after coding.
IT之家 7 月 30 日消息,据外媒 The Verge 报道,当地时间周三(29 日),微软 CEO 萨提亚 · 纳德拉在财报电话会议上透露,微软正在打造一款 AI“超级应用”,计划把 Copilot 的 对话、编程和智能体 功能整合到同一个应用中。这款“超级应用”将于今年发布,同时面向个人用户和企业客户。 纳德拉表示:“Copilot 正在迅速从聊天工…
AI 点评 · 微软将AI聊天、编程与智能体整合为单一应用,或重塑用户与AI的交互方式。
Microsoft is working on an AI "super app" that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the…
AI 点评 · 微软将Copilot升级为超级应用,整合聊天、编程与智能体,重塑AI生态格局。
Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence learning, nonlinear recurrent updates often require repeated circuit evaluations and…

In this tutorial, we configure and operate Kimi CLI as a fully non-interactive AI coding agent. We install the CLI through uv with an isolated Python 3.13 environment, configure Mo…
Fireworks AI has released Fireworks Nexus, an AI management and routing platform aimed at engineering organizations. It connects the coding tools developers already use to a manage…
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program fro…
AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Exis…
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
AI 点评 · 科学家用AI编程代理加速科研,从基因组学到各领域,将颠覆传统计算模式。
Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure life…
AI 点评 · AI Coding普及后,开发者应警惕工具依赖,聚焦核心创造力与问题解决能力。
Perplexity has released pplx, an official command line client for its Search API. The tool exposes two commands — pplx search web and pplx content fetch — and returns exactly one J…

一句话,一个周末,Claude 直接把自家最前沿的模型跑在了一台全新的 AMD MI355X 机架上! 人类工程师全程没动手改过一行代码 。 谁能想到,英伟达花 20 年堆起的 CUDA 护城河,就这样被跨了过去。 一个周末,Claude 把 AMD 新 GPU 调通了 故事是这样的。 前段时间,AMD 给 Anthropic 送来一台搭载 MI355X 的…
An old coder's strategy for the agent era: don't read the code — make it run the gauntlet. Evidence-first development skill for coding agents, inspired by Uncle…
runs anywhere. uses anything
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed…

Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the…
AI 点评 · 用前沿模型做规划,低成本模型执行编码,AI协作效率大幅提升。
The KwaiKAT Team at Kuaishou has published the KAT-Coder-V2.5 technical report, arguing that agentic coding capability is bottlenecked by training infrastructure rather than model…

An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests,…
AI 点评 · AI催生考试变革,68%教师改考法,教育公平与技能评估面临新挑战。

IT之家 7 月 25 日消息,腾讯 WorkBuddy 桌面版现已上架华为鸿蒙电脑 App Gallery 应用商店。 WorkBuddy 是腾讯推出的一款全场景 AI 办公智能体桌面工作台,覆盖日常办公、代码开发与设计创意。应用商店页面显示,这款应用 原生适配 了鸿蒙系统。 据IT之家此前报道,7 月 18 日, WorkBuddy 发布移动端独立 Ap…

Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Op…
The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.
The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.
The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief that how AI assistants interact with users i…
Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effect…
See what your coding agents did and what it cost. Breaks each task down into work steps — tools used, files changed, tests run, time and tokens spent. Local-fir…
给它一张图,还你整个世界
AI 点评 · AWS推出企业级AI安全方案,填补智能体代码防护空白。
**Anthropic** launched the **Claude Opus 5** model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. The model achieved an *…

In this post, we explain how Amazon Bedrock Guardrails can be configured for code generation workflows with coding assistants to overcome these constraints. With these best practic…
AI 点评 · 用护栏策略规范代码生成,平衡AI效率与安全风险,解决企业实际痛点。
AI 点评 · AI应用落地加速,从编程到办公场景的革新值得紧盯。
A practical guide to AI: from running your first local model to building your own agents. 52 files covering LLMs, Ollama, RAG, prompt engineering, machine learn…
A practical guide to AI: from running your first local model to building your own agents. 52 files covering LLMs, Ollama, RAG, prompt engineering, machine learn…
文|周鑫雨 编辑|张雨忻 什么样画像的AI创业者会饱受瞩目?不同投资人心里或许有不同的答案,但其中一个答案一定是:剪映系。 陈冕,Lovart、LibTV等热门创作工具的缔造者;明超平,AI Coding社区YouWare的创始人,2025年一级市场最火的95后;闹闹,其打造的AI视频创作工具OiiOii,被高瓴、锦秋投资...... 他们履历的共同点是:曾…
AI 点评 · 爆款操盘手首次分享AI产品方法论,揭示从工具到生态的进化逻辑。
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its cor…
AI 点评 · 首个抗污染多领域编码智能体评测基准,为评估AI编程能力提供更可靠标准。
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected…
AI 点评 · 评估编码代理从单任务执行转向交互式项目构建能力,标志AI编程工具应用场景的质变。

AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does. Nine in ten Builders use AI coding tools…
AI 点评 · monday.com在亚马逊Bedrock上大规模部署AI队友,实战经验揭示企业级AI代理落地关键。
Over the past few months, our team has been building more and more slidedecks using web frontend technologies with coding harnesses like Claude Code, but a common complaint is to m…
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat su…
Hey HN! This is Divit from Almanac (YC S26). We built CodeAlmanac, a wiki for your coding agents that updates as you talk to them. It is open-source, local, and free. Here’s a demo…

A new type of malware can worm deep into AI coding systems to steal data and logins—and can flip a “death switch” to destroy files and keep out real users.
AI 点评 · 瞄准AI基础设施的隐形黑客工具,能窃取数据并触发“死亡开关”,暴露AI安全盲点。
“2026年,创投圈的浪潮再次翻涌:AI从技术概念走进产业深水区,硬科技创业从“小众赛道” 变成“主流共识”,年轻的创业者们正在用代码和双手,重新定义中国创新的未来坐标。 每一年,由36氪 · 暗涌主办的WAVES大会,都是中国创投圈的年度风向标。 今年的 WAVES 2026以“今年盛夏”为主题,落地广州番禺良仓新造创意园,在两天的时间里,我们汇聚了顶级投…
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. Howeve…

Augment Code's Vinay Perneti talks models, harnesses, and context.
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code…
Databricks has remade its image into an AI company and has published research on the cost savings of open-weight AI models for coding.
AI 点评 · 估值188亿美元,从数据平台成功转型AI公司,开源模型成本优势研究或重塑行业格局。
A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable pl…
36氪获悉,7月17日,在2026世界人工智能大会(WAIC 2026)上,网易智企携全新升级的一站式企业AI应用服务亮相,集中展示AI Agent编排、AI Coding、AI客服、AI私域助理、AI智能数据与AI Agent安全等企业级AI能力,围绕安全治理、组织协作与业务增长三大场景,呈现企业级AI应用实践。
AI 点评 · 展示企业级AI全栈能力,聚焦安全治理与业务增长三大场景,为行业提供可落地的AI应用标杆。
模型、Harness Engineering与产品的持续进化
AI 点评 · AI编码市场格局剧变,阿里Qoder以绝对优势领先,凸显大模型工程化落地的关键转折点。
文 | 周鑫雨 编辑 | 张雨忻 智能涌现从多个独立信源处获悉, 截至 2026 年 7 月, 智谱的 ARR(年度经常性收入)已经达到 10 亿美元 。 截至发稿前,针对上述信息,智谱未回复。 过去一年,AI Coding 和视频生成模型已经成为全球造血能力最强的 AI 赛道。 海外,Anthropic 的 Claude Code 仅发布半年,ARR 就飙…
AI 点评 · 智谱ARR半年飙升15倍达10亿美元,印证AI商业化进入爆发期,行业格局加速重塑。

IT之家 7 月 17 日消息,Google(谷歌)当地时间 16 日宣布,其 AI 笔记助理 NotebookLM 更名为 Gemini Notebook,以实现人工智能产品的品牌统一。 谷歌表示,NotebookLM 仍然是一款独立的产品,但 现在将在包括 Gemini 应用和 Google 搜索的整个 Google 生态系统中发挥更大的作用 。 Gem…
AI 点评 · 品牌统一战略升级,原生代码支持强化AI笔记工具实用性,值得开发者关注。

Creator says he will "very loudly ignore" those arguing for a ban on AI tools.
AI 点评 · Linus强硬表态,AI编码争议升级,Linux社区自由与创新的核心原则面临考验。
AI 点评 · 多智能体协作编程,突破单Agent局限,是AI自主开发的关键一步。
AI 点评 · AI原生开发全流程革新,企业人机协同研发范式从理论走向实践。
Tool: Mermaid to Unicode box art (grok-mermaid) While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs , a "se…
AI 点评 · 将Mermaid图表转为Unicode字符画,适合终端环境,实用且有趣。

IT之家 7 月 16 日消息,马斯克旗下 SpaceXAI 公司昨日(7 月 15 日)宣布开源 Grok Build, 并将源代码发布至 GitHub 平台。 在官方博文中,SpaceXAI 表示: 开源发布源代码,是构建强大、可靠框架的最直接方法。用户可以阅读源代码,了解其从上下文构建到工具调用分发的完整工作原理。 开源也让框架更容易探索和扩展:如果用…
AI 点评 · 开源代码降低使用门槛,推动编程AI智能体技术快速迭代,值得开发者关注。
OpenAI, which is in the middle of a legal battle with Apple over hardware trade theft allegations, just released a light-up keyboard designed to be paired with its agentic coding a…
AI 点评 · 硬件纠纷未平却推高价键盘,OpenAI跨界硬件野心与Codex生态绑定值得关注。
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While p…
The startup has reached a $120 million annualized revenue run rate and more than 200,000 paying customers.
SpaceXAI's Grok Build AI coding tool was spotted uploading users' entire codebases to Google Cloud before it was reported, and the company turned it off. The Register reports that…
AI 点评 · 隐私漏洞暴露AI编程工具安全隐患,用户数据保护机制亟待加强。

MIT students designed, built, and tested a jet engine with AI copilots, assessing AI’s usefulness in developing high-performance aerospace systems.
一手实测这就奉上
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the acti…
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functio…
A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.
A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods…
Hello HN, I don't post on here much, but wanted to get some eyes on a new project I'm just launching. I think we definitely need one more AI code agent.. I'm a long-term C++ dev, a…

IT之家 7 月 12 日消息,Meta 于 7 月 9 日正式发布适用于 AI 智能体的多模态推理模型 Muse Spark 1.1 版本,重点提升了模型在智能体任务中的规划、协同与执行能力,并增强了工具调用、代码开发、应用操作能力。 Meta 表示,Muse Spark 1.1 强化了多智能体协作机制,由主智能体负责收集信息、制定计划,再将任务拆分并分配…
AI 点评 · 多智能体协作机制是AI落地的关键突破,Meta这次强化了任务拆解与分工能力。
AI 点评 · 3D可视化编码过程,直观追踪AI代理如何理解代码库,革新调试与协作方式。
换更大的模型就等于更聪明? 【导读】 换更大的模型就等于更聪明?这可能是Claude Code用户最深的误会。很多人为此一路换到最贵的Fable,近日,Anthropic,亲手澄清了这个误区。 你有没有过这种时刻:Claude Code写代码写砸了,第一反应,就是赶紧换个更强的模型。 但这一招,很多时候并不管用,甚至是在白花钱。 近日,Anthropic官方…
AI 点评 · 揭示模型性能瓶颈不在参数规模,而在于使用方式,颠覆用户认知。

IT之家 7 月 11 日消息,月之暗面官方昨晚宣布, K2.7 Code 高速版结束 Beta,成为常驻可选模式 。 订阅 Allegretto 及以上会员计划的用户,在 Kimi Code CLI 或其他工具中使用 Coding Plan 进行编程时,无需任何申请,可以直接调用 K2.7 Code HighSpeed。 官方表示,K2.7 Code Hi…
AI 点评 · 编程效率再升级,K2.7 Code高速版免费开放,开发者可无缝调用。

Amid live coding sessions and Silicon Valley optimism, the UN’s AI for Good summit wrestled with an urgent question: Can global governance catch up before the technology races beyo…
AI 点评 · 联合国AI峰会展示前沿科技,核心挑战在于全球治理能否跟上技术飞速发展。
OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.
AI 点评 · OpenAI与微软合作稳固,GPT-5.6成Copilot 365核心,预示AI办公生态新阶段。

“IT早报”时间,大家好,现在是 2026 年 7 月 10 日星期五,今天的重要科技资讯有: 1、OpenAI 最强 AI 模型:GPT-5.6 系列正式上线,纳德拉称微软 Copilot 同步接入 OpenAI 公司 7 月 10 日发布公告,宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及…
AI 点评 · OpenAI模型重大升级,微软深度整合,预示AI竞争进入新阶段。

IT之家 7 月 10 日消息,OpenAI 公司今天(7 月 10 日)发布公告, 宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及 API 中上线 GPT-5.6 系列模型。 在模型方面,IT之家援引博文介绍,OpenAI 本次共发布 3 档模型: 旗舰版 Sol(太阳):每 100 万 T…
AI 点评 · 微软同步接入,意味着AI竞争格局突变,企业级应用迎来新拐点。
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations — the kind of automation that enterprises are increasingly…
AI 点评 · Meta携Spark 1.1切入企业级AI编码自动化,专注大型代码迁移和修复,展现巨头竞争新方向。
After reentering the AI race with its first in-house Muse Spark model in April, Meta is now opening up the doors to developers with a new model that can plug into AI coding softwar…
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
AI 点评 · 微软将GPT-5.6深度集成办公套件,标志AI办公进入新阶段,用户体验将显著提升。
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments an…

IT之家 7 月 9 日消息,SpaceXAI 今日正式发布了其 Grok 4.5 模型,这是该公司首个专门针对编程和智能体任务训练的模型。 据介绍,该模型由 SpaceXAI 与 Cursor 联合完成训练,在提供前沿智能水平的同时,兼具领先的速度与成本效率。马斯克将其称为“Opus 级模型”。 Grok 4.5 面向真实工程场景设计,擅长处理大型代码库以…
AI 点评 · 编程智能体成本砍半效率翻倍,马斯克联手Cursor的定价策略才真值得行业关注。
AI 点评 · AI时代软件工程师的门槛被重新定义,挑战传统开发思维。
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces ro…
AI 点评 · AI自动修复漏洞,提升开发效率与安全性的关键一步。
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revisi…
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use thes…
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-f…
AI 点评 · AI冲击初级程序员岗位,暴露行业技能结构转型危机。
AI 点评 · 孤岛编码实验揭示AI自主编程的进化路径。

IT之家 7 月 4 日消息,科技媒体 AppleInsider 昨日(7 月 3 日)发布博文,报道称在 iOS 27 开发者测试版中, 发现相关文字描述,指向具备摄像头功能的苹果 AirPods 耳机产品。 程序员 Sam Henri Gold 昨日在 X 平台发布推文,指出在 iOS 27 开发者测试版中,发现了代号为 B790 的智能眼镜。IT之家附…
AI 点评 · 苹果在耳机上集成摄像头的尝试,或为空间计算与AI视觉交互开辟新入口。
Control plane for AI coding agents: route tasks, reduce token spend, run multi-agent workflows, fallback executors, and track cost per task.
AI 点评 · Agent正突破编程边界,重塑各行各业工作流程,预示AI自主执行任务的未来。
AI 点评 · AI编程工具成瘾性暴露工程师效率与代码质量间的深层矛盾。
Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would lo…
AI 点评 · 首个开源LLM编程代理框架,简化代码生成与迭代流程,开发者可快速上手实验。
AI 点评 · 用短链约束AI编程,突破游戏创作瓶颈,展现人机协作新思路。
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt…
AI 点评 · 用微虚拟机隔离智能体,提升了云计算安全性与灵活性。
Coding agents don't have long-term memory. But you do have months of full-fidelity agent transcripts stored on your machine. A simple solution that goes a long way: ingest those tr…
今日热点导览 OpenAI据悉迎来重大技术突破:系统优化使模型推理成本减半 苹果首次上架iPhone16e官翻机,仅有黑色与白色可选 国内航线燃油附加费7月5日起大幅下调 巴菲特“断供”盖茨基金会 董明珠喊话股东:家电不换成格力,凭什么要分红 詹姆斯发文告别湖人 TOP3大新闻 二手豪华燃油车价格大跳水:宾利近27万、保时捷15万 7月1日,在青岛市的陆海汽…

“IT早报”时间,大家好,现在是 2026 年 7 月 2 日星期四,今天的重要科技资讯有: 1、美光 CEO 发声“甩锅”:内存供应失衡别只怪我们,客户 2023 年曾压价到三分之一导致 AI 爆发前没能扩产 美光 CEO 梅赫罗特拉表示,当前内存芯片供应失衡不应全怪芯片商,部分客户曾将价格压至三分之一,导致行业在 AI 需求爆发前投资不足。他警告短缺可能…
AI 点评 · 内存价格博弈、小红书自我复制、AI合规争议,多条行业动态折射科技产业链关键变局。
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to co…
AI 点评 · AI Coding颠覆传统开发模式,预示大厂工程师角色将根本性重塑。
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current R…
OpenSquilla 上线后数周内 GitHub star 增至数千量级
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized b…
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iterat…
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
Open-source local policy, recovery, and audit layer for explicit Codex execution.
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete re…
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo…
OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stack.
Program verifiers play a central role in training coding agents, including selecting trajectories for supervised fine-tuning (SFT) and providing rewards for reinforcement learning (RL). Standard execu…
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the r…
Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet accurate whole-matrix forecasting remains challenging when prediction must be performe…
Open-source, self-hostable alternative to Claude Tag — a Slack-style workspace where your team and its AI agents (Claude Code, Codex, GitHub Copilot, and more)…
Local coding-agent orchestrator — DAG of auto-approved, git-worktree-isolated sub-sessions across LLM providers (Claude/Kimi/Grok/DeepSeek/local). AGPL-3.0.
UmaDev: A coding agent that works like a real dev team, commanding the Claude Code / Codex / OpenCode you already use.
Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser…
Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcripts included.
The unified agent for long-horizon productivity and coding, launching with Work and Code modes. Plus, a new Vibe VS Code extension.
Related ongoing thread: DeepSeek makes the V4 Pro price discount permanent - https://news.ycombinator.com/item?id=48237663 - May 2026 (384 comments)
Related ongoing thread: DeepSeek makes the V4 Pro price discount permanent - https://news.ycombinator.com/item?id=48237663 - May 2026 (384 comments)
Introducing Mistral Medium 3.5, remote coding agents in Vibe, plus new Work mode in Le Chat for complex tasks.
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…

The artificial intelligence coding revolution comes with a catch: it's expensive. Claude Code , Anthropic's terminal-based AI agent that can write, debug, and deploy code autonomou…

Anthropic released Cowork on Monday, a new AI agent capability that extends the power of its wildly successful Claude Code tool to non-technical users — and according to company in…

Nous Research , the open-source artificial intelligence startup backed by crypto venture firm Paradigm , released a new competitive programming model on Monday that it says matches…
GITHUB HUGGING FACE MODELSCOPE DISCORD Today, we’re announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we’re excited to in…