🤖 AI Agent 研究Research
受过强化学习( RL )训练的多回合代理在每个轨迹上获得单个标量奖励,这激励自我策略提取( OPD ) ,由具有特权任务技能的自学老师提供密集的令牌级监督,让没有技能的学生将它们内化。因此,我们建议RetireOPD (自我退休随机蒸馏) ,它首先优化具有环境奖励的解耦,技能条件的教师,然后与R共同培训无技能的学生
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.
随着编码代理从监督代码完成转变为无人值守、全天候的探索,他们的工作从孤立的预测扩展到推理、工具使用和反馈的长轨迹。因此,令牌效率对于扩展递归自我改进变得非常重要。
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement.
编码线束塑造了自主编码代理如何将模型功能转化为长远的软件工程性能,但现有的工作通常将线束评估为单片系统,使得单个组件的有效性尚不清楚。在SWE-Bench Verified和Terminal-Bench 2.1上评估的四个模型中,我们评估了176个匹配的设置,涵盖五个上下文管理策略、四个上下文窗口预算以及有针对性的规划和行动空间。
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space.
客服代表轨迹记录了客服代表的工作内容以及接下来会发生什么。然而,标准监督微调( SFT )仅将损失应用于座席撰写的操作令牌,使用环境观测作为上下文,而不是作为预测目标。
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets.
⭐ GitHub 热门项目GitHub Trending
【GitHub】适用于Linux的非官方开源Claude桌面客户端。Tauri 2 + React + Rust ,由Claude Code CLI提供支持—无API密钥,无遥测。(⭐ 0 )
【GitHub】An unofficial, open-source Claude desktop client for Linux. Tauri 2 + React + Rust, powered by the Claude Code CLI — no API key, no telemetry. (⭐ 0)
【GitHub】针对Codex和Claude Code的由Jev提供支持的AI代理技能:代码审查、测试、调试、提取和QA证据。开源,麻省理工学院。(⭐ 0 )
【GitHub】Jev-powered AI agent skills for Codex and Claude Code: code review, testing, debugging, extraction and QA evidence. Open source, MIT. (⭐ 0)
【GitHub】用于多语言政策协调和合规治理的闭环、多代理大语言模型( LLM )评估框架。此Python包自动检测“需求dr (⭐ 1 )
【GitHub】A closed-loop, multi-agent Large Language Model (LLM) evaluation framework for multilingual policy alignment and compliance governance. This Python package automates the detection of "requirements dr (⭐ 1)
【GitHub】开源Claude代码修改(⭐ 0 )
【GitHub】Open Source Claude Code Modifications (TPRM-Modified) — a customized open-source build of the Claude Code coding agent, published on GitHub. (⭐ 0)
🚀 模型与行业动态Models & Industry
世界上的每一个模特都坐在一堆现金和大量的嗡嗡声上,但祝你好运,从创始人到自己的数据提供商,任何人都可以告诉你他们实际在建造什么。
Everyone in the world-models space is sitting on a pile of cash and a ton of buzz, but good luck getting anyone — from the founders to their own data suppliers — to tell you what they're actually building.
在股票方面,我们讨论了AI高管是否真的想放慢脚步。
On Equity, we debated whether Ai executives are serious about wanting to slow down.
谷歌表示, Gemini通过立即终止每次黑客攻击“采取了适当的行动”。
Google said Gemini had "acted appropriately" by ending each hack immediately. Google's Gemini is the latest AI model to hack other companies.
人工智能领导者一直承诺,人工智能是治愈人类疾病的关键。人类学研究人员也一直在警告说,人工智能可能会杀死我们所有人。
AI leaders have been promising that AI is the key to curing human disease. Anthropic researchers have also been warning that AI might kill us all.
🔥 社区热议Community
【HN】热度: 607 分 | 320 评论
【HN】热度: 607 分 | 320 评论
【HN】热度: 11 分 | 2 评论
【HN】热度: 11 分 | 2 评论