🤖 AI Agent 研究Research
智能体在实现目标的过程中造成的副作用的基准已经存在,但HarvestBench是第一个对避免副作用定价并将该副作用命名为生物的基准。这是一个农场模拟: LLM子代理与田间动物一起驾驶两辆拖拉机通过合作玉米收获。
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field.
生成的视频的视觉流畅性并不意味着物理可靠性,仅凭标量质量分数无法指示剪辑违反的义务或失败的时刻。我们介绍VeriPhy ,这是一个可审计的物理验证系统,其中纯文本计划程序将提示编译为类型化的物理义务和静态验证的执行计划,然后再观察任何帧。
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed.
当任务具有编程检查器时,来自可验证奖励的强化学习效果很好,但大多数长期代理域都没有。我们在结果盲环境下工作,在这种环境下,无法获得地面真相的成功信号。
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available.
理解代理行为需要能够扩展到数千条轨迹的方法,并在预先构建的分类器不足的长期、通常不熟悉的任务中呈现出新的模式。我们建议将基础理论引入代理轨迹分析:这是一种具有六十年历史的社会科学定性方法,具有原则性的饱和标准和从数据到理论的可审计轨迹。
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, often unfamiliar tasks where pre-built classifiers fall short. We propose to bring grounded theory into agent trajectory analysis: a six-decade-old qualitative method from the social sciences, with a principled saturation criterion and an auditable trail from data to theory.
⭐ GitHub 热门项目GitHub Trending
【GitHub】用于基本因素研究的假设驱动的AI代理,具有证据门限验证和研究记忆。(⭐ 15 )
【GitHub】A hypothesis-driven AI agent for fundamental factor research, with evidence-gated validation and research memory. (⭐ 15)
【GitHub】适用于Claude Code、Codex、GLM和DeepSeek使用的流畅、本机Windows小部件。便携、本地优先、开源。(⭐ 1 )
【GitHub】A sleek, native Windows widget for Claude Code, Codex, GLM and DeepSeek usage. Portable, local-first and open source. (⭐ 1)
【GitHub】您拥有的编码代理的开源控制平面—并行运行Codex、Claude Code、Gemini、Cursor和ACP代理,共享历史记录、权限、交接和审计日志。(⭐ 2 )
【GitHub】Open-source control plane for coding agents you own — run Codex, Claude Code, Gemini, Cursor and ACP agents side by side with shared history, permissions, handoffs and audit logs. (⭐ 2)
【GitHub】从开发人员机器收集AI编码代理会话,清除秘密,随年龄加密,并上传到组织的存储空间。(⭐ 5 )
【GitHub】Collects AI coding-agent sessions from developer machines, scrubs secrets, encrypts with age, and uploads to your organisation's storage. (⭐ 5)
【GitHub】一种开源的反重力技能,迫使Gemini像Claude Models一样编码。它让人工智能读取您的文件,计划其步骤,并在说完成之前实际测试代码。(⭐ 1 )
【GitHub】An open-source Antigravity skill that forces Gemini to code like Claude Models. It makes the AI read your files, plan its steps, and actually test the code before saying it's done. (⭐ 1)
🚀 模型与行业动态Models & Industry
上个月,一位Claude用户注意到他的帐户正在使用代币,即使他没有工作。此后, Anthropic向用户发出了有关黑客的警告。
Last month, a Claude user noticed his account was consuming tokens even though he wasn't working. Anthropic has since warned users about hackers.
Cognition的估值倍数高于Cursor在出售给SpaceX之前的估值倍数。
Cognition's valuation multiple is higher than Cursor's was before selling to SpaceX.
Meta的新个人人工智能代理Muse希望访问用户的电子邮件、日历、付款、医疗服务等,这使该公司成为最大的消费者人工智能赌注,但也是人们是否仍然信任Meta提供数据的主要考验。
Meta's new personal AI agent Muse wants access to users' email, calendars, payments, health services, and more — making the company's biggest consumer AI bet yet a major test of whether people still trust Meta with their data.
谷歌云通过埃森哲扩展其企业人工智能,押注于前沿部署的工程师,以推动采用并克服部署瓶颈。
Google Cloud expands its enterprise AI push with Accenture, betting on forward-deployed engineers to drive adoption and overcome deployment bottlenecks.
🔥 社区热议Community
【HN】热度: 93 分 | 48 评论
【HN】热度: 93 分 | 48 评论
【HN】热度: 323 分 | 256 评论
【HN】热度: 323 分 | 256 评论