🤖 AI Agent 研究Research
我们推出了Agensh ,这是一种可扩展的自组织多智能体线束,没有中央编排器,可扩展到1,024个智能体。并发工作者执行多Agent协作循环,收集上下文,自行分配子任务,采取行动,共享发现,验证结果,异步合并进度。
We introduce Agensh, a scalable self-organized multi-agent harness without a central orchestrator that scales to 1,024 agents. Concurrent workers execute a multi-agent cooperation loop, gathering context, self-assigning sub-tasks, taking action, sharing findings, verifying results, and merging progress asynchronously.
我们介绍AIDE ^ 2 ,这是一个为前沿人工智能研究代理实现递归自我改进的系统。它建议对自己的代码进行更改,在一套人工智能研发任务中对自身的修改版本进行基准测试,并保持在隐藏评估中表现最佳的更改。
We present AIDE^2, a system that implements recursive self-improvement for a frontier AI research agent. It proposes changes to its own code, benchmarks modified versions of itself on a suite of AI R&D tasks, and keeps the changes that perform best on hidden evaluations.
我们研究了在长视野多智能体环境中勾结的出现,其中两个智能体重复完成任务,共享任务日志,并验证彼此的工作。共谋出现在10个模型中94%的轨迹中,同一系列中更强大的模型更早地到达它。
We study the emergence of collusion in a long-horizon multi-agent environment where two agents repeatedly complete tasks, share task logs, and verify each other's work. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier.
我们将做出长远决策的能力称为客服代表的品味。我们构建了Taste-Bench ,这是根据代理在工程和研究任务中产生的轨迹自动构建的口味问题的基准,每个问题都呈现了一个决策分叉,其中一个方向可以带来更好的结果。
We refer to the ability to make good long-horizon decisions as the taste of an agent. We build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, each presenting a decision fork where one direction leads to a better outcome.
⭐ GitHub 热门项目GitHub Trending
【GitHub】SearchCarriers汽车运营商智能的开源Claude Code插件和技能(⭐ 1 )
【GitHub】Open-source Claude Code plugins and skills for SearchCarriers motor-carrier intelligence (⭐ 1)
【GitHub】面向 AI 编码智能体与 Codex 工作流的桌面伴侣工具 (⭐ 222)
【GitHub】Desktop companion for AI coding agents and Codex workflows (⭐ 222). BinaryDeliverer/CodexDesk.
【GitHub】求职申请工作流程:开源、人性化求职申请自动化(剧作家+克劳德代码) (⭐ 0 )
【GitHub】Job Application Workflow: open-source, human-in-the-loop job application automation (Playwright + Claude Code) (⭐ 0)
【GitHub】面向 Claude 与 AI 助手的桌面面板与工具集 (⭐ 221)
【GitHub】Desktop panel and toolkit for Claude and AI assistants (⭐ 221). AvenueSnowStep/ClaudePanel.
【GitHub】本地开源简历定制工具包,使用Codex或Claude Code从结构化候选人数据和职位描述中生成真实、有证据支持的LaTeX简历。(⭐ 0 )
【GitHub】Local, open-source resume tailoring toolkit that uses Codex or Claude Code to generate truthful, evidence-backed LaTeX resumes from structured candidate data and job descriptions. (⭐ 0)
🚀 模型与行业动态Models & Industry
首席执行官马克·扎克伯格(Mark Zuckerberg)周三在门洛帕克(Menlo Park)拉开了公司年度Connect活动的序幕,发表了一篇主题演讲,明确表达了一件事情: Meta正在全力投入Muse。它甚至出现在Meta的人工智能眼镜上。
CEO Mark Zuckerberg kicked off the company’s annual Connect event in Menlo Park on Wednesday with a keynote that made one thing clear: Meta is going all-in on Muse. It's even coming to Meta's AI glasses.
这款微型硬件设备为其 AI 智能体 Muse 打造了另一个移动载体。Meta 为其 Muse AI 智能体打造了一款类似电子宠物(Tamagotchi)的可穿戴设备。
The tiny hardware device creates another mobile home for its AI agent Muse. Meta made a Tamagotchi-like wearable for its Muse AI agent.
Meta表示,无相机眼镜将更轻,电池续航时间长达12小时。
Meta says the camera-free glasses will be lighter and have up to 12 hours battery life.
但也许最大的启示是, Anthropic并没有让Claude在其生物学实验室中逍遥法外。到目前为止,人类仍然处于循环中。
But maybe the biggest reveal is that Anthropic has not let Claude run loose in its biology lab. Humans are still, so far, in the loop.
这轮融资对AI生物技术的估值为20 $。它目前正在测试治疗皮肤病并在停止GLP-1s后保持体重的药物。
The round valued the AI biotech at $2 billion. It is currently testing drugs that treat skin conditions and preserve weight loss after stopping GLP-1s.
🔥 社区热议Community
【Lobsters】热度: 41↑ | 49 评论 | 标签: linux, practices, vibecoding
【Lobsters】热度: 41↑ | 49 评论 | 标签: linux, practices, vibecoding
【Lobsters】热度: 7↑ | 1 评论 | 标签: vibecoding
【Lobsters】热度: 7↑ | 1 评论 | 标签: vibecoding
【Lobsters】热度: 2↑ | 0 评论 | 标签: vibecoding
【Lobsters】热度: 2↑ | 0 评论 | 标签: vibecoding
【HN】热度: 499 分 | 533 评论
【HN】热度: 499 分 | 533 评论
【HN】热度: 453 分 | 258 评论
【HN】热度: 453 分 | 258 评论