🤖 AI Agent 研究Research
为工业部署构建通用代理需要集成在不同培训阶段获得的多种功能,但代理持续学习没有完善的配方。我们介绍了ACLArena ,这是一个用于研究和评估代理持续学习的框架,并提出了一种新的配方,该配方将高质量轨迹的离线重播与LoRA专家的路由网络相结合。
Building general-purpose agents for industrial deployment requires integrating multiple capabilities acquired at distinct training stages, yet no well-established recipe exists for Agent Continual Learning. We introduce ACLArena, a framework for studying and evaluating agent continual learning, and propose a new recipe combining offline replay over high-quality trajectories with a routed network of LoRA experts.
自主代理系统越来越依赖于可重用的技能抽象,但由于意图冲突等微妙的语义不一致,确保其正确性仍然具有挑战性。我们提出了SkillSpec ,这是一个Hoare风格的框架,将技能正确性表述为规范推理,在239个真实世界的技能中识别出763个手动确认的缺陷。
Autonomous agent systems increasingly depend on reusable skill abstractions, but ensuring their correctness remains challenging due to subtle semantic inconsistencies such as intent conflicts. We propose SkillSpec, a Hoare-style framework that formulates skill correctness as specification reasoning, identifying 763 manually confirmed defects across 239 real-world skills.
智能内存对于长期AI智能体来说变得至关重要,但现有系统依赖于自回归LLM ,这些LLM将昂贵的生成置于内存操作的关键路径上。我们引入了Jev-Mem ,这是一种System-One控制的代理内存架构,可提高内存有效性,同时将内存构建时间减少6.6倍,查询延迟减少36.7%。
Agentic memory is becoming essential for long-horizon AI agents, yet existing systems rely on autoregressive LLMs that place expensive generation on the critical path of memory operations. We introduce Jev-Mem, a System-One-controlled agentic memory architecture that improves memory effectiveness while reducing memory construction time by 6.6x and query latency by 36.7%.
工具调用LLM代理越来越多地部署在企业应用程序中,但评估需要高质量、多样化的任务数据集,由于隐私限制,这些数据集很难获得。我们提出了EdgeGen ,这是一个合成任务生成框架,可从代理规范中提取合规性规则,以生成基于数据库的边缘案例任务,从而将微调进度提高多达42%。
Tool-calling LLM agents are increasingly deployed in enterprise applications, but evaluation requires high-quality, diverse task datasets that are difficult to obtain due to privacy constraints. We propose EdgeGen, a synthetic task generation framework that extracts compliance rules from an agent's specification to generate database-grounded edge-case tasks, improving finetuning progress by up to 42%.
⭐ GitHub 热门项目GitHub Trending
【GitHub】AgentJev-0.6B -人工智能代理的快速“系统一”决策模型:向其提供任何非结构化状态(差异、跟踪、日志)和结构化问题,并在一个~ 50毫秒的向前传递中返回校准的概率分布。零输出令牌解码。(⭐ 214 )
【GitHub】AgentJev-0.6B - a fast 'System One' decision model for AI Agents: feed it any unstructured state (diffs, traces, logs) and structured questions, get calibrated probability distributions back in one ~50ms forward pass. Zero output-token decoding. (⭐ 214)
【GitHub】人工智能代理的本地AI网关,通过智能路由和设备隐私保护统一您的订阅和API提供商。(⭐ 41 )
【GitHub】A local AI gateway for AI agents, unifying your subscriptions and API providers with smart routing and on-device privacy protection. (⭐ 41)
【GitHub】Automatech的AI编码代理。使用Claude Code、Ollama和开源AI构建、测试、审查和改进软件。(⭐ 0 )
【GitHub】AI coding agent for Automatech. Building, testing, reviewing, and improving software with Claude Code, Ollama, and open-source AI. (⭐ 0)
【GitHub】求职申请工作流程:开源、人性化求职申请自动化(剧作家+克劳德代码) (⭐ 0 )
【GitHub】Job Application Workflow: open-source, human-in-the-loop job application automation (Playwright + Claude Code) (⭐ 0)
【GitHub】适用于Claude Code 2026的最佳开源Roblox Studio AI工作流程(⭐ 0 )
【GitHub】Best Open-Source Roblox Studio AI Workflow for Claude Code 2026 (⭐ 0). Arther-hup/forge-context-engine.
🚀 模型与行业动态Models & Industry
这家拥有七年历史的初创公司已经筹集了3.5亿美元的E轮融资,以推动其数据即服务的方法。
The seven-year-old startup has raised a $350 million Series E to fuel its data-as-a-service approach.
高通表示,其新的顶级芯片可以在本地运行30B混合专家模型。
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally.
Meta表示, Muse是从头开始构建的,但承认AI助手“深受OpenClaw的启发” —包括其一些工作区文件名和内容。
Meta says Muse was built from scratch, but acknowledges the AI assistant was "heavily inspired" by OpenClaw — down to some of its workspace filenames and content.
OpenAI正在推出两款新型号,据说是从与Astra相同的布料上剪下来的。
OpenAI is launching two new models, which it says are cut from the same cloth as Astra.
🔥 社区热议Community
【Lobsters】热度: 4↑ | 1 评论 | 标签: ai, vibecoding
【Lobsters】热度: 4↑ | 1 评论 | 标签: ai, vibecoding
【HN】热度: 561 分 | 360 评论
【HN】热度: 561 分 | 360 评论
【HN】热度: 425 分 | 218 评论
【HN】热度: 425 分 | 218 评论