02 / 写作 · Agent 与 Skill

Hermes Agent 的 Skill 自进化机制

一、什么是"Skill 自进化机制"

学术界通常将 self-evolving agent 理解为:Agent 能够基于与环境交互过程中获得的经验或反馈,持续调整自身,使后续任务表现得到改善。理解这类机制,可以拆成三个问题:进化什么、什么时候进化、依靠什么信号进化。

  • 进化什么:可以是上下文与记忆、工具、Skill,甚至 Agent 的整体架构
  • 什么时候进化:可以在任务执行过程中实时调整,也可以在任务结束后基于复盘结果更新,或由固定周期、经验积累阈值等条件触发离线进化
  • 依靠什么信号进化:自进化可以依赖奖励信号、成功/失败反馈、模仿已有行为、群体演化等不同机制

基于这个框架,本文所说的 Skill 自进化机制可以进一步定义为:

Agent 在真实使用过程中,根据任务经验自主识别值得复用的知识,并据此创建、修改、合并或淘汰 Skill,使 Skill 库随着使用经验持续改善

基于上述定义,一个完整的 Skill 自进化机制至少需要回答五个问题。后文也会按照这五个环节评价 Hermes:

环节

核心问题

触发

什么时候应该触发一次任务经验沉淀流程?

生成

从完整任务经历中应该抽取什么内容?如何把一次性的经验转化成真正可复用、边界清晰的 Skill?

使用

下一次遇到相似任务时,Agent 能否在正确的时机找到这个 Skill,并且实际使用后带来更好的结果?

验证

Skill 实际使用之后,又如何判断它是否有效、是否导致失败,以及失败后如何修改?

治理

如何处理重复、冲突、过期、低价值和长期不用的 Skill?


二、Hermes 为什么要做 skill 自进化

Hermes 将自己定位为一个能够长期陪伴用户、“The agent that grows with you” 的个人 Agent。对 Hermes 来说,这意味着它不能只依赖基础模型本身的能力,而必须能够把长期使用过程中获得的经验保留下来,并转化为后续可复用的能力。否则,无论用户使用多少次,Hermes 每次面对相似任务仍然需要重新理解、重新探索,也就谈不上真正“随着用户一起成长”。

Hermes 已经通过 Memory 解决了一部分长期积累问题,但 Memory 主要保存的是“用户是谁、偏好什么、发生过什么”这类事实性信息。对于“某类任务应该怎样完成”这类程序性经验,Hermes 需要依赖 Skill。比如一次任务中发现了更有效的执行流程、工具组合、判断规则或避坑方法,只有把这些经验沉淀进 Skill,Hermes 才能在之后的相似任务中直接复用,而不是每次重新推理。

但 Hermes 面临的并不是一个静态环境。随着使用次数增加,它会持续遇到新的成功经验、失败案例和用户纠正;用户习惯、工具能力和外部环境本身也可能发生变化。因此,Hermes 不能只在某个时刻生成一份 Skill 并长期不变,而需要根据后续使用结果不断修改已有 Skill,必要时创建新的 Skill,或者淘汰已经失效的 Skill。这样才能形成 任务经验 → Skill 沉淀 → 后续复用 → 新反馈 → Skill 修正 的持续闭环,以达到进化 Skill 能力的目的。


三、拆解 Hermes 的 skill 自进化机制

3.1 总览

HERMES / SKILL

Hermes Skill 自进化机制总览

任务执行 → 发现值得学习的经验 → 创建或修改 Skill → 后续继续使用 → 定期整理和治理。

任务执行

每一轮对话:模型调用工具、推进任务、交付回复。

01 / 四条学习路径

自动

前台学习

学习规则写入主 Agent 的 System Prompt。模型在执行任务时识别可复用的方法,直接调用 skill_manage 创建或修改 Skill。

自动

后台审查

工具调用迭代累计达到 10 步(可跨轮累积),在最终回复交付后启动独立 /fork 复盘;实际存入后计数清零。/fork 只能看不能改,只允许往记忆和技能库写;根据预定义说明提取“以后怎么做”和“为什么”。

用户触发

/learn

用户说“把这次经验学下来”。阅读相关材料后检查已有 Skill:有对应能力就更新,没有才新建,按统一格式保存。

用户触发

/refine

用户要求重新检查刚发生的对话和执行过程,判断是否需要更新 Memory 或 Skill。Memory 分支直接进入记忆;Skill 分支进入下方写入流程。

02 / 唯一的 Skill 写入口

skill_manage

八道关,按顺序执行。Memory 更新不经过此 Skill 写入口。

  1. 先看这次请求本身写得对不对。
  2. 检查名字、格式,简介不许超过 60 字。
  3. 想改旧技能?必须先完整读一遍才准动。
  4. 安全扫描(默认不开)。
  5. 等人点头确认(默认不开)。
  6. 真正写进文件。
  7. 体检一遍,只提意见,不拒绝写入的操作。
  8. 记一笔账:改动前后各存一份,随时能退回。

03 / Skill 库

写入的终点,也是使用的起点

三层结构;随需要逐步加载。

索引层

常驻系统提示

每条 = 名称 + ≤60 字符描述,每一轮都在上下文里。

正文层

命中时加载

SKILL.md 正文,被选中时才全量加载进上下文。

支持层

按需触达

references/ 读进上下文;templates/ 拷贝后修改;scripts/ 直接执行、不进上下文。

04 / 每一轮的使用回路

从 Skill 回到任务

  1. 每轮把索引注入系统提示。
  2. 模型判断是否命中。
  3. skill_view 全量加载 SKILL.md,使用计数 +1。
  4. 按需读 references、抄 templates、跑 scripts。
  5. 继续执行任务。
↶ 返回任务执行 · 下一轮继续使用

05 / 定期治理

Curator

  • 平时不参与,定期才跑;闲置满 2 小时才动手,大约一周一次。
  • 按上次触发时间算冷热:30 天没碰标为闲置,90 天没碰收进储藏室(随时能取回)。
  • 只整理系统自己攒的和用户明确交给它管的;外面装来的一律不碰。
  • 把内容重复的技能合并成一个(默认不开)。
  • 从不删东西;每次动手前先把整个技能库完整备份一份。
↶ 返回 Skill 库 · 计龄 / 标记 / 归档 / 快照

Hermes 的 Skill 自进化不是由单一流程完成,而是由前台学习、后台审查、/learn、/refine 和 Curator 几套机制共同组成。可以把整体流程理解为:

任务执行 → 发现值得学习的经验 → 创建或修改 Skill → 后续继续使用 → 定期整理和治理

3.2 自动学习:前台学习

Hermes 的前台学习是直接把“发现可复用经验后进行沉淀”的规则写入主 Agent 的 System Prompt,也就意味着前台学习的内容全靠模型自身的判断。当 Agent 在任务中摸索出一个有价值的工作流、发现已有 Skill 存在问题,或者获得了值得以后复用的经验时,它可以直接调用 skill_manage 创建或修改 Skill。因此,前台学习本质上是一种 “将学习意识直接嵌入主 Agent” 的机制,对于模型的判断能力要求极强,而且实际上规则设定的极为模糊,并没有一个确定的标准说明哪些经验是的确有价值的,这一步的粗糙设定让 Hermes 的前台学习实际上产出的 Skill 的质量让人疑虑。

3.3 自动学习:后台审查

除了主动学习外,Hermes 还设计了一套独立的后台审查机制(background review),用于定期回顾已经发生的任务轨迹,从中发现主 Agent 当时没有主动沉淀的经验,主要是通过结合用户反馈、实际任务过程、已有 Skill 的表现来判断哪些经验值得沉淀。它的触发是按照 Agent 实际进行了多少次“工具调用迭代”计算。当累计次数达到阈值后,Hermes 并不会立刻打断当前任务。它会等主 Agent 完成本轮任务并把最终回复交付给用户之后,再复制当前会话的执行轨迹,启动一个独立的后台 Review Agent 来重点寻找对应的学习信号。

它重点寻找四类学习信号:

  • correct:用户纠正 Agent 原来的做法,这类信号优先级很高,会优化对 Skill 进行改动,因为用户已经明确表示“现在这套做法有问题”;
  • improve:任务过程中找到了一种更高效、更可靠的新方法;
  • solve:遇到实际问题或坑之后,经过探索最终找到稳定解决方案;
  • renew:实际使用某个 Skill 后发现其内容错误、缺少步骤或者已经过时。

发现学习信号之后,Hermes 的 Review Prompt 明确要求按照优先级处理:

优先修改当前已经加载的 Skill → 再寻找已有的同类 Skill 并更新 → 必要时给已有 Skill 增加支持文件 → 只有确实不存在合适 Skill 时,才创建新的 class-level Skill

这样做的目的,是尽量把新的经验吸收到已有能力中,而不是让每一次任务都产生一个新的 Skill,导致 Skill 库快速膨胀。

相比于前台的主动学习方法而言,后台最大的提升是:它更多依赖外部可观察信号,而不是单纯依赖模型觉得任务是否复杂;同时对“什么不能存”定义得非常细。 但它依然没有一套真正量化的价值评分,最终还是由 LLM 根据这些规则做判断。

3.4 用户主动触发入口:/learn 与 /refine

Hermes 还提供了两个由用户主动触发的学习入口:/learn 和 /refine。它们解决的是同一个问题:当用户明确知道“这段经验值得学”或“现有经验需要重新检查”时,可以直接要求系统进行学习,而不必等待自动机制触发。

其中,/learn 更偏向把当前经验正式沉淀成 Skill。用户可以基于当前对话、某份材料或一次任务过程,直接要求 Hermes“把这次经验学下来”。系统会读取相关材料,理解其中真正值得复用的方法,再检查现有 Skill:如果已经有对应能力,就更新原有 Skill;如果没有,则新建 Skill,最后按统一格式整理并保存。/refine 则更偏向复盘和修正已有经验。它主要用于重新检查刚刚发生的对话和执行过程,判断其中是否存在应该更新的 Memory 或 Skill。

3.5 长期治理:Curator

当 Skill 持续通过前台学习、后台审查以及用户主动触发不断积累后,新的问题就会出现:Skill 库会越来越大,也可能逐渐出现重复、过期、低价值或相互重叠的内容。Curator 负责的正是这一层问题。它是负责长期维护已经存在的 Skill 库,让其中的内容保持可用、精简和相对一致。它的处理大致分成两个阶段:

  • 第一阶段是基于规则和状态信息的机械处理。系统会根据 Skill 的使用情况、时间等信号,识别长期没有被使用或已经进入低活跃状态的 Skill,并进行标记、降级或归档。
  • 第二阶段是基于 Agent 的内容审查。辅助 Agent 会进一步检查 Skill 本身的内容关系,判断它们应该继续保留、进行修补、与其他 Skill 合并,还是彻底归档。

Curator 主要处理的是 Skill 的四类问题:保留、修补、合并、归档。 Curator 对 Skill 的判断主要依赖 Skill 文本本身及少量使用元数据,而不会重新结合原始任务轨迹、实际成功/失败结果或用户纠正来验证这些 Skill 的真实效果。因此,它更擅长发现“文本上重复、重叠或过时”的内容,却无法可靠判断两个 Skill 在真实执行中是否真的等价、哪个效果更好。

3.6 局限性

3.6.1 经验主要来自单条轨迹,缺少跨轨迹比较

Hermes 的后台审查主要依据当前任务形成的一条执行轨迹。它可以总结“这次是怎么完成的”“这次踩了什么坑”,却无法仅凭这条轨迹判断某项经验是普遍规律,还是只对当前任务成立。系统也没有把多个相似任务的成功与失败轨迹放在一起比较,因此难以发现其他路径为什么更好或更差。

3.6.2 规定了什么经验值得存,但没有验证模型能否判断准确

针对经验筛选,Hermes 现有做法主要是在这类问题发生后补充规则,再用测试确保规则不被删除,这能修补已经发现的问题,但没有证明模型在其他场景下也能正确筛选。测试目前没有设置机制去判断模型的执行是否准确。因此,目前无法回答两个直接影响技能库质量的问题:不值得保存的经验,有多少被存了进去?真正有用的经验,又有多少被漏掉了?导致随着经验不断积累,低价值或错误内容仍可能进入技能库。

3.6.3 Skill 修改后,没有验收它是否真的改善了任务表现

对于需要进行修改的 Skill 而言,Hermes 会检查格式和结构,要求修改前读取原文,也会保留版本记录。后续任务发现错误时,模型仍有机会修正。因此,Hermes 具有一定针对于 Skill 修改质量的核查,但是缺少的是对修改效果的量化验收。 Hermes 并不会用同一组任务比较“不用 Skill、使用旧版、使用新版”的结果,再据此决定是否接受修改,导致 Skill 的更新,更多是机制上的一厢情愿。 这意味着每次自动更新都只能算一次改进尝试,尚不能算已经验证的提升。

3.6.4 Skill 的触发和使用由模型判断,可能漏用,也可能滥用

Hermes 可以根据工具是否可用等条件筛选 Skill,但“当前任务是否需要这个 Skill”“它是否比其他 Skill 更合适”,仍主要由模型判断。Hermes 通过强指令要求模型加载匹配的 Skill,并倾向于在不确定时多加载。这有助于减少漏用,却也可能让模型遵循不适合当前任务的 Skill,造成 Skill 的滥用。

因此,Agent 除了积累“怎么完成任务”的经验,还需要积累“什么时候该用哪个 Skill、什么时候不该用”的经验。Agent 可以选错,但不应在相似情境下反复犯同一种错误。这需要审查实际执行轨迹,而不只是最终结果:模型当时为什么选择这个 Skill,哪些步骤真正有帮助,哪些步骤造成了绕路或返工。只有把这些判断沉淀为后续选用的依据,Agent 才能从拥有更多 Skill,进一步发展到逐渐把 Skill 用对。

3.6.5 Curator 按使用情况治理技能库,而不是按实际价值治理

Hermes 的 Curator 主要根据最近一次使用、查看或修改的时间决定 Skill 是否保持活跃。这个机制能够清理长期无人触碰的内容,却不能判断一个 Skill 是否正确、是否仍然有效,以及它是否真正改善了任务表现。在缺少效果评估的情况下,Curator 治理主要解决的是数量和活跃度问题,而不是质量问题。要转向价值驱动,系统还需要结合任务成功率、返工情况、执行成本、适用范围以及新旧版本对照结果,判断一个 Skill 应该保留、修改、合并还是归档。


四、更好、更完备的方向

4.1 从单条轨迹转向多轨迹归纳

Hermes 当前主要根据一次任务的执行过程生成 Skill。它可以总结“这次怎么完成了”,却难以判断某项经验是普遍规律,还是只适用于当前任务。一次偶然成功、不必要的步骤或环境特例,都可能被写进 Skill。

更可靠的做法,是为同类任务收集多条成功和失败轨迹,比较不同路径的差异,再从中归纳可迁移的经验。Trace2Skill 采用的就是这条路线:并行分析一个广泛的轨迹池,提炼针对性补丁,再合并成没有冲突的 Skill。它说明,经验的价值主要来自对多次执行的比较,而不是对单次成功的复述。

4.2 把独立验证变成入库门槛

生成 Skill 只是一次改进尝试,不代表任务表现已经提高。更完备的机制必须把验证设为硬性门槛:带代码的 Skill 运行测试,纯文本 Skill 在新会话或沙箱中复现任务,验证不通过就不能入库。CoEvoSkills 的消融实验提供了直接证据:去掉代理验证器后,SkillsBench 通过率从 71.1% 降至 41.1%。SkillLearnBench 也发现,在没有新增外部信号时,反复自我改写不会稳定变好,甚至可能造成递归漂移。真正有效的改进依赖测试结果、执行反馈、老师模型或独立验证器等外部信号。因此,验证者不能只是原模型在原上下文里再次检查自己的作品。它需要在隔离环境中重新执行任务,并对失败原因给出具体诊断。

4.3 用反事实对照判断 Skill 是否有贡献

调用次数只能证明一个 Skill 被加载过,不能证明它改善了任务表现。高使用率甚至可能来自过度触发或路由错误。

真正有意义的指标,是同类任务在“不使用 Skill、使用旧版、使用新版”三种情况下,成功率、成本、延迟和回归率分别有什么变化。SkillAudit 已经采用了类似方法:让同一个任务分别在使用和不使用候选 Skill 的条件下执行,通过比较两条轨迹,判断 Skill 究竟提供了帮助,还是引入了干扰。

这种反事实对照也应该成为 Skill 修改的验收标准。只有新版稳定优于旧版,修改才能被视为已经验证的提升。

4.4 不仅要写出 Skill,还要确保它被正确使用

一个有效的 Skill,也可能因为路由失败而没有被调用;一个不适合当前任务的 Skill,也可能因为描述过宽或强制加载而被滥用。因此,系统不仅要评估 Skill 的内容,还要审查它在运行时是否被正确选择和执行。

更完整的使用机制应包括语义检索、选择歧义处理和采用检测:该用时能够找到,找到后真正执行,不该用时不会强行套用。SkillEvolver 的“静默失效”审计,就是专门检查“Skill 内容看起来正确,但运行时从未被调用”的情况。

Agent 需要积累的不只是“怎么完成任务”的经验,还包括“什么时候该用哪个 Skill、什么时候不该用”的经验。这要求系统审查实际执行轨迹,判断哪些步骤提供了帮助,哪些步骤造成了绕路或返工。

4.5 把“自动提出”和“自动生效”分开

运行时 Agent 仍然可以提出新 Skill 或修改建议,但这些建议应该先进入候选队列,而不是直接影响后续任务。候选需要经过测试、独立审计和版本检查,再决定是否生效。自动化程度也不应一刀切。低风险、高频、可重放的任务,例如数据清洗、代码检查和文档格式化,适合自动进行多轨迹验证,验证通过后自动生效。有外部副作用的任务,例如发消息、下单和修改云端数据,重跑本身就可能造成损失,因此需要沙箱、dry-run、人工审批和渐进发布。无论自动化程度多高,系统都需要保留候选、版本、回退、冲突处理和人工审批入口,并能够解释某次行为为什么发生了变化。

4.6 最终要验证的是产品价值

Skill 自进化不能只衡量“生成了多少 Skill”或“调用了多少次”。它还需要计算单位经济:一次审查和验证需要多少 token、多少模型调用和多少费用;这个 Skill 预计会被复用多少次;每次能节省多少时间;错误使用一次又会造成多大损失。只有长期收益高于生成、验证和治理成本,自动进化才有实际价值。否则,系统可能只是越来越会编写 Skill,却没有让任务完成得更准确、更稳定或更便宜。

02 / WRITING · AGENTS & SKILLS

The Skill Self-Evolution Mechanism in Hermes Agent

I. What Is a “Skill Self-Evolution Mechanism”?

In academic research, a self-evolving agent is generally understood as an agent that continually adjusts itself based on experience or feedback gained through interaction with its environment, improving its performance on subsequent tasks. To understand these mechanisms, we can ask three questions: what evolves, when it evolves, and what signals drive that evolution.

  • What evolves: context and memory, tools, Skills, or even the agent’s overall architecture.
  • When it evolves: adjustments can happen in real time during a task, after a task based on a retrospective, or through offline evolution triggered by a fixed schedule, an experience threshold, or other conditions.
  • What signals drive evolution: self-evolution may rely on rewards, success or failure feedback, imitation of existing behavior, population-based evolution, or other mechanisms.

Within this framework, the Skill self-evolution mechanism discussed in this article can be defined more specifically:

During real-world use, an agent independently identifies reusable knowledge from task experience, then creates, modifies, merges, or retires Skills accordingly, allowing the Skill library to improve continually as experience accumulates.

Under this definition, a complete Skill self-evolution mechanism must answer at least five questions. These five stages also provide the framework for evaluating Hermes below:

Stage

Core question

Triggering

When should a process for capturing task experience be triggered?

Generation

What should be extracted from a complete task experience? How can a one-off experience become a genuinely reusable Skill with clearly defined boundaries?

Use

When a similar task comes along, can the agent find the Skill at the right time, and does using it actually produce a better result?

Validation

After a Skill is used, how can we determine whether it was effective or caused a failure, and how should it be revised after failure?

Governance

How should duplicate, conflicting, outdated, low-value, and long-unused Skills be handled?


II. Why Hermes Needs Skill Self-Evolution

Hermes positions itself as a personal agent that accompanies its user over time: “The agent that grows with you.” For Hermes, this means it cannot rely solely on the capabilities of its underlying model. It must retain experience gained through long-term use and turn that experience into capabilities it can reuse. Otherwise, no matter how often someone uses it, Hermes would still need to understand and explore similar tasks from scratch each time. That would hardly amount to growing alongside the user.

Hermes already addresses some of this long-term accumulation through Memory, but Memory primarily stores factual information: who the user is, what they prefer, and what has happened. Procedural experience—how to complete a particular kind of task—requires Skills. For example, a task may reveal a more effective workflow, combination of tools, decision rule, or way to avoid a pitfall. Only by capturing this experience in a Skill can Hermes reuse it directly on similar tasks rather than reasoning everything out again.

Yet Hermes does not operate in a static environment. With continued use, it encounters new successes, failures, and user corrections; user habits, tool capabilities, and the external environment may also change. Hermes therefore cannot generate a Skill once and leave it unchanged indefinitely. It needs to revise existing Skills based on subsequent results, create new ones when necessary, and retire Skills that no longer work. This creates a continuous loop—task experience → capture in a Skill → subsequent reuse → new feedback → Skill revision—through which Skill capabilities can evolve.


III. Unpacking Hermes’s Skill Self-Evolution Mechanism

3.1 Overview

HERMES / SKILL

Hermes Skill Self-Evolution at a Glance

Execute a task → identify reusable experience → create or modify a Skill → reuse it later → periodically organize and govern the library.

Task execution

Every conversation turn: the model calls tools, advances the task, and delivers a reply.

01 / FOUR LEARNING PATHS

Automatic

Foreground learning

Learning rules live in the main agent’s System Prompt. During a task, the model identifies reusable methods and calls skill_manage directly to create or modify a Skill.

Automatic

Background review

When accumulated tool-call iterations reach 10 steps (carried across turns), an independent /fork reviews the trajectory after the final reply. The count resets after saving. The fork is read-only except for Memory and Skill writes; predefined instructions guide it to extract how and why to act next time.

User-triggered

/learn

The user asks to capture this experience. Read the material, then check existing Skills: update a matching capability, or create one if none exists, and save it in the standard format.

User-triggered

/refine

Revisit the recent conversation and execution process to decide whether Memory or a Skill needs updating. Memory updates go to Memory; Skill updates enter the write flow below.

02 / THE SINGLE SKILL WRITE ENTRY POINT

skill_manage

Eight checks and actions, in order. Memory updates do not pass through this Skill write entry point.

  1. Validate the request itself.
  2. Check the name and format; the description must not exceed 60 characters.
  3. Before editing an existing Skill, read it in full.
  4. Run a security scan (off by default).
  5. Wait for human approval (off by default).
  6. Write the file.
  7. Run a health check; it advises but does not reject the write.
  8. Record snapshots before and after the change so it can be rolled back.

03 / SKILL LIBRARY

Where writes end and reuse begins

Three layers, loaded progressively as needed.

Index

Always in the system prompt

Each entry = name + description of ≤60 characters, present in context every turn.

Body

Loaded on a match

The full SKILL.md is loaded into context only when selected.

Support

Accessed on demand

Read references/ into context; copy and edit templates/; execute scripts/ directly without loading them into context.

04 / THE USE LOOP, EVERY TURN

From Skills back to tasks

  1. Inject the index into the system prompt each turn.
  2. The model judges whether a Skill matches.
  3. skill_view loads the full SKILL.md; increment use count by 1.
  4. Read references, copy templates, and run scripts as needed.
  5. Continue executing the task.
↶ Back to task execution · reuse on the next turn

05 / PERIODIC GOVERNANCE

Curator

  • Runs periodically, not during normal work; requires 2 hours of inactivity and runs roughly weekly.
  • Age by last activation: mark idle after 30 days untouched; archive after 90 days, with restoration available.
  • Manage only system-generated Skills and those explicitly entrusted by the user; leave externally installed Skills alone.
  • Merge duplicate Skills (off by default).
  • Never delete; make a complete library backup before every operation.
↶ Back to the Skill library · age / mark / archive / snapshot

Skill self-evolution in Hermes is not a single process. It combines foreground learning, background review, /learn, /refine, and Curator. The overall flow can be understood as:

Execute a task → identify experience worth learning from → create or modify a Skill → reuse it later → periodically organize and govern the library

3.2 Automatic Learning: Foreground Learning

Hermes implements foreground learning by putting the rule “capture reusable experience when you discover it” directly into the main agent’s System Prompt. This means the content of foreground learning depends entirely on the model’s own judgment. When an agent discovers a valuable workflow, finds a problem in an existing Skill, or gains experience worth reusing, it can call skill_manage directly to create or modify a Skill. Foreground learning thus embeds the awareness of learning directly into the main agent. This places very high demands on the model’s judgment. In practice, the rules are extremely vague: there is no definite standard for what experience is genuinely valuable. Such a rough setup raises doubts about the quality of Skills produced through Hermes’s foreground learning.

3.3 Automatic Learning: Background Review

Alongside active learning, Hermes provides an independent background review mechanism. It periodically revisits past task trajectories to identify experience that the main agent did not capture at the time, primarily considering user feedback, the actual task process, and the performance of existing Skills. Its trigger counts the agent’s actual tool-call iterations. Once the accumulated count reaches the threshold, Hermes does not immediately interrupt the task. It waits until the main agent finishes the current task and delivers its final reply, then copies the current session’s execution trajectory and starts an independent background Review Agent to look specifically for learning signals.

It looks for four types of learning signal:

  • correct: the user corrects the agent’s earlier approach. This is a high-priority signal that favors changing the Skill, because the user has explicitly indicated that the current approach is flawed;
  • improve: a more efficient or reliable method is discovered during the task;
  • solve: after encountering a real problem or pitfall, exploration leads to a stable solution;
  • renew: using a Skill reveals errors, missing steps, or outdated content.

After finding learning signals, Hermes’s Review Prompt explicitly requires the following order of priority:

First modify a currently loaded Skill → then find and update an existing Skill of the same class → add supporting files to an existing Skill when necessary → create a new class-level Skill only when no suitable Skill exists

The purpose is to absorb new experience into existing capabilities wherever possible, rather than generating a new Skill for every task and rapidly inflating the library.

Compared with foreground active learning, the main improvement is that background review relies more on externally observable signals than on the model simply deciding whether a task feels complex. It also defines what must not be saved in considerable detail. However, it still has no genuinely quantitative value score; the LLM ultimately makes the judgment based on these rules.

3.4 User-Triggered Entry Points: /learn and /refine

Hermes also offers two user-triggered learning entry points: /learn and /refine. Both address the same need: when a user knows that an experience is worth learning from or that existing experience should be re-examined, they can ask the system to learn directly instead of waiting for an automatic trigger.

/learn focuses on formally capturing current experience in a Skill. Based on the current conversation, a document, or a task process, the user can ask Hermes to “learn from this experience.” The system reads the material, identifies genuinely reusable methods, and checks existing Skills. If the corresponding capability already exists, it updates that Skill; otherwise, it creates one, then organizes and saves it in the standard format. /refine focuses more on reviewing and correcting existing experience. It revisits the conversation and execution process that just took place to determine whether any Memory or Skill should be updated.

3.5 Long-Term Governance: Curator

As foreground learning, background review, and user-triggered learning continue to accumulate Skills, new problems emerge. The library grows and may develop duplicate, outdated, low-value, or overlapping content. Curator addresses this layer of the problem. It maintains the existing Skill library over time, keeping it usable, concise, and reasonably consistent. Its work falls roughly into two stages:

  • The first is mechanical processing based on rules and state information. Using signals such as usage and time, the system identifies Skills that have gone unused for a long time or entered a low-activity state, then marks, demotes, or archives them.
  • The second is agent-based content review. An auxiliary agent examines relationships among the Skills’ contents and decides whether to retain, patch, merge, or fully archive them.

Curator primarily handles four actions: retain, patch, merge, and archive. Its judgments rely mainly on Skill text and a small amount of usage metadata. It does not return to original task trajectories, actual success or failure outcomes, or user corrections to validate the Skills’ real effects. It is therefore better at finding content that is textually redundant, overlapping, or outdated than at reliably determining whether two Skills are truly equivalent in execution or which performs better.

3.6 Limitations

3.6.1 Experience Mainly Comes from a Single Trajectory, Without Cross-Trajectory Comparison

Hermes’s background review is primarily based on a single execution trajectory from the current task. It can summarize how the task was completed and what went wrong, but that trajectory alone cannot establish whether a lesson is a general rule or only applies to this task. The system also does not compare successful and failed trajectories across similar tasks, making it difficult to discover why other approaches are better or worse.

3.6.2 Rules Define What Is Worth Saving, but the Model’s Judgment Has Not Been Validated

For experience selection, Hermes mainly adds rules after a problem is discovered, then uses tests to ensure those rules are not removed. This can patch known issues, but it does not demonstrate that the model can select correctly in other situations. The tests currently provide no mechanism for judging whether the model’s execution is accurate. Two questions that directly affect library quality therefore remain unanswered: how much experience that is not worth saving gets stored, and how much genuinely useful experience is missed? As experience accumulates, low-value or incorrect content may consequently continue to enter the library.

3.6.3 Skill Changes Are Not Accepted on the Basis of Demonstrated Task Improvement

When a Skill is modified, Hermes checks its format and structure, requires the original to be read before editing, and retains version history. The model can still correct errors discovered in later tasks. Hermes thus has some checks on the quality of Skill changes, but it lacks quantitative acceptance criteria for their effects. Hermes does not compare outcomes on the same set of tasks with no Skill, the old version, and the new version, then use that evidence to decide whether to accept a change. As a result, Skill updates rest more on the mechanism’s hopes than on demonstrated results. Every automatic update is therefore an attempt at improvement, not yet a validated gain.

3.6.4 Model-Driven Skill Selection Can Lead to Both Underuse and Overuse

Hermes can filter Skills based on conditions such as tool availability, but whether the current task needs a Skill and whether it is more appropriate than other Skills remain largely matters of model judgment. Hermes strongly instructs the model to load matching Skills and favors loading more when uncertain. This can reduce missed use, but may also cause the model to follow a Skill unsuited to the task, resulting in overuse.

Beyond accumulating experience about how to complete tasks, an agent therefore needs to learn when to use which Skill, and when not to use one. An agent can make the wrong choice, but should not repeatedly make the same mistake in similar circumstances. This requires reviewing execution trajectories, not just final outcomes: why the model chose a particular Skill, which steps actually helped, and which caused detours or rework. Only by retaining these judgments as guidance for future selection can an agent progress from merely having more Skills to using them correctly.

3.6.5 Curator Governs by Usage Rather Than Actual Value

Hermes’s Curator mainly uses the time of a Skill’s last use, viewing, or modification to decide whether it stays active. This can clear away content that has long been untouched, but cannot determine whether a Skill is correct, still effective, or actually improves task performance. Without effect evaluation, Curator primarily addresses quantity and activity rather than quality. To govern by value, the system would also need task success rates, rework, execution costs, scope of applicability, and comparisons between old and new versions to decide whether to retain, revise, merge, or archive a Skill.


IV. Toward a Better, More Complete Mechanism

4.1 From Single-Trajectory Experience to Multi-Trajectory Induction

Hermes currently generates Skills mainly from a single task’s execution process. It can summarize how that task was completed, but struggles to distinguish a general lesson from one applicable only to that task. An accidental success, an unnecessary step, or an environmental exception can all end up embedded in a Skill.

A more reliable approach is to collect multiple successful and failed trajectories for the same class of task, compare the paths, and infer transferable lessons. Trace2Skill follows this approach: it analyzes a broad pool of trajectories in parallel, extracts targeted patches, and combines them into conflict-free Skills. This illustrates that the value of experience comes mainly from comparing multiple executions, rather than retelling a single success.

4.2 Make Independent Validation a Prerequisite for Library Admission

Generating a Skill is only an attempt at improvement; it does not mean task performance has increased. A more complete mechanism must make validation a hard gate: run tests for Skills containing code, reproduce tasks in a fresh session or sandbox for text-only Skills, and reject admission if validation fails. The CoEvoSkills ablation study provides direct evidence: removing the proxy verifier reduced the SkillsBench pass rate from 71.1% to 41.1%. SkillLearnBench likewise found that, without new external signals, repeated self-rewriting does not produce consistent improvement and may even cause recursive drift. Effective improvement depends on external signals such as test results, execution feedback, teacher models, or independent verifiers. A verifier therefore cannot simply be the original model checking its own work again in the original context. It needs to execute the task afresh in an isolated environment and diagnose failures specifically.

4.3 Use Counterfactual Comparisons to Determine a Skill’s Contribution

Invocation counts prove only that a Skill was loaded, not that it improved task performance. High usage can even result from excessive triggering or routing errors.

Meaningful metrics compare success rates, costs, latency, and regression rates for the same class of tasks under three conditions: no Skill, the old version, and the new version. SkillAudit already uses a similar method: executing the same task with and without a candidate Skill, then comparing the two trajectories to determine whether the Skill helped or introduced interference.

This counterfactual comparison should also become the acceptance standard for Skill changes. Only when the new version consistently outperforms the old one should a change count as a validated improvement.

4.4 Do Not Just Write Skills—Ensure They Are Used Correctly

An effective Skill may never be invoked because routing fails, while an unsuitable Skill may be overused because its description is too broad or loading is forced. The system therefore needs to evaluate not only a Skill’s content but also whether it is selected and executed correctly at runtime.

A more complete usage mechanism should include semantic retrieval, ambiguity handling in selection, and adoption detection: find the Skill when needed, actually follow it once found, and avoid forcing it onto inappropriate tasks. SkillEvolver’s “silent failure” audit specifically checks cases where a Skill appears correct in content but is never invoked at runtime.

An agent needs to accumulate experience not just about how to complete a task, but also about when to use which Skill and when not to. This requires reviewing actual execution trajectories to determine which steps helped and which caused detours or rework.

4.5 Separate Automatic Proposals from Automatic Activation

A runtime agent can still propose new Skills or revisions, but those proposals should first enter a candidate queue rather than immediately affect subsequent tasks. Candidates must undergo tests, independent audits, and version checks before a decision is made on activation. The degree of automation should not be uniform either. Low-risk, frequent, replayable tasks—such as data cleaning, code checks, and document formatting—are suited to automatic multi-trajectory validation followed by automatic activation when validation passes. Tasks with external side effects—such as sending messages, placing orders, or modifying cloud data—can cause damage merely by being rerun, so they require sandboxes, dry runs, human approval, and gradual rollout. Regardless of the level of automation, the system must retain candidates, versions, rollback, conflict handling, and human approval entry points, and be able to explain why a particular behavior changed.

4.6 Ultimately, Validate Product Value

Skill self-evolution cannot be measured only by how many Skills are generated or how often they are invoked. It also needs to account for unit economics: how many tokens, model calls, and how much money a review and validation consume; how often the Skill is expected to be reused; how much time each use saves; and how costly a single incorrect use could be. Automatic evolution has practical value only when its long-term benefits exceed the costs of generation, validation, and governance. Otherwise, the system may simply become better at writing Skills without making task completion more accurate, reliable, or inexpensive.