The Skill Self-Evolution Mechanism in Hermes Agent
I. What Is a “Skill Self-Evolution Mechanism”?
In academic research, a self-evolving agent is generally understood as an agent that continually adjusts itself based on experience or feedback gained through interaction with its environment, improving its performance on subsequent tasks. To understand these mechanisms, we can ask three questions: what evolves, when it evolves, and what signals drive that evolution.
What evolves: context and memory, tools, Skills, or even the agent’s overall architecture.
When it evolves: adjustments can happen in real time during a task, after a task based on a retrospective, or through offline evolution triggered by a fixed schedule, an experience threshold, or other conditions.
What signals drive evolution: self-evolution may rely on rewards, success or failure feedback, imitation of existing behavior, population-based evolution, or other mechanisms.
Within this framework, the Skill self-evolution mechanism discussed in this article can be defined more specifically:
During real-world use, an agent independently identifies reusable knowledge from task experience, then creates, modifies, merges, or retires Skills accordingly, allowing the Skill library to improve continually as experience accumulates.
Under this definition, a complete Skill self-evolution mechanism must answer at least five questions. These five stages also provide the framework for evaluating Hermes below:
Stage
Core question
Triggering
When should a process for capturing task experience be triggered?
Generation
What should be extracted from a complete task experience? How can a one-off experience become a genuinely reusable Skill with clearly defined boundaries?
Use
When a similar task comes along, can the agent find the Skill at the right time, and does using it actually produce a better result?
Validation
After a Skill is used, how can we determine whether it was effective or caused a failure, and how should it be revised after failure?
Governance
How should duplicate, conflicting, outdated, low-value, and long-unused Skills be handled?
II. Why Hermes Needs Skill Self-Evolution
Hermes positions itself as a personal agent that accompanies its user over time: “The agent that grows with you.” For Hermes, this means it cannot rely solely on the capabilities of its underlying model. It must retain experience gained through long-term use and turn that experience into capabilities it can reuse. Otherwise, no matter how often someone uses it, Hermes would still need to understand and explore similar tasks from scratch each time. That would hardly amount to growing alongside the user.
Hermes already addresses some of this long-term accumulation through Memory, but Memory primarily stores factual information: who the user is, what they prefer, and what has happened. Procedural experience—how to complete a particular kind of task—requires Skills. For example, a task may reveal a more effective workflow, combination of tools, decision rule, or way to avoid a pitfall. Only by capturing this experience in a Skill can Hermes reuse it directly on similar tasks rather than reasoning everything out again.
Yet Hermes does not operate in a static environment. With continued use, it encounters new successes, failures, and user corrections; user habits, tool capabilities, and the external environment may also change. Hermes therefore cannot generate a Skill once and leave it unchanged indefinitely. It needs to revise existing Skills based on subsequent results, create new ones when necessary, and retire Skills that no longer work. This creates a continuous loop—task experience → capture in a Skill → subsequent reuse → new feedback → Skill revision—through which Skill capabilities can evolve.
III. Unpacking Hermes’s Skill Self-Evolution Mechanism
3.1 Overview
HERMES / SKILL
Hermes Skill Self-Evolution at a Glance
Execute a task → identify reusable experience → create or modify a Skill → reuse it later → periodically organize and govern the library.
Task execution
Every conversation turn: the model calls tools, advances the task, and delivers a reply.
↓
01 / FOUR LEARNING PATHS
Automatic
Foreground learning
Learning rules live in the main agent’s System Prompt. During a task, the model identifies reusable methods and calls skill_manage directly to create or modify a Skill.
Automatic
Background review
When accumulated tool-call iterations reach 10 steps (carried across turns), an independent /fork reviews the trajectory after the final reply. The count resets after saving. The fork is read-only except for Memory and Skill writes; predefined instructions guide it to extract how and why to act next time.
User-triggered
/learn
The user asks to capture this experience. Read the material, then check existing Skills: update a matching capability, or create one if none exists, and save it in the standard format.
User-triggered
/refine
Revisit the recent conversation and execution process to decide whether Memory or a Skill needs updating. Memory updates go to Memory; Skill updates enter the write flow below.
↓
02 / THE SINGLE SKILL WRITE ENTRY POINT
skill_manage
Eight checks and actions, in order. Memory updates do not pass through this Skill write entry point.
Validate the request itself.
Check the name and format; the description must not exceed 60 characters.
Before editing an existing Skill, read it in full.
Run a security scan (off by default).
Wait for human approval (off by default).
Write the file.
Run a health check; it advises but does not reject the write.
Record snapshots before and after the change so it can be rolled back.
↓
03 / SKILL LIBRARY
Where writes end and reuse begins
Three layers, loaded progressively as needed.
Index
Always in the system prompt
Each entry = name + description of ≤60 characters, present in context every turn.
Body
Loaded on a match
The full SKILL.md is loaded into context only when selected.
Support
Accessed on demand
Read references/ into context; copy and edit templates/; execute scripts/ directly without loading them into context.
↓
04 / THE USE LOOP, EVERY TURN
From Skills back to tasks
Inject the index into the system prompt each turn.
The model judges whether a Skill matches.
skill_view loads the full SKILL.md; increment use count by 1.
Read references, copy templates, and run scripts as needed.
Skill self-evolution in Hermes is not a single process. It combines foreground learning, background review, /learn, /refine, and Curator. The overall flow can be understood as:
Execute a task → identify experience worth learning from → create or modify a Skill → reuse it later → periodically organize and govern the library
3.2 Automatic Learning: Foreground Learning
Hermes implements foreground learning by putting the rule “capture reusable experience when you discover it” directly into the main agent’s System Prompt. This means the content of foreground learning depends entirely on the model’s own judgment. When an agent discovers a valuable workflow, finds a problem in an existing Skill, or gains experience worth reusing, it can call skill_manage directly to create or modify a Skill. Foreground learning thus embeds the awareness of learning directly into the main agent. This places very high demands on the model’s judgment. In practice, the rules are extremely vague: there is no definite standard for what experience is genuinely valuable. Such a rough setup raises doubts about the quality of Skills produced through Hermes’s foreground learning.
3.3 Automatic Learning: Background Review
Alongside active learning, Hermes provides an independent background review mechanism. It periodically revisits past task trajectories to identify experience that the main agent did not capture at the time, primarily considering user feedback, the actual task process, and the performance of existing Skills. Its trigger counts the agent’s actual tool-call iterations. Once the accumulated count reaches the threshold, Hermes does not immediately interrupt the task. It waits until the main agent finishes the current task and delivers its final reply, then copies the current session’s execution trajectory and starts an independent background Review Agent to look specifically for learning signals.
It looks for four types of learning signal:
correct: the user corrects the agent’s earlier approach. This is a high-priority signal that favors changing the Skill, because the user has explicitly indicated that the current approach is flawed;
improve: a more efficient or reliable method is discovered during the task;
solve: after encountering a real problem or pitfall, exploration leads to a stable solution;
renew: using a Skill reveals errors, missing steps, or outdated content.
After finding learning signals, Hermes’s Review Prompt explicitly requires the following order of priority:
First modify a currently loaded Skill → then find and update an existing Skill of the same class → add supporting files to an existing Skill when necessary → create a new class-level Skill only when no suitable Skill exists
The purpose is to absorb new experience into existing capabilities wherever possible, rather than generating a new Skill for every task and rapidly inflating the library.
Compared with foreground active learning, the main improvement is that background review relies more on externally observable signals than on the model simply deciding whether a task feels complex. It also defines what must not be saved in considerable detail. However, it still has no genuinely quantitative value score; the LLM ultimately makes the judgment based on these rules.
3.4 User-Triggered Entry Points: /learn and /refine
Hermes also offers two user-triggered learning entry points: /learn and /refine. Both address the same need: when a user knows that an experience is worth learning from or that existing experience should be re-examined, they can ask the system to learn directly instead of waiting for an automatic trigger.
/learn focuses on formally capturing current experience in a Skill. Based on the current conversation, a document, or a task process, the user can ask Hermes to “learn from this experience.” The system reads the material, identifies genuinely reusable methods, and checks existing Skills. If the corresponding capability already exists, it updates that Skill; otherwise, it creates one, then organizes and saves it in the standard format. /refine focuses more on reviewing and correcting existing experience. It revisits the conversation and execution process that just took place to determine whether any Memory or Skill should be updated.
3.5 Long-Term Governance: Curator
As foreground learning, background review, and user-triggered learning continue to accumulate Skills, new problems emerge. The library grows and may develop duplicate, outdated, low-value, or overlapping content. Curator addresses this layer of the problem. It maintains the existing Skill library over time, keeping it usable, concise, and reasonably consistent. Its work falls roughly into two stages:
The first is mechanical processing based on rules and state information. Using signals such as usage and time, the system identifies Skills that have gone unused for a long time or entered a low-activity state, then marks, demotes, or archives them.
The second is agent-based content review. An auxiliary agent examines relationships among the Skills’ contents and decides whether to retain, patch, merge, or fully archive them.
Curator primarily handles four actions: retain, patch, merge, and archive. Its judgments rely mainly on Skill text and a small amount of usage metadata. It does not return to original task trajectories, actual success or failure outcomes, or user corrections to validate the Skills’ real effects. It is therefore better at finding content that is textually redundant, overlapping, or outdated than at reliably determining whether two Skills are truly equivalent in execution or which performs better.
3.6 Limitations
3.6.1 Experience Mainly Comes from a Single Trajectory, Without Cross-Trajectory Comparison
Hermes’s background review is primarily based on a single execution trajectory from the current task. It can summarize how the task was completed and what went wrong, but that trajectory alone cannot establish whether a lesson is a general rule or only applies to this task. The system also does not compare successful and failed trajectories across similar tasks, making it difficult to discover why other approaches are better or worse.
3.6.2 Rules Define What Is Worth Saving, but the Model’s Judgment Has Not Been Validated
For experience selection, Hermes mainly adds rules after a problem is discovered, then uses tests to ensure those rules are not removed. This can patch known issues, but it does not demonstrate that the model can select correctly in other situations. The tests currently provide no mechanism for judging whether the model’s execution is accurate. Two questions that directly affect library quality therefore remain unanswered: how much experience that is not worth saving gets stored, and how much genuinely useful experience is missed? As experience accumulates, low-value or incorrect content may consequently continue to enter the library.
3.6.3 Skill Changes Are Not Accepted on the Basis of Demonstrated Task Improvement
When a Skill is modified, Hermes checks its format and structure, requires the original to be read before editing, and retains version history. The model can still correct errors discovered in later tasks. Hermes thus has some checks on the quality of Skill changes, but it lacks quantitative acceptance criteria for their effects. Hermes does not compare outcomes on the same set of tasks with no Skill, the old version, and the new version, then use that evidence to decide whether to accept a change. As a result, Skill updates rest more on the mechanism’s hopes than on demonstrated results. Every automatic update is therefore an attempt at improvement, not yet a validated gain.
3.6.4 Model-Driven Skill Selection Can Lead to Both Underuse and Overuse
Hermes can filter Skills based on conditions such as tool availability, but whether the current task needs a Skill and whether it is more appropriate than other Skills remain largely matters of model judgment. Hermes strongly instructs the model to load matching Skills and favors loading more when uncertain. This can reduce missed use, but may also cause the model to follow a Skill unsuited to the task, resulting in overuse.
Beyond accumulating experience about how to complete tasks, an agent therefore needs to learn when to use which Skill, and when not to use one. An agent can make the wrong choice, but should not repeatedly make the same mistake in similar circumstances. This requires reviewing execution trajectories, not just final outcomes: why the model chose a particular Skill, which steps actually helped, and which caused detours or rework. Only by retaining these judgments as guidance for future selection can an agent progress from merely having more Skills to using them correctly.
3.6.5 Curator Governs by Usage Rather Than Actual Value
Hermes’s Curator mainly uses the time of a Skill’s last use, viewing, or modification to decide whether it stays active. This can clear away content that has long been untouched, but cannot determine whether a Skill is correct, still effective, or actually improves task performance. Without effect evaluation, Curator primarily addresses quantity and activity rather than quality. To govern by value, the system would also need task success rates, rework, execution costs, scope of applicability, and comparisons between old and new versions to decide whether to retain, revise, merge, or archive a Skill.
IV. Toward a Better, More Complete Mechanism
4.1 From Single-Trajectory Experience to Multi-Trajectory Induction
Hermes currently generates Skills mainly from a single task’s execution process. It can summarize how that task was completed, but struggles to distinguish a general lesson from one applicable only to that task. An accidental success, an unnecessary step, or an environmental exception can all end up embedded in a Skill.
A more reliable approach is to collect multiple successful and failed trajectories for the same class of task, compare the paths, and infer transferable lessons. Trace2Skill follows this approach: it analyzes a broad pool of trajectories in parallel, extracts targeted patches, and combines them into conflict-free Skills. This illustrates that the value of experience comes mainly from comparing multiple executions, rather than retelling a single success.
4.2 Make Independent Validation a Prerequisite for Library Admission
Generating a Skill is only an attempt at improvement; it does not mean task performance has increased. A more complete mechanism must make validation a hard gate: run tests for Skills containing code, reproduce tasks in a fresh session or sandbox for text-only Skills, and reject admission if validation fails. The CoEvoSkills ablation study provides direct evidence: removing the proxy verifier reduced the SkillsBench pass rate from 71.1% to 41.1%. SkillLearnBench likewise found that, without new external signals, repeated self-rewriting does not produce consistent improvement and may even cause recursive drift. Effective improvement depends on external signals such as test results, execution feedback, teacher models, or independent verifiers. A verifier therefore cannot simply be the original model checking its own work again in the original context. It needs to execute the task afresh in an isolated environment and diagnose failures specifically.
4.3 Use Counterfactual Comparisons to Determine a Skill’s Contribution
Invocation counts prove only that a Skill was loaded, not that it improved task performance. High usage can even result from excessive triggering or routing errors.
Meaningful metrics compare success rates, costs, latency, and regression rates for the same class of tasks under three conditions: no Skill, the old version, and the new version. SkillAudit already uses a similar method: executing the same task with and without a candidate Skill, then comparing the two trajectories to determine whether the Skill helped or introduced interference.
This counterfactual comparison should also become the acceptance standard for Skill changes. Only when the new version consistently outperforms the old one should a change count as a validated improvement.
4.4 Do Not Just Write Skills—Ensure They Are Used Correctly
An effective Skill may never be invoked because routing fails, while an unsuitable Skill may be overused because its description is too broad or loading is forced. The system therefore needs to evaluate not only a Skill’s content but also whether it is selected and executed correctly at runtime.
A more complete usage mechanism should include semantic retrieval, ambiguity handling in selection, and adoption detection: find the Skill when needed, actually follow it once found, and avoid forcing it onto inappropriate tasks. SkillEvolver’s “silent failure” audit specifically checks cases where a Skill appears correct in content but is never invoked at runtime.
An agent needs to accumulate experience not just about how to complete a task, but also about when to use which Skill and when not to. This requires reviewing actual execution trajectories to determine which steps helped and which caused detours or rework.
4.5 Separate Automatic Proposals from Automatic Activation
A runtime agent can still propose new Skills or revisions, but those proposals should first enter a candidate queue rather than immediately affect subsequent tasks. Candidates must undergo tests, independent audits, and version checks before a decision is made on activation. The degree of automation should not be uniform either. Low-risk, frequent, replayable tasks—such as data cleaning, code checks, and document formatting—are suited to automatic multi-trajectory validation followed by automatic activation when validation passes. Tasks with external side effects—such as sending messages, placing orders, or modifying cloud data—can cause damage merely by being rerun, so they require sandboxes, dry runs, human approval, and gradual rollout. Regardless of the level of automation, the system must retain candidates, versions, rollback, conflict handling, and human approval entry points, and be able to explain why a particular behavior changed.
4.6 Ultimately, Validate Product Value
Skill self-evolution cannot be measured only by how many Skills are generated or how often they are invoked. It also needs to account for unit economics: how many tokens, model calls, and how much money a review and validation consume; how often the Skill is expected to be reused; how much time each use saves; and how costly a single incorrect use could be. Automatic evolution has practical value only when its long-term benefits exceed the costs of generation, validation, and governance. Otherwise, the system may simply become better at writing Skills without making task completion more accurate, reliable, or inexpensive.