Files
2026-09-04 14:58:42 +08:00

304 lines
15 KiBLFS
Markdown

# Skill Invocation Surfaces — claude-code / opencode / openhands / codex / pi vs. agentskills.io spec
Source-level audit of *how* each harness exposes skills to the model (tool call vs.
prompt injection), what frontmatter it recognises, and how invocations show up in the
trajectory. Goes deeper than per-path discovery on the five harnesses benchflow
currently registers.
**Verified 2026-05-03.** OpenHands rows reflect the unreleased CLI `main` (sdk
1.19+), which has changed the implementation substantively since our 2026-04-30
audit; released CLI `1.15.0` (sdk 1.17.0) does **not** yet have these features.
---
## Sources audited
| Harness | Ref | Files |
|---|---|---|
| **agentskills.io spec** | `agentskills/agentskills@main` | `docs/specification.mdx` |
| **claude-code** | docs at `code.claude.com/docs/en/skills` | reference impl |
| **opencode** | `opencode.ai/docs/skills` (current) | TypeScript runtime, Skills doc rev May 2026 |
| **openhands (new sdk)** | `OpenHands/software-agent-sdk@main` (2026-04-28) | `openhands-sdk/openhands/sdk/tool/builtins/invoke_skill.py`, `openhands-sdk/openhands/sdk/skills/__init__.py` |
| **codex** | `openai/codex@main` | `codex-rs/core-skills/src/injection.rs`, `codex-rs/core/src/context/skill_instructions.rs`, `codex-rs/core/src/skills.rs` |
| **pi** | last-verified 2026-04-30; `inception-ai/pi-mono` no longer publicly accessible | flagged below — **needs reverification** |
---
## Frontmatter recognition vs. spec
The agentskills.io spec defines `name` and `description` as required and
`license`/`compatibility`/`metadata`/`allowed-tools` (experimental) as optional
([spec](https://agentskills.io/specification),
[`docs/specification.mdx`](https://github.com/agentskills/agentskills/blob/main/docs/specification.mdx)).
| Field | Spec | claude-code | opencode | openhands (sdk 1.19+) | codex | pi (Apr 30 — needs reverify) |
|---|---|---|---|---|---|---|
| `name` (req) | required | ✓ | ✓ | ✓ | ✓ | ✓ |
| `description` (req) | required | ✓ | ✓ | ✓ | ✓ | ✓ |
| `license` | optional | ✓ | ✓ | ignored | ignored | ✓ |
| `compatibility` | optional | ✓ | ✓ | ignored | ignored | ✓ |
| `metadata` | optional | ✓ | ✓ | ignored | ignored | ✓ |
| `allowed-tools` | experimental | **enforced** | ignored | ignored | ignored | **enforced** |
| Beyond-spec extensions | n/a | `disable-model-invocation`, `paths`, `context: fork`, `agent`, `model`, `effort`, `hooks`, `arguments` | none (spec-strict) | `triggers` (KeywordTrigger / TaskTrigger) | sibling `agents/openai.yaml`: `policy.allow_implicit_invocation`, `dependencies.tools` (per-skill MCP) | `disable-model-invocation` |
opencode evidence: docs explicitly enumerate the recognised fields and add
"Unknown frontmatter fields are ignored" — `allowed-tools` is not in the list
([opencode.ai/docs/skills](https://opencode.ai/docs/skills)).
openhands evidence: `openhands-sdk/openhands/sdk/skills/__init__.py` exports
`Skill`, `SkillResources`, triggers (`KeywordTrigger`, `TaskTrigger`), and
install lifecycle but no permission frontmatter handling.
---
## Discovery paths
| Harness | Paths searched |
|---|---|
| **claude-code** | `~/.claude/skills/`, `<proj>/.claude/skills/`, plugin dirs, enterprise (live-watched) |
| **opencode** | `.opencode/skills/`, `~/.config/opencode/skills/`, `.claude/skills/`, `~/.claude/skills/`, `.agents/skills/`, `~/.agents/skills/` (project paths walked to git root) |
| **openhands** | `.agents/skills/`, `~/.openhands/skills/`, plus `OpenHands/extensions` public registry via `load_public_skills` |
| **codex** | `.agents/skills/` walked to repo root, `~/.agents/skills/`, `/etc/codex/skills/`, bundled |
| **pi** | `.pi/skills/`, `.agents/skills/` (walked), `~/.pi/agent/skills/`, `~/.agents/skills/`, packages, `--skill <path>` |
opencode evidence: "OpenCode searches these locations: `.opencode/skills/<name>/SKILL.md`,
`~/.config/opencode/skills/<name>/SKILL.md`, `.claude/skills/<name>/SKILL.md`,
`~/.claude/skills/<name>/SKILL.md`, `.agents/skills/<name>/SKILL.md`,
`~/.agents/skills/<name>/SKILL.md`" ([opencode.ai/docs/skills](https://opencode.ai/docs/skills)).
---
## Invocation surface
| | claude-code | opencode | openhands (new sdk) | codex | pi |
|---|---|---|---|---|---|
| **Activation** | model picks from desc list (auto) + `/skill-name` slash | model calls `skill({name})` tool | model calls `invoke_skill(name="…")` tool | `$skill-name` mention OR command-pattern implicit match | model picks from desc list + `/skill:name` |
| **Body delivery** | tool-call observation | tool-call observation | tool-call observation | injected as `<skill>` user-role text fragment | tool-call observation |
| **Trajectory event** | `Skill(name=…)` tool call | `skill(name=…)` tool call | `invoke_skill(name=…)` tool call | **none** — only a user-role text fragment with `<skill>` markers (analytics counter `codex.skill.injected` is out-of-band) | tool call |
| **Progressive disclosure** | desc in ctx, body on invoke | desc in ctx, body on invoke | desc in ctx, body on invoke | nothing in ctx until activation triggers, then full body | desc in ctx, body on invoke |
### Codex evidence — prompt injection, not a tool
`codex-rs/core-skills/src/injection.rs` (excerpted):
```rust
pub async fn build_skill_injections(
mentioned_skills: &[SkillMetadata],
...
) -> SkillInjections {
...
for skill in mentioned_skills {
match fs.read_file_text(&skill.path_to_skills_md, None).await {
Ok(contents) => {
...
invocations.push(SkillInvocation {
skill_name: skill.name.clone(),
skill_scope: skill.scope,
skill_path: skill.path_to_skills_md.to_path_buf(),
invocation_type: InvocationType::Explicit,
});
result.items.push(SkillInjection {
name: skill.name.clone(),
path: skill.path_to_skills_md.to_string_lossy().into_owned(),
contents,
});
}
...
}
}
analytics_client.track_skill_invocations(tracking, invocations);
...
}
```
`codex-rs/core/src/context/skill_instructions.rs`:
```rust
impl ContextualUserFragment for SkillInstructions {
const ROLE: &'static str = "user";
const START_MARKER: &'static str = "<skill>";
const END_MARKER: &'static str = "</skill>";
fn body(&self) -> String {
format!("\n<name>{}</name>\n<path>{}</path>\n{}\n",
self.name, self.path, self.contents)
}
}
```
The `SkillInvocation` analytics record is a separate stream from the trajectory.
Activation entry points: `collect_explicit_skill_mentions` (`$skill-name` token via
`TOOL_MENTION_SIGIL = "$"`) and `detect_implicit_skill_invocation_for_command`
(pattern-match against skill metadata) — both in `codex-rs/core-skills/src/`.
### OpenHands evidence — first-class tool
`openhands-sdk/openhands/sdk/tool/builtins/invoke_skill.py` (excerpted):
```python
class InvokeSkillAction(Action):
name: str = Field(description="Name of the loaded skill to invoke.")
TOOL_DESCRIPTION = """Invoke a skill by name.
This is the only supported way to invoke a skill listed in
`<available_skills>`. Call it with the `<name>` shown in that block; the
skill's full content is rendered (including any dynamic context) and
returned as the tool result.
"""
class InvokeSkillExecutor(ToolExecutor):
@staticmethod
def _record_invocation(conversation, name):
invoked = conversation.state.invoked_skills
if name not in invoked:
invoked.append(name)
def __call__(self, action, conversation=None):
skills, working_dir = self._get_skills_and_working_dir(conversation)
match = next((s for s in skills if s.name == action.name.strip()), None)
...
rendered = render_content_with_commands(match.content, working_dir=working_dir)
rendered = self._append_skill_location_footer(rendered, match.source, working_dir)
self._record_invocation(conversation, action.name.strip())
return InvokeSkillObservation.from_text(text=rendered, skill_name=action.name.strip())
```
Two trajectory-auditable signals: (1) the `invoke_skill` tool call event itself,
(2) `conversation.state.invoked_skills` accumulator.
CLI version pinning:
- CLI `1.15.0` (released 2026-04-24) → `openhands-sdk==1.17.0` → no `invoke_skill`
- CLI `main` (unreleased) → `openhands-sdk==1.19.0` → has `invoke_skill`
- `invoke_skill.py` first appears at `software-agent-sdk@v1.18.0` (verified by 404→200 on raw URL across tags)
### opencode evidence — first-class tool
[opencode.ai/docs/skills](https://opencode.ai/docs/skills): "Agents see available
skills and can load the full content when needed." Tool description format:
```xml
<available_skills>
<skill>
<name>git-release</name>
<description>Create consistent releases and changelogs</description>
</skill>
</available_skills>
```
"The agent loads a skill by calling the tool: `skill({ name: "git-release" })`."
---
## Permission / scoping per skill
| | claude-code | opencode | openhands | codex | pi |
|---|---|---|---|---|---|
| **Per-skill tool gating** | `allowed-tools` frontmatter | none (skill-level allow/deny via `opencode.json` glob patterns: `allow`/`deny`/`ask`) | none | `policy.allow_implicit_invocation` per skill | `allowed-tools` |
| **Disable model invocation** | `disable-model-invocation` frontmatter | global tool disable (`tools.skill: false`) + per-agent permissions | none | n/a | `disable-model-invocation` frontmatter |
| **Sandbox / fork** | `context: fork` → subagent | none | none | inherits session sandbox (landlock+seccomp / seatbelt) | none |
| **Per-skill MCP deps** | none | none | none | `dependencies.tools[type=mcp]` in sibling `agents/openai.yaml` | none |
opencode evidence: `permission.skill` block with glob patterns and three
verbs (`allow`/`deny`/`ask`); per-agent override via agent frontmatter
`permission.skill` map.
---
## Lifecycle & registry
| | claude-code | opencode | openhands | codex | pi |
|---|---|---|---|---|---|
| **Install/enable/disable API** | filesystem + plugin manager | filesystem + permissions | **`install_skill`/`uninstall_skill`/`enable_skill`/`update_skill`** (only one with this) | filesystem + bundled | `--skill <path>` flag |
| **Public registry** | Anthropic-bundled + plugins | none | `OpenHands/extensions` (`load_public_skills`) | bundled in repo | none |
| **Live reload** | yes (filesystem watch) | unknown | unknown | unknown | `/model` reloads; skill watch unknown |
openhands evidence — `openhands-sdk/openhands/sdk/skills/__init__.py` module docstring:
```
**Installed Skills Management:**
- `install_skill` - Install a skill from a source
- `uninstall_skill` - Uninstall a skill
- `list_installed_skills` - List all installed skills
- `load_installed_skills` - Load enabled installed skills
- `enable_skill`, `disable_skill` - Toggle skill enabled state
- `update_skill` - Update an installed skill
```
---
## Spec compliance summary
| Harness | Reads spec frontmatter | Implements `allowed-tools` (experimental) | Beyond-spec extensions |
|---|---|---|---|
| **claude-code** | full | ✓ | richest superset |
| **opencode** | full | ✗ (ignored) | none — spec-strict |
| **openhands** | partial — only `name`+`description` | ✗ | `triggers` |
| **codex** | partial — only `name`+`description` | ✗ | sibling-file MCP/policy |
| **pi (Apr 30)** | full | ✓ | `disable-model-invocation` |
---
## One-line characterisation
- **claude-code** — reference implementation. Richest frontmatter, richest scoping. Skill = first-class tool with permission gates.
- **opencode** — spec-strict. Native `skill` tool. Best discovery breadth (reads `.claude/skills/` and `.agents/skills/` natively). Permissions live in config, not frontmatter.
- **openhands (new sdk)** — first-class `invoke_skill` tool. Only one with a real install/registry lifecycle. No per-skill scoping yet.
- **codex** — *intentionally* not a tool. Skills are prompt injections, but with the richest activation model (mention + command-pattern) and the only per-skill MCP wiring.
- **pi** — closest spiritual clone of claude-code. Same frontmatter superset, same auto+slash invocation. No registry. **Public source no longer accessible — needs reverification.**
---
## Implication for SkillsBench measurement
For "did the agent use the skill?" — counted from the trajectory alone:
| Harness | Counting method | Reliable? |
|---|---|---|
| claude-code | grep `Skill(name=…)` tool calls | ✓ |
| opencode | grep `skill(name=…)` tool calls | ✓ |
| openhands (new sdk) | grep `invoke_skill(name=…)` tool calls **or** read `state.invoked_skills` | ✓ |
| codex | grep user-role text for `<skill>` markers (misses implicit pattern triggers — those land in OTel `codex.skill.injected` only) | partial |
| pi | grep skill tool calls | ✓ (assuming Apr 30 model still holds) |
OpenHands' new SDK flips the harness from worst-case (rebrand of microagents,
keyword-trigger prompt injection, no audit signal) to tied-best for SkillsBench's
measurement model. Codex remains an outlier — its skills are deliberately
not tool calls, so reward-vs-skill-use analysis from trajectories alone
under-counts it.
---
## Re-rank delta vs. earlier (2026-04-30) audit
| Old (Apr 30) | New (May 03) | Movement |
|---|---|---|
| 1. claude-code | 1. claude-code | unchanged |
| 2. opencode | 2. opencode ↔ openhands (tied) | openhands jumps 3 places |
| 3. pi | 3. pi | unchanged |
| 4. codex | 4. codex | unchanged |
| 5. openhands | — | promoted |
The Apr 30 audit placed openhands at #5 because skills were keyword-triggered
prompt extensions (microagents rebrand) with no tool-call audit signal. SDK
v1.18.0 (released bundled with the Apr 28 CLI-main bump to sdk 1.19.0)
introduced `invoke_skill` as a first-class tool, which addresses both the
auditability and progressive-disclosure concerns. Per-skill permission
scoping is still missing, so claude-code retains #1.
---
## Open questions
1. **Pi current state** — repo no longer publicly accessible; if pi-mono moved or
went closed, our claim of `allowed-tools` enforcement is unverifiable. Re-audit
before relying on it for harness ranking.
2. **OpenHands CLI release timing** — `invoke_skill` only ships when CLI cuts a
release pinning sdk ≥1.18.0. As of 2026-05-03, latest released CLI is 1.15.0.
3. **`allowed-tools` adoption** — only claude-code and pi enforce it. The spec
marks it experimental. Worth raising with the spec maintainers whether to
promote or drop, since 3 of 5 major harnesses ignore it.
4. **Codex implicit-trigger counting** — `detect_implicit_skill_invocation_for_command`
never appears in the trajectory. To count codex skill use accurately, SkillsBench
would need to ingest the `codex.skill.injected` OTel counter or the
`track_skill_invocations` analytics stream — not possible in a sandboxed eval
without instrumentation hooks.