15 KiBLFS
Skill Invocation Surfaces — claude-code / opencode / openhands / codex / pi vs. agentskills.io spec
Source-level audit of how each harness exposes skills to the model (tool call vs. prompt injection), what frontmatter it recognises, and how invocations show up in the trajectory. Goes deeper than per-path discovery on the five harnesses benchflow currently registers.
Verified 2026-05-03. OpenHands rows reflect the unreleased CLI main (sdk
1.19+), which has changed the implementation substantively since our 2026-04-30
audit; released CLI 1.15.0 (sdk 1.17.0) does not yet have these features.
Sources audited
| Harness | Ref | Files |
|---|---|---|
| agentskills.io spec | agentskills/agentskills@main |
docs/specification.mdx |
| claude-code | docs at code.claude.com/docs/en/skills |
reference impl |
| opencode | opencode.ai/docs/skills (current) |
TypeScript runtime, Skills doc rev May 2026 |
| openhands (new sdk) | OpenHands/software-agent-sdk@main (2026-04-28) |
openhands-sdk/openhands/sdk/tool/builtins/invoke_skill.py, openhands-sdk/openhands/sdk/skills/__init__.py |
| codex | openai/codex@main |
codex-rs/core-skills/src/injection.rs, codex-rs/core/src/context/skill_instructions.rs, codex-rs/core/src/skills.rs |
| pi | last-verified 2026-04-30; inception-ai/pi-mono no longer publicly accessible |
flagged below — needs reverification |
Frontmatter recognition vs. spec
The agentskills.io spec defines name and description as required and
license/compatibility/metadata/allowed-tools (experimental) as optional
(spec,
docs/specification.mdx).
| Field | Spec | claude-code | opencode | openhands (sdk 1.19+) | codex | pi (Apr 30 — needs reverify) |
|---|---|---|---|---|---|---|
name (req) |
required | ✓ | ✓ | ✓ | ✓ | ✓ |
description (req) |
required | ✓ | ✓ | ✓ | ✓ | ✓ |
license |
optional | ✓ | ✓ | ignored | ignored | ✓ |
compatibility |
optional | ✓ | ✓ | ignored | ignored | ✓ |
metadata |
optional | ✓ | ✓ | ignored | ignored | ✓ |
allowed-tools |
experimental | enforced | ignored | ignored | ignored | enforced |
| Beyond-spec extensions | n/a | disable-model-invocation, paths, context: fork, agent, model, effort, hooks, arguments |
none (spec-strict) | triggers (KeywordTrigger / TaskTrigger) |
sibling agents/openai.yaml: policy.allow_implicit_invocation, dependencies.tools (per-skill MCP) |
disable-model-invocation |
opencode evidence: docs explicitly enumerate the recognised fields and add
"Unknown frontmatter fields are ignored" — allowed-tools is not in the list
(opencode.ai/docs/skills).
openhands evidence: openhands-sdk/openhands/sdk/skills/__init__.py exports
Skill, SkillResources, triggers (KeywordTrigger, TaskTrigger), and
install lifecycle but no permission frontmatter handling.
Discovery paths
| Harness | Paths searched |
|---|---|
| claude-code | ~/.claude/skills/, <proj>/.claude/skills/, plugin dirs, enterprise (live-watched) |
| opencode | .opencode/skills/, ~/.config/opencode/skills/, .claude/skills/, ~/.claude/skills/, .agents/skills/, ~/.agents/skills/ (project paths walked to git root) |
| openhands | .agents/skills/, ~/.openhands/skills/, plus OpenHands/extensions public registry via load_public_skills |
| codex | .agents/skills/ walked to repo root, ~/.agents/skills/, /etc/codex/skills/, bundled |
| pi | .pi/skills/, .agents/skills/ (walked), ~/.pi/agent/skills/, ~/.agents/skills/, packages, --skill <path> |
opencode evidence: "OpenCode searches these locations: .opencode/skills/<name>/SKILL.md,
~/.config/opencode/skills/<name>/SKILL.md, .claude/skills/<name>/SKILL.md,
~/.claude/skills/<name>/SKILL.md, .agents/skills/<name>/SKILL.md,
~/.agents/skills/<name>/SKILL.md" (opencode.ai/docs/skills).
Invocation surface
| claude-code | opencode | openhands (new sdk) | codex | pi | |
|---|---|---|---|---|---|
| Activation | model picks from desc list (auto) + /skill-name slash |
model calls skill({name}) tool |
model calls invoke_skill(name="…") tool |
$skill-name mention OR command-pattern implicit match |
model picks from desc list + /skill:name |
| Body delivery | tool-call observation | tool-call observation | tool-call observation | injected as <skill> user-role text fragment |
tool-call observation |
| Trajectory event | Skill(name=…) tool call |
skill(name=…) tool call |
invoke_skill(name=…) tool call |
none — only a user-role text fragment with <skill> markers (analytics counter codex.skill.injected is out-of-band) |
tool call |
| Progressive disclosure | desc in ctx, body on invoke | desc in ctx, body on invoke | desc in ctx, body on invoke | nothing in ctx until activation triggers, then full body | desc in ctx, body on invoke |
Codex evidence — prompt injection, not a tool
codex-rs/core-skills/src/injection.rs (excerpted):
pub async fn build_skill_injections(
mentioned_skills: &[SkillMetadata],
...
) -> SkillInjections {
...
for skill in mentioned_skills {
match fs.read_file_text(&skill.path_to_skills_md, None).await {
Ok(contents) => {
...
invocations.push(SkillInvocation {
skill_name: skill.name.clone(),
skill_scope: skill.scope,
skill_path: skill.path_to_skills_md.to_path_buf(),
invocation_type: InvocationType::Explicit,
});
result.items.push(SkillInjection {
name: skill.name.clone(),
path: skill.path_to_skills_md.to_string_lossy().into_owned(),
contents,
});
}
...
}
}
analytics_client.track_skill_invocations(tracking, invocations);
...
}
codex-rs/core/src/context/skill_instructions.rs:
impl ContextualUserFragment for SkillInstructions {
const ROLE: &'static str = "user";
const START_MARKER: &'static str = "<skill>";
const END_MARKER: &'static str = "</skill>";
fn body(&self) -> String {
format!("\n<name>{}</name>\n<path>{}</path>\n{}\n",
self.name, self.path, self.contents)
}
}
The SkillInvocation analytics record is a separate stream from the trajectory.
Activation entry points: collect_explicit_skill_mentions ($skill-name token via
TOOL_MENTION_SIGIL = "$") and detect_implicit_skill_invocation_for_command
(pattern-match against skill metadata) — both in codex-rs/core-skills/src/.
OpenHands evidence — first-class tool
openhands-sdk/openhands/sdk/tool/builtins/invoke_skill.py (excerpted):
class InvokeSkillAction(Action):
name: str = Field(description="Name of the loaded skill to invoke.")
TOOL_DESCRIPTION = """Invoke a skill by name.
This is the only supported way to invoke a skill listed in
`<available_skills>`. Call it with the `<name>` shown in that block; the
skill's full content is rendered (including any dynamic context) and
returned as the tool result.
"""
class InvokeSkillExecutor(ToolExecutor):
@staticmethod
def _record_invocation(conversation, name):
invoked = conversation.state.invoked_skills
if name not in invoked:
invoked.append(name)
def __call__(self, action, conversation=None):
skills, working_dir = self._get_skills_and_working_dir(conversation)
match = next((s for s in skills if s.name == action.name.strip()), None)
...
rendered = render_content_with_commands(match.content, working_dir=working_dir)
rendered = self._append_skill_location_footer(rendered, match.source, working_dir)
self._record_invocation(conversation, action.name.strip())
return InvokeSkillObservation.from_text(text=rendered, skill_name=action.name.strip())
Two trajectory-auditable signals: (1) the invoke_skill tool call event itself,
(2) conversation.state.invoked_skills accumulator.
CLI version pinning:
- CLI
1.15.0(released 2026-04-24) →openhands-sdk==1.17.0→ noinvoke_skill - CLI
main(unreleased) →openhands-sdk==1.19.0→ hasinvoke_skill invoke_skill.pyfirst appears atsoftware-agent-sdk@v1.18.0(verified by 404→200 on raw URL across tags)
opencode evidence — first-class tool
opencode.ai/docs/skills: "Agents see available skills and can load the full content when needed." Tool description format:
<available_skills>
<skill>
<name>git-release</name>
<description>Create consistent releases and changelogs</description>
</skill>
</available_skills>
"The agent loads a skill by calling the tool: skill({ name: "git-release" })."
Permission / scoping per skill
| claude-code | opencode | openhands | codex | pi | |
|---|---|---|---|---|---|
| Per-skill tool gating | allowed-tools frontmatter |
none (skill-level allow/deny via opencode.json glob patterns: allow/deny/ask) |
none | policy.allow_implicit_invocation per skill |
allowed-tools |
| Disable model invocation | disable-model-invocation frontmatter |
global tool disable (tools.skill: false) + per-agent permissions |
none | n/a | disable-model-invocation frontmatter |
| Sandbox / fork | context: fork → subagent |
none | none | inherits session sandbox (landlock+seccomp / seatbelt) | none |
| Per-skill MCP deps | none | none | none | dependencies.tools[type=mcp] in sibling agents/openai.yaml |
none |
opencode evidence: permission.skill block with glob patterns and three
verbs (allow/deny/ask); per-agent override via agent frontmatter
permission.skill map.
Lifecycle & registry
| claude-code | opencode | openhands | codex | pi | |
|---|---|---|---|---|---|
| Install/enable/disable API | filesystem + plugin manager | filesystem + permissions | install_skill/uninstall_skill/enable_skill/update_skill (only one with this) |
filesystem + bundled | --skill <path> flag |
| Public registry | Anthropic-bundled + plugins | none | OpenHands/extensions (load_public_skills) |
bundled in repo | none |
| Live reload | yes (filesystem watch) | unknown | unknown | unknown | /model reloads; skill watch unknown |
openhands evidence — openhands-sdk/openhands/sdk/skills/__init__.py module docstring:
**Installed Skills Management:**
- `install_skill` - Install a skill from a source
- `uninstall_skill` - Uninstall a skill
- `list_installed_skills` - List all installed skills
- `load_installed_skills` - Load enabled installed skills
- `enable_skill`, `disable_skill` - Toggle skill enabled state
- `update_skill` - Update an installed skill
Spec compliance summary
| Harness | Reads spec frontmatter | Implements allowed-tools (experimental) |
Beyond-spec extensions |
|---|---|---|---|
| claude-code | full | ✓ | richest superset |
| opencode | full | ✗ (ignored) | none — spec-strict |
| openhands | partial — only name+description |
✗ | triggers |
| codex | partial — only name+description |
✗ | sibling-file MCP/policy |
| pi (Apr 30) | full | ✓ | disable-model-invocation |
One-line characterisation
- claude-code — reference implementation. Richest frontmatter, richest scoping. Skill = first-class tool with permission gates.
- opencode — spec-strict. Native
skilltool. Best discovery breadth (reads.claude/skills/and.agents/skills/natively). Permissions live in config, not frontmatter. - openhands (new sdk) — first-class
invoke_skilltool. Only one with a real install/registry lifecycle. No per-skill scoping yet. - codex — intentionally not a tool. Skills are prompt injections, but with the richest activation model (mention + command-pattern) and the only per-skill MCP wiring.
- pi — closest spiritual clone of claude-code. Same frontmatter superset, same auto+slash invocation. No registry. Public source no longer accessible — needs reverification.
Implication for SkillsBench measurement
For "did the agent use the skill?" — counted from the trajectory alone:
| Harness | Counting method | Reliable? |
|---|---|---|
| claude-code | grep Skill(name=…) tool calls |
✓ |
| opencode | grep skill(name=…) tool calls |
✓ |
| openhands (new sdk) | grep invoke_skill(name=…) tool calls or read state.invoked_skills |
✓ |
| codex | grep user-role text for <skill> markers (misses implicit pattern triggers — those land in OTel codex.skill.injected only) |
partial |
| pi | grep skill tool calls | ✓ (assuming Apr 30 model still holds) |
OpenHands' new SDK flips the harness from worst-case (rebrand of microagents, keyword-trigger prompt injection, no audit signal) to tied-best for SkillsBench's measurement model. Codex remains an outlier — its skills are deliberately not tool calls, so reward-vs-skill-use analysis from trajectories alone under-counts it.
Re-rank delta vs. earlier (2026-04-30) audit
| Old (Apr 30) | New (May 03) | Movement |
|---|---|---|
| 1. claude-code | 1. claude-code | unchanged |
| 2. opencode | 2. opencode ↔ openhands (tied) | openhands jumps 3 places |
| 3. pi | 3. pi | unchanged |
| 4. codex | 4. codex | unchanged |
| 5. openhands | — | promoted |
The Apr 30 audit placed openhands at #5 because skills were keyword-triggered
prompt extensions (microagents rebrand) with no tool-call audit signal. SDK
v1.18.0 (released bundled with the Apr 28 CLI-main bump to sdk 1.19.0)
introduced invoke_skill as a first-class tool, which addresses both the
auditability and progressive-disclosure concerns. Per-skill permission
scoping is still missing, so claude-code retains #1.
Open questions
- Pi current state — repo no longer publicly accessible; if pi-mono moved or
went closed, our claim of
allowed-toolsenforcement is unverifiable. Re-audit before relying on it for harness ranking. - OpenHands CLI release timing —
invoke_skillonly ships when CLI cuts a release pinning sdk ≥1.18.0. As of 2026-05-03, latest released CLI is 1.15.0. allowed-toolsadoption — only claude-code and pi enforce it. The spec marks it experimental. Worth raising with the spec maintainers whether to promote or drop, since 3 of 5 major harnesses ignore it.- Codex implicit-trigger counting —
detect_implicit_skill_invocation_for_commandnever appears in the trajectory. To count codex skill use accurately, SkillsBench would need to ingest thecodex.skill.injectedOTel counter or thetrack_skill_invocationsanalytics stream — not possible in a sandboxed eval without instrumentation hooks.