15 KiBLFS
SkillsBench Research Questions
Analysis and Evidence - January 2026
Executive Summary
SkillsBench addresses two fundamental research questions about AI agent skills. This document synthesizes our ecosystem analysis (47,143 skills, 6,324 repos) with the research goals.
| Research Question | Focus | Current Evidence |
|---|---|---|
| RQ1: Do skills help agents perform better? | Effectiveness | Preliminary - needs benchmark data |
| RQ2: Can agents compose multiple skills? | Composition | Design target: 3+ skills, <39% SOTA |
| RQ3: What is the state of the skills ecosystem? | Ecosystem Analysis | Complete - 47K skills analyzed |
RQ1: Do Skills Help Agents Perform Better?
Question
Comparing agent performance with vs. without Skills on identical tasks
Hypothesis
Skills—procedural knowledge encoded in SKILL.md files—should improve agent performance on specialized tasks by providing:
- Domain-specific patterns and workflows
- Tool usage best practices
- Error handling strategies
- Output format specifications
Evidence Framework
| Metric | With Skills | Without Skills | Delta |
|---|---|---|---|
| Task completion rate | TBD | TBD | TBD |
| Time to completion | TBD | TBD | TBD |
| Error rate | TBD | TBD | TBD |
| Output quality score | TBD | TBD | TBD |
Existing Evidence (External Sources)
Note: Anthropic's official Skills documentation provides architectural guidance but no quantitative effectiveness metrics. This represents a critical gap that SkillsBench aims to fill.
Related Agent Benchmarks (December 2025)
Current benchmarks measure raw model capability, not skill-augmented performance:
| Benchmark | Claude Opus 4.5 | GPT-5.2 | What It Measures |
|---|---|---|---|
| SWE-bench Verified | 77.2% | ~80% | Real-world software engineering |
| Terminal-bench 2.0 | 59.3% | 47.6% | Command-line proficiency |
| ARC-AGI-2 | 37.6% | 52.9% | Abstract reasoning |
| Deep Research | 85.3% | - | Multi-step agentic tasks |
Gap: None of these benchmarks measure whether skills improve agent performance on identical tasks. SkillsBench will be the first to provide this controlled comparison.
Ecosystem Support for RQ1
From our skills analysis, we identified high-impact skill categories for benchmarking:
| Category | Skills Count | Why Measurable |
|---|---|---|
| Testing/TDD | 1,444 (3.1%) | Pass/fail, coverage metrics |
| Database | 976 (2.1%) | Query correctness, data integrity |
| Security | 589 (1.2%) | Vulnerability detection rates |
| Document Processing | 647 (1.4%) | Format compliance, data extraction |
| Code Review | 157 skills | Defect detection, style compliance |
Proposed Methodology
-
Controlled Comparison: Same task, same agent, same environment
- Control: Agent with no skills in
/root/.claude/skills/ - Treatment: Agent with relevant skills installed
- Control: Agent with no skills in
-
Metrics:
- Binary: Task completed successfully (reward = 1 or 0)
- Continuous: Partial credit based on test assertions passed
- Behavioral: Number of tool calls, tokens used, time elapsed
-
Statistical Requirements:
- Minimum 30 tasks per category
- Multiple agent types (Claude Code, Codex, Goose, etc.)
- Multiple model sizes/families
Current Data Gaps
- Need benchmark run data with/without skills
- Need standardized task difficulty ratings
- Need agent behavior logs for analysis
RQ2: Can Agents Compose Multiple Skills?
Question
Measuring how well agents select and combine Skills for complex workflows
Hypothesis
Real-world tasks often require combining multiple specialized skills. Effective skill composition requires:
- Skill discovery (finding relevant skills)
- Skill selection (choosing which to apply)
- Skill ordering (sequencing correctly)
- Skill integration (combining outputs)
Design Targets (from README)
- Composition requirement: 3+ skills together
- SOTA target: <39% pass rate with frontier models
- Distractor tolerance: <10 irrelevant skills
Composition Complexity Levels
| Level | Skills Required | Example Task |
|---|---|---|
| Simple | 1 skill | "Convert PDF to markdown" |
| Moderate | 2-3 skills | "Extract data from PDF, analyze in Python, generate report" |
| Complex | 4-6 skills | "Build dashboard: scrape API, transform data, visualize, deploy" |
Evidence from Ecosystem Analysis
Skill Diversity Supports Composition:
- 47,143 total skills across 20+ categories
- Only ~5% semantic duplication (at 90% threshold)
- 6,324 unique GitHub repositories = diverse implementations
Natural Skill Groupings (from category analysis):
| Workflow | Potential Skills |
|---|---|
| Data Pipeline | Database (976) + Python (2,591) + API (1,020) |
| Code Quality | Testing (1,444) + Code Quality (872) + Git (1,216) |
| Document Processing | PDF/DOCX (647) + Data (976) + Visualization |
| DevOps | Docker/K8s (5,497) + Cloud (398) + Security (589) |
| Scientific | Bio/Chem (142 curated) + ML/AI (2,284) + Visualization |
Composition Metrics
| Metric | Description | Target |
|---|---|---|
| Skill Selection Accuracy | % of relevant skills used | >80% |
| Distractor Rejection | % of irrelevant skills ignored | >90% |
| Ordering Correctness | Skills applied in valid sequence | >70% |
| Integration Success | Outputs correctly combined | >60% |
Challenge: Measuring Composition
Unlike simple task completion, composition quality requires:
- Instrumentation: Log which skills agent reads/uses
- Ground Truth: Define optimal skill set for each task
- Partial Credit: Score based on skill usage patterns
Proposed Task Design
tasks/compose-data-pipeline/
├── task.md # prompt + metadata
├── environment/
│ └── skills/
│ ├── pandas-etl/ # Relevant (ETL)
│ ├── sql-queries/ # Relevant (database)
│ ├── data-viz/ # Relevant (output)
│ ├── git-workflow/ # Distractor
│ ├── docker-deploy/ # Distractor
│ └── testing-tdd/ # Distractor
└── verifier/
└── test_outputs.py # Verify pipeline works
RQ3: What Is the State of the Skills Ecosystem?
Question
Use embeddings to deduplicate skills and characterize the ecosystem, identifying useful skills for benchmark task creation
Status: COMPLETE
Methodology
-
Data Collection: Scraped 47,153 skills from 7 sources
- SkillsMP.com API (46,942 skills, 472 pages)
- K-Dense-AI/claude-scientific-skills (142 curated)
- anthropics/skills (16 official)
- openai/skills (10 official) - NEW
- obra/superpowers (14 curated)
- netresearch/claude-code-marketplace (22 curated)
- awesome-claude-skills (6 curated)
-
Embedding-Based Deduplication: OpenAI
text-embedding-3-large(512 dims)- ALL 47,153 skills embedded (not sampled)
- Cosine similarity thresholds: 95%, 90%, 85%, 80%
- Nearest neighbor search for clustering
-
Registry Overlap Verification: 1,100 random skills validated across Smithery/SkillsMP
Results
Deduplication Analysis (FULL 47K)
| Threshold | Unique | Duplicates | Dup Rate | Clusters |
|---|---|---|---|---|
| 95% | 42,071 | 5,082 | 10.8% | 3,197 |
| 90% | 40,721 | 6,432 | 13.6% | - |
| 85% | 38,495 | 8,658 | 18.4% | - |
| 80% | 34,758 | 12,395 | 26.3% | - |
Finding: ~13.6% semantic duplication at 90% threshold across all 47K skills. ~40,721 semantically unique skills.
Category Distribution
| Category | Count | % | Task Potential |
|---|---|---|---|
| JavaScript/TypeScript | 23,271 | 49.4% | High - primary use case |
| DevOps/Infrastructure | 5,497 | 11.7% | High - measurable |
| Python | 2,591 | 5.5% | High - common |
| AI/ML | 2,284 | 4.8% | Medium - specialized |
| Frontend/UI | 1,476 | 3.1% | High - visual verification |
| Testing | 1,444 | 3.1% | High - pass/fail metrics |
| Git/Version Control | 1,216 | 2.6% | High - verifiable |
| API | 1,020 | 2.2% | High - response validation |
| Database | 976 | 2.1% | High - query correctness |
| Documentation | 913 | 1.9% | Medium - format compliance |
| Security | 749 | 1.6% | High - critical |
Top Duplicated Skills (High Demand Indicators)
| Skill Name | Copies | Implication for Tasks |
|---|---|---|
| skill-creator | 223 | Meta-skill for creating skills |
| code-review | 157 | PR review automation tasks |
| frontend-design | 151 | UI implementation tasks |
| brainstorming | 92 | Planning/ideation tasks |
| git-workflow | 86 | Version control tasks |
| testing | 78 | Test generation tasks |
| systematic-debugging | 65 | Bug diagnosis tasks |
| test-driven-development | 59 | TDD workflow tasks |
Semantic Duplicate Clusters (Examples)
| Cluster | Skills | Similarity |
|---|---|---|
| Visualization | Matplotlib, Plotly, Seaborn | 0.77-0.83 |
| Browser Automation | Playwright, Playwriter, playwright-skill | 0.76-0.86 |
| Scientific DB | PubChem, ChEMBL, DrugBank | 0.81-0.81 |
| Clinical | ClinicalTrials.gov, ClinVar, FDA Databases | 0.81-0.83 |
Skills for Benchmark Task Creation
Based on analysis, high-value skills for tasks:
Tier 1: Most Useful (High duplication + Measurable)
| Skill | Why Useful | Verification Method |
|---|---|---|
| code-review | 157 copies, high demand | Defect detection rate |
| testing/tdd | 78+59 copies | Test pass/fail, coverage |
| git-workflow | 86 copies | Commit history, merge success |
| frontend-design | 151 copies | Visual diff, accessibility |
| systematic-debugging | 65 copies | Bug resolution success |
Tier 2: Domain-Specific (Curated + Specialized)
| Skill Category | Source | Example Tasks |
|---|---|---|
| Scientific (142) | K-Dense-AI | Bioinformatics pipeline, protein analysis |
| Document Processing | anthropics/skills | PDF/XLSX/DOCX manipulation |
| Enterprise (22) | netresearch | TYPO3, PHP modernization |
Tier 3: Composition Candidates (Multi-skill workflows)
| Workflow | Required Skills | Complexity |
|---|---|---|
| Data Pipeline | pandas-etl + sql-queries + data-viz | 3 skills |
| Code Quality | testing + code-review + git-workflow | 3 skills |
| Scientific Report | literature-review + matplotlib + scientific-writing | 3 skills |
| Security Audit | semgrep + security-audit + reporting | 3 skills |
Key Insights for SkillsBench
- JavaScript/TypeScript dominates (49%) → Focus benchmark tasks here
- Low duplication (~5%) → Genuine diversity in implementations
- Testing underrepresented (3.1%) → Opportunity for benchmark tasks
- code-review most duplicated → Highest demand, priority for tasks
- Scientific skills well-curated (142) → High-quality domain tasks
- Community:Official ratio 2,773:1 → Rich ecosystem for task variety
Registry Overlap Analysis
Key Finding
We verified that Smithery.ai skills are a subset of SkillsMP.com:
| Test | Result |
|---|---|
| Sample size | 1,100 random skills |
| Found on both registries | 1,100 (100%) |
| Conclusion | Single source sufficient |
Implication for Research
- SkillsMP's 47,143 skills with GitHub URLs provides complete coverage
- No need to scrape Smithery separately
- All skills traceable to source code for analysis
Data Assets for Research
| File | Contents | Use |
|---|---|---|
all_skills_combined.json |
47,143 skills | RQ1/RQ2 skill selection |
curated_skills.json |
201 verified skills | High-quality subset |
full_analysis.json |
Category distribution | Task design |
embedding_analysis_full.json |
Semantic clusters | Deduplication |
skillsmp_smithery_verification.json |
Registry overlap | Methodology |
Next Steps
For RQ1 (Effectiveness)
- Design 30+ tasks with clear success criteria
- Run each task with/without skills
- Collect metrics: pass rate, time, token usage
- Statistical analysis of skill impact
For RQ2 (Composition)
- Design tasks requiring 3-6 skills
- Include 5-10 distractor skills per task
- Instrument skill usage logging
- Define composition quality metrics
Priority Tasks to Create
Based on ecosystem analysis, highest-value benchmark tasks:
| Task Idea | Required Skills | Measurability |
|---|---|---|
| Automated code review | git-workflow + code-review + testing | PR quality metrics |
| Data pipeline ETL | pandas + sql + data-viz | Output correctness |
| API integration | http-client + auth + error-handling | Response validation |
| Document conversion | pdf + docx + data-extraction | Format compliance |
| Security audit | semgrep + security-patterns + reporting | Vuln detection rate |
Appendix: Skills Ecosystem Statistics
Overall Numbers
- Total skills: 47,143
- Unique names: 32,222
- Semantic unique (90%): ~44,750
- GitHub repositories: 6,324
Top Categories
- JavaScript/TypeScript: 49.4%
- DevOps/Infrastructure: 11.7%
- Python: 5.5%
- AI/ML: 4.8%
- Testing: 3.1%
Official vs Community
- Anthropic official: 17 skills
- Community created: 47,126 skills
- Ratio: 1:2,773
Appendix B: Skills Ecosystem Timeline & Adoption
Timeline
| Date | Event |
|---|---|
| Nov 2024 | MCP (Model Context Protocol) ships |
| Mar 2025 | OpenAI adopts MCP |
| Oct 2025 | Anthropic releases Agent Skills beta |
| Dec 2025 | Skills spec released as open standard; 9 agents adopt |
| Jan 2026 | SkillsBench research begins |
Agent Adoption (December 2025)
- Claude Code (Anthropic)
- Codex CLI (OpenAI)
- ChatGPT (OpenAI)
- Cursor
- Goose
- VS Code agents
- And 3+ others
Key Differentiators: Skills vs MCP
| Aspect | Skills | MCP Servers |
|---|---|---|
| Purpose | Teach how to work | Provide power to work |
| Format | Markdown + YAML | JSON-RPC protocol |
| Mechanism | Context injection | Tool API calls |
| Portability | File-based, copy anywhere | Server configuration |
External Resources
- Claude Skills Documentation
- Claude Code Skills Guide
- Anthropic Engineering Blog on Skills
- SkillsMP Marketplace
- Simon Willison's Analysis
- Deep Dive Technical Analysis
Research framework for SkillsBench paper - January 2026