Files
SkillCompiler/data/skills-bench/docs/paper-figures/README.md
T
2026-09-04 14:58:42 +08:00

80 lines
6.3 KiBLFS
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `docs/paper-figures/` — NeurIPS 2026 paper figures
This directory contains the figure-rendering pipeline for the SkillsBench
NeurIPS 2026 paper. It is **self-contained** and does not import or depend on
the rest of the SkillsBench codebase (no benchmark infra, no harness, no real
trajectory I/O).
```
docs/paper-figures/
├── README.md ← this file
├── scripts/ ← 9 matplotlib scripts + shared utils.py
└── figures/ ← rendered PDFs (one per script + method.pdf)
```
## ⚠️ Synthetic data
**Every script in `scripts/` uses synthetic / placeholder data.** No external
CSVs, manifests, or trajectory dumps are read. The fake-data shape is
documented at the top of each script in a `FAKE DATA FORMAT` docstring:
- which module-level constant defines the input (`MODELS`, `SYNTHETIC_MODELS`,
`DOMAIN_PALETTE`, etc.),
- the schema of each row / record (column names, types, value ranges),
- and the noise model used to derive the synthetic numbers.
This makes the pipeline easy to swap in real data later: replace the
fake-data dict / list with values loaded from your real run, keep the
plotting code, regenerate the PDFs.
## How to regenerate
Requirements: Python ≥ 3.10, `matplotlib`, `numpy`, `pandas`. From this
directory:
```bash
cd scripts
for f in *.py; do [ "$f" = "utils.py" ] || python3 "$f"; done
```
Each script writes its PDF to `../figures/<name>.pdf`, overwriting the
existing copy. `utils.py` is the shared style + palette module imported by
every script — do not run it directly.
## What each figure shows
The numbering follows the script filenames; the paper-section labels and
in-paper figure numbers are listed for cross-reference. (Paper figure
numbering can shift; the script number is the stable identifier.)
| Script / PDF | Paper § | Paper Fig | Purpose |
|---|---|---|---|
| `method.pdf` | §3 SkillsBench | Fig 2 | **Pipeline overview**: Phase 1 (Construction — Skills + 322 candidate tasks from 105 contributors), Phase 2 (Filtering — automated checks + human review), Phase 3 (Evaluation — 3 conditions × 4 harnesses, 7,308 trajectories). Static diagram, not generated by a script in this directory. |
| `01_pareto_vector.py` | §1 Intro / teaser | Fig 1 | **Cost–performance Pareto with skill-effect arrows.** Per model, two markers (hollow circle = no skills, filled triangle = with skills) connected by an arrow; overlaid no-skill and with-skill Pareto frontiers. Demonstrates that curated Skills shift the Pareto envelope upward at modest cost. |
| `02_condition_bars.py` | §4.1 Same Harness, Different Models | Fig 3 | **Per-model pass rate under 3 skill conditions on OpenCode (12 models).** Vertical dumbbell per model with hollow circle / filled triangle / hollow square = no-skill / curated / self-gen, plus Wilson 95 % CI whiskers and a curated−baseline Δpp annotation. Isolates model effect on skill loading and execution. |
| `06_harness_parity.py` | §4.2 Same Models, Different Harnesses | Fig 4 | **Top-4 provider-native models evaluated on all 4 harnesses** (Claude Code / Codex CLI / Gemini CLI / OpenCode), 1×4 panel grid. Shows Codex CLI consistently produces curated ≈ no-skill (the harness, not the model, governs skill loading), while OpenCode delivers a clean lift across families. |
| `14_skill_creator_slopes.py` | §4.3 Different Skill Creators | Fig 5 | **12 models under 4 skill-source conditions** (Distractor → No Skills → Self-Generated → Curated), one slope-line per model colored by provider, plus an aggregate-mean diamond overlay with headline Δpp annotations (−6.3 / +1.3 / +14.4 pp). Isolates skill content quality from model and harness. |
| `05_reasoning_smallmultiples.py` | §4.4 Reasoning Effort | Fig 6 | **Top-4 reasoning × skills small-multiples** sorted by Δslope. Each panel: dashed = no skill, solid = curated, x-axis = reasoning effort (Low/Mid/High/Max). Shows that curated skills compose multiplicatively with reasoning for flagship/thinking models and additively for lite-tier. |
| `03_tsapr_attribution_domain.py` | §5.1 Skill's Functions | Fig 7 | **TSAPR attribution counts across 11 domains** for successful skill trajectories (multi-label: column sums exceed n_success). P (Procedure-prior) and A (Action-augmentation) dominate; R (Runtime-verification) concentrates in Security and Multimodal; S (State-abstraction) in Scientific Computing and Graph/Dialog. Predictive of which authoring effort yields benefit per domain. |
| `15_failure_count_by_model.py` | §5.2 Failure Analysis | Fig 8 | **Absolute failed-trajectory counts for 12 OpenCode models under curated skills**, single bar per model stacked by failure category (color + hatch encode category; bars sorted ascending by total failures). Task Reasoning is the dominant residual; Execution Timeout the second-largest. |
| `04_task_heterogeneity.py` | §5.3 Domain-Level Heterogeneity | Fig 9 | **84-task scatter: baseline pass rate vs with-curated-skills pass rate.** Dashed y = x reference + median-baseline vertical split four quadrants: skill-rescued, stuck-low, skill-amplified, context-burden. Marker size = difficulty tier; marker color = domain. Quantifies that skills are not Pareto-improving across tasks. |
| `09_domain_failure_panels.py` | Appendix | App. Fig | **Per-domain failure burden (left panel) + failure composition (right panel) across 11 domains.** Companion to Fig 8 at domain granularity, used in the appendix to motivate domain-specific instrumentation choices. |
## Replacing fake data with real data
For each script, the workflow is:
1. Open the script and read the `FAKE DATA FORMAT` docstring at the top.
2. Replace the module-level constant (e.g. `MODELS`, `SYNTHETIC_MODELS`,
`DOMAIN_PALETTE`) and / or the `_synthetic*()` / `_fake_tasks()` /
`fake_data()` helper with code that loads your real measurements.
3. Keep the row / column schema identical to the docstring — every plotting
helper downstream assumes that shape.
4. Re-run the script; the PDF in `../figures/` is overwritten.
The shared style (`utils.py`) is read-only for the plotting scripts — do not
add data-loading helpers there. If a future paper figure needs a new
script, follow the existing pattern: a top-level `FAKE DATA FORMAT`
docstring, a single `OUTPUT_PATH = … / "figures" / "<n>_<name>.pdf"`, and a
`main()` that ends with `save_figure(fig, OUTPUT_PATH)`.