80 lines
6.3 KiBLFS
Markdown
80 lines
6.3 KiBLFS
Markdown
# `docs/paper-figures/` — NeurIPS 2026 paper figures
|
||
|
||
This directory contains the figure-rendering pipeline for the SkillsBench
|
||
NeurIPS 2026 paper. It is **self-contained** and does not import or depend on
|
||
the rest of the SkillsBench codebase (no benchmark infra, no harness, no real
|
||
trajectory I/O).
|
||
|
||
```
|
||
docs/paper-figures/
|
||
├── README.md ← this file
|
||
├── scripts/ ← 9 matplotlib scripts + shared utils.py
|
||
└── figures/ ← rendered PDFs (one per script + method.pdf)
|
||
```
|
||
|
||
## ⚠️ Synthetic data
|
||
|
||
**Every script in `scripts/` uses synthetic / placeholder data.** No external
|
||
CSVs, manifests, or trajectory dumps are read. The fake-data shape is
|
||
documented at the top of each script in a `FAKE DATA FORMAT` docstring:
|
||
|
||
- which module-level constant defines the input (`MODELS`, `SYNTHETIC_MODELS`,
|
||
`DOMAIN_PALETTE`, etc.),
|
||
- the schema of each row / record (column names, types, value ranges),
|
||
- and the noise model used to derive the synthetic numbers.
|
||
|
||
This makes the pipeline easy to swap in real data later: replace the
|
||
fake-data dict / list with values loaded from your real run, keep the
|
||
plotting code, regenerate the PDFs.
|
||
|
||
## How to regenerate
|
||
|
||
Requirements: Python ≥ 3.10, `matplotlib`, `numpy`, `pandas`. From this
|
||
directory:
|
||
|
||
```bash
|
||
cd scripts
|
||
for f in *.py; do [ "$f" = "utils.py" ] || python3 "$f"; done
|
||
```
|
||
|
||
Each script writes its PDF to `../figures/<name>.pdf`, overwriting the
|
||
existing copy. `utils.py` is the shared style + palette module imported by
|
||
every script — do not run it directly.
|
||
|
||
## What each figure shows
|
||
|
||
The numbering follows the script filenames; the paper-section labels and
|
||
in-paper figure numbers are listed for cross-reference. (Paper figure
|
||
numbering can shift; the script number is the stable identifier.)
|
||
|
||
| Script / PDF | Paper § | Paper Fig | Purpose |
|
||
|---|---|---|---|
|
||
| `method.pdf` | §3 SkillsBench | Fig 2 | **Pipeline overview**: Phase 1 (Construction — Skills + 322 candidate tasks from 105 contributors), Phase 2 (Filtering — automated checks + human review), Phase 3 (Evaluation — 3 conditions × 4 harnesses, 7,308 trajectories). Static diagram, not generated by a script in this directory. |
|
||
| `01_pareto_vector.py` | §1 Intro / teaser | Fig 1 | **Cost–performance Pareto with skill-effect arrows.** Per model, two markers (hollow circle = no skills, filled triangle = with skills) connected by an arrow; overlaid no-skill and with-skill Pareto frontiers. Demonstrates that curated Skills shift the Pareto envelope upward at modest cost. |
|
||
| `02_condition_bars.py` | §4.1 Same Harness, Different Models | Fig 3 | **Per-model pass rate under 3 skill conditions on OpenCode (12 models).** Vertical dumbbell per model with hollow circle / filled triangle / hollow square = no-skill / curated / self-gen, plus Wilson 95 % CI whiskers and a curated−baseline Δpp annotation. Isolates model effect on skill loading and execution. |
|
||
| `06_harness_parity.py` | §4.2 Same Models, Different Harnesses | Fig 4 | **Top-4 provider-native models evaluated on all 4 harnesses** (Claude Code / Codex CLI / Gemini CLI / OpenCode), 1×4 panel grid. Shows Codex CLI consistently produces curated ≈ no-skill (the harness, not the model, governs skill loading), while OpenCode delivers a clean lift across families. |
|
||
| `14_skill_creator_slopes.py` | §4.3 Different Skill Creators | Fig 5 | **12 models under 4 skill-source conditions** (Distractor → No Skills → Self-Generated → Curated), one slope-line per model colored by provider, plus an aggregate-mean diamond overlay with headline Δpp annotations (−6.3 / +1.3 / +14.4 pp). Isolates skill content quality from model and harness. |
|
||
| `05_reasoning_smallmultiples.py` | §4.4 Reasoning Effort | Fig 6 | **Top-4 reasoning × skills small-multiples** sorted by Δslope. Each panel: dashed = no skill, solid = curated, x-axis = reasoning effort (Low/Mid/High/Max). Shows that curated skills compose multiplicatively with reasoning for flagship/thinking models and additively for lite-tier. |
|
||
| `03_tsapr_attribution_domain.py` | §5.1 Skill's Functions | Fig 7 | **TSAPR attribution counts across 11 domains** for successful skill trajectories (multi-label: column sums exceed n_success). P (Procedure-prior) and A (Action-augmentation) dominate; R (Runtime-verification) concentrates in Security and Multimodal; S (State-abstraction) in Scientific Computing and Graph/Dialog. Predictive of which authoring effort yields benefit per domain. |
|
||
| `15_failure_count_by_model.py` | §5.2 Failure Analysis | Fig 8 | **Absolute failed-trajectory counts for 12 OpenCode models under curated skills**, single bar per model stacked by failure category (color + hatch encode category; bars sorted ascending by total failures). Task Reasoning is the dominant residual; Execution Timeout the second-largest. |
|
||
| `04_task_heterogeneity.py` | §5.3 Domain-Level Heterogeneity | Fig 9 | **84-task scatter: baseline pass rate vs with-curated-skills pass rate.** Dashed y = x reference + median-baseline vertical split four quadrants: skill-rescued, stuck-low, skill-amplified, context-burden. Marker size = difficulty tier; marker color = domain. Quantifies that skills are not Pareto-improving across tasks. |
|
||
| `09_domain_failure_panels.py` | Appendix | App. Fig | **Per-domain failure burden (left panel) + failure composition (right panel) across 11 domains.** Companion to Fig 8 at domain granularity, used in the appendix to motivate domain-specific instrumentation choices. |
|
||
|
||
## Replacing fake data with real data
|
||
|
||
For each script, the workflow is:
|
||
|
||
1. Open the script and read the `FAKE DATA FORMAT` docstring at the top.
|
||
2. Replace the module-level constant (e.g. `MODELS`, `SYNTHETIC_MODELS`,
|
||
`DOMAIN_PALETTE`) and / or the `_synthetic*()` / `_fake_tasks()` /
|
||
`fake_data()` helper with code that loads your real measurements.
|
||
3. Keep the row / column schema identical to the docstring — every plotting
|
||
helper downstream assumes that shape.
|
||
4. Re-run the script; the PDF in `../figures/` is overwritten.
|
||
|
||
The shared style (`utils.py`) is read-only for the plotting scripts — do not
|
||
add data-loading helpers there. If a future paper figure needs a new
|
||
script, follow the existing pattern: a top-level `FAKE DATA FORMAT`
|
||
docstring, a single `OUTPUT_PATH = … / "figures" / "<n>_<name>.pdf"`, and a
|
||
`main()` that ends with `save_figure(fig, OUTPUT_PATH)`.
|