docs/paper-figures/ — NeurIPS 2026 paper figures
This directory contains the figure-rendering pipeline for the SkillsBench NeurIPS 2026 paper. It is self-contained and does not import or depend on the rest of the SkillsBench codebase (no benchmark infra, no harness, no real trajectory I/O).
docs/paper-figures/
├── README.md ← this file
├── scripts/ ← 9 matplotlib scripts + shared utils.py
└── figures/ ← rendered PDFs (one per script + method.pdf)
⚠️ Synthetic data
Every script in scripts/ uses synthetic / placeholder data. No external
CSVs, manifests, or trajectory dumps are read. The fake-data shape is
documented at the top of each script in a FAKE DATA FORMAT docstring:
- which module-level constant defines the input (
MODELS,SYNTHETIC_MODELS,DOMAIN_PALETTE, etc.), - the schema of each row / record (column names, types, value ranges),
- and the noise model used to derive the synthetic numbers.
This makes the pipeline easy to swap in real data later: replace the fake-data dict / list with values loaded from your real run, keep the plotting code, regenerate the PDFs.
How to regenerate
Requirements: Python ≥ 3.10, matplotlib, numpy, pandas. From this
directory:
cd scripts
for f in *.py; do [ "$f" = "utils.py" ] || python3 "$f"; done
Each script writes its PDF to ../figures/<name>.pdf, overwriting the
existing copy. utils.py is the shared style + palette module imported by
every script — do not run it directly.
What each figure shows
The numbering follows the script filenames; the paper-section labels and in-paper figure numbers are listed for cross-reference. (Paper figure numbering can shift; the script number is the stable identifier.)
| Script / PDF | Paper § | Paper Fig | Purpose |
|---|---|---|---|
method.pdf |
§3 SkillsBench | Fig 2 | Pipeline overview: Phase 1 (Construction — Skills + 322 candidate tasks from 105 contributors), Phase 2 (Filtering — automated checks + human review), Phase 3 (Evaluation — 3 conditions × 4 harnesses, 7,308 trajectories). Static diagram, not generated by a script in this directory. |
01_pareto_vector.py |
§1 Intro / teaser | Fig 1 | Cost–performance Pareto with skill-effect arrows. Per model, two markers (hollow circle = no skills, filled triangle = with skills) connected by an arrow; overlaid no-skill and with-skill Pareto frontiers. Demonstrates that curated Skills shift the Pareto envelope upward at modest cost. |
02_condition_bars.py |
§4.1 Same Harness, Different Models | Fig 3 | Per-model pass rate under 3 skill conditions on OpenCode (12 models). Vertical dumbbell per model with hollow circle / filled triangle / hollow square = no-skill / curated / self-gen, plus Wilson 95 % CI whiskers and a curated−baseline Δpp annotation. Isolates model effect on skill loading and execution. |
06_harness_parity.py |
§4.2 Same Models, Different Harnesses | Fig 4 | Top-4 provider-native models evaluated on all 4 harnesses (Claude Code / Codex CLI / Gemini CLI / OpenCode), 1×4 panel grid. Shows Codex CLI consistently produces curated ≈ no-skill (the harness, not the model, governs skill loading), while OpenCode delivers a clean lift across families. |
14_skill_creator_slopes.py |
§4.3 Different Skill Creators | Fig 5 | 12 models under 4 skill-source conditions (Distractor → No Skills → Self-Generated → Curated), one slope-line per model colored by provider, plus an aggregate-mean diamond overlay with headline Δpp annotations (−6.3 / +1.3 / +14.4 pp). Isolates skill content quality from model and harness. |
05_reasoning_smallmultiples.py |
§4.4 Reasoning Effort | Fig 6 | Top-4 reasoning × skills small-multiples sorted by Δslope. Each panel: dashed = no skill, solid = curated, x-axis = reasoning effort (Low/Mid/High/Max). Shows that curated skills compose multiplicatively with reasoning for flagship/thinking models and additively for lite-tier. |
03_tsapr_attribution_domain.py |
§5.1 Skill's Functions | Fig 7 | TSAPR attribution counts across 11 domains for successful skill trajectories (multi-label: column sums exceed n_success). P (Procedure-prior) and A (Action-augmentation) dominate; R (Runtime-verification) concentrates in Security and Multimodal; S (State-abstraction) in Scientific Computing and Graph/Dialog. Predictive of which authoring effort yields benefit per domain. |
15_failure_count_by_model.py |
§5.2 Failure Analysis | Fig 8 | Absolute failed-trajectory counts for 12 OpenCode models under curated skills, single bar per model stacked by failure category (color + hatch encode category; bars sorted ascending by total failures). Task Reasoning is the dominant residual; Execution Timeout the second-largest. |
04_task_heterogeneity.py |
§5.3 Domain-Level Heterogeneity | Fig 9 | 84-task scatter: baseline pass rate vs with-curated-skills pass rate. Dashed y = x reference + median-baseline vertical split four quadrants: skill-rescued, stuck-low, skill-amplified, context-burden. Marker size = difficulty tier; marker color = domain. Quantifies that skills are not Pareto-improving across tasks. |
09_domain_failure_panels.py |
Appendix | App. Fig | Per-domain failure burden (left panel) + failure composition (right panel) across 11 domains. Companion to Fig 8 at domain granularity, used in the appendix to motivate domain-specific instrumentation choices. |
Replacing fake data with real data
For each script, the workflow is:
- Open the script and read the
FAKE DATA FORMATdocstring at the top. - Replace the module-level constant (e.g.
MODELS,SYNTHETIC_MODELS,DOMAIN_PALETTE) and / or the_synthetic*()/_fake_tasks()/fake_data()helper with code that loads your real measurements. - Keep the row / column schema identical to the docstring — every plotting helper downstream assumes that shape.
- Re-run the script; the PDF in
../figures/is overwritten.
The shared style (utils.py) is read-only for the plotting scripts — do not
add data-loading helpers there. If a future paper figure needs a new
script, follow the existing pattern: a top-level FAKE DATA FORMAT
docstring, a single OUTPUT_PATH = … / "figures" / "<n>_<name>.pdf", and a
main() that ends with save_figure(fig, OUTPUT_PATH).