# `docs/paper-figures/` — NeurIPS 2026 paper figures This directory contains the figure-rendering pipeline for the SkillsBench NeurIPS 2026 paper. It is **self-contained** and does not import or depend on the rest of the SkillsBench codebase (no benchmark infra, no harness, no real trajectory I/O). ``` docs/paper-figures/ ├── README.md ← this file ├── scripts/ ← 9 matplotlib scripts + shared utils.py └── figures/ ← rendered PDFs (one per script + method.pdf) ``` ## ⚠️ Synthetic data **Every script in `scripts/` uses synthetic / placeholder data.** No external CSVs, manifests, or trajectory dumps are read. The fake-data shape is documented at the top of each script in a `FAKE DATA FORMAT` docstring: - which module-level constant defines the input (`MODELS`, `SYNTHETIC_MODELS`, `DOMAIN_PALETTE`, etc.), - the schema of each row / record (column names, types, value ranges), - and the noise model used to derive the synthetic numbers. This makes the pipeline easy to swap in real data later: replace the fake-data dict / list with values loaded from your real run, keep the plotting code, regenerate the PDFs. ## How to regenerate Requirements: Python ≥ 3.10, `matplotlib`, `numpy`, `pandas`. From this directory: ```bash cd scripts for f in *.py; do [ "$f" = "utils.py" ] || python3 "$f"; done ``` Each script writes its PDF to `../figures/.pdf`, overwriting the existing copy. `utils.py` is the shared style + palette module imported by every script — do not run it directly. ## What each figure shows The numbering follows the script filenames; the paper-section labels and in-paper figure numbers are listed for cross-reference. (Paper figure numbering can shift; the script number is the stable identifier.) | Script / PDF | Paper § | Paper Fig | Purpose | |---|---|---|---| | `method.pdf` | §3 SkillsBench | Fig 2 | **Pipeline overview**: Phase 1 (Construction — Skills + 322 candidate tasks from 105 contributors), Phase 2 (Filtering — automated checks + human review), Phase 3 (Evaluation — 3 conditions × 4 harnesses, 7,308 trajectories). Static diagram, not generated by a script in this directory. | | `01_pareto_vector.py` | §1 Intro / teaser | Fig 1 | **Cost–performance Pareto with skill-effect arrows.** Per model, two markers (hollow circle = no skills, filled triangle = with skills) connected by an arrow; overlaid no-skill and with-skill Pareto frontiers. Demonstrates that curated Skills shift the Pareto envelope upward at modest cost. | | `02_condition_bars.py` | §4.1 Same Harness, Different Models | Fig 3 | **Per-model pass rate under 3 skill conditions on OpenCode (12 models).** Vertical dumbbell per model with hollow circle / filled triangle / hollow square = no-skill / curated / self-gen, plus Wilson 95 % CI whiskers and a curated−baseline Δpp annotation. Isolates model effect on skill loading and execution. | | `06_harness_parity.py` | §4.2 Same Models, Different Harnesses | Fig 4 | **Top-4 provider-native models evaluated on all 4 harnesses** (Claude Code / Codex CLI / Gemini CLI / OpenCode), 1×4 panel grid. Shows Codex CLI consistently produces curated ≈ no-skill (the harness, not the model, governs skill loading), while OpenCode delivers a clean lift across families. | | `14_skill_creator_slopes.py` | §4.3 Different Skill Creators | Fig 5 | **12 models under 4 skill-source conditions** (Distractor → No Skills → Self-Generated → Curated), one slope-line per model colored by provider, plus an aggregate-mean diamond overlay with headline Δpp annotations (−6.3 / +1.3 / +14.4 pp). Isolates skill content quality from model and harness. | | `05_reasoning_smallmultiples.py` | §4.4 Reasoning Effort | Fig 6 | **Top-4 reasoning × skills small-multiples** sorted by Δslope. Each panel: dashed = no skill, solid = curated, x-axis = reasoning effort (Low/Mid/High/Max). Shows that curated skills compose multiplicatively with reasoning for flagship/thinking models and additively for lite-tier. | | `03_tsapr_attribution_domain.py` | §5.1 Skill's Functions | Fig 7 | **TSAPR attribution counts across 11 domains** for successful skill trajectories (multi-label: column sums exceed n_success). P (Procedure-prior) and A (Action-augmentation) dominate; R (Runtime-verification) concentrates in Security and Multimodal; S (State-abstraction) in Scientific Computing and Graph/Dialog. Predictive of which authoring effort yields benefit per domain. | | `15_failure_count_by_model.py` | §5.2 Failure Analysis | Fig 8 | **Absolute failed-trajectory counts for 12 OpenCode models under curated skills**, single bar per model stacked by failure category (color + hatch encode category; bars sorted ascending by total failures). Task Reasoning is the dominant residual; Execution Timeout the second-largest. | | `04_task_heterogeneity.py` | §5.3 Domain-Level Heterogeneity | Fig 9 | **84-task scatter: baseline pass rate vs with-curated-skills pass rate.** Dashed y = x reference + median-baseline vertical split four quadrants: skill-rescued, stuck-low, skill-amplified, context-burden. Marker size = difficulty tier; marker color = domain. Quantifies that skills are not Pareto-improving across tasks. | | `09_domain_failure_panels.py` | Appendix | App. Fig | **Per-domain failure burden (left panel) + failure composition (right panel) across 11 domains.** Companion to Fig 8 at domain granularity, used in the appendix to motivate domain-specific instrumentation choices. | ## Replacing fake data with real data For each script, the workflow is: 1. Open the script and read the `FAKE DATA FORMAT` docstring at the top. 2. Replace the module-level constant (e.g. `MODELS`, `SYNTHETIC_MODELS`, `DOMAIN_PALETTE`) and / or the `_synthetic*()` / `_fake_tasks()` / `fake_data()` helper with code that loads your real measurements. 3. Keep the row / column schema identical to the docstring — every plotting helper downstream assumes that shape. 4. Re-run the script; the PDF in `../figures/` is overwritten. The shared style (`utils.py`) is read-only for the plotting scripts — do not add data-loading helpers there. If a future paper figure needs a new script, follow the existing pattern: a top-level `FAKE DATA FORMAT` docstring, a single `OUTPUT_PATH = … / "figures" / "_.pdf"`, and a `main()` that ends with `save_figure(fig, OUTPUT_PATH)`.