Files
2026-09-04 14:58:42 +08:00

6.3 KiBLFS
Raw Permalink Blame History

docs/paper-figures/ — NeurIPS 2026 paper figures

This directory contains the figure-rendering pipeline for the SkillsBench NeurIPS 2026 paper. It is self-contained and does not import or depend on the rest of the SkillsBench codebase (no benchmark infra, no harness, no real trajectory I/O).

docs/paper-figures/
├── README.md          ← this file
├── scripts/           ← 9 matplotlib scripts + shared utils.py
└── figures/           ← rendered PDFs (one per script + method.pdf)

⚠️ Synthetic data

Every script in scripts/ uses synthetic / placeholder data. No external CSVs, manifests, or trajectory dumps are read. The fake-data shape is documented at the top of each script in a FAKE DATA FORMAT docstring:

  • which module-level constant defines the input (MODELS, SYNTHETIC_MODELS, DOMAIN_PALETTE, etc.),
  • the schema of each row / record (column names, types, value ranges),
  • and the noise model used to derive the synthetic numbers.

This makes the pipeline easy to swap in real data later: replace the fake-data dict / list with values loaded from your real run, keep the plotting code, regenerate the PDFs.

How to regenerate

Requirements: Python ≥ 3.10, matplotlib, numpy, pandas. From this directory:

cd scripts
for f in *.py; do [ "$f" = "utils.py" ] || python3 "$f"; done

Each script writes its PDF to ../figures/<name>.pdf, overwriting the existing copy. utils.py is the shared style + palette module imported by every script — do not run it directly.

What each figure shows

The numbering follows the script filenames; the paper-section labels and in-paper figure numbers are listed for cross-reference. (Paper figure numbering can shift; the script number is the stable identifier.)

Script / PDF Paper § Paper Fig Purpose
method.pdf §3 SkillsBench Fig 2 Pipeline overview: Phase 1 (Construction — Skills + 322 candidate tasks from 105 contributors), Phase 2 (Filtering — automated checks + human review), Phase 3 (Evaluation — 3 conditions × 4 harnesses, 7,308 trajectories). Static diagram, not generated by a script in this directory.
01_pareto_vector.py §1 Intro / teaser Fig 1 Cost–performance Pareto with skill-effect arrows. Per model, two markers (hollow circle = no skills, filled triangle = with skills) connected by an arrow; overlaid no-skill and with-skill Pareto frontiers. Demonstrates that curated Skills shift the Pareto envelope upward at modest cost.
02_condition_bars.py §4.1 Same Harness, Different Models Fig 3 Per-model pass rate under 3 skill conditions on OpenCode (12 models). Vertical dumbbell per model with hollow circle / filled triangle / hollow square = no-skill / curated / self-gen, plus Wilson 95 % CI whiskers and a curated−baseline Δpp annotation. Isolates model effect on skill loading and execution.
06_harness_parity.py §4.2 Same Models, Different Harnesses Fig 4 Top-4 provider-native models evaluated on all 4 harnesses (Claude Code / Codex CLI / Gemini CLI / OpenCode), 1×4 panel grid. Shows Codex CLI consistently produces curated ≈ no-skill (the harness, not the model, governs skill loading), while OpenCode delivers a clean lift across families.
14_skill_creator_slopes.py §4.3 Different Skill Creators Fig 5 12 models under 4 skill-source conditions (Distractor → No Skills → Self-Generated → Curated), one slope-line per model colored by provider, plus an aggregate-mean diamond overlay with headline Δpp annotations (−6.3 / +1.3 / +14.4 pp). Isolates skill content quality from model and harness.
05_reasoning_smallmultiples.py §4.4 Reasoning Effort Fig 6 Top-4 reasoning × skills small-multiples sorted by Δslope. Each panel: dashed = no skill, solid = curated, x-axis = reasoning effort (Low/Mid/High/Max). Shows that curated skills compose multiplicatively with reasoning for flagship/thinking models and additively for lite-tier.
03_tsapr_attribution_domain.py §5.1 Skill's Functions Fig 7 TSAPR attribution counts across 11 domains for successful skill trajectories (multi-label: column sums exceed n_success). P (Procedure-prior) and A (Action-augmentation) dominate; R (Runtime-verification) concentrates in Security and Multimodal; S (State-abstraction) in Scientific Computing and Graph/Dialog. Predictive of which authoring effort yields benefit per domain.
15_failure_count_by_model.py §5.2 Failure Analysis Fig 8 Absolute failed-trajectory counts for 12 OpenCode models under curated skills, single bar per model stacked by failure category (color + hatch encode category; bars sorted ascending by total failures). Task Reasoning is the dominant residual; Execution Timeout the second-largest.
04_task_heterogeneity.py §5.3 Domain-Level Heterogeneity Fig 9 84-task scatter: baseline pass rate vs with-curated-skills pass rate. Dashed y = x reference + median-baseline vertical split four quadrants: skill-rescued, stuck-low, skill-amplified, context-burden. Marker size = difficulty tier; marker color = domain. Quantifies that skills are not Pareto-improving across tasks.
09_domain_failure_panels.py Appendix App. Fig Per-domain failure burden (left panel) + failure composition (right panel) across 11 domains. Companion to Fig 8 at domain granularity, used in the appendix to motivate domain-specific instrumentation choices.

Replacing fake data with real data

For each script, the workflow is:

  1. Open the script and read the FAKE DATA FORMAT docstring at the top.
  2. Replace the module-level constant (e.g. MODELS, SYNTHETIC_MODELS, DOMAIN_PALETTE) and / or the _synthetic*() / _fake_tasks() / fake_data() helper with code that loads your real measurements.
  3. Keep the row / column schema identical to the docstring — every plotting helper downstream assumes that shape.
  4. Re-run the script; the PDF in ../figures/ is overwritten.

The shared style (utils.py) is read-only for the plotting scripts — do not add data-loading helpers there. If a future paper figure needs a new script, follow the existing pattern: a top-level FAKE DATA FORMAT docstring, a single OUTPUT_PATH = … / "figures" / "<n>_<name>.pdf", and a main() that ends with save_figure(fig, OUTPUT_PATH).