# Oracle Patterns
The oracle (`oracle/solve.sh`) is human-written, demonstrates the answer is producible, and runs without skills (so anything from a skill must be copied into `oracle/` first). Five patterns recur across SkillsBench task types.
## 1. Derive vs. copy — the unavoidable trade-off
The implementation rubric says "derive through computation; skeptical of over-engineered solutions." That cuts two ways:
- **Derive**: Python+library oracle that reconstructs the answer step by step. Demonstrates the workflow.
- **Copy**: place a pre-authored reference artifact at the output path. Demonstrates the artifact is producible.
Pick **derive** when the workflow is procedural: compute a number, run a query, transform a CSV, generate a report from data, run a static analyzer on code, parse a file. Examples: `econ-detrending-correlation` computes a correlation; `dialogue-parser` parses chat logs; `find-topk-similiar-chemicals` does ranking. The oracle naturally rewrites the same pipeline a smart agent would.
Pick **copy** when the answer is an artifact a human authored in a tool that the toolchain can't fully reproduce. Examples:
- **Excel with array formulas.** LibreOffice's `calculateAll()` evaluates `_xlfn.XLOOKUP` but leaves `` empty in the saved XML; openpyxl reads back `None`.
- **PowerPoint with embedded charts / equations.** python-pptx writes the structure; the rendered visuals only get cached when PowerPoint or Keynote opens and saves.
- **PDF with hand-laid-out forms.** Reflowing a fillable PDF in Python loses the original visual layout.
- **Audio with mastering chain.** A waveform generated by SoX won't match an artifact mastered in Logic.
- **Hand-drawn vector diagrams.** SVG produced by D3 won't match a designer's Figma export.
In all the copy cases, the "derivation" already happened in the authoring tool. A Python re-derivation is itself the over-engineering the rubric warns against.
When you copy: ship the reference in `oracle/` (a byte copy of the corresponding `verifier/`) so the oracle has access without crossing the verifier lockdown:
```bash
#!/bin/bash
set -e
cp /oracle/expected.xlsx /root/output.xlsx
```
Note the trade-off in the PR description so reviewers can weigh in.
## 2. Common derivation bugs (across formats)
### Loop range exceeds populated data range
When formulas / commands / queries depend on a key column populated only for a sub-range, but the loop runs over the full range, every row past the populated range produces `None` / `#N/A` / a stack trace.
Excel example: column R has 2025 dates only on rows 2..366; writing `=VLOOKUP($R, ...)` for r=2..2558 produces `#N/A` in 13,152 cells.
CSV / SQL example: a JOIN / lookup against an empty key column drops all rows past the populated range silently.
Fix: gate the loop to the populated range explicitly. Don't trust `max_row` / `len(...)` blindly.
### Unpopulated scaffolding rows
When the input file is meant to have anchor cells (country code, region label, primary key, header) seeded but the agent leaves them empty, formulas / queries that reference them skip silently or return null.
Excel example: Chart 1 sheet expects B/C anchors; if oracle skips rows where C is None, every D/E formula goes empty.
JSON example: a template file has `{"id": null, "value": ...}`; the agent populates `value` but never the `id`, so the test that joins on `id` fails for all rows.
Fix: the oracle (or instruction) must populate scaffolding before writing dependent values. Don't assume the input has them.
### Type-mismatched keys
Lookups that compare `datetime` to `"2019-01-01"` (string) silently miss every row. Same trap with `int` vs `str`, `Decimal` vs `float`, timezone-aware vs naive datetimes.
Fix: normalize key types before the lookup. Excel: `dt.datetime.fromisoformat(s)` before writing date columns. JSON: cast IDs explicitly. SQL: rely on the schema's declared types, not Python's.
### Locale / encoding drift
Tools that interpret numeric strings differently across locales: `1,5` → `1.5` (DE) vs `15` (US). UTF-8 vs Windows-1252 in CSV exports. CRLF vs LF in code diffs.
Fix: pin `LC_ALL=C.UTF-8` in the Dockerfile or test.sh. Set `encoding="utf-8"` explicitly when reading text. Never assume the host locale matches the container's.
## 3. Format-specific quirks
### Excel (openpyxl + LibreOffice)
The dominant SkillsBench oracle pattern when expected.xlsx contains formulas: write formula strings with openpyxl, then run LibreOffice via the bundled `recalc.py` to populate cached values:
```bash
cp tasks/xlsx-recover-data/environment/skills/xlsx/recalc.py tasks//oracle/recalc.py
```
```bash
python3 /oracle/recalc.py /root/output.xlsx 300
```
The `xlsx` skill is included with several existing tasks — reuse it directly, don't reimplement.
**Hard limit**: `_xlfn.XLOOKUP` array formulas. `calculateAll()` evaluates them but doesn't serialize cached values. Two paths: (a) use the `cp expected.xlsx` oracle, or (b) post-process with `soffice --headless --calc --convert-to xlsx --outdir /tmp/lo /root/output.xlsx && mv /tmp/lo/output.xlsx /root/output.xlsx` — the second open/save sometimes serializes values the macro doesn't.
### PowerPoint (python-pptx)
For tasks like `exceltable-in-ppt`, edit XML inside the embedded xlsx via the zip, never via python-pptx's chart API (it loses structure on save). See that task's `solve.sh` for the canonical pattern.
### PDF (pypdf / reportlab / pdfplumber)
For form-filling tasks (`court-form-filling`, `edit-pdf`): pypdf preserves layout; reportlab regenerates from scratch and loses original styling. Pick based on whether the test verifies content (pypdf is fine) vs visual layout (need a tool that doesn't reflow).
### Code (build / test / static analysis)
For tasks like `fix-druid-loophole-cve`, `fix-build-agentops`, the oracle is usually `cp /oracle/.patch /root/code && cd /root && git apply .patch`. Don't write the patch in the oracle — the patch IS the oracle.
### 3D / scientific (binary STL, NetCDF, FITS, CIF)
For `3d-scan-calc`, `crystallographic-wyckoff-position-analysis`: format-specific Python libs (numpy-stl, xarray, astropy, gemmi). Bundle test data; don't fetch live.
## 4. External APIs (research-track) — Playwright over urllib
ArcGIS REST, government open-data portals, search APIs have caching, rate-limiting, and pagination quirks that bite raw HTTP clients but pass through a real browser context cleanly.
```bash
# Dockerfile addition (drop --with-deps to avoid Debian font issues)
RUN pip install --no-cache-dir playwright==1.49.1 && \
playwright install chromium
```
```python
import time
from playwright.sync_api import sync_playwright
def get_with_retry(ctx, url, params, retries=4):
for attempt in range(retries):
response = ctx.request.get(url, params=params, timeout=120000)
if response.status == 200:
data = response.json()
if not data.get("error"):
return data
err = data.get("error", {}).get("code")
if err and 500 <= int(err) < 600:
time.sleep(2 ** attempt); continue
raise RuntimeError(f"API error: {data['error']}")
if 500 <= response.status < 600 or response.status in (304, 429):
time.sleep(2 ** attempt); continue
raise RuntimeError(f"HTTP {response.status}")
raise RuntimeError("retries exhausted")
with sync_playwright() as p:
browser = p.chromium.launch()
ctx = browser.new_context()
data = get_with_retry(ctx, ENDPOINT, params)
browser.close()
```
Always wrap external HTTP in retry-with-backoff. Government / ArcGIS / portal endpoints occasionally return 504 / 304 / 429 even when nominally up. Four-attempt exponential backoff (1s, 2s, 4s, 8s) recovers from most blips without slowing the happy path noticeably.
### Pagination — split keys, don't trust deep offsets
ArcGIS FeatureServer caps at ~1000 features per response and rejects `resultOffset` past 6k–23k for some queries. Split by a stable key (country×year, date×region, etc.) so each query's offset stays under the limit. Pattern applies broadly: many SaaS APIs have similar deep-pagination thresholds.
### Live internet is acceptable for stable government endpoints
The contributing guide prefers no-internet tasks but explicitly allows internet when the source is the canonical workflow. Document the source URL in the skill's `references/`. Government / standards-body endpoints (IMF, USGS, ECB, ICANN) are stable; private SaaS endpoints (Pinecone, Slack, Discord) are not — for those, bundle a snapshot.
Anchor a snapshot date in the instruction either way (see [time-invariance.md](time-invariance.md)).
## 5. Long-running oracles — budget and benchflow gotchas
Oracles that download multi-megabyte datasets, run LibreOffice recalc on tens of thousands of formulas, or invoke browser-based testing can take 5–10 minutes. Set `[agent].timeout_sec = 1800` and `[verifier].timeout_sec = 900`. Build timeout 600s is enough for Playwright + chromium install.
**Watch for benchflow's idle-600s detector.** If the oracle (or agent) launches a multi-minute Bash subprocess that doesn't emit ACP events, benchflow kills the trial mid-execution. Workarounds:
- Break the long subprocess into chunks with periodic status prints (`fetched 10000 rows…`).
- Use Python loops with intermediate `print(..., flush=True)` between iterations.
- File benchflow issues for new manifestations; track #211 for the canonical fix.
## Self-check before declaring the oracle done
```bash
bench eval run --tasks-dir tasks/ --agent oracle --sandbox docker --jobs-dir jobs/oracle-check
cat jobs/oracle-check/*/__*/result.json | python3 -c \
"import json,sys; r=json.load(sys.stdin)['rewards']; print(r)"
```
`{"reward": 1.0}` and nothing else. If you see fractional reward, fix it before the agent runs — agent runs on a half-passing oracle waste compute and produce uninterpretable signal.
## Three things people forget
1. **Verifier is locked, oracle is not.** `oracle/` is mounted at `/oracle/` for oracle runs and blocked from agent runs. Helper scripts and reference artifacts go there.
2. **`/root/`, not `/app/`.** The Dockerfile's `WORKDIR /root` is the agent's cwd; the verifier's `--rootdir` is `/app`. Tests should reference absolute paths regardless.
3. **For Excel: `recalc.py` reads cached values; openpyxl with `data_only=True` doesn't re-evaluate.** If the agent saves the workbook without running recalc, every cached value is `None`. Tests that check `cached_value is not None` are how you catch this.