Files
2026-09-04 14:58:42 +08:00

112 lines
5.6 KiBLFS
Markdown

# Dataset versioning
How SkillsBench versions its task set, modeled on Terminal-Bench's 2.0 → 2.1
practice (immutable, separately published dataset versions, decoupled from the
harness) — with the conventions written down, which Terminal-Bench never did.
## Principles
1. **A dataset version's content identity is immutable.** The task roster,
per-task digests, and git snapshot pins of a published entry are never
edited, re-pointed, or yanked — a content fix ships as a *new* version
with its own registry entry and digests. Results, validation runs, and
paper numbers attach to a version forever. Compat/metadata fields
(`bench_version`, `description`, `notes`) are *not* part of that
identity and may be corrected in place, with the correction recorded in
`notes`: they change which harnesses may run a version, never what its
results mean.
2. **Dataset and harness are versioned independently.** The harness
([benchflow](https://github.com/benchflow-ai/benchflow), the `bench` CLI)
is a Python package with its own releases; each registry entry declares the
compatible range in `bench_version`. A run declares the dataset it used.
3. **Supersession is annotated, never silent.** When a new version replaces an
old one, the old version's release notes and leaderboard get a
"superseded by X" banner. Old leaderboards stay up — flagged, not deleted.
4. **One leaderboard namespace per dataset version.** Results from different
versions are never merged by default; cross-version comparison is a
deliberate analysis.
## Version semantics
| Bump | Meaning | Example |
| ----- | ------------------------------------- | ---------------------------------------- |
| minor | in-place task fixes; roster unchanged | verifier/path fixes on the same task set |
| major | roster changes | adding/removing/swapping tasks |
There is no patch tier: compat/metadata fields are corrected in place on
the existing entry (principle 1), so version numbers move only when task
content does — one current version stays one leaderboard.
## Registry
[`registry.json`](../registry.json) at the repo root lists every published
dataset version. Each entry pins:
- `git_tag` / `git_commit_id` — the exact repo snapshot,
- per task: `path` and a content `digest`,
- `bench_version` — the harness range the version was validated against
(`null` for versions predating the `bench` CLI, with the actual harness
recorded in `notes`).
## Task digest
`digest` pins task *content*, not just a commit:
```
task_digest = sha256( for each regular file under the task directory,
sorted by POSIX relative path:
update( path_utf8 + b"\x00" + sha256(file_bytes).digest() ) )
```
Symlinks and file modes are excluded. The hex digest is prefixed `sha256:`.
The digest is reproducible from a plain checkout — it does not depend on git
object formats. (`bench tasks digest` is the planned CLI home for this
computation; until it lands, the reference implementation lives in the release
tooling.)
## Published versions
| Version | Git tag | Snapshot | Contents |
| ------- | ------- | -------- | -------- |
| `skillsbench@1.0` | [`v1.0`](https://github.com/benchflow-ai/skillsbench/releases/tag/v1.0) | `14f33967` (2026-06-12) | 87-task benchmark set before the native package cutover. Superseded by 1.1 (same roster, native `task.md` packages) for current evaluation. |
| `skillsbench@1.1` | [`v1.1`](https://github.com/benchflow-ai/skillsbench/releases/tag/v1.1) | `27738384` (2026-06-14) | Native `task.md` package release (task.md + environment/ + oracle/ + verifier/) for the 87-task roster. Compatible with BenchFlow `>=0.6.2,<0.7` (latest at release: 0.6.2). |
Each release ships a `skillsbench-<version>-task-manifest.json` asset with the
full per-task archive (category, difficulty, tags, last substantive commit,
digest, archived-folder URL).
### Renumbering note (2026-06-15)
The release line was consolidated on 2026-06-15: the former `v1.2` (native
`task.md` packages) was renamed to **`v1.1`**, the former `v1.1` (87-task set)
to **`v1.0`**, and the original `v1.0` — the arXiv-v1 paper snapshot (86 tasks,
84 evaluated) — was retired as a published release. That paper snapshot is
still reachable at commit `adbaf4d7` for reproducing the paper, but is no
longer tagged or listed here. This was a one-time pre-stable consolidation; the
Principle 1 immutability guarantee applies from this renumbered `v1.0`/`v1.1`
line forward.
## Run plumbing (planned)
Tracked as follow-up engineering work in the harness and website:
- `bench eval run -d skillsbench@1.1 ...` resolves the registry entry
(publishable runs); `--tasks-dir ./local` stays as visibly-distinct dev mode.
- Every `result.json` gets stamped with `dataset_name`, `dataset_version`, and
per-task `task_digest`.
- The leaderboard groups results by `dataset@version` and shows a superseded
banner on outdated boards.
## Roadmap
- **`skillsbench@1.0` referent (resolved 2026-06-12):** it pins the git tag
`v1.0` at snapshot `14f33967`, the 87-task set before the native package
cutover. The registry entry records `git_tag`, `git_commit_id`, and per-task
content digests.
- **`skillsbench@1.1` referent (resolved 2026-06-14):** it pins the git tag
`v1.1` at snapshot `27738384`, the native `task.md` package release for the
same 87-task roster as 1.0. The registry entry records `git_tag`,
`git_commit_id`, and per-task content digests computed by the release
tooling's reference implementation. No in-repo digest CLI exists yet
(`bench tasks digest` is planned).