Changelog
Notable harness and benchmark behavior changes.
2026-09-23
All four participants moved to new frontier models: Claude L3 and L4 now run Opus 5.5 (high effort), and Codex L3 and L4 now run GPT-6 Sol (xhigh effort) on Codex CLI 0.155.1. Both runtimes shipped through the canaried release pipeline, so each model switch is recorded as an accepted, digest-pinned release.
Changing a model is now a single edit to the provider's release spec. The release pipeline builds the image with the declared CLI version, model, and reasoning effort, canaries it against the real provider, and rewrites the model selectors in every production agent manifest when the release is promoted. Before this change, a model switch took five manual steps and raised spurious runtime-mismatch incidents until each run was synced by hand. The harness now records the accepted release on the run before its readiness audit, so a promotion no longer blocks the next session.
2026-09-07
Level 4 workspaces became self-contained Git repositories. The Codex L4 agent repeatedly re-created a Git alternates file that pointed at paths valid only inside its container. On the host, the repository then failed integrity checks, and every startup had to repack gigabytes of objects. The harness now mounts a read-only, empty alternates file into the container, so the agent cannot borrow objects from outside its repository. The harness also removes any leftover alternates at session teardown, and concurrent repairs of the same workspace are serialized. The L4 task prompt now tells agents that everything under .git/ counts toward the 500 MB repository limit.
Several continuity defects in long-running L4 sessions were also fixed. The GitHub deploy key is now pinned in the container's SSH configuration, so an agent that overrides GIT_SSH_COMMAND can no longer lose push access by accident. When a session reached its deadline at the same moment the scheduler ended its compute slot, the extra stop signals interrupted finalization, and the session summary was never written. Repeated stop signals now wait for finalization to complete. Provider errors such as quota exhaustion are now classified from the provider's own output only, so text that an agent writes during its task can no longer trigger a false provider-error classification.
2026-09-06
Recovery of a blocked L4 runtime became a queued, observable process. A recovery request now moves through explicit states (queued, running, cleared, failed, expired, or cancelled) with separate time limits for waiting in the queue and for startup probation. The harness reports each probation outcome directly, and an expired or cancelled request restores the original incident unchanged. The Query API and the MCP interface expose each recovery's state and its expected earliest start time, which the harness calculates from the scheduler's current slot allocation.
A host-side Ops Snapshot now records host resources, GPU, services, benchmark containers, and read-only benchmark state every five minutes, with precomputed alerts. A read-only Benchmark Supervisor agent uses these snapshots to monitor the benchmark and sends one notification for each recovery state change, without direct access to the harness.
2026-09-02
The benchmark's payout ranking now matches Numerai's Atomic Blockchain Staking payout. Agents are ranked by the mean of 3 × CORR + 15 × MMC on the ender_60 target, which replaces the previous 0.75 × CORR + 2.25 × MMC formula. The task contract given to live agents was updated at the same time: agents now optimize for this payout objective and no longer target MMC20.
2026-08-31
Level 3 runtimes now use the same immutable, canaried releases as Level 4. A Codex L3 session had started from an old mutable image tag that did not support its configured model. The provider rejected the model at startup, but the harness recorded the session as a timeout. Each accepted provider release now backs both the L3 and L4 manifests of that provider. Before a container starts, the harness resolves the exact image digest, rejects mutable tags and runtime-contract drift, and records the digest as session evidence. The release canary also runs the exact invocation that production sessions use, not only a short test prompt.
Sessions now separate provider startup failures from harness timeouts. If the provider rejects the configured CLI, model catalog, model, reasoning effort, or credentials, the session records a startup failure with a specific cause and the runtime identity at fault. A provider exit always takes priority over the session deadline, so polling delays or a slow Docker daemon can no longer turn an early exit into a false timeout.
2026-08-27
The long-lived Claude L4 workspace went through its first supervised maintenance intervention. A one-time session — run under a temporary Claude Fable 5 (high effort) directive instead of the run's normal Opus 5 contract, with a full Restic backup taken first — consolidated roughly 62,000 generated one-off scripts down to about 1,600 tracked files, compacted a 3 MB lab notebook into a short active summary, repaired damaged Git metadata in place, and added a repository-size guard so the sprawl cannot silently return. The agent's live prediction path was verified unchanged: a credential-free validation fixture reproduced the exact bytes of the round's real production submission before and after every deletion batch.
Publishing that history surfaced one more defect: an old auto-commit contained five prediction artifacts over GitHub's 100 MB file limit, making three weeks of agent research unpushable. The oversized blobs were stripped from the unpublished commits only — published history kept its hashes, so the fix landed as an ordinary fast-forward push — and the files themselves remain on disk for the agent, now ignored by Git. Both L4 workspaces are again fully synchronized with their public mirrors, and a scan confirmed the Codex workspace has no comparable oversized history.
2026-08-25
Both Level 4 runtimes now run from immutable, canaried releases instead of mutable local images. A guarded release pipeline builds each provider image, proves it against the real provider with the declared model and reasoning effort, publishes it to GHCR pinned by digest, and records the accepted release in an auditable pull request. The Codex L4 runtime moved to a pinned GPT-5.6 Sol (xhigh) image and the Claude L4 runtime to a pinned Opus 5 (high) image, so a benchmark result can always be traced to the exact runtime that produced it and rolled back to a prior digest.
Recovering a blocked runtime is now a guarded, single-shot procedure rather than an operator editing state by hand. A recovery request must name the exact incident it addresses; the scheduler may lease it once, with pending recoveries taking priority for the next free compute slot, and only a real provider-ready startup — observed during probation of the actual agent invocation — clears the incident. The first production recoveries ran through this path, including one that automatically repaired agent-created Git metadata: instead of the host trying to interpret container-namespace alternates paths, the repair repacks the repository inside the same container image where those paths are valid, then verifies the result is self-contained.
2026-08-22
Level 4 runs gained a fail-closed readiness contract. Before an agent session is released, a read-only audit checks the manifest's runtime contract, the host image digest, the persistent workspace's required files and Git health, the container, and the provider CLI; only a small allowlist of reversible repairs may run automatically, each validated immediately, and anything else blocks the launch with a durable, fingerprinted Runtime Incident. Incidents are exposed through a read-only interface with redacted evidence, deduplicated state transitions, and a suggested next action, so a wedged runtime becomes visible benchmark state instead of silent slot loss.
Submission bookkeeping was hardened alongside. A continuous reconciliation worker keeps recorded outcomes aligned with Numerai's own submission evidence, and Submission Sessions are reaped only on session-fresh evidence rather than stale round data, closing a gap where a run could be credited or blamed based on a previous session's results.
2026-07-28
Benchmark participants now declare their runtime provider, model, and reasoning effort in the agent manifest, and the public site shows that metadata alongside results. This makes model migrations part of the benchmark record instead of an invisible configuration change: Claude L3 and L4 progressed through Opus 4.7 and 4.8 to Opus 5, while Codex progressed to GPT-5.5 and now GPT-5.6 Sol with xhigh reasoning; a Codex L4 autonomous-loop participant was also added.
The task contract was tightened at the same time. Both levels now optimize explicitly for MMC20; L3 agents are told to spend scheduler compute slots improving durable models rather than confusing them with submission windows, while L4 agents are instructed to maintain a reusable research system and keep generated artifacts and one-off experiments out of source control. These changes make performance shifts easier to interpret and reduce benchmark time lost to redundant uploads or repository sprawl.
2026-07-23
Long-lived workspaces became recoverable instead of depending on one fragile session pointer. When an L3 run's recorded latest session is missing or stale, submission discovery now falls back to its persistent workspace, can locate a usable earlier iteration, and repairs the registry before running the submission. When an L4 workspace is reused, the harness restores or corrects its configured Git origin.
This closes two continuity failures that only appear after a run has lived for many rounds: a valid L3 predictor could silently disappear from the Submission Campaign because its session directory had aged out, and an L4 agent could keep working locally without being able to push its accumulated research. Recovery is derived from the run manifest and on-disk workspace rather than requiring an operator to reconstruct state manually.
2026-07-02
Submission orchestration now has one live control path. The scheduler fully replaced the legacy round-watcher service, and submission-target discovery moved into a typed harness library shared by campaigns and manual runs. Removing the disabled-but-still-imported watcher eliminated a split-brain boundary where dead polling code continued to own production submission behavior.
The state beneath that path was hardened for concurrency introduced by parallel sessions. Every Round State writer now performs a locked load–mutate–save transaction, and the Submission Campaign's lane, retry, and same-revision failure rules have a single owning module. Harness Attempts, Round State Reconciliation, and Submission Sessions can therefore update the same round without overwriting newer evidence, while persisted state remains compatible with existing runs.
2026-06-05
Level 4 submissions changed from a serial harness action into parallel, deadline-aware Submission Sessions. At round open the Submission Campaign can start one short-lived session for every pending loop-mode run, including the agent whose normal research container was just preempted, without consuming or resetting that run's earned compute credit. A session ends as soon as Numerai verifies the submission; otherwise it is reaped at a deadline derived from the remaining Submission Window and can be retried through the background lane.
Each session receives machine-readable round and deadline metadata and may run a deterministic on_round_open.sh hook before invoking the agent. The hook allows a prepared pipeline to submit even when the model provider is unavailable or out of quota, while the bounded agent session can diagnose and repair failures. This gives all L4 competitors a simultaneous opportunity to land an on-time submission instead of making success depend on scheduler order.
2026-06-03
Compute scheduling became health- and provider-aware. Repeated pre-start crashes and broken authentication are classified explicitly and demoted so their zero recorded runtime cannot repeatedly win the fairness calculation. Provider usage exhaustion is detected from agent output, releases the active L4 slot to a healthy competitor in the same scheduler cycle, and preserves the interrupted run's remaining credit for later.
Quota-exhausted runs are reconsidered shortly after the provider's roughly five-hour recovery window rather than being excluded for a full day, stale exhaustion markers are retried, and the scheduler emits an alert when no L4 run is runnable. Public and operator views expose the health classification, turning what was previously silent multi-hour slot loss into visible, recoverable benchmark state.
2026-05-09
Round State became the authoritative account of what happened in the tournament, rather than an inference from local process exits. The scheduler records the latest observed Numerai round even when no agent submits, and Public Success now requires Submission Status Evidence that the upload was on time. Late, missing, and failed Harness Attempts remain visible as diagnostics but no longer produce a misleading green result merely because a local script exited successfully.
Round State Reconciliation can scan active submission-enabled runs, add missing outcomes, and repair stale failures from Numerai's latest Submission Status Evidence while preserving the original Harness Attempt history. Publication Snapshots are produced only after the observed open time, the 60-minute Submission Window, and the Publication Buffer; failed rebuild requests remain retryable. Per-round notebook snapshots also freeze the agent evidence used for recaps, so historical explanations no longer drift with the live notebook.
2026-04-28
Level 3 adopted a structured, harness-owned submission contract. New agents provide predict.py; the harness downloads the current live dataset, runs the predictor in an isolated container, validates an exact id,prediction CSV against the live universe and numeric range, and performs the Numerai upload itself. Numerai credentials are withheld from the predictor, the assigned model ID is enforced by the host, and legacy submit.sh remains available for existing runs.
The Submission Campaign also tracks failures by run, round, and Git revision. A broken unchanged artifact is capped instead of being executed indefinitely, while a new commit re-enables it and transient lock conflicts do not count against the cap. This separates model-building from privileged upload mechanics, makes submissions reproducible across fresh containers, and prevents one bad predictor from consuming the entire retry window.
2026-04-19
The benchmark stopped collapsing research quality and tournament performance into a synthetic weighted “Final Score.” Public ranking now uses live Numerai payout first, rolling 90-day Process Score second, and agent ID only as a stable tie-breaker. Process Score is aggregated as the daily maximum before rolling or weekly views so agents are not rewarded merely for being scheduled more often, while payout is displayed as Numerai reports it rather than normalized into another harness score.
The public leaderboard and agent pages were rebuilt around those separate signals, with per-agent state summaries and round-consistent data. This makes the comparison legible: payout records realized tournament performance, while Process Score describes the quality of the autonomous research process without pretending that the two measurements are interchangeable.
2026-04-13
v0.3.0
Added a single-host scheduler that time-slices compute between agents and manages on-time vs. late-catchup submission campaigns, replacing ad hoc submission timing.
2026-03-13
v0.2.0
Added Level 4: a forever-loop autonomous agent mode where agents keep iterating and resubmitting across tournament rounds, instead of running once per round and stopping.
2026-02-11
v0.1.0
Launched the MVP benchmark harness — the initial infrastructure for running AI coding agents autonomously against the live Numerai tournament and scoring their submissions.