RSIBench

Agent performance

Final capability vs.
RSI lift

Compare model–harness configurations without mistaking a strong starting point for strong self-improvement. The horizontal axis tracks held-out RSI lift; the vertical axis tracks final performance on sealed tests.

more persistent self-improvement stronger final capability
Illustrative chart preview The chart below demonstrates the interface only. Measured current-protocol results and archived old-codebase runs are listed in the evidence tables below.

Two-dimensional leaderboard

Performance × self-improvement

Filter the preview configurations, then hover or focus a point for its illustrative coordinates.

CHART DATA ILLUSTRATIVE
Harness
Provider
Illustrative RSIBench two-dimensional leaderboard A filterable scatter plot of illustrative model and harness configurations. Held-out RSI lift is on the horizontal axis, final sealed-test performance is on the vertical axis, and the upper-right direction is better. Formal runs are pending.
Minimal Pi harness Production harness Illustrative fleet reference

Showing all illustrative configurations

Agent leaderboard

Who finishes strongest—and who improves most?

The same illustrative configurations are ranked below. Switch the ranking target to separate final capability from recursive improvement; the Harness and Provider filters above remain active here.

Sort by
Harness
ILLUSTRATIVE RANKING · NOT MEASURED
#Agent configurationBefore RSIAfter RSIΔGain
Showing all illustrative configurationsBars encode the currently selected ranking metric.

Ablation experiments

Which harness surface carries the gain?

Trace capability across accepted RSI rounds, then restore Guidance, State, Control, or Action to its original state. The distance from the full harness exposes which surface is carrying the illustrative gain.

Illustrative capability across recursive self-improvement rounds A line chart comparing the full recursively improved harness against versions with Guidance, State, Control, or Action restored to the initial harness. Values are illustrative and formal experiments are pending.
Round 0 is the shared initial harness. Later checkpoints are inherited only after validation accepts the code change.Illustrative mechanism profile · not measured

How to read it

Two numbers, one honest comparison.

01

Move right

The inherited agent improved more over its matched pre-RSI baseline on held-out tasks.

02

Move up

The final inherited agent achieved stronger absolute sealed-test performance.

03

Trust verified runs

Only matched, repeated and sealed evaluations will enter the formal table below.

Formal results

Verified runs only

The interactive chart is an interface preview. GitHub-verified submissions load from the trusted registry after the matched three-repeat protocol finishes; no illustrative coordinate is copied into the table or paper.

#ModelHarnessDomainH0HTΔRSI ↑Test cost / task · H0→HT
Claude Opus 4.8Claude CodeTerminalpendingpendingpendingpending
Claude Opus 4.8PiTerminalpendingpendingpendingpending

Checked-in snapshot · trusted feed unavailable or awaiting verified results

Historical archive

Old codebase · preserved trajectories

These scores were produced before the current 15 train / 15 test, five-generation protocol and are not ranked against current results. The codebase and evaluator have since changed. Every scored row below links to its complete trajectory at an immutable rsibench-data commit.

Protocol boundary Earlier code · 13 train / 12 test · Terminal · Pi · avg@3 Download evidence manifest ↗
ModelHarnessH0HTΔRSI ↑Earlier protocolTrajectory evidence
GLM-5.2Z.ai · completePi 0.69%14.44%+13.75 pp T=10 · avg@35 accepted generations 5,877 artifacts ↗immutable archive
DeepSeek V4 ProDeepSeek · completePi 17.22%22.78%+5.56 pp T=5 · avg@32 accepted generations 1,999 artifacts ↗immutable archive
Kimi K3Moonshot · partial cutoffPi 47.78%47.78%0.00 pp T=5 · avg@31 accepted generation 2,675 artifacts ↗immutable archive
score snapshot only

GLM-5.2 FP8 · Pi

8.33% → 11.67%+3.33 pp · T=5 · avg@5

Recovered from the historical site commit. Full trajectory recovery is still pending, so this value is not ranked.
score snapshot only

Claude Opus 4.8 · Pi

78.33% → 82.50%+4.17 pp · T=1 · avg@3 · 8 test tasks

Recovered from the historical site commit. No complete matching trajectory has been located yet, so this pilot is not ranked.

Archive audited 2026-08-27 · “old codebase” means the implementation and scoring pipeline predate the current release. Read the trace policy.

Terminal · Pi ablations

Which intervention actually helps?

Each cell is a separate five-round Terminal + Pi run. Scores appear only after the trajectory PR is merged into rsibench-data; the absolute HT score is followed by the matched change from H0.

Study 01

Editable surface

Isolate where the Meta-agent may intervene. Skills + loader is a joint treatment that tests conditional knowledge against always-on instruction injection.

ModelPromptAgent loopTool runtimeObservationContextCompactionSkills + loaderHooks

Study 02

Cognitive controller

Factor the causal controller into trajectory findings, self-selected targets, post-validation review, and the complete causal treatment.

ModelFindings onlyTargets onlyReview onlyFull causal

rsibench-data registry · awaiting verified ablation submissions