Two-dimensional leaderboard
Performance × self-improvement
Filter the preview configurations, then hover or focus a point for its illustrative coordinates.
Agent performance
Compare model–harness configurations without mistaking a strong starting point for strong self-improvement. The horizontal axis tracks held-out RSI lift; the vertical axis tracks final performance on sealed tests.
Two-dimensional leaderboard
Filter the preview configurations, then hover or focus a point for its illustrative coordinates.
Agent leaderboard
The same illustrative configurations are ranked below. Switch the ranking target to separate final capability from recursive improvement; the Harness and Provider filters above remain active here.
| # | Agent configuration | Before RSI | After RSI | Δ | Gain |
|---|
Ablation experiments
Trace capability across accepted RSI rounds, then restore Guidance, State, Control, or Action to its original state. The distance from the full harness exposes which surface is carrying the illustrative gain.
How to read it
The inherited agent improved more over its matched pre-RSI baseline on held-out tasks.
The final inherited agent achieved stronger absolute sealed-test performance.
Only matched, repeated and sealed evaluations will enter the formal table below.
Formal results
The interactive chart is an interface preview. GitHub-verified submissions load from the trusted registry after the matched three-repeat protocol finishes; no illustrative coordinate is copied into the table or paper.
| # | Model | Harness | Domain | H0 | HT | ΔRSI ↑ | Test cost / task · H0→HT |
|---|---|---|---|---|---|---|---|
| — | Claude Opus 4.8 | Claude Code | Terminal | pending | pending | pending | pending |
| — | Claude Opus 4.8 | Pi | Terminal | pending | pending | pending | pending |
Checked-in snapshot · trusted feed unavailable or awaiting verified results
Historical archive
These scores were produced before the current 15 train / 15 test, five-generation protocol and are not ranked against current results. The codebase and evaluator have since changed. Every scored row below links to its complete trajectory at an immutable rsibench-data commit.
| Model | Harness | H0 | HT | ΔRSI ↑ | Earlier protocol | Trajectory evidence |
|---|---|---|---|---|---|---|
| GLM-5.2Z.ai · complete | Pi | 0.69% | 14.44% | +13.75 pp | T=10 · avg@35 accepted generations | 5,877 artifacts ↗immutable archive |
| DeepSeek V4 ProDeepSeek · complete | Pi | 17.22% | 22.78% | +5.56 pp | T=5 · avg@32 accepted generations | 1,999 artifacts ↗immutable archive |
| Kimi K3Moonshot · partial cutoff | Pi | 47.78% | 47.78% | 0.00 pp | T=5 · avg@31 accepted generation | 2,675 artifacts ↗immutable archive |
8.33% → 11.67%+3.33 pp · T=5 · avg@5
Recovered from the historical site commit. Full trajectory recovery is still pending, so this value is not ranked.78.33% → 82.50%+4.17 pp · T=1 · avg@3 · 8 test tasks
Recovered from the historical site commit. No complete matching trajectory has been located yet, so this pilot is not ranked.Archive audited 2026-08-27 · “old codebase” means the implementation and scoring pipeline predate the current release. Read the trace policy.
Terminal · Pi ablations
Each cell is a separate five-round Terminal + Pi run. Scores appear only after the trajectory PR is merged into rsibench-data; the absolute HT score is followed by the matched change from H0.
Study 01
Isolate where the Meta-agent may intervene. Skills + loader is a joint treatment that tests conditional knowledge against always-on instruction injection.
| Model | Prompt | Agent loop | Tool runtime | Observation | Context | Compaction | Skills + loader | Hooks |
|---|
Study 02
Factor the causal controller into trajectory findings, self-selected targets, post-validation review, and the complete causal treatment.
| Model | Findings only | Targets only | Review only | Full causal |
|---|
rsibench-data registry · awaiting verified ablation submissions