Leaderboard
MLS-Bench-Lite Score. The evaluation is based on Harbor with a 5-hour exploration budget for each agent.
MLS-Bench Lite
Human SOTA
| # | Model | Harness | Performance |
|---|---|---|---|
| 1 | Claude Codemax(with fallback) | 49.9 | |
| 2 | Kimi-Codemax | 48.3 | |
| 3 | Codexmax | 46.2 | |
| 4 | Claude Codemax | 42.8 | |
| 5 | Claude Code | 41.0 | |
| 6 | Claude Codemax | 40.4 | |
| 7 | Codexxhigh | 35.5 | |
| 8 | Kimi-Code | 35.1 | |
| 9 | Claude Code | 31.7 | |
| 10 | Kimi-Code | 26.7 |