
Medal rate by difficulty split, against R&D-Agent, AIRA-dojo, ML-Master and AIDE
These results were submitted as an official submission to MLE-Bench.
Usage
CLI options
Configuration modes
Both modes run the
generic search strategy, but the MLE handler supplies the official evaluation: solutions are scored by mlebench grading rather than agent-built evaluation.- MLE_GENERIC
- MINIMAL
The campaign standard. Ideation on Claude, implementation sessions on the Codex CLI.
Stages
The handler automatically adjusts strategy based on budget progress:Output structure
The agent generates:Code requirements
Generated code must:- Support
--debugflag for fast testing - Write
final_submission.csvin the output directory - Print progress and metrics
- Handle GPU efficiently (batch size, device selection)
- Use early stopping and learning rate scheduling
Competition types
Environment variables
Related
Installation
Installing the harness
CLI reference
Running a campaign
pip install leeroo-kapso · Every page as plain text: llms.txt.