{"article":{"slug":"frontis-ma1-training-an-ai4ai-model-towards-recursive-self-improvement-in-machine-learning-engineering","title":"Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering","subtitle":null,"summary":"Frontis.AI / Horizon Research open-source OpenMLE (gym, RL, Evo) and Frontis-MA1-35B, lifting MLE-Bench Lite medal average to 71.21% under a single RTX 4090 budget toward executable RSI research.","content_type":"research","language":"en","canonical_url":"https://frontisai.github.io/OpenRSI/","author":{"name":"Junlin Yang et al.","url":"https://frontisai.github.io/OpenRSI/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Frontis.AI","url":"https://frontisai.github.io/OpenRSI/","listing_slug":null,"listing":null},"topics":[{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":385,"reading_minutes":2,"published_at":"2026-09-01T00:00:00.000Z","added_at":"2026-09-25T06:19:32.364Z","updated_at":"2026-09-25T06:19:32.364Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/frontis-ma1-training-an-ai4ai-model-towards-recursive-self-improvement-in-machine-learning-engineering","markdown_url":"https://listedarticles.com/articles/frontis-ma1-training-an-ai4ai-model-towards-recursive-self-improvement-in-machine-learning-engineering.md","example":false,"citation":"Junlin Yang et al., Frontis.AI. \"Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering.\" 1 Sept 2026. https://frontisai.github.io/OpenRSI/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://frontisai.github.io/OpenRSI/"},"body_markdown":"# Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering\n\nHorizon Research, Frontis.AI · Tsinghua University\n\nOpen weights · open gym · open search — the full OpenMLE stack, released\n\n## Abstract\n\nRecursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop.\n\nOn MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI.\n\n## Stack overview\n\n- **OpenMLE-Gym** — a gym, not a dataset: thousands of executable tasks with structured sandbox feedback modes (MLE-Bench excluded from training).\n- **OpenMLE-ERL** — execution-grounded SFT + RL with asynchronous rollouts.\n- **OpenMLE-Evo** — test-time scaling toward test-time learning with experience cards and operator-conditioned memory.\n- **Frontis-MA1 (30B / 35B)** — trained by OpenMLE, driving OpenMLE, evaluated on third-party benchmarks.\n\n## Four operators\n\nDraft (generate from scratch), Improve (refine a parent), Debug (repair failing code), and Crossover (recombine two parents) form a unified action space for code evolution, invoked thousands of times per task.\n\n## Results snapshot\n\n| System | Medal Average (MLE-Bench Lite) |\n| --- | --- |\n| Qwen3.6-35B-A3B base · OpenMLE-Evo | 39.39 |\n| Frontis-MA1-35B post-trained · OpenMLE-Evo | 60.61 |\n| Claude Opus 4.8 Claude Code | 63.64 |\n| GPT-5.5 Codex | 68.18 |\n| Frontis-MA1-35B OpenMLE-Evo-Max | 71.21 |\n| GPT-5.6 Sol / Kimi K3 | 72.73 |\n\n## Release\n\nWeights, gym, sandbox, training, search, and evaluation harness are released for reproducible AI4AI / RSI research. Paper: arXiv:2607.28568.","body_html":"<h1 id=\"frontis-ma1-training-an-ai4ai-model-towards-recursive-self-impro\">Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h1>\n<p>Horizon Research, Frontis.AI · Tsinghua University</p>\n<p>Open weights · open gym · open search — the full OpenMLE stack, released</p>\n<h2 id=\"abstract\">Abstract</h2>\n<p>Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop.</p>\n<p>On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI.</p>\n<h2 id=\"stack-overview\">Stack overview</h2>\n<ul><li><strong>OpenMLE-Gym</strong> — a gym, not a dataset: thousands of executable tasks with structured sandbox feedback modes (MLE-Bench excluded from training).</li><li><strong>OpenMLE-ERL</strong> — execution-grounded SFT + RL with asynchronous rollouts.</li><li><strong>OpenMLE-Evo</strong> — test-time scaling toward test-time learning with experience cards and operator-conditioned memory.</li><li><strong>Frontis-MA1 (30B / 35B)</strong> — trained by OpenMLE, driving OpenMLE, evaluated on third-party benchmarks.</li></ul>\n<h2 id=\"four-operators\">Four operators</h2>\n<p>Draft (generate from scratch), Improve (refine a parent), Debug (repair failing code), and Crossover (recombine two parents) form a unified action space for code evolution, invoked thousands of times per task.</p>\n<h2 id=\"results-snapshot\">Results snapshot</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>System</th><th>Medal Average (MLE-Bench Lite)</th></tr></thead><tbody><tr><td>Qwen3.6-35B-A3B base · OpenMLE-Evo</td><td>39.39</td></tr><tr><td>Frontis-MA1-35B post-trained · OpenMLE-Evo</td><td>60.61</td></tr><tr><td>Claude Opus 4.8 Claude Code</td><td>63.64</td></tr><tr><td>GPT-5.5 Codex</td><td>68.18</td></tr><tr><td>Frontis-MA1-35B OpenMLE-Evo-Max</td><td>71.21</td></tr><tr><td>GPT-5.6 Sol / Kimi K3</td><td>72.73</td></tr></tbody></table></div>\n<h2 id=\"release\">Release</h2>\n<p>Weights, gym, sandbox, training, search, and evaluation harness are released for reproducible AI4AI / RSI research. Paper: arXiv:2607.28568.</p>","headings":[{"level":1,"text":"Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering","id":"frontis-ma1-training-an-ai4ai-model-towards-recursive-self-impro"},{"level":2,"text":"Abstract","id":"abstract"},{"level":2,"text":"Stack overview","id":"stack-overview"},{"level":2,"text":"Four operators","id":"four-operators"},{"level":2,"text":"Results snapshot","id":"results-snapshot"},{"level":2,"text":"Release","id":"release"}]}}