---
title: "Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering"
slug: frontis-ma1-training-an-ai4ai-model-towards-recursive-self-improvement-in-machine-learning-engineering
url: https://listedarticles.com/articles/frontis-ma1-training-an-ai4ai-model-towards-recursive-self-improvement-in-machine-learning-engineering
canonical_url: https://frontisai.github.io/OpenRSI/
content_type: research
language: en
published_at: 2026-09-01T00:00:00.000Z
updated_at: 2026-09-25T06:19:32.364Z
author: "Junlin Yang et al."
author_url: https://frontisai.github.io/OpenRSI/
authored_by: human
publisher: "Frontis.AI"
publisher_url: https://frontisai.github.io/OpenRSI/
topics: ["Machine Learning", "Research", "AI Agents", "Open Source", "Benchmarks"]
license: all-rights-reserved
word_count: 385
reading_minutes: 2
citation: "Junlin Yang et al., Frontis.AI. \"Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering.\" 1 Sept 2026. https://frontisai.github.io/OpenRSI/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

> Frontis.AI / Horizon Research open-source OpenMLE (gym, RL, Evo) and Frontis-MA1-35B, lifting MLE-Bench Lite medal average to 71.21% under a single RTX 4090 budget toward executable RSI research.

# Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Horizon Research, Frontis.AI · Tsinghua University

Open weights · open gym · open search — the full OpenMLE stack, released

## Abstract

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop.

On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI.

## Stack overview

- **OpenMLE-Gym** — a gym, not a dataset: thousands of executable tasks with structured sandbox feedback modes (MLE-Bench excluded from training).
- **OpenMLE-ERL** — execution-grounded SFT + RL with asynchronous rollouts.
- **OpenMLE-Evo** — test-time scaling toward test-time learning with experience cards and operator-conditioned memory.
- **Frontis-MA1 (30B / 35B)** — trained by OpenMLE, driving OpenMLE, evaluated on third-party benchmarks.

## Four operators

Draft (generate from scratch), Improve (refine a parent), Debug (repair failing code), and Crossover (recombine two parents) form a unified action space for code evolution, invoked thousands of times per task.

## Results snapshot

| System | Medal Average (MLE-Bench Lite) |
| --- | --- |
| Qwen3.6-35B-A3B base · OpenMLE-Evo | 39.39 |
| Frontis-MA1-35B post-trained · OpenMLE-Evo | 60.61 |
| Claude Opus 4.8 Claude Code | 63.64 |
| GPT-5.5 Codex | 68.18 |
| Frontis-MA1-35B OpenMLE-Evo-Max | 71.21 |
| GPT-5.6 Sol / Kimi K3 | 72.73 |

## Release

Weights, gym, sandbox, training, search, and evaluation harness are released for reproducible AI4AI / RSI research. Paper: arXiv:2607.28568.
