Can a Model Learn New Skills as Add-Ons?
Each new skill is trained as a small, separate expert with its own router, and experts trained independently can be merged into one model in seconds without losing what each one learned.
If you want a large model to get better at something new, say medicine or code, the usual answer is to fine-tune it.
That works, but it has two costs. Fine-tuning edits a network that already supports everything else the model knows, so a new skill can quietly damage old ones. And every new skill means another pass over the model, run by whoever owns the whole model.
What if a new skill could be built separately, by someone else, and simply plugged in?
What if a new skill were an add-on, not a rewrite?
Our latest experiments test exactly that on a 15.7B-parameter Mixture-of-Experts (MoE) model. In an MoE model, each layer holds many small "expert" networks, and a router picks a handful of them for every token. Our base model picks 6.
Instead of retraining any of the existing experts, we append one new expert to each layer and give it its own router. The original model stays frozen, exactly as it was. At the start of training the new expert contributes nothing, so the model begins identical to the original.
The new router doesn't compete with the old one. It has one question to answer: does this token need the new skill? If its score passes a threshold, the new expert joins in on top of the usual six; if not, it stays silent.
In symbols, the layer's output is the original output plus a gated contribution from the new expert:
h(x) = h_original(x) + g(x) · E_new(x)
where g(x) is close to 0 when the new router's score is below the threshold and close to 1 above it.
How does the new expert learn when to stay quiet?
The router is what makes the add-on safe. It is trained on two kinds of data at once:
| Data | Example | What the router is taught |
|---|---|---|
| General data | ordinary web text | Stay off. Any activation is penalised. |
| Expertise data | math problems, for a math expert | Switch on, at a chosen target rate (20% of tokens in these runs). |
The intuition fits in one line:
Stay out of the way on what the model already knows. Show up on the new capability.
The model's existing knowledge is never edited. The new expert only adds to it, and only where it is needed.
Does each expert learn its skill?
We trained five experts separately: math, code, medicine, law and finance. Each ran for 500 steps on its own data, without any knowledge of the others. Here is each one next to the untouched base model, on a benchmark from its own domain. Higher is better.
| Expert | Benchmark | Base model | Base + expert |
|---|---|---|---|
| Math | GSM8K | 37.9 | 61.7 |
| Code | HumanEval | 27.4 | 35.4 |
| Medical | MedQA | 42.3 | 46.7 |
| Law | LegalBench | 48.6 | 53.7 |
| Finance | FinQA | 1.2 | 26.6 |
Every expert lifts its own domain, and none of the original model was changed to get there.
Do independently trained experts still work together?
This is the real test. An add-on that works alone is useful; add-ons that can be stacked are a different kind of asset.
So we took experts from those separate runs and merged them into the same base model, two or three at a time. Merging copies each expert and its router into place, side by side. Nothing is averaged, and nothing is retrained.
Math + Code
| Model | GSM8K (math) | HumanEval (code) | MBPP (code) |
|---|---|---|---|
| Base model | 37.9 | 27.4 | 42.8 |
| Math expert alone | 61.7 | 33.5 | 41.6 |
| Code expert alone | 41.2 | 35.4 | 42.2 |
| Math + Code, merged | 59.5 | 37.8 | 44.6 |
Math + Medical
| Model | GSM8K (math) | MedQA (medical) | MedMCQA (medical) |
|---|---|---|---|
| Base model | 37.9 | 42.3 | 40.4 |
| Math + Medical, merged | 60.6 | 47.1 | 46.3 |
Math + Code + Law
| Model | GSM8K (math) | HumanEval (code) | LegalBench (law) |
|---|---|---|---|
| Base model | 37.9 | 27.4 | 48.6 |
| Math + Code + Law, merged | 53.0 | 35.4 | 52.0 |
In every merged model, every score is above the base model. Each run is a single seed, so differences of a point or two are within noise. The pattern is the point: skills trained apart, combined afterwards, and none was lost.
- Size of one new skill: ~1.4% of the model's parameters (~225M on a 15.7B base)
- Cost of combining skills: seconds on a CPU, no retraining
Why does this matter as models grow?
Because it changes what a unit of AI improvement is.
Today, improving a large model is one big job: one owner, one training run, and one model at the end. With add-on experts, a new capability becomes a small, separate component:
- It can be trained in parallel by different contributors, on different data and hardware, without waiting for each other.
- It is cheap relative to the model: about 1.4% of the parameters here, and a smaller share still on a larger base.
- It can be combined on demand, so a customer gets exactly the set of skills their workload needs.
The same recipe does not depend on the base model's size. As the base grows toward 100B parameters and beyond, a new skill still means training one small expert and its router, not retraining the whole. That is how model improvement becomes fast enough, and cheap enough, to deliver to market one capability at a time.
Train a skill once, reuse it everywhere. That is the building block Connito's training network is designed around.