{"article":{"slug":"can-a-model-learn-new-skills-as-add-ons","title":"Can a Model Learn New Skills as Add-Ons?","subtitle":null,"summary":"Connito Research trains residual MoE experts with their own routers on a frozen DeepSeek-V2-Lite base, then merges independently trained math, code, medical, law, and finance experts in seconds without retraining—lifting domain benchmarks while leaving the original model untouched.","content_type":"research","language":"en","canonical_url":"https://connito.ai/blog/can-a-model-learn-new-skills-as-add-ons","author":{"name":"Connito Research","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Connito","url":"https://connito.ai/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":913,"reading_minutes":4,"published_at":"2026-09-28T12:00:00.000Z","added_at":"2026-09-29T09:18:23.062Z","updated_at":"2026-09-29T09:18:23.062Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/can-a-model-learn-new-skills-as-add-ons","markdown_url":"https://listedarticles.com/articles/can-a-model-learn-new-skills-as-add-ons.md","example":false,"citation":"Connito Research, Connito. \"Can a Model Learn New Skills as Add-Ons?.\" 28 Sept 2026. https://connito.ai/blog/can-a-model-learn-new-skills-as-add-ons (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://connito.ai/blog/can-a-model-learn-new-skills-as-add-ons"},"body_markdown":"# Can a Model Learn New Skills as Add-Ons?\n\nEach new skill is trained as a small, separate expert with its own router, and experts trained independently can be merged into one model in seconds without losing what each one learned.\n\nIf you want a large model to get better at something new, say medicine or code, the usual answer is to fine-tune it.\n\nThat works, but it has two costs. Fine-tuning edits a network that already supports everything else the model knows, so a new skill can quietly damage old ones. And every new skill means another pass over the model, run by whoever owns the whole model.\n\nWhat if a new skill could be built separately, by someone else, and simply plugged in?\n\n## What if a new skill were an add-on, not a rewrite?\n\nOur latest experiments test exactly that on a 15.7B-parameter Mixture-of-Experts (MoE) model. In an MoE model, each layer holds many small \"expert\" networks, and a router picks a handful of them for every token. Our base model picks 6.\n\nInstead of retraining any of the existing experts, we append one new expert to each layer and give it its own router. The original model stays frozen, exactly as it was. At the start of training the new expert contributes nothing, so the model begins identical to the original.\n\nThe new router doesn't compete with the old one. It has one question to answer: does this token need the new skill? If its score passes a threshold, the new expert joins in on top of the usual six; if not, it stays silent.\n\nIn symbols, the layer's output is the original output plus a gated contribution from the new expert:\n\n`h(x) = h_original(x) + g(x) · E_new(x)`\n\nwhere `g(x)` is close to 0 when the new router's score is below the threshold and close to 1 above it.\n\n## How does the new expert learn when to stay quiet?\n\nThe router is what makes the add-on safe. It is trained on two kinds of data at once:\n\n| Data | Example | What the router is taught |\n| --- | --- | --- |\n| General data | ordinary web text | Stay off. Any activation is penalised. |\n| Expertise data | math problems, for a math expert | Switch on, at a chosen target rate (20% of tokens in these runs). |\n\nThe intuition fits in one line:\n\n> Stay out of the way on what the model already knows. Show up on the new capability.\n\nThe model's existing knowledge is never edited. The new expert only adds to it, and only where it is needed.\n\n## Does each expert learn its skill?\n\nWe trained five experts separately: math, code, medicine, law and finance. Each ran for 500 steps on its own data, without any knowledge of the others. Here is each one next to the untouched base model, on a benchmark from its own domain. Higher is better.\n\n| Expert | Benchmark | Base model | Base + expert |\n| --- | --- | --- | --- |\n| Math | GSM8K | 37.9 | 61.7 |\n| Code | HumanEval | 27.4 | 35.4 |\n| Medical | MedQA | 42.3 | 46.7 |\n| Law | LegalBench | 48.6 | 53.7 |\n| Finance | FinQA | 1.2 | 26.6 |\n\nEvery expert lifts its own domain, and none of the original model was changed to get there.\n\n## Do independently trained experts still work together?\n\nThis is the real test. An add-on that works alone is useful; add-ons that can be stacked are a different kind of asset.\n\nSo we took experts from those separate runs and merged them into the same base model, two or three at a time. Merging copies each expert and its router into place, side by side. Nothing is averaged, and nothing is retrained.\n\n**Math + Code**\n\n| Model | GSM8K (math) | HumanEval (code) | MBPP (code) |\n| --- | --- | --- | --- |\n| Base model | 37.9 | 27.4 | 42.8 |\n| Math expert alone | 61.7 | 33.5 | 41.6 |\n| Code expert alone | 41.2 | 35.4 | 42.2 |\n| Math + Code, merged | 59.5 | 37.8 | 44.6 |\n\n**Math + Medical**\n\n| Model | GSM8K (math) | MedQA (medical) | MedMCQA (medical) |\n| --- | --- | --- | --- |\n| Base model | 37.9 | 42.3 | 40.4 |\n| Math + Medical, merged | 60.6 | 47.1 | 46.3 |\n\n**Math + Code + Law**\n\n| Model | GSM8K (math) | HumanEval (code) | LegalBench (law) |\n| --- | --- | --- | --- |\n| Base model | 37.9 | 27.4 | 48.6 |\n| Math + Code + Law, merged | 53.0 | 35.4 | 52.0 |\n\nIn every merged model, every score is above the base model. Each run is a single seed, so differences of a point or two are within noise. The pattern is the point: skills trained apart, combined afterwards, and none was lost.\n\n- Size of one new skill: ~1.4% of the model's parameters (~225M on a 15.7B base)\n- Cost of combining skills: seconds on a CPU, no retraining\n\n## Why does this matter as models grow?\n\nBecause it changes what a unit of AI improvement is.\n\nToday, improving a large model is one big job: one owner, one training run, and one model at the end. With add-on experts, a new capability becomes a small, separate component:\n\n- It can be trained in parallel by different contributors, on different data and hardware, without waiting for each other.\n- It is cheap relative to the model: about 1.4% of the parameters here, and a smaller share still on a larger base.\n- It can be combined on demand, so a customer gets exactly the set of skills their workload needs.\n\nThe same recipe does not depend on the base model's size. As the base grows toward 100B parameters and beyond, a new skill still means training one small expert and its router, not retraining the whole. That is how model improvement becomes fast enough, and cheap enough, to deliver to market one capability at a time.\n\nTrain a skill once, reuse it everywhere. That is the building block Connito's training network is designed around.\n","body_html":"<h1 id=\"can-a-model-learn-new-skills-as-add-ons\">Can a Model Learn New Skills as Add-Ons?</h1>\n<p>Each new skill is trained as a small, separate expert with its own router, and experts trained independently can be merged into one model in seconds without losing what each one learned.</p>\n<p>If you want a large model to get better at something new, say medicine or code, the usual answer is to fine-tune it.</p>\n<p>That works, but it has two costs. Fine-tuning edits a network that already supports everything else the model knows, so a new skill can quietly damage old ones. And every new skill means another pass over the model, run by whoever owns the whole model.</p>\n<p>What if a new skill could be built separately, by someone else, and simply plugged in?</p>\n<h2 id=\"what-if-a-new-skill-were-an-add-on-not-a-rewrite\">What if a new skill were an add-on, not a rewrite?</h2>\n<p>Our latest experiments test exactly that on a 15.7B-parameter Mixture-of-Experts (MoE) model. In an MoE model, each layer holds many small &quot;expert&quot; networks, and a router picks a handful of them for every token. Our base model picks 6.</p>\n<p>Instead of retraining any of the existing experts, we append one new expert to each layer and give it its own router. The original model stays frozen, exactly as it was. At the start of training the new expert contributes nothing, so the model begins identical to the original.</p>\n<p>The new router doesn&#39;t compete with the old one. It has one question to answer: does this token need the new skill? If its score passes a threshold, the new expert joins in on top of the usual six; if not, it stays silent.</p>\n<p>In symbols, the layer&#39;s output is the original output plus a gated contribution from the new expert:</p>\n<p><code>h(x) = h_original(x) + g(x) · E_new(x)</code></p>\n<p>where <code>g(x)</code> is close to 0 when the new router&#39;s score is below the threshold and close to 1 above it.</p>\n<h2 id=\"how-does-the-new-expert-learn-when-to-stay-quiet\">How does the new expert learn when to stay quiet?</h2>\n<p>The router is what makes the add-on safe. It is trained on two kinds of data at once:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Data</th><th>Example</th><th>What the router is taught</th></tr></thead><tbody><tr><td>General data</td><td>ordinary web text</td><td>Stay off. Any activation is penalised.</td></tr><tr><td>Expertise data</td><td>math problems, for a math expert</td><td>Switch on, at a chosen target rate (20% of tokens in these runs).</td></tr></tbody></table></div>\n<p>The intuition fits in one line:</p>\n<blockquote><p>Stay out of the way on what the model already knows. Show up on the new capability.</p></blockquote>\n<p>The model&#39;s existing knowledge is never edited. The new expert only adds to it, and only where it is needed.</p>\n<h2 id=\"does-each-expert-learn-its-skill\">Does each expert learn its skill?</h2>\n<p>We trained five experts separately: math, code, medicine, law and finance. Each ran for 500 steps on its own data, without any knowledge of the others. Here is each one next to the untouched base model, on a benchmark from its own domain. Higher is better.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Expert</th><th>Benchmark</th><th>Base model</th><th>Base + expert</th></tr></thead><tbody><tr><td>Math</td><td>GSM8K</td><td>37.9</td><td>61.7</td></tr><tr><td>Code</td><td>HumanEval</td><td>27.4</td><td>35.4</td></tr><tr><td>Medical</td><td>MedQA</td><td>42.3</td><td>46.7</td></tr><tr><td>Law</td><td>LegalBench</td><td>48.6</td><td>53.7</td></tr><tr><td>Finance</td><td>FinQA</td><td>1.2</td><td>26.6</td></tr></tbody></table></div>\n<p>Every expert lifts its own domain, and none of the original model was changed to get there.</p>\n<h2 id=\"do-independently-trained-experts-still-work-together\">Do independently trained experts still work together?</h2>\n<p>This is the real test. An add-on that works alone is useful; add-ons that can be stacked are a different kind of asset.</p>\n<p>So we took experts from those separate runs and merged them into the same base model, two or three at a time. Merging copies each expert and its router into place, side by side. Nothing is averaged, and nothing is retrained.</p>\n<p><strong>Math + Code</strong></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>GSM8K (math)</th><th>HumanEval (code)</th><th>MBPP (code)</th></tr></thead><tbody><tr><td>Base model</td><td>37.9</td><td>27.4</td><td>42.8</td></tr><tr><td>Math expert alone</td><td>61.7</td><td>33.5</td><td>41.6</td></tr><tr><td>Code expert alone</td><td>41.2</td><td>35.4</td><td>42.2</td></tr><tr><td>Math + Code, merged</td><td>59.5</td><td>37.8</td><td>44.6</td></tr></tbody></table></div>\n<p><strong>Math + Medical</strong></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>GSM8K (math)</th><th>MedQA (medical)</th><th>MedMCQA (medical)</th></tr></thead><tbody><tr><td>Base model</td><td>37.9</td><td>42.3</td><td>40.4</td></tr><tr><td>Math + Medical, merged</td><td>60.6</td><td>47.1</td><td>46.3</td></tr></tbody></table></div>\n<p><strong>Math + Code + Law</strong></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>GSM8K (math)</th><th>HumanEval (code)</th><th>LegalBench (law)</th></tr></thead><tbody><tr><td>Base model</td><td>37.9</td><td>27.4</td><td>48.6</td></tr><tr><td>Math + Code + Law, merged</td><td>53.0</td><td>35.4</td><td>52.0</td></tr></tbody></table></div>\n<p>In every merged model, every score is above the base model. Each run is a single seed, so differences of a point or two are within noise. The pattern is the point: skills trained apart, combined afterwards, and none was lost.</p>\n<ul><li>Size of one new skill: ~1.4% of the model&#39;s parameters (~225M on a 15.7B base)</li><li>Cost of combining skills: seconds on a CPU, no retraining</li></ul>\n<h2 id=\"why-does-this-matter-as-models-grow\">Why does this matter as models grow?</h2>\n<p>Because it changes what a unit of AI improvement is.</p>\n<p>Today, improving a large model is one big job: one owner, one training run, and one model at the end. With add-on experts, a new capability becomes a small, separate component:</p>\n<ul><li>It can be trained in parallel by different contributors, on different data and hardware, without waiting for each other.</li><li>It is cheap relative to the model: about 1.4% of the parameters here, and a smaller share still on a larger base.</li><li>It can be combined on demand, so a customer gets exactly the set of skills their workload needs.</li></ul>\n<p>The same recipe does not depend on the base model&#39;s size. As the base grows toward 100B parameters and beyond, a new skill still means training one small expert and its router, not retraining the whole. That is how model improvement becomes fast enough, and cheap enough, to deliver to market one capability at a time.</p>\n<p>Train a skill once, reuse it everywhere. That is the building block Connito&#39;s training network is designed around.</p>","headings":[{"level":1,"text":"Can a Model Learn New Skills as Add-Ons?","id":"can-a-model-learn-new-skills-as-add-ons"},{"level":2,"text":"What if a new skill were an add-on, not a rewrite?","id":"what-if-a-new-skill-were-an-add-on-not-a-rewrite"},{"level":2,"text":"How does the new expert learn when to stay quiet?","id":"how-does-the-new-expert-learn-when-to-stay-quiet"},{"level":2,"text":"Does each expert learn its skill?","id":"does-each-expert-learn-its-skill"},{"level":2,"text":"Do independently trained experts still work together?","id":"do-independently-trained-experts-still-work-together"},{"level":2,"text":"Why does this matter as models grow?","id":"why-does-this-matter-as-models-grow"}]}}