Can a Model Learn New Skills as Add-Ons?

Each new skill is trained as a small, separate expert with its own router, and experts trained independently can be merged into one model in seconds without losing what each one learned.

If you want a large model to get better at something new, say medicine or code, the usual answer is to fine-tune it.

That works, but it has two costs. Fine-tuning edits a network that already supports everything else the model knows, so a new skill can quietly damage old ones. And every new skill means another pass over the model, run by whoever owns the whole model.

What if a new skill could be built separately, by someone else, and simply plugged in?