Halo: Frontier-Lab Training for Everyone

Today we are open-sourcing Halo, the training framework we use to train every model at White Circle. Halo adds distributed training directly to HuggingFace models, and it is deliberately minimal. There is no porting step, no new checkpoint format, and significantly less code to write: adding a new model family takes ~100 lines instead of a full re-implementation. On the same hardware, it delivers 2.3–2.8x the throughput of stock TRL at lower peak memory. If you fine-tune open models, you have probably hit the wall that made us build Halo: your models are too large for stock HuggingFace tooling like TRL, and not large enough for Megatron to make sense. Training models with gradients, optimizer states, and activations stacked on top of the weights eventually pushes the setup past what a single GPU can hold. And if you want to train a Mixture-of-Experts (MoE) model, the frameworks that most people go for like TRL, Axolotl, or LLaMA-Factory do not have the parallelism strategies MoE training requires. The alternative so far has been to shift to a framework like Megatron, which comes at quite a significant cost. You leave the HuggingFace ecosystem behind, re-implement the model in a new codebase, have the checkpoints converted to a new format and write thousands of lines of porting and bridging code for every model family trained. Megatron was built for massive scale, 500+ GPU clusters, and only makes sense when operating at that scale. Most teams find themselves somewhere between the two current solutions, and this is what Halo is built for. It brings the parallelism strategies of the big frameworks to the models HuggingFace already defines, along with the components that make that parallelism usable in practice: fused kernels, bf16 optimizer, packing, and an async RL loop. None of it requires new algorithms, a new checkpoint format, or a new mental model. The result is the fastest training we measured among the frameworks we benchmarked, and it scales well past plain fine-tuning: agentic and async RL, on-policy distillation, embeddings learning, and more. 2.3-2.8x Halo throughput compared to stock TRL, with the model kept in plain Hugging Face format. Throughput against peak memory · B300B300, sequence 4,096, GC off. Solid points: batch 1, each framework at a representative configuration. Up and to the left is faster at less memory. Dashed curves add the batch-2 Expert-Parallel sweep (EP1→EP8) for Halo and Megatron-LM; Halo's stays up and to the left at every width, and Megatron OOMs at EP1. ## So what is Halo? Halo is a distributed-training layer for Hugging Face models. Currently, most distributed training frameworks work by owning the model. They define their own version of each architecture, tuned for their parallelism implementation, and expect you to write the conversion scripts that move weights in and out. Every new model family means a new port with model code, config mapping, a checkpoint bridge, conversion tooling. Halo inverts this. It adds four ways of splitting work across multiple GPUs directly onto Hugging Face models, without requiring a separate model definition or checkpoint format. A raw Hugging Face model goes in, and a raw Hugging Face model comes back out, loadable through the standard from_pretrained. The four splitting strategies target different bottlenecks: Expert Parallelism (EP) splits an MoE model's experts across GPUs. Context Parallelism (CP) splits the input sequence itself across GPUs rather than the model. CP chops long sequence into chunks distributed across GPUs and coordinates them to still compute attention correctly. Tensor Parallelism (TP) slices a single matrix multiplication across GPUs, with each GPU computing its portion and the results combined afterward. This applies to the dense, non-expert parts of the model. Expert-Tensor Parallelism (ETP) applies TP's slicing technique inside an individual expert, for cases where even one expert is too large for a single GPU. All of Halo's parallelism degrees are set through config rather than requiring manual intervention or a kernel support matrix to navigate by hand. They can also be combined together: EP+CP and EP+TP run in production, including EP across nodes, while EP+ETP works but is still experimental. The rest of Halo makes each step faster: custom Triton kernels alongside Liger (the open-source fused-kernel library), a fused bf16 AdamW (AdamWBF16) that halves optimizer memory by keeping the master weights and Adam moments in bf16 instead of the fp32 stock TRL holds, and an asynchronous RL loop. ## Our design choices for Halo Two things set Halo apart from how these frameworks usually get built: it works with the open-source ecosystem instead of replacing it, and it's built for the scale most teams operate at day to day. The ing choices cascaded from that: 1 HuggingFace is the training representation Every trainer subclasses a stock Hugging Face or TRL trainer, and checkpoints stay plain SafeTensors. New architectures and community methods come for free, and the methods we add ship back through one public repo, with none of the private checkpoint formats the Megatron route needs. 2 Distributed behavior is added as wrappers Frameworks like Megatron require you to re-define the model inside the framework, which means a custom class rebuilt from the framework's own building blocks before any training can start. Halo skips that step: EP wraps the existing MoE block, CP wraps attention, and dense and expert-FFN weights are sharded in place. No need to rewrite anything. 3 PyTorch-native parallelism Tensor and data parallelism ride PyTorch's own FSDP2, DTensor, and DeviceMesh; only Expert Parallelism needs hand-built process groups, since DeepEP's data-dependent all-to-all is more than DTensor can express. We also deliberately left out pipeline parallelism. Pipeline parallelism splits a model into sequential stages that live on different GPUs, and its known weakness is the "bubble": GPUs sitting idle while they wait for the stage before them to finish. On a modern rack-scale machine, where every GPU shares a fast NVLink connection, that trade-off no longer pays off. FSDP2 plus Expert Parallelism achieves the same memory savings with no idle time, and the two can be scaled independently of each other. The EP/CP/TP/ETP degrees are set in YAML and resolved at startup — no manual sharding, no kernel support matrix to navigate. 4 Every method shares one distributed core Every method is one trainer subclass over the same distributed layer, so it inherits the parallelism, kernels, and bf16 optimizer for free. Wrap any TRL method (SFT, DPO, GRPO) or add one of our own (SMPO, Offline GRPO, online distillation, reward modeling) — either way it's a small, self-contained change, and switching methods never means re-architecting the parallelism. 5 Tuned for the newest GPUs We train on B200/B300 (full-bf16 AdamWBF16, FlashAttention-4), but the stack steps down to FlashAttention-3 on Hopper, FA2 on anything else, and LoRA fine-tuning on a 24 GB consumer card. Whichever GPU you run on, the software underneath is the same and tracks the latest versions of the ecosystem Halo builds on: FSDP2, DeepEP V2, Transformers v5, CUDA 13. 6 Rollout generation stays outside the training image Online and multi-turn RL need an inference engine generating rollouts while the trainer runs. In Halo that engine is SGLang and it lives in its own isolated serving environment, so the training stack's Transformers/PyTorch versions stay untouched. 7 Halo is modular Each speedup is a separate unit that can be added or dropped independently. DeepEP, Liger kernels, and Ulysses-style attention wrappers all attach through the same mechanism, so a new kernel or attention variant is one more wrapper, and not a stack rework. 8 Built to be driven by agents The Docker images ship a CLI and Claude Code skills covering Halo's configs and common failure modes, so a coding agent can set up and launch runs without a human reading the docs first. It's how we run Halo ourselves. ## Running a job In practice, a run is one Docker image and one YAML file. Images are built per GPU architecture with FlashAttention and DeepEP V2 already compiled in, so there is no host Python environment to assemble. Pull the image for your GPU type from our registry and clone the repo to see examples and cookbooks.Bash # Clone to see cookbooks and examples git clone --recurse-submodules https://github.com/whitecircle/halo # Pull the image for B200 / B300 docker pull public.ecr.aws/whitecircle/halo:blackwell # Pull the image for H100 / H200 docker pull public.ecr.aws/whitecircle/halo:hopper Everything about the run then exists within a single YAML: the model, the data, the training method, and the parallelism degrees (EP, CP, TP, ETP), all resolved at startup. There is no kernel-support matrix to navigate and no manual sharding to reason about. The standard TrainingArguments / SFTConfig fields work alongside Halo's own blocks:YAML # examples/sft/gptoss/oss-20b-multinode-ep.yaml (abridged) model_name_or_path: openai/gpt-oss-20b dataset: - ... expert_parallel_size: 16 bf16: true # AdamWBF16 engages automatically learning_rate: 5.0e-6 per_device_train_batch_size: 2 Then launch it:Bash docker run --rm --gpus=all --ipc=host -it public.ecr.aws/whitecircle/halo:blackwell \ halo launch sft examples/sft/gptoss/oss-20b-multinode-ep.yaml \ --nproc=8 -- --expert_parallel_size=8 One thing to know about this example: the YAML above is written for two nodes. expert_parallel_size: 16 spreads the experts across 16 GPUs, with DeepEP's all-to-all running over AWS EFA or InfiniBand between them. The --expert_parallel_size=8 at the end of the launch command overrides that so the same config runs on a single 8-GPU node. Any YAML field can be overridden this way — whatever s -- on the command line takes precedence over the file. Switching training methods means changing the launch target and its config: halo launch smpo examples/preference/smpo-qwen3-4b-ultrafeedback.yaml runs SMPO (Smooth Margin Preference Optimization), our reference-free preference method. ## One trainer architecture, one parallelism stack Every training method runs through the same trainer stack, over the same distributed setup: Supervised fine-tuning Preference optimization: DPO, KTO, and our SMPO The GRPO family: offline, online with verifiable rewards (RLVR), and environmental Distillation: teacher, self, and online (Self-Distilled Policy Gradient, SDPG) Reward modeling, classification, and embedding1Pick training method Start with the trainer you already use — stock Hugging Face or TRL.Every trainer gets the sameHalo layer2Halo makes it distributed DistributedTrainerMixin: One abstract layer mixed into every trainer. Everything below it already works by definition — you don't wire any of it up.OptimizersAdamWBF16, distributed-readyLossesevery method, untouchedKernelsFlashAttention · DeepEP V2Data & checkpointsparallel-awareJust one parallelism config in your YAML3Choose parallelism Pick one or more dimensions. Halo creates the process groups.EPExpertsCPContextTPTensorsETPExpert Tensors Whatever data-parallel dimension is left falls through to FSDP2 fully_shard — native PyTorch. Each method's trainer is just its upstream HuggingFace or TRL trainer (SFTTrainer, GRPOTrainer, DPOTrainer) with one shared DistributedTrainerMixin mixed in ahead of it, so the mixin's overrides win and everything else falls through to the trainer you already know. The mixin is where the distributed work actually happens, and it's ours: we strip Accelerate's DDP wrap and synchronize gradients ourselves, keep gradient clipping correct when experts are sharded across GPUs, swap in the AdamWBF16 optimizer, and make data loading and checkpointing parallelism-aware. Underneath, the parallelism you set in the YAML (EP, CP, TP, ETP) partitions GPU groups first, and FSDP2 shards whatever data-parallel dimension is left. Adding a method is mostly subclassing its HuggingFace base and declaring which splits it supports. Not all support every split: Context Parallelism, for instance, is currently limited to the SFT and SMPO trainers, since the others work over whole sequences or paired completions. Everything else rides the same EP/TP/FSDP2 engine. Staying inside the Transformers ecosystem means new architectures and community methods come for free. And because the trainer stack is ours, we drop our own algorithms in right next to them. LoRA runs under all of Halo's data-parallel modes and under EP and CP. Under Expert Parallelism, Halo applies LoRA directly to the fused MoE experts. Stock PEFT can't reach these at all because fused experts are stored as single 3-D weight tensors rather than the nn.Linear layers PEFT knows how to adapt. This is why we built a grouped LoRA that adapts them natively. The payoff is that, with the experts frozen, they carry no optimizer state and skip the expert gradient all-to-all entirely, so LoRA trains about 1.25x faster than full fine-tuning — 9,600–9,900 vs 7,896 tok/s/GPU on gpt-oss-20b at EP2 (8x B300, 4k sequence, GC on) — at about a third of the memory. On dense models the effect reverses: attention-only LoRA stays close to full fine-tuning while all-linear costs more; those numbers are in the appendix, along with the full LoRA/QLoRA support matrix across parallelism modes. We also added quantization-aware training and a real fp8/fp4 path on DeepGEMM (actual low-precision kernels, not simulated quantization that is common in most stacks), though both are still experimental and aimed at producing quantized models rather than faster training. ## The numbers: faster at less memory across the board Setup for the runs below, unless a caption says otherwise: gpt-oss-20b in bf16 on an 8-GPU NVIDIA B300 node (274 GB usable per GPU), synthetic exact-length sequences, GC on. The TRL baseline is the strongest stock configuration we could build: it runs the same FlashAttention-4, Liger kernels (with fused-linear cross-entropy), and grouped-GEMM experts as Halo, so only the trainer, the sharding, and the optimizer differ. The harness behind the TRL comparisons ships in the repo, so those numbers are reproducible end to end. The four figures test one claim — faster at less memory — from four angles: whether it fits at all, how it compares to the tools already used, whether it holds without MoE, and what are the results at extreme length. Same convergence Halo wraps the same HuggingFace modules TRL trains, so a bf16 run reproduces the stock TRL loss curve step for step. The throughput and memory wins below are free, with no quality tradeoff. Convergence · Halo tracks stock TRLstock TRLHalo dense · EP1Halo EP2Halo EP8gpt-oss-20b, same seed and data: training loss over 200 steps for stock TRL and Halo at EP1 (no expert sharding), EP2, and EP8. ### Halo vs. stock TRL This is the comparison that matters for most teams, as stock TRL's SFTTrainer is where most fine-tuning happens today. Halo EP1 reaches roughly 3x the throughput of the stock TRL SFTTrainer at 4k, 16k, and 32k (top figure), while staying below it on peak memory the whole way. Peak measured throughput on this model: 24,456 tok/s/GPU (EP1, 4k, batch 4, GC off). Throughput by sequence-and-batch shape4k·b14k·b24k·b416k·b116k·b2Seq · batchstock TRL · ZeRO-3Halo EP1 · ZeRO-2Halo EP1 · ZeRO-3Halo EP2 · ZeRO-2Halo EP8 · ZeRO-24k·b13,8859,0095,56010,4798,320011,00022,000 tokens/s/GPUTok/s/GPU by sequence-and-batch shape: stock TRL ZeRO-3 against Halo EP1 (both ZeRO-2 and the matched ZeRO-3), EP2, and EP8, GC on. Matched sharding · Halo EP1 vs stock TRL4k·b116k·b116k·b432k·b164k·b1Seq · batchHalo EP1 · ZeRO-2Halo EP1 · ZeRO-3stock TRL · ZeRO-3 Throughput · multiplier vs stock TRL4k·b19,009 · 2.3×5,5603,885011,00022,000 tok/s/GPU Peak memory · lower is better4k·b160 GB29 GB48 GB085170 GB per GPUThe matched-sharding view: same EP1 (DP=8) layout and identical kernels on both sides, so the trainer and AdamWBF16 account for the whole gap. Qwen3-30B-A3B has 128 experts to gpt-oss-20b's 32, and that narrows the gap. It runs from a near-tie at EP8, batch 2 up to 1.8x at EP2, batch 1. Qwen3-30B-A3B throughputstock TRLHalo EP8Halo EP24k·b13,7876,3486,8984k·b28,2528,38211,34306,00012,000 tokens/s/GPUQwen3-30B-A3B, 8x B300, 4k, batch 1 and 2, GC on: stock TRL against Halo EP8 and EP2. ### Fitting in less memory The fastest configurations above spend memory to buy speed. Halo's default ZeRO-2 (FSDP2 shards the optimizer and gradients but replicates parameters) keeps more state resident than TRL's fully-sharded ZeRO-3. When memory matters more than speed, the opposite configuration wins. Expert Parallelism shards the experts, which are most of a MoE's weights, across GPUs, and AdamWBF16 holds optimizer state at half the bytes of fp32-master AdamW. At a 4k, batch-1 shape, Halo at EP8 trains gpt-oss-20b in 26 GB per GPU against stock TRL ZeRO-3's 47.6 GB — under the 40 GB of an older-generation card, though all measurements here are on B300. Matched ZeRO-3 lands at 28.7 GB and stays faster. Peak memory by sequence-and-batch shape4k·b14k·b24k·b416k·b116k·b2Seq · batchstock TRL · ZeRO-3Halo EP1 · ZeRO-2Halo EP1 · ZeRO-3Halo EP2 · ZeRO-2Halo EP8 · ZeRO-24k·b148 GB60 GB29 GB77 GB26 GB065130 GB per GPUPeak GB per GPU at the same shapes as the throughput figure; lower is better. EP8 at 4k, batch 1 is the 26 GB configuration. ### Still faster, even without MoE Everything so far involved Expert Parallelism. We therefore decided to run tests on Qwen3.5-4B, a dense VLM with no experts to distribute. We conducted the run in the same single-GPU FSDP2 route as every other framework, and Halo still leads at 16k. This means the other optimizations behind Halo, like kernel stack and the bf16 optimizer state, win on their own. ### Training at 256k tokens Finally, we compared gpt-oss-20b at up to 256k tokens on 8x B300. The gains are more modest here but still present; at these lengths, moving data is as important as computing on it. Halo runs 2.1x stock TRL's throughput at 64k, and then to 1.3x at 256k. We find that splitting the sequence across GPUs with Context Parallelism roughly halves the footprint per device, which is what makes 256k trainable at all. Long context · up to 256k tokens64k128k256kCtx sizestock TRL · ZeRO-3Halo EP1 · ZeRO-3Halo EP8+CP8Halo CP-only · ZeRO-3 Throughput64k6,87614,6545,2724,00707,50015,000 tok/s/GPU Peak memory · lower is better64k66 GB100 GB32 GB29 GB080160 GB per GPUgpt-oss-20b, 8x B300, batch 1, GC on, fused-linear cross-entropy on both sides. Left: throughput — Halo EP1 (sequence kept whole) leads TRL at every length. Right: peak memory — EP8+CP8 and CP-only (no EP) land at roughly half TRL's footprint. ## End-to-end example: GLM-4.7-Flash-Coder We used Halo to train GLM-4.7-Flash-Coder, a 30B MoE model, with multi-node SFT on a 177M-token agentic dataset. It lifted the base model's resolved rate on SWE-rebench-V2 from 33.15% to 41.65%, a 26% relative gain on a live agentic coding benchmark. The config, evaluation, and agentic setup are documented on the model card. SWE-rebench-V2 resolved ratebefore Halo SFT33.15%after Halo SFT41.65%0%25%50% resolvedSWE-rebench-V2 resolved rate before and after the Halo SFT run. ## Agentic and multi-turn RL, natively Online and multi-turn RL require a separate inference engine for rollout generation. The model talks to an environment — calling tools, running code, searching — before any reward exists, while a separate serving fleet generates those conversations. Halo folds this into the same trainer stack as SFT, still on top of TRL's GRPO trainer, with a runnable code-contests example on gpt-oss-20b in the repo. SGLang is the primary rollout engine. Halo uses SGLang as the primary backend: SGLang runs in an isolated serving environment, returns sampled token IDs, log probabilities, and optional MoE routing decisions, while Halo synchronizes updated model parameters over NCCL. With more than one server, the next batch generates while the current one trains, and a truncated importance-sampling ratio corrects the small off-policy lag. Halo supports selective full-parameter fine-tuning, where only a subset of transformer blocks is optimized. Instead of synchronizing the full policy after every optimizer step, Halo performs an element-wise comparison against the parameter state currently loaded by SGLang and transmits only modified BF16 elements as (flat_index, value) pairs. In our LFM R3 run, transformer blocks 7, 14, and 21 were trainable, totaling 1.1B parameters. The median synchronization contained 19.46M modified BF16 elements, or ~116 MB of indexed parameter updates. This is lossless and includes all modified parameter types, not only MoE experts. Compared with Halo's standard 16.94 GB full-policy dense synchronization on LFM, this reduces the median transfer volume by 145x. MoE routing replay is independent of parameter synchronization and adds only ~44 KB of metadata per rollout request. Environments are built in: code contests graded by their tests, SWE tasks, tool use over MCP, and search QA, with a registry for your own and SandboxFusion for sandboxed execution. We handle the multi-turn objective exactly, too. A model's own tokens and a tool's injected outputs land in one flat sequence with no marker between them, so training naively would put loss on the tool's text; instead we recover each assistant turn's exact token span by diffing tokenized chat-template prefixes, and train on nothing else. ## Model-family integration is cheap Adding a model family to Halo takes a few hundred lines at most, orders of magnitude less than Megatron. The integration is only a wrapper around the model's MoE or attention block. Once registered, the family reuses Halo's existing trainer, optimizer, checkpoint, and launcher code. Our assessment of LOC needed across Halo and Megatron-Bridges counts the model, config, bridge, and conversion files needed to connect the same family to a Megatron-Core stack. Lines of code per integrationHalo · EP wrapper127 LOCMegatron-Bridge · Qwen3.5627 LOCMegatron-Bridge · Bailing/Ling1,925 LOC01,0002,000 LOCHalo: one wrapper file around the Hugging Face block, counted from the repo. Megatron-Bridge: a full upstreamed integration for a recent decoder family — bridge, config, recipe, and conversion tests, ~900–1,800 lines; the weight-mapping file alone is smaller. ## Novel model integration examplePoolside · Laguna-S-2.1 Laguna-S-2.1 is a non-common architecture model released by Poolside. Its official implementation is not yet merged into the Transformers repository. Nevertheless, Halo provides support for Laguna in just 44 LoC. Laguna S 2.1 118B-A8B · throughput and peak memory · 4x B300 ThroughputHalo · EP411,298 · 2.0×Axolotl5,61706,00012,000 tok/s Peak memory · lower is betterHalo · EP4200.5 · −15%Axolotl237.10125250 GiBFull-parameter BF16 SFT, sequence 2,048, batch 1 per GPU, mean of 20 post-warmup steps. Halo runs EP4 with grouped GEMM and fused SwiGLU; Axolotl runs FSDP2 with grouped MM. Lower peak memory is better. Laguna support is 44 lines of Halo code. ## LiquidAI LFM supportLiquid AI · LFM2.5-8B-A1B Liquid's LFM2.5-8B-A1B is a hybrid MoE with 8.3B total and 1.5B active parameters, 18 double-gated LIV convolutional blocks, six GQA blocks, and a 128k context window. Halo trains it in native Hugging Face format. Roughly 120 lines wrap the original MoE block with EP, ETP, and combined EP+ETP, after which every Halo trainer and capability works unchanged. That includes dedicated embedding-model fine-tuning. There is no model reimplementation or checkpoint conversion. Among the frameworks we benchmarked, Halo is the only one with model-aware LFM2.5 support for both EP and ETP, including combined EP+ETP configurations. On one B300, Halo was 10–20% faster than Axolotl across the two working single-GPU comparison runs. SequenceHalo tok/s/GPUAxolotl tok/s/GPU 8,19231,26228,330 16,38435,51729,520 ## Supported families Parallelism support is per family. Families not listed fall back to the standard FSDP2 route, and any causal LM or VLM that Transformers loads trains there. Cond. () marks support with a condition attached: Zaya's EP and ETP require gradient checkpointing off. Context Parallelism currently applies to SFT and SMPO. The full matrix, including the combined modes, is in the appendix. FamilyEPCPTPETP GLM-4 MoE LiteYesYesYesYes GPT-OSSYesYesYesYes Mistral4 MoEYesYesYesYes Qwen3 MoEYesYesYesYes LFM2.5-8B-A1BYesNo—YesYes Qwen3.5/Qwen3.6 MoEYesNo—YesYes InclusionAI LingYesNo—No—Yes Gemma 4 MoEYesNo—No—Yes Kimi LinearYesNo—No—Yes Poolside LagunaYesNo—No—Yes ZyZyphra ZayaNo—No— Cohere Command A+YesNo—No—No— Thinking Machines InklingYesNo—No—No— The table tracks parallelism only; the full method-by-family matrix is in the repo's agent-docs/models/index.md. ## Some of the hard problems along the way Making an unmodified Hugging Face MoE train fast and reliably at frontier scale surfaced a run of non-obvious problems. The step is almost all communication. In an MoE layer, each token is routed to a few chosen experts, computed there, and gathered back. Even with DeepEP moving the tokens as fast as the hardware allows, that round trip takes up 88–93% of a training step. The expert matmul itself is only 6.6%. This shaped what we optimized. Rewriting the routing's backward pass to be atomic-free (PyTorch's default serializes on atomic adds) improved end-to-end throughput by up to 65% on a 256-expert model. Quantizing the experts to fp8 or fp4, on the other hand, changed nothing: the matmul already runs as fast as memory can feed it, so making it cheaper to compute doesn't help. A crash three files from its cause. For weeks, EP8 training died once sequences reached 64k tokens, with a CUDA fault deep in DeepEP's network transport. The actual cause was elsewhere. A batch holds many sequences, and MoE routing is uneven — in the failing runs, the router had concentrated 745k tokens onto the experts of a single GPU. At that size, a fused SwiGLU kernel in a different part of the codebase overflowed