Ziyue (Alvin) Liu
MLSys 2026Artifact evaluated

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models

Zhengyang Wang*1, Ziyue Liu*1, Ruijie Zhang1, Avinash Maurya2, Bogdan Nicolae2, Paul Hovland2, Franck Cappello2, Zheng Zhang1

1University of California, Santa Barbara2Argonne National Laboratory*Equal contribution

In short

Low-rank models such as CoLA should train faster than full-rank models, but under standard tensor parallelism they end up slower, because it adds all-reduces on full-width activations. BOOST places every all-reduce at the low-rank bottleneck instead, and adds Online RMSNorm, linear layer grouping and low-rank activation checkpointing on top of it. Low-rank models then train 1.46–1.91× faster than full-rank models, and 1.87–2.27× faster than with standard tensor parallelism.

Average time per training iteration (s), lower is better
LLaMA-2-style models; CoLA as the low-rank model
Full-rank TPVanilla low-rank TPBOOST
0123
0.59
0.78
0.72
1.30
1.27
1.52
1B1 GPU3B2 GPUs7B4 GPUs13B8 GPUs30B16 GPUs40B32 GPUs
Show the numbers
ModelFull-rank TPVanilla TPBOOSTBOOST speedupvs full-rankvs vanilla
1B1 GPU0.850.560.591.44×0.95×
3B2 GPUs1.141.410.781.46×1.81×
7B4 GPUs1.061.640.721.47×2.28×
13B8 GPUs2.072.421.301.59×1.86×
30B16 GPUs2.432.581.271.91×2.03×
40B32 GPUs2.823.391.521.86×2.23×
1B runs on one GPU without tensor parallelism. 3B and 7B use tensor parallelism of 2 and 4; 13B, 30B and 40B add 2, 4 and 8 pipeline stages across nodes. Micro-batch 4, sequence length 4,096.
1.46–1.91×
faster than full-rank training with standard tensor parallelism
1.87–2.27×
faster than the same low-rank models with standard tensor parallelism
5.7×
less tensor-parallel traffic than standard low-rank TP, and 1.14× less than full-rank TP

Abstract

The scale of transformer model pre-training is constrained by the increasing computation and communication cost. Low-rank bottleneck architectures offer a promising solution to significantly reduce the training time and memory footprint with minimum impact on accuracy. Despite algorithmic efficiency, bottleneck architectures scale poorly under standard tensor parallelism. Simply applying 3D parallelism designed for full-rank methods leads to excessive communication and poor GPU utilization. To address this limitation, we propose BOOST, an efficient training framework tailored for large-scale low-rank bottleneck architectures. BOOST introduces a novel Bottleneck-aware Tensor Parallelism, and combines optimizations such as online-RMSNorm, linear layer grouping, and low-rank activation checkpointing to achieve end-to-end training speedup. Evaluations on different low-rank bottleneck architectures demonstrate that BOOST achieves 1.46–1.91× speedup over full-rank model baselines and 1.87–2.27× speedup over low-rank model with naively integrated 3D parallelism, with improved GPU utilization and reduced communication overhead.

Move the chunk boundary to the bottleneck

A bottleneck model replaces each weight with a down-projection to a small rank r and an up-projection back; the paper uses r = d/4. CoLA, SVD-style factorization and LaX all have this shape. Standard Megatron-style tensor parallelism treats each factor pair as one chunk: the down-projection is split by columns, the up-projection by rows, and an all-reduce closes the chunk. For a bottleneck model this hurts in two ways. An all-reduce now follows every factorized layer and carries a full-width activation, 5 to 6.5× the traffic of full-rank TP per block. The split also cuts the small rank into smaller pieces, so the matrix multiplies do little work per byte they move: in LLaMA-7B MLP blocks, 0.2× the arithmetic intensity of full-rank TP. In the paper’s runs, this vanilla low-rank TP is slower than full-rank TP from 3B parameters up.

RMSNormAqkvBqkvAttnAoBoRMSNormAupBupActAdownBdownxl−1xlOnlineRMSNormAqkvBqkvAttnAoBoOnlineRMSNormAupBupActAdownBdownxl−1xlVanilla low-rank TPall-reduce width per block: 5d + 2dff 3dd2dff dBOOST: bottleneck-aware TPall-reduce width per block: 7r3rr2rrNormAqkvBqkvAttnAoBoxl−1NormAupBupActAdownBdownxlNormAqkvBqkvAttnAoBoxl−1NormAupBupActAdownBdownxlVanilla low-rank TPper block: 5d + 2dff 3dd2dff dBOOST: bottleneck-aware TPper block: 7r3rr2rr
vanilla TP chunkBOOST chunkOnline RMSNormall-reduce, width per token
One LLaMA decoder block, left to right. A is a down-projection to rank r and B an up-projection back to d or dff. Each shaded box is one tensor-parallel chunk and ends with one all-reduce, labeled with the number of values it sends per token. BOOST’s first and last chunks continue into the neighboring blocks.

BOOST shifts every chunk by one layer. A chunk now starts at an up-projection, split by columns, and ends at the next down-projection, split by rows, so its only all-reduce carries an r-wide activation. The weights are split along the large dimension d or dff instead of r, which gives 2.5× the arithmetic intensity of vanilla TP in LLaMA-7B MLP blocks.

VolumeLLaMA-7Bvs full-rank TP
Full-rank TP2bsd1×
Vanilla low-rank TP5bsd + 2bsdff5.19×
BOOST7bsr0.88×
Tensor-parallel all-reduce volume per decoder block and pass. b is the micro-batch size and s the sequence length; LLaMA-7B has d = 4096, dff = 11008 and r = 1024.

Data and pipeline parallelism need no change. Gradient all-reduces shrink with the parameter count, by about 2.5× at r = d/4, and pipeline stages pass the same d-wide tensors as before.

Three supporting changes

  1. Online RMSNorm. After the shift, RMSNorm falls inside a chunk, where each GPU holds only a slice of the hidden vector. A separate all-reduce for its statistic would be tiny but slow. Each GPU instead normalizes with its local RMS, scales the chunk’s output back by it, and sends its sum of squares along with the chunk’s all-reduce; dividing by the global RMS then gives the exact result. Section 4.2
  2. Linear layer grouping. The Q, K and V down-projections read the same input, so they run as one matrix multiply followed by one all-reduce of width 3r, and their up-projections run as one batched multiply. The MLP’s gate and up projections are grouped the same way. Section 4.3
  3. Checkpointing without communication. Only the r-wide activations at chunk boundaries are stored. Recomputing a chunk in the backward pass never crosses an all-reduce, while under vanilla TP the recomputation needs an extra full-width one. Section 4.4

Results

LLaMA-2-style models run on NERSC Perlmutter, whose nodes each have four A100 80GB GPUs. Tensor parallelism runs inside a node and pipeline parallelism across nodes. CoLA is the default bottleneck model, sequences are 4,096 tokens long, and each time is the average of 8 iterations after 2 warm-up iterations.

Faster than full-rank training from 3B up

The 1B model fits on one GPU without tensor parallelism, and its low-rank version is already 1.4× faster than full rank. From 3B on, where tensor parallelism is needed, vanilla TP falls behind full-rank TP and BOOST is the fastest of the three: up to 1.91× faster than full-rank TP (30B) and 2.28× faster than vanilla TP (7B). The gain holds at 13B to 40B, where the pipeline spans 2 to 8 nodes. The chart at the top of this page shows these runs.

Larger batches and other bottleneck models

On LLaMA-7B with four GPUs, BOOST is 1.3×, 1.42× and 1.48× faster than full-rank TP at micro-batch sizes 1, 2 and 4, and it is the only one of the three that fits a micro-batch of 8. SVD, CoLA and LaX models all train about 2.2× faster with BOOST than with vanilla TP, and about 1.5× faster than the full-rank model. SVD is the fastest because nothing sits between its two factors; LaX is the slowest because of its extra residual path.

Iteration time (s) by micro-batch size
LLaMA-7B on four GPUs
Full-rank TPVanilla low-rank TPBOOST
00.511.52
0.27
0.42
0.72
OOM
1.32
1248
micro-batch size
Show the numbers
Micro-batchFull-rank TPVanilla TPBOOST
10.360.460.27
20.600.800.42
41.061.640.72
8OOMOOM1.32
Iteration time (s) by bottleneck model
LLaMA-7B on four GPUs; dashed: full-rank TP, 1.06 s
Vanilla low-rank TPBOOST
00.511.52
1.57
0.70
1.64
0.72
1.72
0.75
SVDCoLALaX
bottleneck model
Show the numbers
ModelVanilla TPBOOSTBOOST speedupvs vanillavs full-rank
SVD1.570.702.24×1.51×
CoLA1.640.722.28×1.47×
LaX1.720.752.29×1.41×

Less communication than full-rank TP

At micro-batch 4, BOOST’s all-reduces take up to 8% less time per decoder block than full-rank TP’s, and about 5.3× less than vanilla TP’s. Its linear layers also run at higher hardware utilization than under vanilla TP at every model size and batch size measured. On CoLA LLaMA-7B at micro-batch 4, BOOST needs 27.1 GB per GPU against 35.7 GB for vanilla TP. Weights, gradients and optimizer states are the same; the difference is activations and communication buffers.

All-reduce time per decoder block and pass (ms)
tensor parallelism of 4, micro-batch 4
Full-rank TPVanilla low-rank TPBOOST
04812
1.57
7.61
1.47
2.01
9.87
1.86
2.47
12.09
2.27
3B7B13B
Show the numbers
ModelFull-rank TPVanilla TPBOOST
3B1.577.611.47
7B2.019.871.86
13B2.4712.092.27

What each change adds

Grouping speeds up a CoLA LLaMA-7B decoder block by 1.16× at micro-batch 1 and 1.04× at 4; it helps most when the matrix multiplies are small. Checkpointing under BOOST saves 1.70× more memory per millisecond of recomputation than under vanilla TP at micro-batch 4, and 1.56× at 8. Online RMSNorm matches the one-GPU result to 7×10−7 in FP32, and a tiny LLaMA trained with it follows the same loss curve as the one-GPU baseline.

Ablations
Micro-batch 1time per decoder block (µs)
DefaultGroupedSpeedup
Gate and up, compute3552921.22×
Gate and up, all-reduce2662181.22×
QKV, compute3912551.53×
QKV, all-reduce4062881.41×
Whole block2,7732,3951.16×
Micro-batch 4time per decoder block (µs)
DefaultGroupedSpeedup
Gate and up, compute1,1151,0821.03×
Gate and up, all-reduce6205801.07×
QKV, compute9398771.07×
QKV, all-reduce9818061.22×
Whole block7,5777,2661.04×

CoLA LLaMA-7B. Gate and up are the MLP’s first two projections; the whole block also includes kernels not listed here.

Scope

Citation

BibTeX
@inproceedings{MLSYS2026_127e7093,
 author = {Wang, Zhengyang and Liu, Ziyue and Zhang, Ruijie and Maurya, Avinash and Nicolae, Bogdan and Hovland, Paul and Cappello, Franck and Zhang, Zheng},
 booktitle = {Proceedings of Machine Learning and Systems},
 editor = {A. Chowdhery and Z. Jia},
 pages = {1350--1368},
 publisher = {MLSys},
 title = {BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models},
 url = {https://proceedings.mlsys.org/paper_files/paper/2026/file/127e7093c38a45290524237be8eb39c5-Paper-Conference.pdf},
 volume = {8},
 year = {2026}
}