About
I work on making LLM training cheaper, faster, and more reliable.
I am a Ph.D. student in Computer Science at the University of California, Santa Barbara, advised by Prof. Zheng Zhang. Before my Ph.D., I received an M.A. in Statistics from UCSB and a B.S. in Statistics from Nankai University.
My research covers the training stack end to end, particularly for pre-training: low-rank model architectures that reduce compute, Muon-style optimizers that train faster, and scalable, fault-tolerant training systems designed for clusters of 100k+ GPUs.
Most recently, I was a Software Engineer Intern at Google Cloud. Before that, I was a visiting student at Argonne National Laboratory and a research intern at Cadence Design Systems.
I am on the job market for full-time industry research and engineering roles starting in 2027. Feel free to reach out at ziyueliu@ucsb.edu.
News
Scroll for earlier updatesReCoVer was accepted to NeurIPS 2026.
Muon² was accepted to EMNLP 2026 as an oral presentation.
MuonQ was accepted to COLM 2026.
Started my internship at Google Cloud, on MoE adapter tuning and serving.
BOOST was accepted to MLSys 2026.
DeepOHeat-v1 was published in IEEE TCPMT.
LaX was accepted to NeurIPS 2025.
Visiting student at Argonne National Laboratory, on fault-tolerant LLM pre-training.
SepONet was published in TMLR.
CoMERA was accepted to NeurIPS 2024.
Third research internship at Cadence, on data center thermal and CFD modeling.
Started my Ph.D. in Computer Science at UCSB.
Back at Cadence as a research intern, on automotive aerodynamic simulation.
Received my M.A. in Statistics from UCSB.
DeepOHeat was accepted to DAC 2023.
TT-PINN was accepted to the ICML 2022 HAET workshop.
First research internship at Cadence, on 3D-IC thermal simulation.
Joined Prof. Zheng Zhang's group at UCSB.
Came to UCSB to start my M.A. in Statistics.
Research
My current research is on efficient LLM training, across three layers of the stack. Earlier in my graduate studies, I worked on scientific machine learning, building efficient neural PDE solvers with applications in 3D-IC thermal design.
LLM training stack
Architecture
Low-rank and tensor-compressed models that cut the compute and memory cost of training while preserving quality.
Optimizer
Muon-based optimizers that further speed up convergence, reduce orthogonalization cost, and improve quality.
System
Scalable training for low-rank models, and fault-tolerant training that avoids system stalling on frequent restarts at O(100k+) scale.
load checkpoint
restore
restore
no rework
Scientific machine learning
PDE solvers
Physics-informed neural networks and neural operators for fast PDE solving, applied to 3D-IC thermal simulation and design.
Selected Publications
Full list on Google ScholarFirst or co-first author* Equal contribution
-
NeurIPS 2026
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
Keeps pre-training on track through GPU failures without restarting the job.
-
EMNLP 2026Oral
Muon²: Boosting Muon via Adaptive Second-Moment Preconditioning
Adaptive second-moment preconditioning that makes Muon converge faster at lower cost.
-
EMNLP 2025Oral
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
Low-rank activations that cut model size and compute while matching full-rank quality.
-
MLSys 2026
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
Bottleneck-aware parallelism that makes large-scale low-rank training faster than full-rank.
-
NeurIPS 2025
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
A plug-and-play module that helps low-rank models match full-rank quality.
-
NeurIPS 2024
CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
Rank-adaptive tensor-compressed training that reduces both memory and runtime on GPUs.
-
DAC 2023
DeepOHeat: Operator Learning-based Ultra-fast Thermal Simulation in 3D-IC Design
Operator learning that makes 3D-IC thermal simulation orders of magnitude faster.
-
COLM 2026
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
Quantizes Muon's optimizer states to low bits while closely matching full-precision training.
-
Under review
Muon+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
One extra normalization step after orthogonalization that consistently improves Muon.
-
TCPMT 2026
DeepOHeat-v1: Efficient Operator Learning for Fast and Trustworthy Thermal Simulation and Optimization in 3D-IC Design
A more accurate, efficient, and trustworthy DeepOHeat for 3D-IC thermal design.
-
Under review
DeepOHeat-v2: Self-Improving Operator Learning for Fast and Trustworthy Thermal Optimization in 3D-IC Design
A self-improving DeepOHeat for thermal optimization of high-contrast multi-die 3D-ICs.
-
Under review
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
Extends Muon's orthogonalization from single layers to a tensor across layers.
-
EMNLP 2025
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
Memory-efficient fine-tuning of quantized LLMs using only low-bit forward passes.
-
TMLR 2024
Separable Operator Networks
Faster, more memory-efficient physics-informed operator learning with separable networks.
-
ICML 2022Workshop
TT-PINN: A Tensor-Compressed Neural PDE Solver for Edge Computing
Tensor-train compressed PINNs that can be trained on edge devices.
Experience
Industry and research
-
Jun 2026 – Sep 2026
Software Engineer Intern
Google Cloud · Sunnyvale, CA
Efficient MoE adapter tuning and multi-tenant serving.
-
Apr 2025 – Sep 2025
Visiting Student
Argonne National Laboratory · Lemont, IL
Fault-tolerant LLM pre-training.
-
Jun 2022 – Sep 2024 · 3 terms
Research Intern
Cadence Design Systems · Austin, TX
- 2024
Data-driven modeling of real-world data center thermal and CFD simulations.
- 2023
Data-driven modeling of large-scale automotive aerodynamic simulations.
- 2022
Physics-informed operator learning for 3D-IC thermal simulations.
- 2024
Education
-
2023 – 2027 (expected)
Ph.D. in Computer Science
University of California, Santa Barbara
Advisor: Prof. Zheng Zhang
-
2021 – 2023
M.A. in Statistics
University of California, Santa Barbara
-
2016 – 2020
B.S. in Statistics
Nankai University, School of Mathematical Sciences
Teaching
Teaching assistant at the University of California, Santa Barbara.
- CMPSC 130AData Structures and AlgorithmsWinter 2025, Winter 2026
- CMPSC 165BMachine LearningWinter 2024
- PSTAT 231Statistical Machine LearningFall 2022, Winter 2023
- PSTAT 120BProbability and StatisticsFall 2021, Spring 2022
- PSTAT 5LSStatistics for Life SciencesWinter 2022