DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
Name
3689031.3717468.pdf
Size
1.43 MB
Format
Adobe PDF
Checksum (MD5)
64da8939a58ad4ed3791e778543d7a85
Author(s) • •
Yao, Xiaozhe
Hu, Qinghao
Klimovic, Ana
Date Issued
March 30, 2025
Publisher
ACM|Twentieth European Conference on Computer Systems
Citation
Xiaozhe Yao, Qinghao Hu, and Ana Klimovic. 2025. DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs. In Proceedings of the Twentieth European Conference on Computer Systems (EuroSys '25). Association for Computing Machinery, New York, NY, USA, 110–127.
Version
Final published version
Abstract
Fine-tuning large language models (LLMs) greatly improves model quality for downstream tasks. However, serving many fine-tuned LLMs concurrently is challenging due to the sporadic, bursty, and varying request patterns of different LLMs. To bridge this gap, we present DeltaZip, an LLM serving system that efficiently serves multiple full-parameter fine-tuned models concurrently by aggressively compressing model deltas by up to 10× while maintaining high model quality. The key insight behind this design is that fine-tuning results in small-magnitude changes to the pre-trained model. By co-designing the serving system with the compression algorithm, DeltaZip achieves 2× to 12× improvement in throughput compared to the state-of-the-art systems.
Description
EuroSys ’25, March 30–April 3, 2025, Rotterdam, Netherlands
MIT Department
Massachusetts Institute of Technology. Research Laboratory of Electronics
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3689031.3717468