Xuejian Gong

dblp:256/8229 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0006-1678-0365ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
abstract
Large Language Models (LLMs) have significantly advanced natural language processing with exceptional task generalization capabilities. Low-Rank Adaption (LoRA) offers a cost-effective fine-tuning solution, freezing the original model parameters and training only lightweight, low-rank adapter matrices. However, the memory footprint of LoRA is largely dominated by the original model parameters. To mitigate this, we propose LoRAM, a memory-efficient LoRA training scheme founded on the intuition that many neurons in over-parameterized LLMs have low training utility but are essential for inference. LoRAM presents a unique twist: it trains on a pruned (small) model to obtain pruned low-rank matrices, which are then recovered and utilized with the original (large) model for inference. Additionally, minimal-cost continual pre-training, performed by the model publishers in advance, aligns the knowledge discrepancy between pruned and original models. Our extensive experiments demonstrate the efficacy of LoRAM across various pruning strategies and downstream tasks. For a model with 70 billion parameters, LoRAM enables training on a GPU with only 20G HBM, replacing an A100-80G GPU for LoRA training and 15 GPUs for full fine-tuning. Specifically, QLoRAM implemented by structured pruning combined with 4-bit quantization, for LLaMA-3.1-70B (LLaMA-2-70B), reduces the parameter storage cost that dominates the memory usage in low-rank matrix training by 15.81× (16.95×), while achieving dominant performance gains over both the original LLaMA-3.1-70B (LLaMA-2-70B) and LoRA-trained LLaMA-3.1-8B (LLaMA-2-13B). Code is available at https://github.com/junzhang-zj/LoRAM.
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Guiming Xie, Xuejian Gong, Kunlong Zhou
ICLR8
2025 HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Guiming Xie, Xuejian Gong
VLDB J.9
2023 Collaborative contracting for Manufacturing-as-a-Service (MaaS) by information content measurement and decision tree learning
Xuejian Gong, Roger Jianxin Jiao, Nagi Gebraeel
Adv. Eng. Informatics1
2022 Quantum entanglement inspired hard constraint handling for operations engineering optimization with an application to airport shift planning
Pan Zou, Xuejian Gong, Roger Jianxin Jiao, Feng Zhou 0003
Expert Syst. Appl.3
2021 Smart dispatching and optimal elevator group control through real-time occupancy-aware deep learning of usage patterns
Xuejian Gong, Mulang Song, Cindy Y. Fei, Stefan Quaadgras, Jianyuan Peng, Pan Zou, Jerred Chen, Roger Jianxin Jiao
Adv. Eng. Informatics2