Jue Wang 0019

dblp:69/393-19 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-6712-1929ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Efficient and distributed learning · 62% Information extraction and text analysis · 16% Language models and text generation · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
3.042025
FloE: On-the-Fly MoE Inference on Memory-constrained GPU · ICML 2025
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025
Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023
Machine learning › Efficient and distributed learning
federated learning
1.422025
${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025
Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022
Natural language and speech › Information extraction and text analysis
text classification
1.022024
Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023
Machine learning › Efficient and distributed learning › federated learning › heterogeneous federated learning
client heterogeneity
0.912025
${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025
Machine learning › Efficient and distributed learning › federated learning › label-efficient federated learning
federated active learning
0.912025
${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025
Natural language and speech › Language models and text generation
large language model fine-tuning
0.912025
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
0.912025
FloE: On-the-Fly MoE Inference on Memory-constrained GPU · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation › large language model inference
pre-trained language model inference
0.912025
HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models · VLDB J. 2025
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.822023
Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023
Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022
Machine learning › Efficient and distributed learning
inference acceleration
0.812024
Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification
0.812024
Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Machine learning › Efficient and distributed learning › inference acceleration › speculative decoding
self-speculative decoding
0.812024
Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.812024
Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
0.812024
Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Machine learning › Learning paradigms
continual learning
0.712023
Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
lightweight adapters
0.712023
Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023
Machine learning › Efficient and distributed learning › federated learning
federated continual learning
0.612022
Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.612022
Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022
Machine learning › Efficient and distributed learning › adaptive computation
layer skipping
0.612022
SkipBERT: Efficient Inference with Shallow Layer Skipping · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
slot filling
0.512021
Effective Slot Filling via Weakly-Supervised Dual-Model Learning · AAAI 2021
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction
0.412020
Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence Encoders · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Pyramid: A Layered Model for Nested Named Entity Recognition · ACL 2020
Natural language and speech › Information extraction and text analysis › named entity recognition
nested named entity recognition
0.412020
Pyramid: A Layered Model for Nested Named Entity Recognition · ACL 2020
Machine learning › Efficient and distributed learning
active learning
0.312025
${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025
Machine learning › Efficient and distributed learning
data selection
0.312025
${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.312025
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
large language model inference
0.212022
SkipBERT: Efficient Inference with Shallow Layer Skipping · ACL (1) 2022
Machine learning › Learning paradigms
weakly supervised learning
0.112021
Effective Slot Filling via Weakly-Supervised Dual-Model Learning · AAAI 2021

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 3.0sparse prediction · 1.7offloading · 1.7hierarchical indexing · 1.7expert parameter compression · 1.7structured pruning · 0.9quantization · 0.9low-rank adaptation · 0.9epistemic uncertainty · 0.9active learning · 0.9
YearPublicationVenuePosition
2025 Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
abstract
Large Language Models (LLMs) have significantly advanced natural language processing with exceptional task generalization capabilities. Low-Rank Adaption (LoRA) offers a cost-effective fine-tuning solution, freezing the original model parameters and training only lightweight, low-rank adapter matrices. However, the memory footprint of LoRA is largely dominated by the original model parameters. To mitigate this, we propose LoRAM, a memory-efficient LoRA training scheme founded on the intuition that many neurons in over-parameterized LLMs have low training utility but are essential for inference. LoRAM presents a unique twist: it trains on a pruned (small) model to obtain pruned low-rank matrices, which are then recovered and utilized with the original (large) model for inference. Additionally, minimal-cost continual pre-training, performed by the model publishers in advance, aligns the knowledge discrepancy between pruned and original models. Our extensive experiments demonstrate the efficacy of LoRAM across various pruning strategies and downstream tasks. For a model with 70 billion parameters, LoRAM enables training on a GPU with only 20G HBM, replacing an A100-80G GPU for LoRA training and 15 GPUs for full fine-tuning. Specifically, QLoRAM implemented by structured pruning combined with 4-bit quantization, for LLaMA-3.1-70B (LLaMA-2-70B), reduces the parameter storage cost that dominates the memory usage in low-rank matrix training by 15.81× (16.95×), while achieving dominant performance gains over both the original LLaMA-3.1-70B (LLaMA-2-70B) and LoRA-trained LLaMA-3.1-8B (LLaMA-2-13B). Code is available at https://github.com/junzhang-zj/LoRAM.
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Guiming Xie, Xuejian Gong, Kunlong Zhou
ICLR2
2025 FloE: On-the-Fly MoE Inference on Memory-constrained GPU
abstract
With the widespread adoption of Mixture-of-Experts (MoE) models, there is a growing demand for efficient inference on memory-constrained devices. While offloading expert parameters to CPU memory and loading activated experts on demand has emerged as a potential solution, the large size of activated experts overburdens the limited PCIe bandwidth, hindering the effectiveness in latency-sensitive scenarios. To mitigate this, we propose FloE, an on-the-fly MoE inference system on memory-constrained GPUs. FloE is built on the insight that there exists substantial untapped redundancy within sparsely activated experts. It employs various compression techniques on the expert's internal parameter matrices to reduce the data movement load, combined with low-cost sparse prediction, achieving perceptible inference acceleration in wall-clock time on resource-constrained devices. Empirically, FloE achieves a 9.3$\times$ compression of parameters per expert in Mixtral-8$\times$7B; enables deployment on a GPU with only 11GB VRAM, reducing the memory footprint by up to 8.5$\times$; and delivers a 48.7$\times$ inference speedup compared to DeepSpeed-MII on a single GeForce RTX 3090—all with only a 4.4\% $\sim$ 7.6\% average performance degradation.
Zheng Li 0006, Jun Zhang 0069, Jue Wang 0019, Yiping Wang 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou
ICML4
2025 ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning
abstract
Active learning (AL) reduces human annotation costs for machine learning systems by strategically selecting the most informative unlabeled data for annotation, but performing it individually may still be insufficient due to restricted data diversity and annotation budget. Federated Active Learning (FAL) addresses this by facilitating collaborative data selection and model training, while preserving the confidentiality of raw data samples. Yet, existing FAL methods fail to account for the heterogeneity of data distribution across clients and the associated fluctuations in global and local model parameters, adversely affecting model accuracy. To overcome these challenges, we propose${\sf CHASe}$(Client Heterogeneity-Aware Data Selection), specifically designed for FAL.${\sf CHASe}$focuses on identifying those unlabeled samples with high epistemic variations (EVs), which notably oscillate around the decision boundaries during training. To achieve both effectiveness and efficiency,${\sf CHASe}$encompasses techniques for 1) tracking EVs by analyzing inference inconsistencies across training epochs, 2) calibrating decision boundaries of inaccurate models with a new alignment loss, and 3) enhancing data selection efficiency via a data freeze and awaken mechanism with subset sampling. Experiments show that${\sf CHASe}$surpasses various established baselines in terms of effectiveness and efficiency, validated across diverse datasets, model complexities, and heterogeneous federation settings.
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou
IEEE Trans. Knowl. Data Eng.2
2025 HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Guiming Xie, Xuejian Gong
VLDB J.2
2024 Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
abstract
We present a novel inference scheme, selfspeculative decoding, for accelerating Large Language Models (LLMs) without the need for an auxiliary model.This approach is characterized by a two-stage process: drafting and verification.The drafting stage generates draft tokens at a slightly lower quality but more quickly, which is achieved by selectively skipping certain intermediate layers during drafting.Subsequently, the verification stage employs the original LLM to validate those draft output tokens in one forward pass.This process ensures the final output remains identical to that produced by the unaltered LLM.Moreover, the proposed method requires no additional neural network training and no extra memory footprint, making it a plug-and-play and cost-effective solution for inference acceleration.Benchmarks with LLaMA-2 and its variants demonstrated a speedup up to 1.99×. 1 * Huan Li and Lidan Shou are the corresponding authors. 1 Code is available at https://github.com/dilab-zju/ self-speculative-decoding.
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Sharad Mehrotra
ACL (1)2
2024 Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification
abstract
Large-scale multi-label text classification (LMTC) aims at tagging each text with multiple relevant labels from a large label space, which typically demonstrates high sparsity, diversity, and skewness. To learn text representations in LMTC, a straightforward strategy is to learn a single vector to represent the whole text, yet limiting good generalization to diverse labels; another popular one is to learn specific representation per label via attention weighting, but excessively emphasizing tail labels restricts the overall performance. To cope with these limitations, we propose a novel LMTC framework, dubbed LADAR, which learns label-adaptive text representations to ensure high performance on large-scale labels. Specifically, we construct a representation pool for each text by collecting multi-layer features of the deep model as well as multi-granularity features of the text. Furthermore, all labels are adaptively matched to their most relevant representations to predict the final scores. Experiments over five benchmark datasets demonstrate the LADAR achieves highly superior results to state-of-the-art LMTC approaches. In particular, LADAR achieves significantly better performance on tail labels, e.g., 5.09% relative improvement on PSP@5 on the Amazon-670K dataset than the best baseline.
Cheng Peng 0011, Haobo Wang 0001, Jue Wang 0019, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Chang Yao 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Effective Continual Learning for Text Classification with Lightweight Snapshots
abstract
Continual learning is known for suffering from catastrophic forgetting, a phenomenon where previously learned concepts are forgotten upon learning new tasks. A natural remedy is to use trained models for old tasks as ‘teachers’ to regularize the update of the current model to prevent such forgetting. However, this requires storing all past models, which is very space-consuming for large models, e.g. BERT, thus impractical in real-world applications. To tackle this issue, we propose to construct snapshots of seen tasks whose key knowledge is captured in lightweight adapters. During continual learning, we transfer knowledge from past snapshots to the current model through knowledge distillation, allowing the current model to review previously learned knowledge while learning new tasks. We also design representation recalibration to better handle the class-incremental setting. Experiments over various task sequences show that our approach effectively mitigates catastrophic forgetting and outperforms all baselines.
Jue Wang 0019, Dajie Dong, Lidan Shou, Ke Chen 0005, Gang Chen 0001
AAAI1
2022 SkipBERT: Efficient Inference with Shallow Layer Skipping
abstract
In this paper, we propose SkipBERT to accelerate BERT inference by skipping the computation of shallow layers.To achieve this, our approach encodes small text chunks into independent representations, which are then materialized to approximate the shallow representation of BERT.Since the use of such approximation is inexpensive compared with transformer calculations, we leverage it to replace the shallow layers of BERT to skip their runtime overhead.With off-the-shelf early exit mechanisms, we also skip redundant computation from the highest few layers to further improve inference efficiency.Results on GLUE show that our approach can reduce latency by 65% without sacrificing performance.By using only two-layer transformer calculations, we can still maintain 95% accuracy of BERT. 1
Jue Wang 0019, Ke Chen 0005, Gang Chen 0001, Lidan Shou, Julian J. McAuley
ACL (1)1
2022 Continual Federated Learning Based on Knowledge Distillation
abstract
Federated learning (FL) is a promising approach for learning a shared global model on decentralized data owned by multiple clients without exposing their privacy. In real-world scenarios, data accumulated at the client-side varies in distribution over time. As a consequence, the global model tends to forget the knowledge obtained from previous tasks while learning new tasks, showing signs of "catastrophic forgetting". Previous studies in centralized learning use techniques such as data replay and parameter regularization to mitigate catastrophic forgetting. Unfortunately, these techniques cannot adequately solve the non-trivial problem in FL. We propose Continual Federated Learning with Distillation (CFeD) to address catastrophic forgetting under FL. CFeD performs knowledge distillation on both the clients and the server, with each party independently having an unlabeled surrogate dataset, to mitigate forgetting. Moreover, CFeD assigns different learning objectives, namely learning the new task and reviewing old tasks, to different clients, aiming to improve the learning ability of the model. The results show that our method performs well in mitigating catastrophic forgetting and achieves a good trade-off between the two objectives.
Zhongle Xie, Jue Wang 0019, Ke Chen 0005, Lidan Shou
IJCAI3
2021 Effective Slot Filling via Weakly-Supervised Dual-Model Learning
Jue Wang 0019, Ke Chen 0005, Lidan Shou, Sai Wu, Gang Chen 0001
AAAI1
2020 Pyramid: A Layered Model for Nested Named Entity Recognition
abstract
This paper presents Pyramid, a novel layered model for Nested Named Entity Recognition (nested NER).In our approach, token or text region embeddings are recursively inputted into L flat NER layers, from bottom to top, stacked in a pyramid shape.Each time an embedding passes through a layer of the pyramid, its length is reduced by one.Its hidden state at layer l represents an l-gram in the input text, which is labeled only if its corresponding text region represents a complete entity mention.We also design an inverse pyramid to allow bidirectional interaction between layers.The proposed method achieves state-of-the-art F1 scores in nested NER on ACE-2004, ACE-2005, GENIA, and NNE, which are 80.27, 79.42, 77.78, and 93.70 with conventional embeddings, and 87.74, 86.34, 79.31, and 94.68 with pre-trained contextualized embeddings.In addition, our model can be used for the more general task of Overlapping Named Entity Recognition.A preliminary experiment confirms the effectiveness of our method in overlapping NER.
Jue Wang 0019, Lidan Shou, Ke Chen 0005, Gang Chen 0001
ACL1
2020 Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence Encoders
abstract
Named entity recognition and relation extraction are two important fundamental problems.Joint learning algorithms have been proposed to solve both tasks simultaneously, and many of them cast the joint task as a table-filling problem.However, they typically focused on learning a single encoder (usually learning representation in the form of a table) to capture information required for both tasks within the same space.We argue that it can be beneficial to design two distinct encoders to capture such two different types of information in the learning process.In this work, we propose the novel table-sequence encoders where two different encoders -a table encoder and a sequence encoder are designed to help each other in the representation learning process.Our experiments confirm the advantages of having two encoders over one encoder.On several standard datasets, our model shows significant improvements over existing approaches. 1
Jue Wang 0019, Wei Lu 0011
EMNLP (1)1