EDBT 2026 Demo / reviewers in the wild / expert
Jue Wang 0019
dblp:69/393-19
· DBLP profile ↗
12ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-6712-1929ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Efficient and distributed learning · 62% Information extraction and text analysis · 16% Language models and text generation · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
3.0 | 4 | 2025 | FloE: On-the-Fly MoE Inference on Memory-constrained GPU · ICML 2025 Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025 Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023 |
Machine learning › Efficient and distributed learning
federated learning |
1.4 | 2 | 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025 Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022 |
Natural language and speech › Information extraction and text analysis
text classification |
1.0 | 2 | 2024 | Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023 |
Machine learning › Efficient and distributed learning › federated learning › heterogeneous federated learning
client heterogeneity |
0.9 | 1 | 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Efficient and distributed learning › federated learning › label-efficient federated learning
federated active learning |
0.9 | 1 | 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.9 | 1 | 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
0.9 | 1 | 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference |
0.9 | 1 | 2025 | FloE: On-the-Fly MoE Inference on Memory-constrained GPU · ICML 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
pre-trained language model inference |
0.9 | 1 | 2025 | HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models · VLDB J. 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 2 | 2023 | Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023 Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.8 | 1 | 2024 | Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024 |
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification |
0.8 | 1 | 2024 | Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Machine learning › Efficient and distributed learning › inference acceleration › speculative decoding
self-speculative decoding |
0.8 | 1 | 2024 | Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.8 | 1 | 2024 | Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › text embedding
text representation learning |
0.8 | 1 | 2024 | Learning Label-Adaptive Representation for Large-Scale Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Machine learning › Learning paradigms
continual learning |
0.7 | 1 | 2023 | Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
lightweight adapters |
0.7 | 1 | 2023 | Effective Continual Learning for Text Classification with Lightweight Snapshots · AAAI 2023 |
Machine learning › Efficient and distributed learning › federated learning
federated continual learning |
0.6 | 1 | 2022 | Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.6 | 1 | 2022 | Continual Federated Learning Based on Knowledge Distillation · IJCAI 2022 |
Machine learning › Efficient and distributed learning › adaptive computation
layer skipping |
0.6 | 1 | 2022 | SkipBERT: Efficient Inference with Shallow Layer Skipping · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis
slot filling |
0.5 | 1 | 2021 | Effective Slot Filling via Weakly-Supervised Dual-Model Learning · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.4 | 1 | 2020 | Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence Encoders · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.4 | 1 | 2020 | Pyramid: A Layered Model for Nested Named Entity Recognition · ACL 2020 |
Natural language and speech › Information extraction and text analysis › named entity recognition
nested named entity recognition |
0.4 | 1 | 2020 | Pyramid: A Layered Model for Nested Named Entity Recognition · ACL 2020 |
Machine learning › Efficient and distributed learning
active learning |
0.3 | 1 | 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.3 | 1 | 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.3 | 1 | 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.2 | 1 | 2022 | SkipBERT: Efficient Inference with Shallow Layer Skipping · ACL (1) 2022 |
Machine learning › Learning paradigms
weakly supervised learning |
0.1 | 1 | 2021 | Effective Slot Filling via Weakly-Supervised Dual-Model Learning · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 3.0sparse prediction · 1.7offloading · 1.7hierarchical indexing · 1.7expert parameter compression · 1.7structured pruning · 0.9quantization · 0.9low-rank adaptation · 0.9epistemic uncertainty · 0.9active learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language ModelsabstractLarge Language Models (LLMs) have significantly advanced natural language processing with exceptional task generalization capabilities. Low-Rank Adaption (LoRA) offers a cost-effective fine-tuning solution, freezing the original model parameters and training only lightweight, low-rank adapter matrices. However, the memory footprint of LoRA is largely dominated by the original model parameters. To mitigate this, we propose LoRAM, a memory-efficient LoRA training scheme founded on the intuition that many neurons in over-parameterized LLMs have low training utility but are essential for inference. LoRAM presents a unique twist: it trains on a pruned (small) model to obtain pruned low-rank matrices, which are then recovered and utilized with the original (large) model for inference. Additionally, minimal-cost continual pre-training, performed by the model publishers in advance, aligns the knowledge discrepancy between pruned and original models. Our extensive experiments demonstrate the efficacy of LoRAM across various pruning strategies and downstream tasks. For a model with 70 billion parameters, LoRAM enables training on a GPU with only 20G HBM, replacing an A100-80G GPU for LoRA training and 15 GPUs for full fine-tuning. Specifically, QLoRAM implemented by structured pruning combined with 4-bit quantization, for LLaMA-3.1-70B (LLaMA-2-70B), reduces the parameter storage cost that dominates the memory usage in low-rank matrix training by 15.81× (16.95×), while achieving dominant performance gains over both the original LLaMA-3.1-70B (LLaMA-2-70B) and LoRA-trained LLaMA-3.1-8B (LLaMA-2-13B). Code is available at https://github.com/junzhang-zj/LoRAM. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Guiming Xie, Xuejian Gong, Kunlong Zhou |
ICLR | 2 |
| 2025 | FloE: On-the-Fly MoE Inference on Memory-constrained GPUabstractWith the widespread adoption of Mixture-of-Experts (MoE) models, there is a growing demand for efficient inference on memory-constrained devices.
While offloading expert parameters to CPU memory and loading activated experts on demand has emerged as a potential solution, the large size of activated experts overburdens the limited PCIe bandwidth, hindering the effectiveness in latency-sensitive scenarios.
To mitigate this, we propose FloE, an on-the-fly MoE inference system on memory-constrained GPUs.
FloE is built on the insight that there exists substantial untapped redundancy within sparsely activated experts.
It employs various compression techniques on the expert's internal parameter matrices to reduce the data movement load, combined with low-cost sparse prediction, achieving perceptible inference acceleration in wall-clock time on resource-constrained devices.
Empirically, FloE achieves a 9.3$\times$ compression of parameters per expert in Mixtral-8$\times$7B; enables deployment on a GPU with only 11GB VRAM, reducing the memory footprint by up to 8.5$\times$; and delivers a 48.7$\times$ inference speedup compared to DeepSpeed-MII on a single GeForce RTX 3090—all with only a 4.4\% $\sim$ 7.6\% average performance degradation. Zheng Li 0006, Jun Zhang 0069, Jue Wang 0019, Yiping Wang 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou |
ICML | 4 |
| 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active LearningabstractActive learning (AL) reduces human annotation costs for machine learning systems by strategically selecting the most informative unlabeled data for annotation, but performing it individually may still be insufficient due to restricted data diversity and annotation budget. Federated Active Learning (FAL) addresses this by facilitating collaborative data selection and model training, while preserving the confidentiality of raw data samples. Yet, existing FAL methods fail to account for the heterogeneity of data distribution across clients and the associated fluctuations in global and local model parameters, adversely affecting model accuracy. To overcome these challenges, we propose${\sf CHASe}$(Client Heterogeneity-Aware Data Selection), specifically designed for FAL.${\sf CHASe}$focuses on identifying those unlabeled samples with high epistemic variations (EVs), which notably oscillate around the decision boundaries during training. To achieve both effectiveness and efficiency,${\sf CHASe}$encompasses techniques for 1) tracking EVs by analyzing inference inconsistencies across training epochs, 2) calibrating decision boundaries of inaccurate models with a new alignment loss, and 3) enhancing data selection efficiency via a data freeze and awaken mechanism with subset sampling. Experiments show that${\sf CHASe}$surpasses various established baselines in terms of effectiveness and efficiency, validated across diverse datasets, model complexities, and heterogeneous federation settings. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Guiming Xie, Xuejian Gong |
VLDB J. | 2 |
| 2024 | Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingabstractWe present a novel inference scheme, selfspeculative decoding, for accelerating Large Language Models (LLMs) without the need for an auxiliary model.This approach is characterized by a two-stage process: drafting and verification.The drafting stage generates draft tokens at a slightly lower quality but more quickly, which is achieved by selectively skipping certain intermediate layers during drafting.Subsequently, the verification stage employs the original LLM to validate those draft output tokens in one forward pass.This process ensures the final output remains identical to that produced by the unaltered LLM.Moreover, the proposed method requires no additional neural network training and no extra memory footprint, making it a plug-and-play and cost-effective solution for inference acceleration.Benchmarks with LLaMA-2 and its variants demonstrated a speedup up to 1.99×. 1 * Huan Li and Lidan Shou are the corresponding authors. 1 Code is available at https://github.com/dilab-zju/ self-speculative-decoding. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Sharad Mehrotra |
ACL (1) | 2 |
| 2024 | Learning Label-Adaptive Representation for Large-Scale Multi-Label Text ClassificationabstractLarge-scale multi-label text classification (LMTC) aims at tagging each text with multiple relevant labels from a large label space, which typically demonstrates high sparsity, diversity, and skewness. To learn text representations in LMTC, a straightforward strategy is to learn a single vector to represent the whole text, yet limiting good generalization to diverse labels; another popular one is to learn specific representation per label via attention weighting, but excessively emphasizing tail labels restricts the overall performance. To cope with these limitations, we propose a novel LMTC framework, dubbed LADAR, which learns label-adaptive text representations to ensure high performance on large-scale labels. Specifically, we construct a representation pool for each text by collecting multi-layer features of the deep model as well as multi-granularity features of the text. Furthermore, all labels are adaptively matched to their most relevant representations to predict the final scores. Experiments over five benchmark datasets demonstrate the LADAR achieves highly superior results to state-of-the-art LMTC approaches. In particular, LADAR achieves significantly better performance on tail labels, e.g., 5.09% relative improvement on PSP@5 on the Amazon-670K dataset than the best baseline. Cheng Peng 0011, Haobo Wang 0001, Jue Wang 0019, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Chang Yao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Effective Continual Learning for Text Classification with Lightweight SnapshotsabstractContinual learning is known for suffering from catastrophic forgetting, a phenomenon where previously learned concepts are forgotten upon learning new tasks. A natural remedy is to use trained models for old tasks as ‘teachers’ to regularize the update of the current model to prevent such forgetting. However, this requires storing all past models, which is very space-consuming for large models, e.g. BERT, thus impractical in real-world applications. To tackle this issue, we propose to construct snapshots of seen tasks whose key knowledge is captured in lightweight adapters. During continual learning, we transfer knowledge from past snapshots to the current model through knowledge distillation, allowing the current model to review previously learned knowledge while learning new tasks. We also design representation recalibration to better handle the class-incremental setting. Experiments over various task sequences show that our approach effectively mitigates catastrophic forgetting and outperforms all baselines. Jue Wang 0019, Dajie Dong, Lidan Shou, Ke Chen 0005, Gang Chen 0001 |
AAAI | 1 |
| 2022 | SkipBERT: Efficient Inference with Shallow Layer SkippingabstractIn this paper, we propose SkipBERT to accelerate BERT inference by skipping the computation of shallow layers.To achieve this, our approach encodes small text chunks into independent representations, which are then materialized to approximate the shallow representation of BERT.Since the use of such approximation is inexpensive compared with transformer calculations, we leverage it to replace the shallow layers of BERT to skip their runtime overhead.With off-the-shelf early exit mechanisms, we also skip redundant computation from the highest few layers to further improve inference efficiency.Results on GLUE show that our approach can reduce latency by 65% without sacrificing performance.By using only two-layer transformer calculations, we can still maintain 95% accuracy of BERT. 1 Jue Wang 0019, Ke Chen 0005, Gang Chen 0001, Lidan Shou, Julian J. McAuley |
ACL (1) | 1 |
| 2022 | Continual Federated Learning Based on Knowledge DistillationabstractFederated learning (FL) is a promising approach for learning a shared global model on decentralized data owned by multiple clients without exposing their privacy. In real-world scenarios, data accumulated at the client-side varies in distribution over time. As a consequence, the global model tends to forget the knowledge obtained from previous tasks while learning new tasks, showing signs of "catastrophic forgetting". Previous studies in centralized learning use techniques such as data replay and parameter regularization to mitigate catastrophic forgetting. Unfortunately, these techniques cannot adequately solve the non-trivial problem in FL. We propose Continual Federated Learning with Distillation (CFeD) to address catastrophic forgetting under FL. CFeD performs knowledge distillation on both the clients and the server, with each party independently having an unlabeled surrogate dataset, to mitigate forgetting. Moreover, CFeD assigns different learning objectives, namely learning the new task and reviewing old tasks, to different clients, aiming to improve the learning ability of the model. The results show that our method performs well in mitigating catastrophic forgetting and achieves a good trade-off between the two objectives. Zhongle Xie, Jue Wang 0019, Ke Chen 0005, Lidan Shou |
IJCAI | 3 |
| 2021 | Effective Slot Filling via Weakly-Supervised Dual-Model Learning
Jue Wang 0019, Ke Chen 0005, Lidan Shou, Sai Wu, Gang Chen 0001 |
AAAI | 1 |
| 2020 | Pyramid: A Layered Model for Nested Named Entity RecognitionabstractThis paper presents Pyramid, a novel layered model for Nested Named Entity Recognition (nested NER).In our approach, token or text region embeddings are recursively inputted into L flat NER layers, from bottom to top, stacked in a pyramid shape.Each time an embedding passes through a layer of the pyramid, its length is reduced by one.Its hidden state at layer l represents an l-gram in the input text, which is labeled only if its corresponding text region represents a complete entity mention.We also design an inverse pyramid to allow bidirectional interaction between layers.The proposed method achieves state-of-the-art F1 scores in nested NER on ACE-2004, ACE-2005, GENIA, and NNE, which are 80.27, 79.42, 77.78, and 93.70 with conventional embeddings, and 87.74, 86.34, 79.31, and 94.68 with pre-trained contextualized embeddings.In addition, our model can be used for the more general task of Overlapping Named Entity Recognition.A preliminary experiment confirms the effectiveness of our method in overlapping NER. Jue Wang 0019, Lidan Shou, Ke Chen 0005, Gang Chen 0001 |
ACL | 1 |
| 2020 | Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence EncodersabstractNamed entity recognition and relation extraction are two important fundamental problems.Joint learning algorithms have been proposed to solve both tasks simultaneously, and many of them cast the joint task as a table-filling problem.However, they typically focused on learning a single encoder (usually learning representation in the form of a table) to capture information required for both tasks within the same space.We argue that it can be beneficial to design two distinct encoders to capture such two different types of information in the learning process.In this work, we propose the novel table-sequence encoders where two different encoders -a table encoder and a sequence encoder are designed to help each other in the representation learning process.Our experiments confirm the advantages of having two encoders over one encoder.On several standard datasets, our model shows significant improvements over existing approaches. 1 Jue Wang 0019, Wei Lu 0011 |
EMNLP (1) | 1 |