VLDB 2026 Research / reviewers in the wild / expert
Jun Zhang 0069
dblp:29/4190-69
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-5245-7956ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interleaved Tool-Call Reasoning for Protein Function UnderstandingabstractChuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Ziqiang Cao, Guohong Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chuanliu Fan, Zicheng Ma, Huanran Meng, Jun Zhang 0069, Ziqiang Cao, Guohong Fu |
ACL (1) | 6 |
| 2026 | See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMsabstractVideo Large Language Models (Video-LLMs) excel in video understanding but suffer from high inference latency during autoregressive generation.Speculative Decoding (SD) mitigates this by applying a draft-and-verify paradigm, yet existing methods are constrained by rigid exact-match rules, severely limiting the acceleration potential.To bridge this gap, we propose LVSPEC, the first training-free loosely SD framework tailored for Video-LLMs.Grounded in the insight that generation is governed by sparse visual-relevant anchors (mandating strictness) amidst abundant visualirrelevant fillers (permitting loose verification), LVSPEC employs a lightweight visual-relevant token identification scheme to accurately pinpoint the former.To further maximize acceptance, we augment this with a position-shift tolerant mechanism that effectively salvages positionally mismatched but semantically equivalent tokens.Experiments demonstrate that LVSPEC achieves high fidelity and speed: it preserves >99.8% of target performance while accelerating Qwen2.5-VL-32B by 2.70× and LLaVA-OneVision-72B by 2.94×.Notably, it boosts the mean accepted length and speedup ratio by 136% and 35% compared to SOTA training-free SD methods for Video-LLMs.Code is at https://github.com/zju-jiyicheng/LVSpec. Yicheng Ji, Jun Zhang 0069, Jinpeng Chen 0001, Lidan Shou, Gang Chen 0001, Huan Li 0003 |
ACL (1) | 2 |
| 2026 | HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceabstractMultimodal Large Language Models (MLLMs) have advanced unified reasoning over text, images, and videos, but their inference is hindered by the rapid growth of key-value (KV) caches.Each visual input expands into thousands of tokens, causing caches to scale linearly with context length and remain resident in GPU memory throughout decoding, which leads to prohibitive memory overhead and latency even on high-end GPUs.A common solution is to compress caches under a fixed allocated budget at different granularities: tokenlevel uniformly discards less important tokens, layer-level varies retention across layers, and head-level redistributes budgets across heads.Yet these approaches stop at allocation and overlook the heterogeneous behaviors of attention heads that require distinct compression strategies.We propose HYBRIDKV, a hybrid KV cache compression framework that integrates complementary strategies in three stages: heads are first classified into static or dynamic types using text-centric attention; then a top-down budget allocation scheme hierarchically assigns KV budgets; finally, static heads are compressed by text-prior pruning and dynamic heads by chunk-wise retrieval.Experiments on 11 multimodal benchmarks with Qwen2.5-VL-7Bshow that HYBRIDKV reduces KV cache memory by up to 7.9× and achieves 1.52× faster decoding, with almost no performance drop or even higher relative to the full-cache MLLM. Feiyang Ren, Jun Zhang 0069, Xiaoling Gu, Ke Chen 0005, Lidan Shou, Huan Li 0003 |
ACL (1) | 3 |
| 2025 | LlaMol: A Unified Molecule Designer via Preference Ranking and Numerical EnhancementabstractGoal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints from scratch, is a crucial yet challenging task in drug discovery. Existing research often relies on separate predictors for distinct properties and struggles with integrating substructure constraints due to the complexities involved in modeling structural information via multitask learning. This separation necessitates a dedicated prediction model for each constraint, limiting the flexibility and posing challenges for realworld applications. To address these limitations, we propose a unified framework for molecular design that incorporates multiple property and substructure constraints, leveraging LLMs to handle diverse constraint settings within a single model. We first integrate feedback learning derived from preference ranking to eliminate the need for separate property predictors. Then, we enhance the model's ability to follow numerical instructions by introducing a unified numerical encoding into the prompt. We conduct extensive experiments across single-property, substructureproperty, and multi-property constrained tasks. Experimental results demonstrate that LlaMol consistently outperforms state-of-the-art baselines across various constraint settings. Notably, in the multi-objective binding affinity maximization task, LlaMol achieves a significantly lower$\mathrm{K}_{\mathrm{D}}$value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76 %. These results underscore the effectiveness and versatility of LLM-based frameworks for molecule generation under complex constraints. Chuanliu Fan, Zicheng Ma, Jun Zhang 0069, Ziqiang Cao, Yiqin Gao, Guohong Fu |
BIBM | 5 |
| 2025 | SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token PruningabstractVideo large language models (Vid-LLMs) have shown strong capabilities in understanding video content.However, their reliance on dense video token representations introduces substantial memory and computational overhead in both prefilling and decoding.To mitigate the information loss of recent video token reduction methods and accelerate the decoding stage of Vid-LLMs losslessly, we introduce SPECVLM, a training-free speculative decoding (SD) framework tailored for Vid-LLMs that incorporates staged video token pruning.Building on our novel finding that the draft model's speculation exhibits low sensitivity to video token pruning, SPECVLM prunes up to 90% of video tokens to enable efficient speculation without sacrificing accuracy.To achieve this, we perform a two-stage pruning process: Stage I selects highly informative tokens guided by attention signals from the verifier (target model), while Stage II prunes the remaining redundant ones in a spatially uniform manner.Extensive experiments on four video understanding benchmarks demonstrate the effectiveness and robustness of SPECVLM, which achieves up to 2.68× decoding speedup for LLaVA-OneVision-72B and 2.11× speedup for Qwen2.5-VL-32B.Code is available at https: //github.com/zju-jiyicheng/SpecVLM. Yicheng Ji, Jun Zhang 0069, Heming Xia, Jinpeng Chen 0001, Lidan Shou, Gang Chen 0001, Huan Li 0003 |
EMNLP | 2 |
| 2025 | SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationabstractSpeculative decoding (SD) has emerged as a widely used paradigm to accelerate LLM inference without compromising quality. It works by first employing a compact model to draft multiple tokens efficiently and then using the target LLM to verify them in parallel. While this technique has achieved notable speedups, most existing approaches necessitate either additional parameters or extensive training to construct effective draft models, thereby restricting their applicability across different LLMs and tasks. To address this limitation, we explore a novel plug-and-play SD solution with layer-skipping, which skips intermediate layers of the target LLM as the compact draft model. Our analysis reveals that LLMs exhibit great potential for self-acceleration through layer sparsity and the task-specific nature of this sparsity. Building on these insights, we introduce SWIFT, an on-the-fly self-speculative decoding algorithm that adaptively selects intermediate layers of LLMs to skip during inference. SWIFT does not require auxiliary models or additional training, making it a plug-and-play solution for accelerating LLM inference across diverse input data streams. Our extensive experiments across a wide range of models and downstream tasks demonstrate that SWIFT can achieve over a $1.3\times$$\sim$$1.6\times$ speedup while preserving the original distribution of the generated text. We release our code in https://github.com/hemingkx/SWIFT. Heming Xia, Yongqi Li 0001, Jun Zhang 0069, Cunxiao Du, Wenjie Li 0002 |
ICLR | 3 |
| 2025 | Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language ModelsabstractLarge Language Models (LLMs) have significantly advanced natural language processing with exceptional task generalization capabilities. Low-Rank Adaption (LoRA) offers a cost-effective fine-tuning solution, freezing the original model parameters and training only lightweight, low-rank adapter matrices. However, the memory footprint of LoRA is largely dominated by the original model parameters. To mitigate this, we propose LoRAM, a memory-efficient LoRA training scheme founded on the intuition that many neurons in over-parameterized LLMs have low training utility but are essential for inference. LoRAM presents a unique twist: it trains on a pruned (small) model to obtain pruned low-rank matrices, which are then recovered and utilized with the original (large) model for inference. Additionally, minimal-cost continual pre-training, performed by the model publishers in advance, aligns the knowledge discrepancy between pruned and original models. Our extensive experiments demonstrate the efficacy of LoRAM across various pruning strategies and downstream tasks. For a model with 70 billion parameters, LoRAM enables training on a GPU with only 20G HBM, replacing an A100-80G GPU for LoRA training and 15 GPUs for full fine-tuning. Specifically, QLoRAM implemented by structured pruning combined with 4-bit quantization, for LLaMA-3.1-70B (LLaMA-2-70B), reduces the parameter storage cost that dominates the memory usage in low-rank matrix training by 15.81× (16.95×), while achieving dominant performance gains over both the original LLaMA-3.1-70B (LLaMA-2-70B) and LoRA-trained LLaMA-3.1-8B (LLaMA-2-13B). Code is available at https://github.com/junzhang-zj/LoRAM. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Guiming Xie, Xuejian Gong, Kunlong Zhou |
ICLR | 1 |
| 2025 | FloE: On-the-Fly MoE Inference on Memory-constrained GPUabstractWith the widespread adoption of Mixture-of-Experts (MoE) models, there is a growing demand for efficient inference on memory-constrained devices.
While offloading expert parameters to CPU memory and loading activated experts on demand has emerged as a potential solution, the large size of activated experts overburdens the limited PCIe bandwidth, hindering the effectiveness in latency-sensitive scenarios.
To mitigate this, we propose FloE, an on-the-fly MoE inference system on memory-constrained GPUs.
FloE is built on the insight that there exists substantial untapped redundancy within sparsely activated experts.
It employs various compression techniques on the expert's internal parameter matrices to reduce the data movement load, combined with low-cost sparse prediction, achieving perceptible inference acceleration in wall-clock time on resource-constrained devices.
Empirically, FloE achieves a 9.3$\times$ compression of parameters per expert in Mixtral-8$\times$7B; enables deployment on a GPU with only 11GB VRAM, reducing the memory footprint by up to 8.5$\times$; and delivers a 48.7$\times$ inference speedup compared to DeepSpeed-MII on a single GeForce RTX 3090—all with only a 4.4\% $\sim$ 7.6\% average performance degradation. Zheng Li 0006, Jun Zhang 0069, Jue Wang 0019, Yiping Wang 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou |
ICML | 3 |
| 2025 | ${\sf CHASe}$CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active LearningabstractActive learning (AL) reduces human annotation costs for machine learning systems by strategically selecting the most informative unlabeled data for annotation, but performing it individually may still be insufficient due to restricted data diversity and annotation budget. Federated Active Learning (FAL) addresses this by facilitating collaborative data selection and model training, while preserving the confidentiality of raw data samples. Yet, existing FAL methods fail to account for the heterogeneity of data distribution across clients and the associated fluctuations in global and local model parameters, adversely affecting model accuracy. To overcome these challenges, we propose${\sf CHASe}$(Client Heterogeneity-Aware Data Selection), specifically designed for FAL.${\sf CHASe}$focuses on identifying those unlabeled samples with high epistemic variations (EVs), which notably oscillate around the decision boundaries during training. To achieve both effectiveness and efficiency,${\sf CHASe}$encompasses techniques for 1) tracking EVs by analyzing inference inconsistencies across training epochs, 2) calibrating decision boundaries of inaccurate models with a new alignment loss, and 3) enhancing data selection efficiency via a data freeze and awaken mechanism with subset sampling. Experiments show that${\sf CHASe}$surpasses various established baselines in terms of effectiveness and efficiency, validated across diverse datasets, model complexities, and heterogeneous federation settings. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Zhongle Xie, Ke Chen 0005, Lidan Shou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | HMI: hierarchical knowledge management for efficient multi-tenant inference in pretrained language models
Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Guiming Xie, Xuejian Gong |
VLDB J. | 1 |
| 2024 | Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?abstractZhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaochen Su, Juntao Li 0005, Jun Zhang 0069, Tong Zhu 0002, Xiaoye Qu, Pan Zhou 0001, Yan Bowen, Yu Cheng 0001, Min Zhang 0005 |
ACL (1) | 3 |
| 2024 | Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingabstractWe present a novel inference scheme, selfspeculative decoding, for accelerating Large Language Models (LLMs) without the need for an auxiliary model.This approach is characterized by a two-stage process: drafting and verification.The drafting stage generates draft tokens at a slightly lower quality but more quickly, which is achieved by selectively skipping certain intermediate layers during drafting.Subsequently, the verification stage employs the original LLM to validate those draft output tokens in one forward pass.This process ensures the final output remains identical to that produced by the unaltered LLM.Moreover, the proposed method requires no additional neural network training and no extra memory footprint, making it a plug-and-play and cost-effective solution for inference acceleration.Benchmarks with LLaMA-2 and its variants demonstrated a speedup up to 1.99×. 1 * Huan Li and Lidan Shou are the corresponding authors. 1 Code is available at https://github.com/dilab-zju/ self-speculative-decoding. Jun Zhang 0069, Jue Wang 0019, Huan Li 0003, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Sharad Mehrotra |
ACL (1) | 1 |
| 2024 | TabMedBERT: A Tabular Knowledge Enhanced Biomedical Pretrained Language ModelabstractMost existing biomedical language models are trained on plain text with general learning goals such as random word infilling, failing to capture the knowledge in the biomedical corpus sufficiently. Since biomedical articles usually contain many tables summarising the main entities and their relations, in the paper, we propose a Tabular knowledge enhanced bioMedical pretrained language model, called TabMedBERT. Specifically, we align entities between table cells, and article text spans with pre-defined rules. Then we add two table-related self-supervised tasks to integrate tabular knowledge into the language model: Entity Infilling (EI) and Table Cloze Test (TCT). While EI masks tokens within aligned entities in the article, TCT converts aligned entities in the table layout into a cloze text by erasing one entity and prompts the model to extract the appropriate span to fill in the blank. Experimental results demonstrate that TabMedBERT surpasses all competing language models without adding additional parameters, establishing a new state-of-the-art performance of 85.59% (+1.29%) on the BLURB biomedical NLP benchmark and 7 additional information extraction datasets. Moreover, the model architecture for TCT provides a straightforward solution to revise information extraction with paired entities. Lei Geng, Ziqiang Cao, Juntao Li 0005, Wenjie Li 0002, Sujian Li, Yang Yang 0074, Jun Zhang 0069 |
ECAI | 9 |
| 2024 | Contrastive Learning with High-Quality and Low-Quality Augmented Data for Query-Focused SummarizationabstractUnlike general text summarization, Query-focused summarization (QFS) is severely limited by insufficient datasets, forcing previous research to transform datasets from other tasks into QFS format for data augmentation. However, this approach has resulted in two problems: the task and traintest gaps. To alleviate these gaps, we propose QFS-CL, a novel in-place data augmentation framework equipped with contrastive learning. Firstly, we design diverse prompts for ChatGPT to paraphrase the original QFS data into high-quality/low-quality document-summary pairs, filling the task gap. Then, instead of directly incorporating the augmented data into the training set, we train the QFS baseline model in a contrastive learning scheme. Specifically, our approach encourages the model to imitate high-quality pairs and distinguish itself from low-quality pairs, enabling the model to learn how to acquire reliable information and avoid extracting invalid information. Our method achieves state-of-the-art performance on Debatepedia and DUC datasets in ROUGE scores, GPT-4, and human evaluations. Shaoyao Huang, Ziqiang Cao, Luozheng Qin, Jun Zhang 0069 |
ICASSP | 5 |
| 2024 | Preventing the Popular Item Embedding Based Attack in Federated RecommendationsabstractPrivacy concerns have led to the rise of federated recommender systems (FRS), which can create personalized models across distributed clients. However, FRS is vulnerable to poisoning attacks, where malicious users manipulate gradients to promote their target items intentionally. Existing attacks against FRS have limitations, as they depend on specific models and prior knowledge, restricting their real-world applicability. In our exploration of practical FRS vulnerabilities, we devise a model-agnostic and prior-knowledge-free attack, named PIECK (Popular Item Embedding based Attack). The core module of PIECK is popular item mining, which leverages embedding changes during FRS training to effectively identify the popular items. Built upon the core module, PIECK branches into two diverse solutions: The PIECKIPE solution employs an item popularity enhancement module, which aligns the embeddings of targeted items with the mined popular items to increase item exposure. The PIECKUEA further enhances the robustness of the attack by using a user embedding approximation module, which approximates private user embeddings using mined popular items. Upon identifying PIECK, we evaluate existing federated defense methods and find them ineffective against PIECK, as poisonous gradients inevitably overwhelm the cold target items. We then propose a novel defense method by introducing two regularization terms during user training, which constrain item popularity enhancement and user embedding approximation while preserving FRS performance. We evaluate PIECK and its defense across two base models, three real datasets, four top-tier attacks, and six general defense methods, affirming the efficacy of both PIECK and its defense. Jun Zhang 0069, Huan Li 0003, Dazhong Rong, Yan Zhao 0008, Ke Chen 0005, Lidan Shou |
ICDE | 1 |
| 2024 | ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001 |
NeurIPS | 2 |