Muquan Li

dblp:388/3240 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0009-3417-2007ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 48% Deep learning architectures and training · 14% Generative modeling · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph reasoning
1.012026
GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding · WWW 2026
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding · WWW 2026
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
1.012026
GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding · WWW 2026
Information retrieval › retrieval-augmented generation
graph-based retrieval-augmented generation
1.012026
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation · WWW 2026
Information retrieval
retrieval-augmented generation
1.012026
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation · WWW 2026
Security and privacy of machine learning › retrieval-augmented generation security
knowledge poisoning
1.012026
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation · WWW 2026
Security and privacy of machine learning
poisoning attack
1.012026
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation · WWW 2026
Machine learning › Deep learning architectures and training › backpropagation
backpropagation through time
0.912025
Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation · NeurIPS 2025
Machine learning › Efficient and distributed learning › data selection
coreset selection
0.912025
Adaptive Dataset Quantization · AAAI 2025
Machine learning › Efficient and distributed learning › data reduction
dataset compression
0.912025
Adaptive Dataset Quantization · AAAI 2025
Machine learning › Efficient and distributed learning
dataset distillation
0.912025
Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation · NeurIPS 2025
Machine learning › Efficient and distributed learning › data reduction
dataset quantization
0.912025
Adaptive Dataset Quantization · AAAI 2025
Machine learning › Deep learning architectures and training
training optimization
0.912025
Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
data-free knowledge distillation
0.812024
Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation · ACM Multimedia 2024
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation
0.812024
Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation · ACM Multimedia 2024
Machine learning › Generative modeling
diffusion model
0.812024
Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation · ACM Multimedia 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation · ACM Multimedia 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

tool calling · 1.0subgraph sampling · 1.0multi-agent framework · 1.0truncated backpropagation through time · 0.9low-rank hessian approximation · 0.9dataset distillation · 0.9contrastive learning · 0.9adaptive sampling · 0.9diffusion model · 0.8cosine similarity filtering · 0.8
YearPublicationVenuePosition
2026 KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
Qizhi Chen 0001, Muquan Li, Rongzheng Wang, Dongyang Zhang 0001, Ke Qin, Shuang Liang 0002
WWW4
2026 GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding
abstract
Large language models (LLMs) show promising performance on small-scale graph reasoning tasks but fail when handling real-world graphs with complex queries. This phenomenon arises from LLMs' working memory constraints, which result in their inability to retain long-range graph topology over extended contexts while sustaining coherent multi-step reasoning. However, real-world graphs are often structurally complex, such as Web, Transportation, Social, and Citation networks. To address these limitations, we propose GraphCogent, a collaborative agent framework inspired by human Working Memory Model that decomposes graph reasoning into specialized cognitive processes: sense, buffer, and execute. The framework consists of three modules: Sensory Module standardizes diverse graph text representations via subgraph sampling, Buffer Module integrates and indexes graph data across multiple formats, and Execution Module combines tool calling and tool creation for efficient reasoning. We also introduce Graph4real, a comprehensive benchmark that contains four domains of real-world graphs (Web, Transportation, Social, and Citation) to evaluate LLMs' graph reasoning capabilities. Our Graph4real covers 21 different graph reasoning tasks, categorized into three types (Structural Querying, Algorithmic Reasoning, and Predictive Modeling tasks), with graph scales up to 10 times larger than existing benchmarks. Experiments show that Llama3.1-8B based GraphCogent achieves a 50% improvement over massive-scale LLMs like DeepSeek-R1 (671B). Compared to state-of-the-art code-based baseline, our framework outperforms by 20% in accuracy while reducing token usage by 80% for in-toolset tasks and 30% for out-toolset tasks.
Rongzheng Wang, Shuang Liang 0002, Qizhi Chen 0001, Muquan Li, Yizhuo Ma, Dongyang Zhang 0001, Ke Qin, Man-Fai Leung
WWW5
2026 STPE-map: Multimodal alignment and spatio-temporal priors for online HD map construction
Keke Tian, Muquan Li, Ke Qin
Knowl. Based Syst.2
2026 Efficient Dataset Distillation via Generative Pruning
abstract
Dataset distillation (DD) has demonstrated the promise of synthesizing smaller datasets that enable competitive performance. While most of DD methods operate in pixel space, they often suffer from poor scalability on high-resolution datasets. Recent works have thus shifted towards parameterizing synthetic data using deep generative priors. However, these approaches apply latent updates to all generator layers, leading to substantial computational overhead. To address this limitation, we propose Generative Lightweight Distillation (GLiD), a unified framework that jointly compresses the generator and accelerates latent optimization for efficient distillation. GLiD introduces two key components: (1) We introduce a Classification-Diversity Sensitivity Pruning mechanism that quantifies both discriminative utility and semantic diversity of each output channel to guide structural pruning and selective layer-wise optimization. (2) We also present a Layer-Adaptive Scheduling strategy that dynamically allocates latent update steps across generator stages based on convergence behavior. Extensive experiments on CIFAR-10 and ImageNet-1K and its subsets demonstrate that our GLiD achieves up to 10× acceleration across different datasets, while maintaining performance competitive with state-of-the-art methods.
Yingyi Ma, Muquan Li, Guiduo Duan, Ke Qin, Shuang Liang 0002, Dongyang Zhang 0001
IEEE Trans. Big Data2
2026 Efficient Industrial Dataset Distillation With Textual Trajectory Matching
abstract
Modern industrial environments generate massive streams of discrete textual logs that record device status, error codes, and operational events. Deploying models on resource-constrained edge devices demands extreme data compression and fast inference, yet existing dataset distillation (DD) methods are designed for images and capture only short-term training dynamics via single-step gradient or distribution matching. To address these limitations, we propose textual trajectory matching (TTM), a novel textual DD framework that aligns student trajectories with expert trajectories derived from industrial data training. We introduce a Mask-and-Fill initialization to enhance trajectory diversity, expanding semantic representations. We further propose a manifold distribution-based initialization to preserve original semantic features through low-dimensional manifold analysis. To address model computational costs, we design a subset matching strategy to reducing by aligning critical structural components. Evaluated on SST-2, MNLI-m, AGNews, and BlueGene/L (BGL) with$\text{BERT}_{\text{BASE}}$,$\text{RoBERTa}_{\text{BASE}}$,$\text{XLNet}_{\text{BASE}}$, and LLaMA 3, TTM achieves 3.6% higher accuracy than state-of-the-art methods, enabling efficient industrial dataset and model deployment in smart factories.
Muquan Li, Dongyang Zhang 0001, Ke Qin, Guangchun Luo
IEEE Trans. Ind. Informatics1
2025 Adaptive Dataset Quantization
abstract
Contemporary deep learning, characterized by the training of cumbersome neural networks on massive datasets, confronts substantial computational hurdles. To alleviate heavy data storage burdens on limited hardware resources, numerous dataset compression methods such as dataset distillation (DD) and coreset selection have emerged to obtain a compact but informative dataset through synthesis or selection for efficient training. However, DD involves an expensive optimization procedure and exhibits limited generalization across unseen architectures, while coreset selection is limited by its low data keep ratio and reliance on heuristics, hindering its practicality and feasibility. To address these limitations, we introduce a newly versatile framework for dataset compression, namely Adaptive Dataset Quantization (ADQ). Specifically, we first identify the sub-optimal performance of naive Dataset Quantization (DQ), which relies on uniform sampling and overlooks the varying importance of each generated bin. Subsequently, we propose a novel adaptive sampling strategy through the evaluation of generated bins' representativeness score, diversity score and importance score, where the former two scores are quantified by the texture level and contrastive learning-based techniques, respectively. Extensive experiments demonstrate that our method not only exhibits superior generalization capability across different architectures, but also attains state-of-the-art results.
Muquan Li, Dongyang Zhang 0001, Qiang Dong, Xiurui Xie, Ke Qin
AAAI1
2025 Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
abstract
The growing demand for efficient deep learning has positioned dataset distillation as a pivotal technique for compressing training dataset while preserving model performance. However, existing inner-loop optimization methods for dataset distillation typically rely on random truncation strategies, which lack flexibility and often yield suboptimal results. In this work, we observe that neural networks exhibit distinct learning dynamics across different training stages—early, middle, and late—making random truncation ineffective. To address this limitation, we propose Automatic Truncated Backpropagation Through Time (AT-BPTT), a novel framework that dynamically adapts both truncation positions and window sizes according to intrinsic gradient behavior. AT-BPTT introduces three key components: (1) a probabilistic mechanism for stage-aware timestep selection, (2) an adaptive window sizing strategy based on gradient variation, and (3) a low-rank Hessian approximation to reduce computational overhead. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-1K show that AT-BPTT achieves state-of-the-art performance, improving accuracy by an average of 6.16\% over baseline methods. Moreover, our approach accelerates inner-loop optimization by 3.9 × while saving 63\% memory cost.
Muquan Li, Hang Gou, Dongyang Zhang 0001, Shuang Liang 0002, Xiurui Xie, Deqiang Ouyang, Ke Qin
NeurIPS1
2024 Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation
abstract
Data-free knowledge distillation (DFKD) has emerged as a pivotal technique in the domain of model compression, substantially reducing the dependency on the original training data. Nonetheless, conventional DFKD methods that employ synthesized training data are prone to the limitations of inadequate diversity and discrepancies in distribution between the synthesized and original datasets. To address these challenges, this paper introduces an innovative approach to DFKD through diverse diffusion augmentation (DDA). Specifically, we revise the paradigm of common data synthesis in DFKD to a composite process through leveraging diffusion models subsequent to data synthesis for self-supervised augmentation, which generates a spectrum of data samples with similar distributions while retaining controlled variations. Furthermore, to mitigate excessive deviation in the embedding space, we introduce an image filtering technique grounded in cosine similarity to maintain fidelity during the knowledge distillation process. Comprehensive experiments conducted on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets showcase the superior performance of our method across various teacher-student network configurations, outperforming the contemporary state-of-the-art DFKD methods. Code will be available at: https://github.com/SLGSP/DDA.
Muquan Li, Dongyang Zhang 0001, Tao He 0007, Xiurui Xie, Yuan-Fang Li, Ke Qin
ACM Multimedia1