VLDB 2026 Research / reviewers in the wild / expert
Huixia Li
dblp:140/1219
· DBLP profile ↗
18ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts TrainingabstractExpert parallelism is vital for effectively training Mixture-of-Experts (MoE) models, enabling different devices to host distinct experts, with each device processing different input data. However, during expert parallel training, dynamic routing results in significant load imbalance among experts: a handful of overloaded experts hinder overall iteration, emerging as a training bottleneck. In this paper, we introduce LAER-MoE, an efficient MoE training framework. The core of LAER-MoE is a novel parallel paradigm, Fully Sharded Expert Parallel (FSEP), which fully partitions each expert parameter by the number of devices and restores partial experts at expert granularity through All-to-All communication during training. This allows for flexible re-layout of expert parameters during training to enhance load balancing. In particular, we perform fine-grained scheduling of communication operations to minimize communication overhead. Additionally, we develop a load balancing planner to formulate re-layout strategies of experts and routing schemes for tokens during training. We perform experiments on an A100 cluster, and the results indicate that our system achieves up to 1.69x acceleration compared to the current state-of-the-art training systems. Source code available at https://github.com/PKU-DAIR/Hetu-Galvatron/tree/laer-moe. Fangcheng Fu, Xuefeng Xiao 0001, Huixia Li, Jiashi Li, Bin Cui 0001 |
ASPLOS (2) | 5 |
| 2026 | Research on road damage detection method based on multi-modal information collaborative representation and recursive reflux mechanism
Huixia Li, Ruohan Chen, Nyirandayisabye Ritha |
Adv. Eng. Informatics | 1 |
| 2026 | An automatic detection method for road damage in UAV images based on multi-level perception and feature aggregation
Huixia Li, Ruohan Chen, Nyirandayisabye Ritha |
Adv. Eng. Informatics | 1 |
| 2025 | ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsabstractRecent advancement in text-to-image models and corresponding personalized technologies enables individuals to generate high-quality and imaginative images. However, they often suffer from limitations when generating images with resolutions outside of their trained domain. To overcome this limitation, we present the resolution adapter \textbf{(ResAdapter)}, a domain-consistent adapter designed for diffusion models to generate images with unrestricted resolutions and aspect ratios. Unlike other multi-resolution generation methods that process images of static resolution with complex post-process operations, ResAdapter directly generates images with the dynamical resolution. Especially, after learning a deep understanding of pure resolution priors, ResAdapter trained on the general dataset, generates resolution-free images with personalized diffusion models while preserving their original style domain. Comprehensive experiments demonstrate that ResAdapter with only 0.5M can process images with flexible resolutions for arbitrary diffusion models. More extended experiments demonstrate that ResAdapter is compatible with other modules for image generation across a broad range of resolutions, and can be integrated into other multi-resolution model for efficiently generating higher-resolution images. Jiaxiang Cheng, Pan Xie, Xin Xia 0005, Jiashi Li, Yuxi Ren, Huixia Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu |
AAAI | 7 |
| 2025 | FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismabstractExtending the context length (i.e., the maximum supported sequence length) of LLMs is of paramount significance. To facilitate long context training of LLMs, sequence parallelism has emerged as an essential technique, which scatters each input sequence across multiple devices and necessitates communication to process the sequence. In essence, existing sequence parallelism methods assume homogeneous sequence lengths (i.e., all input sequences are equal in length) and therefore leverages a single, static scattering strategy for all input sequences. However, in reality, the sequence lengths in LLM training corpora exhibit substantial variability, often following a long-tail distribution, which leads to workload heterogeneity. Shiju Wang, Shenhan Zhu, Fangcheng Fu, Xuefeng Xiao 0001, Huixia Li, Jiashi Li, Faming Wu, Bin Cui 0001 |
ASPLOS (2) | 7 |
| 2025 | Training-Free and Adaptive Sparse Attention for Efficient Long Video Generation
Suhan Ling, Fangcheng Fu, Huixia Li, Xuefeng Xiao 0001, Bin Cui 0001 |
ICCV | 5 |
| 2025 | polybasic Speculative Decoding Through a Theoretical PerspectiveabstractInference latency stands as a critical bottleneck in the large-scale deployment of Large Language Models (LLMs). Speculative decoding methods have recently shown promise in accelerating inference without compromising the output distribution. However, existing work typically relies on a dualistic draft-verify framework and lacks rigorous theoretical grounding. In this paper, we introduce a novel \emph{polybasic} speculative decoding framework, underpinned by a comprehensive theoretical analysis. Specifically, we prove a fundamental theorem that characterizes the optimal inference time for multi-model speculative decoding systems, shedding light on how to extend beyond the dualistic approach to a more general polybasic paradigm. Through our theoretical investigation of multi-model token generation, we expose and optimize the interplay between model capabilities, acceptance lengths, and overall computational cost. Our framework supports both standalone implementation and integration with existing speculative techniques, leading to accelerated performance in practice. Experimental results across multiple model families demonstrate that our approach yields speedup ratios ranging from $3.31\times$ to $4.01\times$ for LLaMA2-Chat 7B, up to $3.87 \times$ for LLaMA3-8B, up to $4.43 \times$ for Vicuna-7B and up to $3.85 \times$ for Qwen2-7B---all while preserving the original output distribution. We release our theoretical proofs and implementation code to facilitate further investigation into polybasic speculative decoding. Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao 0001, Xuefeng Xiao 0001, Rongrong Ji |
ICML | 2 |
| 2025 | Breaking Static Barriers: Dynamic Post-Training Quantization for Diffusion ModelsabstractCurrent Post-Training Quantization (PTQ) schemes have been extensively studied for traditional convolutional neural networks and language models; however, PTQ application in diffusion models has shown significant performance degradation due to static settings of PTQ. Existing methods only uniformly and statically sample during each denoising step to construct calibration sets, neglecting the different importance of different steps in diffusion models. Furthermore, diffusion models exhibit a large number of activations with skewed distributions, and maintaining a static zero-point during the reconstruction process causes the model to converge only to local optima. To solve these limitations, it is necessary to dynamically design calibration dataset construction methods for different quantization scenarios and develop specialized optimization strategies tailored to specific activation distributions. Thus we proposed a unified framework, termed Dynamic PTQ, to achieve the aforementioned purposes. The framework first applies an evolutionary search algorithm to dynamically construct calibration sets for different quantization scenarios. Then, we design a dynamic zero-point update strategy for the quantizer, significantly reducing the loss during the reconstruction process. Extensive experiments demonstrate that our method outperforms current PTQ methods for diffusion models in generating high-quality samples. In particular, for the LSUN-bedrooms 256×256 task, our method quantizes the corresponding full-precision LDM-4 to W4A6 with only a 0.84 increase in FID. Huixia Li, Lijiang Li, Xiawu Zheng, Yuexiao Ma, Jie Wu 0001, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001 |
IJCNN | 2 |
| 2025 | PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsabstractIn visual generation, the quadratic complexity of attention mechanisms results in high memory and computational costs, especially for longer token sequences required in high-resolution image or multi-frame video generation. To address this, prior research has explored techniques such as sparsification and quantization.
However, these techniques face significant challenges under low density and reduced bitwidths. Through systematic analysis, we identify that the core difficulty stems from the dispersed and irregular characteristics of visual attention patterns. Therefore, instead of introducing specialized sparsification and quantization design to accommodate such patterns, we propose an alternative strategy: "reorganizing" the attention pattern to alleviate the challenges.
Inspired by the local aggregatin nature of visual feature extraction, we design a novel **P**attern-**A**ware token **R**e**O**rdering (**PARO**) technique, which unifies the diverse attention patterns into a hardware-friendly block-wise pattern. This unification substantially simplifies and enhances both sparsification and quantization.
We evaluate the performance-efficiency trade-offs of various design choices and finalize a methodology tailored for the unified pattern.
Our approach, **PAROAttention**, achieves video and image generation with lossless metrics, and nearly identical results from full-precision (FP) baselines, while operating at notably lower density (**20%-30%**) and bitwidth (**INT8/INT4**), achieving a **1.9 - 2.7x** end-to-end latency speedup. Tianchen Zhao, Ke Hong, Xuefeng Xiao 0001, Huixia Li, Ruiqi Xie, Yichong Zhang, Yu Wang 0002 |
NeurIPS | 5 |
| 2024 | AffineQuant: Affine Transformation Quantization for Large Language ModelsabstractThe significant resource requirements associated with Large-scale Language Models (LLMs) have generated considerable interest in the development of techniques aimed at compressing and accelerating neural networks.
Among these techniques, Post-Training Quantization (PTQ) has emerged as a subject of considerable interest due to its noteworthy compression efficiency and cost-effectiveness in the context of training.
Existing PTQ methods for LLMs limit the optimization scope to scaling transformations between pre- and post-quantization weights.
This constraint results in significant errors after quantization, particularly in low-bit configurations.
In this paper, we advocate for the direct optimization using equivalent Affine transformations in PTQ (AffineQuant).
This approach extends the optimization scope and thus significantly minimizing quantization errors.
Additionally, by employing the corresponding inverse matrix, we can ensure equivalence between the pre- and post-quantization outputs of PTQ, thereby maintaining its efficiency and generalization capabilities.
To ensure the invertibility of the transformation during optimization, we further introduce a gradual mask optimization method.
This method initially focuses on optimizing the diagonal elements and gradually extends to the other elements.
Such an approach aligns with the Levy-Desplanques theorem, theoretically ensuring invertibility of the transformation.
As a result, significant performance improvements are evident across different LLMs on diverse datasets.
Notably, these improvements are most pronounced when using very low-bit quantization, enabling the deployment of large models on edge devices.
To illustrate, we attain a C4 perplexity of $15.76$ (2.26$\downarrow$ vs $18.02$ in OmniQuant) on the LLaMA2-$7$B model of W$4$A$4$ quantization without overhead.
On zero-shot tasks, AffineQuant achieves an average of $58.61\%$ accuracy ( $1.98\%\uparrow$ vs $56.63$ in OmniQuant) when using $4$/$4$-bit quantization for LLaMA-$30$B, which setting a new state-of-the-art benchmark for PTQ in LLMs.
Codes are available at: https://github.com/bytedance/AffineQuant. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
ICLR | 2 |
| 2024 | Outlier-aware Slicing for Post-Training Quantization in Vision TransformerabstractPost-Training Quantization (PTQ) is a vital technique for network compression and acceleration, gaining prominence as model sizes increase. This paper addresses a critical challenge in PTQ: the severe impact of outliers on the accuracy of quantized transformer architectures. Specifically, we introduce the concept of ‘reconstruction granularity’ as a novel solution to this issue, which has been overlooked in previous works. Our work provides theoretical insights into the role of reconstruction granularity in mitigating the outlier problem in transformer models. This theoretical framework is supported by empirical analysis, demonstrating that varying reconstruction granularities significantly influence quantization performance. Our findings indicate that different architectural designs necessitate distinct optimal reconstruction granularities. For instance, the multi-stage Swin Transformer architecture benefits from finer granularity, a deviation from the trends observed in ViT and DeiT models. We further develop an algorithm for determining the optimal reconstruction granularity for various ViT models, achieving state-of-the-art (SOTA) performance in PTQ. For example, applying our method to $4$-bit quantization, the Swin-Base model achieves a Top-1 accuracy of $82.24%$ on the ImageNet classification task. This result surpasses the RepQ-ViT by $3.92%$ ($82.24%$ VS $78.32%$). Similarly, our approach elevates the ViT-Small to a Top-1 accuracy of $80.50%$, outperforming NoisyQuant by $3.64%$ ($80.50%$ VS $76.86%$). Codes are available in Supplementary Materials. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
ICML | 2 |
| 2023 | OMPQ: Orthogonal Mixed Precision QuantizationabstractTo bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash the full potential of network quantization. However, existing approaches rely heavily on an extremely time-consuming search process and various relaxations when seeking the optimal bit configuration. To address this issue, we propose to optimize a proxy metric of network orthogonality that can be efficiently solved with linear programming, which proves to be highly correlated with quantized model accuracy and bit-width. Our approach significantly reduces the search time and the required data amount by orders of magnitude, but without a compromise on quantization accuracy. Specifically, we achieve 72.08% Top-1 accuracy on ResNet-18 with 6.7Mb parameters, which does not require any searching iterations. Given the high efficiency and low data dependency of our algorithm, we use it for the post-training quantization, which achieves 71.27% Top-1 accuracy on MobileNetV2 with only 1.5Mb parameters. Yuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang 0059, Huixia Li, Yongjian Wu 0001, Guannan Jiang, Wei Zhang 0217, Rongrong Ji |
AAAI | 5 |
| 2023 | Solving Oscillation Problem in Post-Training Quantization Through a Theoretical PerspectiveabstractPost-training quantization (PTQ) is widely regarded as one of the most efficient compression methods practically, benefitting from its data privacy and low computation costs. We argue that an overlooked problem of oscillation is in the PTQ methods. In this paper, we take the initiative to explore and present a theoretical proof to explain why such a problem is essential in PTQ. And then, we try to solve this problem by introducing a principled and generalized frame-work theoretically. In particular, we first formulate the oscillation in PTQ and prove the problem is caused by the difference in module capacity. To this end, we define the module capacity (ModCap) under data-dependent and data-free scenarios, where the differentials between adjacent modules are used to measure the degree of oscillation. The problem is then solved by selecting top-k differentials, in which the corresponding modules are jointly optimized and quantized. Extensive experiments demonstrate that our method successfully reduces the performance drop and is generalized to different neural networks and PTQ methods. For example, with 2/4 bit ResNet-50 quantization, our method surpasses the previous state-of-the-art method by 1.9%. It becomes more significant on small model quantization, e.g. surpasses BRECQ method by 6.61% on MobileNetV2 × 0.5. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
CVPR | 2 |
| 2023 | AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model AccelerationabstractDiffusion models are emerging expressive generative models, in which a large number of time steps (inference steps) are required for a single image generation. To accelerate such tedious process, reducing steps uniformly is considered as an undisputed principle of diffusion models. We consider that such a uniform assumption is not the optimal solution in practice; i.e., we can find different optimal time steps for different models. Therefore, we propose to search the optimal time steps sequence and compressed model architecture in a unified framework to achieve effective image generation for diffusion models without any further training. Specifically, we first design a unified search space that consists of all possible time steps and various architectures. Then, a two stage evolutionary algorithm is introduced to find the optimal solution in the designed search space. To further accelerate the search process, we employ FID score between generated and real samples to estimate the performance of the sampled examples. As a result, the proposed method is (i).training-free, obtaining the optimal time steps and model architecture without any training process; (ii). orthogonal to most advanced diffusion samplers and can be integrated to gain better sample quality. (iii). generalized, where the searched time steps and architectures can be directly applied on different diffusion models with the same guidance scale. Experimental results show that our method achieves excellent performance by using only a few time steps, e.g. 17.86 FID score on ImageNet 64 × 64 with only four steps, compared to 138.66 with DDIM. Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu 0032, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001, Rongrong Ji |
ICCV | 2 |
| 2021 | Evolving Fully Automated Machine Learning via Life-Long Knowledge AnchorsabstractAutomated machine learning (AutoML) has achieved remarkable progress on various tasks, which is attributed to its minimal involvement of manual feature and model designs. However, most of existing AutoML pipelines only touch parts of the full machine learning pipeline, e.g., neural architecture search or optimizer selection. This leaves potentially important components such as data cleaning and model ensemble out of the optimization, and still results in considerable human involvement and suboptimal performance. The main challenges lie in the huge search space assembling all possibilities over all components, as well as the generalization ability over different tasks like image, text, and tabular etc. In this paper, we present a first-of-its-kind fully AutoML pipeline, to comprehensively automate data preprocessing, feature engineering, model generation/selection/training and ensemble for an arbitrary dataset and evaluation metric. Our innovation lies in the comprehensive scope of a learning pipeline, with a novel "life-long" knowledge anchor design to fundamentally accelerate the search over the full search space. Such knowledge anchors record detailed information of pipelines and integrates them with an evolutionary algorithm for joint optimization across components. Experiments demonstrate that the result pipeline achieves state-of-the-art performance on multiple datasets and modalities. Specifically, the proposed framework was extensively evaluated in the NeurIPS 2019 AutoDL challenge, and won the only champion with a significant gap against other approaches, on all the image, video, speech, text and tabular tracks. Xiawu Zheng, Yang Zhang 0079, Sirui Hong, Huixia Li, Lang Tang, Youcheng Xiong, Yan Wang 0059, Xiaoshuai Sun, Pengfei Zhu 0001, Chenglin Wu 0001, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | PAMS: Quantized Super-Resolution via Parameterized Max Scale
Huixia Li, Chenqian Yan, Shaohui Lin, Xiawu Zheng, Baochang Zhang 0001, Fan Yang 0016, Rongrong Ji |
ECCV (25) | 1 |
| 2017 | A Novel Ionospheric Sounding Radar Based on USRPabstractIonospheric sounding is a technique that provides real-time data on high-frequency ionospheric-dependent radio propagation. This letter presents a Universal Software Radio Peripheral-based ionospheric sounding radar, which relies on a basic system consisting of a synchronized transmitter and receiver. The radar has the advantages of miniaturization, modularization, low power, and low cost. The three most significant features of the radar system are that it is software-defined and universal platform-based and that it has low transmitting power. This novel software-defined vertical-incidence radar system can probe the ionosphere and obtain real-time plasma parameters according to the simulation. Ionograms that directly express probe results are generated by MATLAB after data processing and simulation. Successful development of such an ionospheric sounding software radar will allow universalization and miniaturization of an ionosonde radar system. This letter introduces the implementation of the novel ionospheric sounding radar. Ziyang Zhao, Ming Yao 0001, Xiaohua Deng, Huixia Li |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | System Design of the Prototype Incoherent Scatter Radar at Nanchang UniversityabstractIncoherent scatter radar (ISR) is the only ground-based instrument that can detect space plasma parameters from tens to thousands of kilometers above the earth. The Prototype Incoherent Scatter Radar System of Nanchang University, a versatile space detection radar, will be built at Nanchang University to detect the ionosphere over China in the end of 2013. Software-defined radio (SDR)-based and active phased antenna array are the two most significant features of this radar system. It can probe real time plasma parameters in the altitude range of the ionosphere (80-400 km) according to simulation. The novel ISR introduced in this letter has the advantages of low power, low cost, and modularization. When properly configurated, this system can also carry out coherent scatter detection and wind profile detection. Successful development of such an ISR will resolve the problem of continuous detection of ionospheric electric field and other parameters. This letter introduces the design of this software defined ISR prototype system. Ming Yao 0001, Liangmin Zhang, Xiaohua Deng, Bo Bai 0003, Huixia Li, Ziqing Lu |
IEEE Geosci. Remote. Sens. Lett. | 7 |