VLDB 2026 Research / reviewers in the wild / expert
Xuefeng Xiao 0001
dblp:245/9547-1 · also XueFeng Xiao 0001
· DBLP profile ↗
40ranked-venue papers
2as first author
36since 2021 · last 2026
0009-0009-8258-1243ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 2 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 21 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts TrainingabstractExpert parallelism is vital for effectively training Mixture-of-Experts (MoE) models, enabling different devices to host distinct experts, with each device processing different input data. However, during expert parallel training, dynamic routing results in significant load imbalance among experts: a handful of overloaded experts hinder overall iteration, emerging as a training bottleneck. In this paper, we introduce LAER-MoE, an efficient MoE training framework. The core of LAER-MoE is a novel parallel paradigm, Fully Sharded Expert Parallel (FSEP), which fully partitions each expert parameter by the number of devices and restores partial experts at expert granularity through All-to-All communication during training. This allows for flexible re-layout of expert parameters during training to enhance load balancing. In particular, we perform fine-grained scheduling of communication operations to minimize communication overhead. Additionally, we develop a load balancing planner to formulate re-layout strategies of experts and routing schemes for tokens during training. We perform experiments on an A100 cluster, and the results indicate that our system achieves up to 1.69x acceleration compared to the current state-of-the-art training systems. Source code available at https://github.com/PKU-DAIR/Hetu-Galvatron/tree/laer-moe. Fangcheng Fu, Xuefeng Xiao 0001, Huixia Li, Jiashi Li, Bin Cui 0001 |
ASPLOS (2) | 4 |
| 2026 | Align Video Diffusion Model with Online Video-Centric Preference OptimizationabstractVideo diffusion models (VDMs) have demonstrated remarkable capabilities in text-to-video (T2V) generation. Despite their success, VDMs still suffer from degraded image quality and flickering artifacts. To address these issues, some approaches have introduced preference learning to exploit human feedback to enhance the video generation. However, these methods primarily adopt the routine in the image domain without an in-depth investigation into video-specific preference optimization. In this paper, we reexamine the design of the video preference learning from two key aspects: feedback source and feedback tuning methodology, and present OnlineVPO, a more efficient preference learning framework tailored specifically for VDMs. On the feedback source, we found that the image-level reward model commonly used in existing methods fails to provide a human-aligned video preference signal due to the modality gap. In contrast, video quality assessment (VQA) models show superior alignment with human perception of video quality. Building on this insight, we propose leveraging VQA models as a proxy of humans to provide more modality-aligned feedback for VDMs. Regarding the preference tuning methodology, we introduce an online DPO algorithm tailored for VDMs. It not only enjoys the benefits of superior scalability in optimizing videos with higher resolution and longer duration compared with the existing method, but also mitigates the insufficient optimization issue caused by off-policy learning via online preference generation and curriculum preference update designs. Extensive experiments on the open-source video-diffusion model demonstrate OnlineVPO as a simple yet effective and, more importantly, scalable preference learning algorithm for video diffusion models. Jie Wu 0001, Yatai Ji, Xuefeng Xiao 0001, Kai Han 0003 |
WACV | 5 |
| 2025 | ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsabstractRecent advancement in text-to-image models and corresponding personalized technologies enables individuals to generate high-quality and imaginative images. However, they often suffer from limitations when generating images with resolutions outside of their trained domain. To overcome this limitation, we present the resolution adapter \textbf{(ResAdapter)}, a domain-consistent adapter designed for diffusion models to generate images with unrestricted resolutions and aspect ratios. Unlike other multi-resolution generation methods that process images of static resolution with complex post-process operations, ResAdapter directly generates images with the dynamical resolution. Especially, after learning a deep understanding of pure resolution priors, ResAdapter trained on the general dataset, generates resolution-free images with personalized diffusion models while preserving their original style domain. Comprehensive experiments demonstrate that ResAdapter with only 0.5M can process images with flexible resolutions for arbitrary diffusion models. More extended experiments demonstrate that ResAdapter is compatible with other modules for image generation across a broad range of resolutions, and can be integrated into other multi-resolution model for efficiently generating higher-resolution images. Jiaxiang Cheng, Pan Xie, Xin Xia 0005, Jiashi Li, Yuxi Ren, Huixia Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu |
AAAI | 8 |
| 2025 | FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismabstractExtending the context length (i.e., the maximum supported sequence length) of LLMs is of paramount significance. To facilitate long context training of LLMs, sequence parallelism has emerged as an essential technique, which scatters each input sequence across multiple devices and necessitates communication to process the sequence. In essence, existing sequence parallelism methods assume homogeneous sequence lengths (i.e., all input sequences are equal in length) and therefore leverages a single, static scattering strategy for all input sequences. However, in reality, the sequence lengths in LLM training corpora exhibit substantial variability, often following a long-tail distribution, which leads to workload heterogeneity. Shiju Wang, Shenhan Zhu, Fangcheng Fu, Xuefeng Xiao 0001, Huixia Li, Jiashi Li, Faming Wu, Bin Cui 0001 |
ASPLOS (2) | 6 |
| 2025 | RayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow TrajectoriesabstractDiffusion models have achieved remarkable success across various domains. However, their slow generation speed remains a critical challenge. Existing acceleration methods, while aiming to reduce steps, often compromise sample quality, controllability, or introduce training complexities. Therefore, we propose RayFlow, a novel diffusion framework that addresses these limitations. Unlike previous methods, RayFlow guides each sample along a unique path towards an instance-specific target distribution. This method minimizes sampling steps while preserving generation diversity and stability. Furthermore, we introduce Time Sampler, an importance sampling technique to enhance training efficiency by focusing on crucial timesteps. Extensive experiments demonstrate RayFlow’s superiority in generating high-quality images with improved speed, control, and training efficiency compared to existing acceleration techniques. Huiyang Shao, Xin Xia 0005, Yuhong Yang 0010, Yuxi Ren, Xuefeng Xiao 0001 |
CVPR | 6 |
| 2025 | Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMabstractText-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation. Yatai Ji, Jie Wu 0001, Shoufa Chen, Chongjian Ge, Peize Sun, Wenqi Shao, Xuefeng Xiao 0001, Ping Luo 0002 |
ICCV | 10 |
| 2025 | Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video SynthesisabstractDistribution Matching Distillation (DMD) is a promising score distillation technique that compresses pre-trained teacher diffusion models into efficient one-step or multi-step student generators. Nevertheless, its reliance on the reverse Kullback-Leibler (KL) divergence minimization potentially induces mode collapse (or mode-seeking) in certain applications. To circumvent this inherent drawback, we propose Adversarial Distribution Matching (ADM), a novel framework that leverages diffusion-based discriminators to align the latent predictions between real and fake score estimators for score distillation in an adversarial manner. In the context of extremely challenging one-step distillation, we further improve the pre-trained generator by adversarial distillation with hybrid discriminators in both latent and pixel spaces. Different from the mean squared error used in DMD2 pre-training, our method incorporates the distributional loss on ODE pairs collected from the teacher model, and thus providing a better initialization for score distillation fine-tuning in the next stage. By combining the adversarial distillation pre-training with ADM fine-tuning into a unified pipeline termed DMDX, our proposed method achieves superior one-step performance on SDXL compared to DMD2 while consuming less GPU time. Additional experiments that apply multi-step ADM distillation on SD3-Medium, SD3.5-Large, and CogVideoX set a new benchmark towards efficient image and video synthesis. Yanzuo Lu, Yuxi Ren, Xin Xia 0005, Shanchuan Lin, Xuefeng Xiao 0001, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai |
ICCV | 6 |
| 2025 | Training-Free and Adaptive Sparse Attention for Efficient Long Video Generation
Suhan Ling, Fangcheng Fu, Huixia Li, Xuefeng Xiao 0001, Bin Cui 0001 |
ICCV | 6 |
| 2025 | Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationabstractDiffusion Transformer (DiT) has demonstrated remarkable performance in text-to-image generation; however, its large parameter size results in substantial inference overhead. Existing parameter compression methods primarily focus on pruning, but aggressive pruning often leads to severe performance degradation due to reduced model capacity. To address this limitation, we pioneer the transformation of a dense DiT into a Mixture of Experts (MoE) for structured sparsification, reducing the number of activated parameters while preserving model capacity. Specifically, we replace the Feed-Forward Networks (FFNs) in DiT Blocks with MoE layers, reducing the number of activated parameters in the FFNs by 62.5\%. Furthermore, we propose the Mixture of Blocks (MoB) to selectively activate DiT blocks, thereby further enhancing sparsity. To ensure an effective dense-to-MoE conversion, we design a multi-step distillation pipeline, incorporating Taylor metric-based expert initialization, knowledge distillation with load balancing, and group feature loss for MoB optimization. We transform large diffusion transformers (e.g., FLUX.1 [dev]) into an MoE structure, reducing activated parameters by 60\% while maintaining original performance and surpassing pruning-based approaches in extensive experiments. Overall, Dense2MoE establishes a new paradigm for efficient text-to-image generation. Youwei Zheng, Yuxi Ren, Xin Xia 0005, Xuefeng Xiao 0001, Xiaohua Xie |
ICCV | 4 |
| 2025 | IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language ModelabstractThe rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of a single scenario, while their ability to associate instances across different scenes has not yet been explored, which is essential for understanding complex visual content, such as movies with multiple characters and intricate plots. Towards movie understanding, a critical initial step for LVLMs is to unleash the potential of character identities memory and recognition across multiple visual scenarios. To achieve the goal, we propose visual instruction tuning with ID reference and develop an ID-Aware Large Vision-Language Model, IDA-VLM. Furthermore, our research introduces a novel benchmark MM-ID, to examine LVLMs on instance IDs memory and recognition across four dimensions: matching, location, question-answering, and captioning. Our findings highlight the limitations of existing LVLMs in recognizing and associating instance identities with ID reference. This paper paves the way for future artificial intelligence systems to possess multi-identity visual inputs, thereby facilitating the comprehension of complex visual narratives like movies. Yatai Ji, Jie Wu 0001, Peize Sun, Xuefeng Xiao 0001, Sidi Yang, Yujiu Yang 0001, Ping Luo 0002 |
ICLR | 6 |
| 2025 | Diffusion Adversarial Post-Training for One-Step Video GenerationabstractThe diffusion models are widely used for image and video generation, but their iterative generation process is slow and expansive. While existing distillation approaches have demonstrated the potential for one-step generation in the image domain, they still suffer from significant quality degradation. In this work, we propose Adversarial Post-Training (APT) against real data following diffusion pre-training for one-step video generation. To improve the training stability and quality, we introduce several improvements to the model architecture and training procedures, along with an approximated R1 regularization objective. Empirically, our experiments show that our adversarial post-trained model can generate two-second, 1280x720, 24fps videos in real-time using a single forward evaluation step. Additionally, our model is capable of generating 1024px images in a single step, achieving quality comparable to state-of-the-art methods. Shanchuan Lin, Xin Xia 0005, Yuxi Ren, Ceyuan Yang, Xuefeng Xiao 0001 |
ICML | 5 |
| 2025 | polybasic Speculative Decoding Through a Theoretical PerspectiveabstractInference latency stands as a critical bottleneck in the large-scale deployment of Large Language Models (LLMs). Speculative decoding methods have recently shown promise in accelerating inference without compromising the output distribution. However, existing work typically relies on a dualistic draft-verify framework and lacks rigorous theoretical grounding. In this paper, we introduce a novel \emph{polybasic} speculative decoding framework, underpinned by a comprehensive theoretical analysis. Specifically, we prove a fundamental theorem that characterizes the optimal inference time for multi-model speculative decoding systems, shedding light on how to extend beyond the dualistic approach to a more general polybasic paradigm. Through our theoretical investigation of multi-model token generation, we expose and optimize the interplay between model capabilities, acceptance lengths, and overall computational cost. Our framework supports both standalone implementation and integration with existing speculative techniques, leading to accelerated performance in practice. Experimental results across multiple model families demonstrate that our approach yields speedup ratios ranging from $3.31\times$ to $4.01\times$ for LLaMA2-Chat 7B, up to $3.87 \times$ for LLaMA3-8B, up to $4.43 \times$ for Vicuna-7B and up to $3.85 \times$ for Qwen2-7B---all while preserving the original output distribution. We release our theoretical proofs and implementation code to facilitate further investigation into polybasic speculative decoding. Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao 0001, Xuefeng Xiao 0001, Rongrong Ji |
ICML | 6 |
| 2025 | Breaking Static Barriers: Dynamic Post-Training Quantization for Diffusion ModelsabstractCurrent Post-Training Quantization (PTQ) schemes have been extensively studied for traditional convolutional neural networks and language models; however, PTQ application in diffusion models has shown significant performance degradation due to static settings of PTQ. Existing methods only uniformly and statically sample during each denoising step to construct calibration sets, neglecting the different importance of different steps in diffusion models. Furthermore, diffusion models exhibit a large number of activations with skewed distributions, and maintaining a static zero-point during the reconstruction process causes the model to converge only to local optima. To solve these limitations, it is necessary to dynamically design calibration dataset construction methods for different quantization scenarios and develop specialized optimization strategies tailored to specific activation distributions. Thus we proposed a unified framework, termed Dynamic PTQ, to achieve the aforementioned purposes. The framework first applies an evolutionary search algorithm to dynamically construct calibration sets for different quantization scenarios. Then, we design a dynamic zero-point update strategy for the quantizer, significantly reducing the loss during the reconstruction process. Extensive experiments demonstrate that our method outperforms current PTQ methods for diffusion models in generating high-quality samples. In particular, for the LSUN-bedrooms 256×256 task, our method quantizes the corresponding full-precision LDM-4 to W4A6 with only a 0.84 increase in FID. Huixia Li, Lijiang Li, Xiawu Zheng, Yuexiao Ma, Jie Wu 0001, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001 |
IJCNN | 8 |
| 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationabstractExisting large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to turn a pre-trained latent video diffusion model into
a real-time, interactive, streaming video generator. Our model autoregressively generates a latent frame at a time using a single neural function evaluation (1NFE). The model can stream the result to the user in real time and receive interactive responses as control to generate the next latent frame. Unlike existing approaches, our method explores adversarial training as an effective paradigm for autoregressive generation. This allows us to design a more efficient architecture for one-step generation and to train the model in a student-forcing way to mitigate error accumulation. The adversarial approach also enables us to train the model for long-duration generation fully utilizing the KV cache. As a result, our 8B model achieves real-time, 24fps, nonstop, streaming video generation at 736x416 resolution on a single H100, or 1280x720 on 8xH100 up to a minute long (1440 frames). Shanchuan Lin, Ceyuan Yang, Jianwen Jiang, Yuxi Ren, Xin Xia 0005, Yang Zhao 0003, Xuefeng Xiao 0001 |
NeurIPS | 8 |
| 2025 | LABridge: Text-Image Latent Alignment Framework via Mean-Conditioned OU ProcessabstractDiffusion models have emerged as state‑of‑the‑art in image synthesis.However, it often suffer from semantic instability and slow iterative denoising. We introduce Latent Alignment Framework (LABridge), a novel Text–Image Latent Alignment Framework via an Ornstein–Uhlenbeck (OU) Process, which explicitly preserves and aligns textual and visual semantics in an aligned latent space. LABridge employs a Text-Image Alignment Encoder (TIAE) to encode text prompts into structured priors that are directly aligned with image latents. Instead of a homogeneous Gaussian, Mean-Conditioned OU process smoothly interpolates between these text‑conditioned priors and image latents, improving stability and reducing the number of denoising steps. Extensive experiments on standard text-to-image benchmarks show that LABridge achieves better text–image alignment metric and competitive FID scores compared to leading diffusion baselines. By unifying text and image representations through principled latent alignment, LABridge paves the way for more efficient, semantically consistent, and high‑fidelity text to image generation. Huiyang Shao, Xin Xia 0005, Yuxi Ren, Xuefeng Xiao 0001 |
NeurIPS | 5 |
| 2025 | VarFlow: Proper Scoring-Rule Diffusion Distillation via Energy Matchingabstract**Diffusion models** achieve remarkable generative performance but are hampered by slow, iterative inference. Model distillation seeks to train a fast student generator. **Variational Score Distillation (VSD)** offers a principled KL-divergence minimization framework for this task. This method cleverly avoids computing the teacher model's Jacobian, but its student gradient relies on the score of the student's own noisy marginal distribution, $\nabla\_{\mathbf{x}\_t} \log p\_{\phi,t}(\mathbf{x}\_t)$. VSD thus requires approximations, such as training an auxiliary network to estimate this score. These approximations can introduce biases, cause training instability, or lead to an incomplete match of the target distribution, potentially focusing on conditional means rather than broader distributional features.
We introduce **VarFlow**, a method based on a **Score-Rule Variational Distillation (SRVD)** framework. VarFlow trains a one-step generator $g_{\phi}(\mathbf{z})$ by directly minimizing an energy distance (derived from the strictly proper energy score) between the student's induced noisy data distribution $p_{\phi,t}(\mathbf{x}_t)$ and the teacher's target noisy distribution $q_t(\mathbf{x}_t)$. This objective is estimated entirely using samples from these two distributions. Crucially, VarFlow bypasses the need to compute or approximate the intractable student score. By directly matching the full noisy marginal distributions, VarFlow aims for a more comprehensive and robust alignment between student and teacher, offering an efficient and theoretically grounded path to high-fidelity one-step generation. Huiyang Shao, Xin Xia 0005, Yuxi Ren, Xuefeng Xiao 0001 |
NeurIPS | 5 |
| 2025 | PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsabstractIn visual generation, the quadratic complexity of attention mechanisms results in high memory and computational costs, especially for longer token sequences required in high-resolution image or multi-frame video generation. To address this, prior research has explored techniques such as sparsification and quantization.
However, these techniques face significant challenges under low density and reduced bitwidths. Through systematic analysis, we identify that the core difficulty stems from the dispersed and irregular characteristics of visual attention patterns. Therefore, instead of introducing specialized sparsification and quantization design to accommodate such patterns, we propose an alternative strategy: "reorganizing" the attention pattern to alleviate the challenges.
Inspired by the local aggregatin nature of visual feature extraction, we design a novel **P**attern-**A**ware token **R**e**O**rdering (**PARO**) technique, which unifies the diverse attention patterns into a hardware-friendly block-wise pattern. This unification substantially simplifies and enhances both sparsification and quantization.
We evaluate the performance-efficiency trade-offs of various design choices and finalize a methodology tailored for the unified pattern.
Our approach, **PAROAttention**, achieves video and image generation with lossless metrics, and nearly identical results from full-precision (FP) baselines, while operating at notably lower density (**20%-30%**) and bitwidth (**INT8/INT4**), achieving a **1.9 - 2.7x** end-to-end latency speedup. Tianchen Zhao, Ke Hong, Xuefeng Xiao 0001, Huixia Li, Ruiqi Xie, Yichong Zhang, Yu Wang 0002 |
NeurIPS | 4 |
| 2025 | VmambaIR: Visual State Space Model for Image RestorationabstractImage restoration is a critical task in low-level computer vision, aiming to restore high-quality images from degraded inputs. Various models, such as convolutional neural networks (CNNs), generative adversarial networks (GANs), transformers, and diffusion models (DMs), have been employed to address this problem with significant impact. However, CNNs have limitations in capturing long-range dependencies. DMs require large prior models and computationally intensive denoising steps. Transformers have powerful modeling capabilities but face challenges due to quadratic complexity with input image size. To tackle these challenges, we propose VmambaIR, one of the first works to introduce State Space Models (SSMs) with linear complexity into comprehensive image restoration tasks. Specifically, we utilize a Unet architecture to stack our proposed Omni Selective Scan (OSS) blocks, consisting of an OSS module and an Efficient Feed-Forward Network (EFFN). Our proposed omni selective scan mechanism overcomes the unidirectional modeling limitation of SSMs by efficiently modeling image information flows in all six directions to better exploit surrounding restoration information. Furthermore, we conducted a comprehensive evaluation of our VmambaIR across multiple image restoration tasks, including image deraining, single image super-resolution, and real-world image super-resolution. Extensive experimental results demonstrate that our proposed VmambaIR achieves state-of-the-art (SOTA) performance with much fewer computational resources and parameters. Our research highlights the potential of state space models as promising alternatives to the transformer and CNN architectures in serving as foundational frameworks for next-generation low-level visual tasks. Bin Xia 0014, Xiaoyu Jin, Xin Xia 0005, Xuefeng Xiao 0001, Wenming Yang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
Ming Li 0010, Taojiannan Yang, Huafeng Kuang, Jie Wu 0032, Zhaoning Wang, Xuefeng Xiao 0001, Chen Chen 0001 |
ECCV (7) | 6 |
| 2024 | ByteEdit: Boost, Comply and Accelerate Generative Image Editing
Yuxi Ren, Jie Wu 0030, Yanzuo Lu, Huafeng Kuang, Xin Xia 0005, Xionghui Wang, Yixing Zhu, Pan Xie, Shiyin Wang, Xuefeng Xiao 0001, Lean Fu |
ECCV (3) | 11 |
| 2024 | AffineQuant: Affine Transformation Quantization for Large Language ModelsabstractThe significant resource requirements associated with Large-scale Language Models (LLMs) have generated considerable interest in the development of techniques aimed at compressing and accelerating neural networks.
Among these techniques, Post-Training Quantization (PTQ) has emerged as a subject of considerable interest due to its noteworthy compression efficiency and cost-effectiveness in the context of training.
Existing PTQ methods for LLMs limit the optimization scope to scaling transformations between pre- and post-quantization weights.
This constraint results in significant errors after quantization, particularly in low-bit configurations.
In this paper, we advocate for the direct optimization using equivalent Affine transformations in PTQ (AffineQuant).
This approach extends the optimization scope and thus significantly minimizing quantization errors.
Additionally, by employing the corresponding inverse matrix, we can ensure equivalence between the pre- and post-quantization outputs of PTQ, thereby maintaining its efficiency and generalization capabilities.
To ensure the invertibility of the transformation during optimization, we further introduce a gradual mask optimization method.
This method initially focuses on optimizing the diagonal elements and gradually extends to the other elements.
Such an approach aligns with the Levy-Desplanques theorem, theoretically ensuring invertibility of the transformation.
As a result, significant performance improvements are evident across different LLMs on diverse datasets.
Notably, these improvements are most pronounced when using very low-bit quantization, enabling the deployment of large models on edge devices.
To illustrate, we attain a C4 perplexity of $15.76$ (2.26$\downarrow$ vs $18.02$ in OmniQuant) on the LLaMA2-$7$B model of W$4$A$4$ quantization without overhead.
On zero-shot tasks, AffineQuant achieves an average of $58.61\%$ accuracy ( $1.98\%\uparrow$ vs $56.63$ in OmniQuant) when using $4$/$4$-bit quantization for LLaMA-$30$B, which setting a new state-of-the-art benchmark for PTQ in LLMs.
Codes are available at: https://github.com/bytedance/AffineQuant. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
ICLR | 5 |
| 2024 | Outlier-aware Slicing for Post-Training Quantization in Vision TransformerabstractPost-Training Quantization (PTQ) is a vital technique for network compression and acceleration, gaining prominence as model sizes increase. This paper addresses a critical challenge in PTQ: the severe impact of outliers on the accuracy of quantized transformer architectures. Specifically, we introduce the concept of ‘reconstruction granularity’ as a novel solution to this issue, which has been overlooked in previous works. Our work provides theoretical insights into the role of reconstruction granularity in mitigating the outlier problem in transformer models. This theoretical framework is supported by empirical analysis, demonstrating that varying reconstruction granularities significantly influence quantization performance. Our findings indicate that different architectural designs necessitate distinct optimal reconstruction granularities. For instance, the multi-stage Swin Transformer architecture benefits from finer granularity, a deviation from the trends observed in ViT and DeiT models. We further develop an algorithm for determining the optimal reconstruction granularity for various ViT models, achieving state-of-the-art (SOTA) performance in PTQ. For example, applying our method to $4$-bit quantization, the Swin-Base model achieves a Top-1 accuracy of $82.24%$ on the ImageNet classification task. This result surpasses the RepQ-ViT by $3.92%$ ($82.24%$ VS $78.32%$). Similarly, our approach elevates the ViT-Small to a Top-1 accuracy of $80.50%$, outperforming NoisyQuant by $3.64%$ ($80.50%$ VS $76.86%$). Codes are available in Supplementary Materials. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
ICML | 5 |
| 2024 | TreeReward: Improve Diffusion Model via Tree-Structured Feedback LearningabstractRecently, there has been significant progress in leveraging human feedback to enhance diffusion-based image generation, garnering considerable interest and attention. However, existing methods fail to achieve a fine-grained performance boost for the following challenges: i) insufficient amount of fine-grained feedback data; ii) lack of effective fine-grained feedback learning framework; To tackle these challenges, we present TreeReward to facilitate the fine-grained feedback optimization for diffusion models. Specifically, to address the limitation of the fine-grained feedback data, we first design a novel "AI + Expert" feedback data construction pipeline, yielding about 2.2M high-quality feedback dataset encompassing six fine-grained dimensions at a relatively low cost. Built upon this dataset, we introduce a tree-structure reward model to exploit the fine-grained feedback data efficiently and provide tailored optimization during feedback learning. We validate the feedback learning performance of our method across different fine-grained dimensions and various downstream tasks. Extensive experiments on both Stable Diffusion v1.5 (SD1.5) and Stable Diffusion XL (SDXL) demonstrate the effectiveness of our method in enhancing the general and fine-grained generation and downstream tasks generalization. Jie Wu 0030, Huafeng Kuang, Haiming Zhang 0001, Yuxi Ren, Manlin Zhang, Xuefeng Xiao 0001, Guanbin Li |
ACM Multimedia | 8 |
| 2024 | Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image SynthesisabstractRecently, a series of diffusion-aware distillation algorithms have emerged to alleviate the computational overhead associated with the multi-step inference process of Diffusion Models (DMs). Current distillation techniques often dichotomize into two distinct aspects: i) ODE Trajectory Preservation; and ii) ODE Trajectory Reformulation. However, these approaches suffer from severe performance degradation or domain shifts. To address these limitations, we propose Hyper-SD, a novel framework that synergistically amalgamates the advantages of ODE Trajectory Preservation and Reformulation, while maintaining near-lossless performance during step compression. Firstly, we introduce Trajectory Segmented Consistency Distillation to progressively perform consistent distillation within pre-defined time-step segments, which facilitates the preservation of the original ODE trajectory from a higher-order perspective. Secondly, we incorporate human feedback learning to boost the performance of the model in a low-step regime and mitigate the performance loss incurred by the distillation process. Thirdly, we integrate score distillation to further improve the low-step generation capability of the model and offer the first attempt to leverage a unified LoRA to support the inference process at all steps. Extensive experiments and user studies demonstrate that Hyper-SD achieves SOTA performance from 1 to 8 inference steps for both SDXL and SD1.5. For example, Hyper-SDXL surpasses SDXL-Lightning by +0.68 in CLIP Score and +0.51 in Aes Score in the 1-step inference. Yuxi Ren, Xin Xia 0005, Yanzuo Lu, Jie Wu 0032, Pan Xie, Xuefeng Xiao 0001 |
NeurIPS | 8 |
| 2024 | UniFL: Improve Latent Diffusion Model via Unified Feedback LearningabstractLatent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion models still suffer from several limitations, including inferior visual quality, inadequate aesthetic appeal, and inefficient inference, without a comprehensive solution in sight. To address these challenges, we present **UniFL**, a unified framework that leverages feedback learning to enhance diffusion models comprehensively. UniFL stands out as a universal, effective, and generalizable solution applicable to various diffusion models, such as SD1.5 and SDXL.
Notably, UniFL consists of three key components: perceptual feedback learning, which enhances visual quality; decoupled feedback learning, which improves aesthetic appeal; and adversarial feedback learning, which accelerates inference.
In-depth experiments and extensive user studies validate the superior performance of our method in enhancing generation quality and inference acceleration. For instance, UniFL surpasses ImageReward by 17\% user preference in terms of generation quality and outperforms LCM and SDXL Turbo by 57\% and 20\% general preference with 4-step inference. Jie Wu 0030, Yuxi Ren, Xin Xia 0005, Huafeng Kuang, Pan Xie, Jiashi Li, Xuefeng Xiao 0001, Shilei Wen, Lean Fu, Guanbin Li |
NeurIPS | 8 |
| 2023 | Solving Oscillation Problem in Post-Training Quantization Through a Theoretical PerspectiveabstractPost-training quantization (PTQ) is widely regarded as one of the most efficient compression methods practically, benefitting from its data privacy and low computation costs. We argue that an overlooked problem of oscillation is in the PTQ methods. In this paper, we take the initiative to explore and present a theoretical proof to explain why such a problem is essential in PTQ. And then, we try to solve this problem by introducing a principled and generalized frame-work theoretically. In particular, we first formulate the oscillation in PTQ and prove the problem is caused by the difference in module capacity. To this end, we define the module capacity (ModCap) under data-dependent and data-free scenarios, where the differentials between adjacent modules are used to measure the degree of oscillation. The problem is then solved by selecting top-k differentials, in which the corresponding modules are jointly optimized and quantized. Extensive experiments demonstrate that our method successfully reduces the performance drop and is generalized to different neural networks and PTQ methods. For example, with 2/4 bit ResNet-50 quantization, our method surpasses the previous state-of-the-art method by 1.9%. It becomes more significant on small model quantization, e.g. surpasses BRECQ method by 6.61% on MobileNetV2 × 0.5. Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji |
CVPR | 4 |
| 2023 | FreeSeg: Unified, Universal and Open-Vocabulary Image SegmentationabstractRecently, open-vocabulary learning has emerged to accomplish segmentation for arbitrary categories of text-based descriptions, which popularizes the segmentation system to more general-purpose application scenarios. However, existing methods devote to designing specialized architectures or parameters for specific segmentation tasks. These customized design paradigms lead to fragmentation between various segmentation tasks, thus hindering the uniformity of segmentation models. Hence in this paper, we propose FreeSeg, a generic framework to accomplish Unified, Universal and Open-Vocabulary Image Segmentation. FreeSeg optimizes an all-in-one network via one-shot training and employs the same architecture and parameters to handle diverse segmentation tasks seamlessly in the inference procedure. Additionally, adaptive prompt learning facilitates the unified model to capture task-aware and category-sensitive concepts, improving model robustness in multi-task and varied scenarios. Extensive experimental results demonstrate that FreeSeg establishes new state-of-the-art results in performance and generalization on three segmentation tasks, which outperforms the best task-specific architectures by a large margin: 5.5% mIoU on semantic segmentation, 17.6% mAP on instance segmentation, 20.1% PQ on panoptic segmentation for the unseen class on COCO. Project page: https://FreeSeg.github.io. Jie Qin 0004, Jie Wu 0032, Pengxiang Yan, Ming Li 0010, Yuxi Ren, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Xingang Wang 0003 |
CVPR | 6 |
| 2023 | AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model AccelerationabstractDiffusion models are emerging expressive generative models, in which a large number of time steps (inference steps) are required for a single image generation. To accelerate such tedious process, reducing steps uniformly is considered as an undisputed principle of diffusion models. We consider that such a uniform assumption is not the optimal solution in practice; i.e., we can find different optimal time steps for different models. Therefore, we propose to search the optimal time steps sequence and compressed model architecture in a unified framework to achieve effective image generation for diffusion models without any further training. Specifically, we first design a unified search space that consists of all possible time steps and various architectures. Then, a two stage evolutionary algorithm is introduced to find the optimal solution in the designed search space. To further accelerate the search process, we employ FID score between generated and real samples to estimate the performance of the sampled examples. As a result, the proposed method is (i).training-free, obtaining the optimal time steps and model architecture without any training process; (ii). orthogonal to most advanced diffusion samplers and can be integrated to gain better sample quality. (iii). generalized, where the searched time steps and architectures can be directly applied on different diffusion models with the same guidance scale. Experimental results show that our method achieves excellent performance by using only a few time steps, e.g. 17.86 FID score on ImageNet 64 × 64 with only four steps, compared to 138.66 with DDIM. Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu 0032, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001, Rongrong Ji |
ICCV | 5 |
| 2023 | AlignDet: Aligning Pre-training and Fine-tuning in Object DetectionabstractThe paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limit the detector’s performance, generalization ability, and convergence speed. To this end, we propose AlignDet, a unified pre-training framework that can be adapted to various existing detectors to alleviate the discrepancies. AlignDet decouples the pre-training process into two stages, i.e., image-domain and box-domain pre-training. The image-domain pre-training optimizes the detection backbone to capture holistic visual abstraction, and box-domain pre-training learns instance-level semantics and task-aware concepts to initialize the parts out of the backbone. By incorporating the self-supervised pretrained backbones, we can pre-train all modules for various detectors in an unsupervised paradigm. As depicted in Figure 1, extensive experiments demonstrate that AlignDet can achieve significant improvements across diverse protocols, such as ${\color{Green}\text{detection algorithms}}, {\color{Blue}\text{model backbones}}, {\color{Red}\text{data settings}}$, and ${\color{SkyBlue}\text{training schedules}}$. For example, AlignDet improves FCOS by 5.3 mAP, RetinaNet by 2.1 mAP, Faster R-CNN by 3.3 mAP, and DETR by 2.3 mAP under fewer epochs. Ming Li 0010, Jie Wu 0032, Xionghui Wang, Chen Chen 0001, Jie Qin 0004, Xuefeng Xiao 0001, Rui Wang 0089 |
ICCV | 6 |
| 2023 | UGC: Unified GAN Compression for Efficient Image-to-Image TranslationabstractRecent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on ponderous computational costs and labor-expensive training data. Current efficient GAN learning techniques often fall into two orthogonal aspects: i) model slimming via reduced calculation costs; ii) data/label-efficient learning with fewer training data/labels. To combine the best of both worlds, we propose a new learning paradigm, Unified GAN Compression (UGC), with a unified optimization objective to seamlessly prompt the synergy of model-efficient and label-efficient learning. UGC sets up semi-supervised-driven network architecture search and adaptive online semi-supervised distillation stages sequentially, which formulates a heterogeneous mutual learning scheme to obtain an architecture-flexible, label-efficient, and performance-excellent model. Extensive experiments demonstrate that UGC obtains state-of-the-art lightweight models even with less than 50% labels. UGC that compresses 40× MACs can achieve 21.43 FID on edges→shoes with 25% labels, which even outperforms the original model with 100% labels by 2.75 FID. Yuxi Ren, Jie Wu 0032, Manlin Zhang, Xuefeng Xiao 0001, Rui Wang 0089 |
ICCV | 5 |
| 2022 | Activation Modulation and Recalibration Scheme for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation (WSSS) is a fundamental yet challenging computer vision task facilitating scene understanding and automatic driving. Most existing methods resort to classification-based Class Activation Maps (CAMs) to play as the initial pseudo labels, which tend to focus on the discriminative image regions and lack customized characteristics for the segmentation task. To alleviate this issue, we propose a novel activation modulation and recalibration (AMR) scheme, which leverages a spotlight branch and a compensation branch to obtain weighted CAMs that can provide recalibration supervision and task-specific concepts. Specifically, an attention modulation module (AMM) is employed to rearrange the distribution of feature importance from the channel-spatial sequential perspective, which helps to explicitly model channel-wise interdependencies and spatial encodings to adaptively modulate segmentation-oriented activation responses. Furthermore, we introduce a cross pseudo supervision for dual branches, which can be regarded as a semantic similar regularization to mutually refine two branches. Extensive experiments show that AMR establishes a new state-of-the-art performance on the PASCAL VOC 2012 dataset, surpassing not only current methods trained with the image-level of supervision but also some methods relying on stronger supervision, such as saliency label. Experiments also reveal that our scheme is plug-and-play and can be incorporated with other approaches to boost their performance. Our code is available at: https://github.com/jieqin-ai/AMR. Jie Qin 0004, Jie Wu 0032, Xuefeng Xiao 0001, Lujun Li 0001, Xingang Wang 0003 |
AAAI | 3 |
| 2022 | Multi-granularity Distillation Scheme Towards Lightweight Semi-supervised Semantic Segmentation
Jie Qin 0004, Jie Wu 0032, Ming Li 0010, Xuefeng Xiao 0001, Xingang Wang 0003 |
ECCV (30) | 4 |
| 2022 | ScalableViT: Rethinking the Context-Oriented Generalization of Vision Transformer
Rui Yang 0010, Jie Wu 0032, Yansong Tang, Xuefeng Xiao 0001, Xiu Li 0001 |
ECCV (24) | 5 |
| 2022 | Progressive Automatic Design of Search Space for One-Shot Neural Architecture SearchabstractNeural Architecture Search (NAS) has attracted growing interest. To reduce the search cost, recent work has explored weight sharing across models and made major progress in One-Shot NAS. However, it has been observed that a model with higher one-shot model accuracy does not necessarily perform better when stand-alone trained. To address this issue, in this paper, we propose Progressive Automatic Design of search space, named PAD-NAS. Un-like previous approaches where the same operation search space is shared by all the layers in the supernet, we formulate a progressive search strategy based on operation pruning and build a layer-wise operation search space. In this way, PAD-NAS can automatically design the operations for each layer and achieve a trade-off between search space quality and model diversity. During the search, we also take the hardware platform constraints into consideration for efficient neural network model deployment. Extensive experiments on ImageNet show that our method can achieve state-of-the-art performance. Xin Xia 0005, Xuefeng Xiao 0001 |
WACV | 2 |
| 2021 | Online Multi-Granularity Distillation for GAN CompressionabstractGenerative Adversarial Networks (GANs) have witnessed prevailing success in yielding outstanding images, however, they are burdensome to deploy on resource-constrained devices due to ponderous computational costs and hulking memory usage. Although recent efforts on compressing GANs have acquired remarkable results, they still exist potential model redundancies and can be further compressed. To solve this issue, we propose a novel online multi-granularity distillation (OMGD) scheme to obtain lightweight GANs, which contributes to generating highfidelity images with low computational demands. We offer the first attempt to popularize single-stage online distillation for GAN-oriented compression, where the progressively promoted teacher generator helps to refine the discriminator-free based student generator. Complementary teacher generators and network layers provide comprehensive and multi-granularity concepts to enhance visual fidelity from diverse dimensions. Experimental results on four benchmark datasets demonstrate that OMGD successes to compress 40× MACs and 82.5× parameters on Pix2Pix and CycleGAN, without loss of image quality. It reveals that OMGD provides a feasible solution for the deployment of real-time image translation on resource-constrained devices. Our code and models are made public at: https://github.com/bytedance/OMGD Yuxi Ren, Jie Wu 0032, Xuefeng Xiao 0001, Jianchao Yang |
ICCV | 3 |
| 2021 | Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression SchemeabstractRecently, a series of algorithms have been explored for GAN compression, which aims to reduce tremendous computational overhead and memory usages when deploying GANs on resource-constrained edge devices. However, most of the existing GAN compression work only focuses on how to compress the generator, while fails to take the discriminator into account. In this work, we revisit the role of discriminator in GAN compression and design a novel generator-discriminator cooperative compression scheme for GAN compression, termed GCC. Within GCC, a selective activation discriminator automatically selects and activates convolutional channels according to a local capacity constraint and a global coordination constraint, which help maintain the Nash equilibrium with the lightweight generator during the adversarial training and avoid mode collapse. The original generator and discriminator are also optimized from scratch, to play as a teacher model to progressively refine the pruned generator and the selective activation discriminator. A novel online collaborative distillation scheme is designed to take full advantage of the intermediate feature of the teacher generator and discriminator to further boost the performance of the lightweight generator. Extensive experiments on various GAN-based generation tasks demonstrate the effectiveness and generalization of GCC. Among them, GCC contributes to reducing 80% computational costs while maintains comparable performance in image translation tasks. Jie Wu 0032, Xuefeng Xiao 0001, Fei Chao 0001, Xudong Mao, Rongrong Ji |
NeurIPS | 3 |
| 2018 | Accelerating and Compressing LSTM Based Model for Online Handwritten Chinese Character RecognitionabstractWith the development of deep learning tools, the online handwritten Chinese character recognition (HCCR) performance has been greatly improved by using deep neural networks (DNNs) especially for long short-term memory (LSTM). However, DNNs suffer from large consumption of computation and storage resources, which may cause problems for service providers, such as server pressure, longer service latency and higher energy consumption. To solve these problems, we propose a framework that combines singular value decomposition (SVD) and adaptive drop weight (ADW) to accelerate and compress LSTM based models. We first build an LSTM based model that achieves an accuracy of 97.83% on the ICDAR2013 online HCCR competition dataset. After restructuring the model with SVD and ADW, it can reduce the FLOPs (floating point operations per second) of the forward process by approximately 10 times and compress the model with 1/30 of the original size with only a 0.5% decrease in accuracy. Finally, integrated with our efficient forward implementation, the recognition of an online character requires only 2.7 ms in average on a CPU with a single thread, while requiring only 0.45 MB for model storage. Yafeng Yang, Kaihuan Liang, Xuefeng Xiao 0001, Zecheng Xie, Jun Sun 0004, Weiying Zhou |
ICFHR | 3 |
| 2017 | A Comprehensive Analysis of Misclassified Handwritten Chinese Character Samples by Incorporating Human RecognitionabstractThe development of convolutional neural networks (CNN) has led to revolutionary progress in the resolution of the offline handwritten Chinese character recognition (HCCR) problem. As the recognition rate on a standard offline HCCR testbed is outstanding, a few samples that remain misclassified have kindled our interest. In this paper, with the help of human recognition results, we present a comprehensive analysis of the samples misclassified by a state-of-the-art CNN model. We performed the analysis based on the top-1-votes, which are obtained from the statistical analysis of human recognition results, and derived the following conclusions: (1) the majority of samples with high top-1-votes were mis-labeled. Besides, by comparing the results of human recognition with that of CNN, some limitations of CNN that provide scope for further improvement are presented; (2) in the samples with medium top- 1-votes, it is shown that the samples with different confidence level have different characteristics. Specifically, some samples could be regarded as multi-label samples; (3) the samples with low top-1- votes are either wrongly written or written extensively in cursive style, which are difficult to match their given ground-truths; (4)the relationship between writing styles and misclassifications are also introduced in the paper. We believe this work should provide some insights and brings new clues on designing new classification methods to deal with these challenging samples. Kaihuan Liang, Zecheng Xie, Xuefeng Xiao 0001, Weiguo Huang |
ICDAR | 4 |
| 2017 | Design of a Very Compact CNN Classifier for Online Handwritten Chinese Character Recognition Using DropWeight and Global PoolingabstractCurrently, owing to the ubiquity of mobile devices, online handwritten Chinese character recognition (HCCR) has become one of the suitable choice for feeding input to cell phones and tablet devices. Over the past few years, larger and deeper convolutional neural networks (CNNs) have extensively been employed for improving character recognition performance. However, its substantial storage requirement is a significant obstacle in deploying such networks into portable electronic devices. To circumvent this problem, we use a novel technique called DropWeight for pruning redundant connections in the CNN architecture. It is revealed that the method not only treats streamlined architectures such as AlexNet and VGGNet well but exhibits remarkable performance for deep residual network and inception network. We also demonstrate that global pooling is a better choice for building very compact online HCCR systems. Experiments were performed on the ICDAR-2013 online HCCR competition dataset using our proposed network, and it is found that the proposed approach requires only 0.57 MB for storage, whereas state-of-the-art CNN-based methods require up to 135 MB; meanwhile the performance is decreased only by 0.91%. Xuefeng Xiao 0001, Yafeng Yang, Tasweer Ahmad, Tianhai Chang |
ICDAR | 1 |
| 2017 | Building fast and compact convolutional neural networks for offline handwritten Chinese character recognition
Xuefeng Xiao 0001, Yafeng Yang, Jun Sun 0004, Tianhai Chang |
Pattern Recognit. | 1 |