EDBT 2026 Demo / reviewers in the wild / expert
Yunshan Zhong
dblp:239/4066
· DBLP profile ↗
19ranked-venue papers
10as first author
17since 2021 · last 2027
0000-0003-0268-4672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Unified-width adaptive dynamic network for all-in-one image restoration
Yimin Xu, Chunmei Yuan, Yunshan Zhong, Fei Chao 0001 |
Inf. Sci. | 3 |
| 2026 | Dynamic trajectory diffusion model for all-in-one image restoration
Yimin Xu, Yunshan Zhong, Fei Chao 0001 |
Expert Syst. Appl. | 2 |
| 2026 | I&S-ViT: An Inclusive & Stable Method for Post-Training ViTs QuantizationabstractAlbeit the scalable performance of vision transformers (ViTs), the dense computational costs undermine their position in industrial applications. Post-training quantization (PTQ), tuning ViTs with a tiny dataset and running in a low-bit format, well addresses the cost issue but unluckily bears more performance drops in lower-bit cases. In this paper, we introduce I&S-ViT, a novel method that regulates the PTQ of ViTs in an inclusive and stable fashion. I&S-ViT first identifies two issues in the PTQ of ViTs: (1) Quantization inefficiency in the prevalent log2 quantizer for post-Softmax activations; (2) Rugged and magnified loss landscape in coarse-grained quantization granularity for post-LayerNorm activations. Then, I&S-ViT addresses these issues by introducing: (1) A novel shift-uniform-log2 quantizer (SULQ) that incorporates a shift mechanism followed by uniform quantization to achieve both an inclusive domain representation and accurate distribution approximation; (2) A three-stage smooth optimization strategy (SOS) that amalgamates the strengths of channel-wise and layer-wise quantization to enable stable learning. Comprehensive evaluations across diverse vision tasks validate I&S-ViT's superiority over existing PTQ of ViTs methods, particularly in low-bit scenarios. For instance, I&S-ViT elevates the performance of W3A3 ViT-B by an impressive 50.68%. Yunshan Zhong, Mingbao Lin, Mengzhao Chen, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything Model
Wenlun Zhang, Yunshan Zhong, Shimpei Ando, Kentaro Yoshioka |
ICCV | 2 |
| 2025 | Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers
Yunshan Zhong, Yuyao Zhou, Yuxin Zhang 0002, Wanchen Sui, Fei Chao 0001, Rongrong Ji |
ICCV | 1 |
| 2025 | Distribution-flexible subset quantization for post-quantizing super-resolution networks
Yunshan Zhong, Mingbao Lin, Jingjing Xie, Yuxin Zhang 0002, Fei Chao 0001, Rongrong Ji |
Sci. China Inf. Sci. | 1 |
| 2025 | Towards Accurate Post-Training Quantization of Vision Transformers via Error ReductionabstractPost-training quantization (PTQ) for vision transformers (ViTs) has received increasing attention from both academic and industrial communities due to its minimal data needs and high time efficiency. However, many current methods fail to account for the complex interactions between quantized weights and activations, resulting in significant quantization errors and suboptimal performance. This paper presents ERQ, an innovative two-step PTQ method specifically crafted to reduce quantization errors arising from activation and weight quantization sequentially. The first step, Activation quantization error reduction (Aqer), first applies Reparameterization Initialization aimed at mitigating initial quantization errors in high-variance activations. Then, it further mitigates the errors by formulating a Ridge Regression problem, which updates the weights maintained at full-precision using a closed-form solution. The second step, Weight quantization error reduction (Wqer), first applies Dual Uniform Quantization to handle weights with numerous outliers, which arise from adjustments made during Reparameterization Initialization, thereby reducing initial weight quantization errors. Then, it employs an iterative approach to further tackle the errors. In each iteration, it adopts Rounding Refinement that uses an empirically derived, efficient proxy to refine the rounding directions of quantized weights, complemented by a Ridge Regression solver to reduce the errors. Comprehensive experimental results demonstrate ERQ's superior performance across various ViTs variants and tasks. For example, ERQ surpasses the state-of-the-art GPTQ by a notable 36.81% in accuracy for W3A4 ViT-S. Yunshan Zhong, You Huang, Yuxin Zhang 0002, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | MBQuant: A novel multi-branch topology method for arbitrary bit-width network quantization
Yunshan Zhong, Yuyao Zhou, Fei Chao 0001, Rongrong Ji |
Pattern Recognit. | 1 |
| 2025 | ARF: Arbitrary Routing Framework for All-in-One Image RestorationabstractAll-in-one image restoration methods, as opposed to conventional image restoration methods, reconstruct images impaired by various degradations within a unified model, eliminating the need for separate network parameters for each task. However, current all-in-one image restoration approaches tackle various types of image degradation using an identical underlying model, neglecting the inherent variability in complexity across different image restoration tasks, resulting in inefficient allocation of computational resources. To address this limitation, this article introduces the arbitrary routing framework (ARF), designed to effectively assess the difficulty of image restoration tasks and identify the most suitable network structure based on these complexities. This framework can be integrated with existing all-in-one image restoration models, enabling efficient inference by activating various proportions of the entire network, that is subnetworks, based on their task-specific complexities. More specifically, the ARF comprises two principal components: 1) the arbitrary routing backbone (ARB) and 2) a task-specific neural architecture search (T-NAS). The ARB incorporates a routing layer between consecutive convolutional groups, offering a wide array of potential subnetwork configurations while adding only negligible extra parameters. Concurrently, T-NAS autonomously identifies the most effective subnetworks for each image restoration task, optimizing both performance and efficiency through an efficiency-aware reward function. Comprehensive experiments across various image restoration tasks demonstrate that the ARF significantly improves performance metrics, that is, an increase of 0.31 in reconstruction PSNR, while also achieving a notable reduction in computational demands by 37.1% compared with the benchmark AirNet method. The code has been made available in the supplementary materials. Yimin Xu, Nanxi Gao, Yunshan Zhong, Fei Chao 0001, Rongrong Ji |
IEEE Trans. Cybern. | 3 |
| 2024 | Learning Image Demoiréing from Unpaired Real DataabstractThis paper focuses on addressing the issue of image demoiréing. Unlike the large volume of existing studies that rely on learning from paired real data, we attempt to learn a demoiréing model from unpaired real data, i.e., moiré images associated with irrelevant clean images. The proposed method, referred to as Unpaired Demoiréing(UnDeM), synthesizes pseudo moiré images from unpaired datasets, generating pairs with clean images for training demoiréing models. To achieve this, we divide real moiré images into patches and group them in compliance with their moiré complexity. We introduce a novel moiré generation framework to synthesize moiré images with diverse moiré features, resembling real moiré patches, and details akin to real moiré-free images. Additionally, we introduce an adaptive denoise method to eliminate the low-quality pseudo moiré images that adversely impact the learning of demoiréing models. We conduct extensive experiments on the commonly-used FHDMi and UHDM datasets. Results manifest that our UnDeM performs better than existing methods when using existing demoiréing models such as MBCNN and ESDNet-L. Code: https://github.com/zysxmu/UnDeM. Yunshan Zhong, Yuyao Zhou, Yuxin Zhang 0002, Fei Chao 0001, Rongrong Ji |
AAAI | 1 |
| 2024 | CaM: Cache Merging for Memory-efficient LLMs InferenceabstractDespite the exceptional performance of Large Language Models (LLMs), the substantial volume of key-value (KV) pairs cached during inference presents a barrier to their efficient deployment. To ameliorate this, recent works have aimed to selectively eliminate these caches, informed by the attention scores of associated tokens. However, such cache eviction invariably leads to output perturbation, regardless of the token choice. This perturbation escalates with the compression ratio, which can precipitate a marked deterioration in LLM inference performance. This paper introduces Cache Merging (CaM) as a solution to mitigate this challenge. CaM adaptively merges to-be-evicted caches into the remaining ones, employing a novel sampling strategy governed by the prominence of attention scores within discarded locations. In this manner, CaM enables memory-efficient LLMs to preserve critical token information, even obviating the need to maintain their corresponding caches. Extensive experiments utilizing LLaMA, OPT, and GPT-NeoX across various benchmarks corroborate CaM’s proficiency in bolstering the performance of memory-efficient LLMs. Code is released at https://github.com/zyxxmu/cam. Yuxin Zhang 0002, Gen Luo, Yunshan Zhong, Zhenyu Zhang 0015, Shiwei Liu 0003, Rongrong Ji |
ICML | 4 |
| 2024 | ERQ: Error Reduction for Post-Training Quantization of Vision TransformersabstractPost-training quantization (PTQ) for vision transformers (ViTs) has garnered significant attention due to its efficiency in compressing models. However, existing methods typically overlook the intricate interdependence between quantized weight and activation, leading to considerable quantization error. In this paper, we propose ERQ, a two-step PTQ approach meticulously crafted to sequentially reduce the quantization error arising from activation and weight quantization. ERQ first introduces Activation quantization error reduction (Aqer) that strategically formulates the minimization of activation quantization error as a Ridge Regression problem, tackling it by updating weights with full-precision. Subsequently, ERQ introduces Weight quantization error reduction (Wqer) that adopts an iterative approach to mitigate the quantization error induced by weight quantization. In each iteration, an empirically derived, efficient proxy is employed to refine the rounding directions of quantized weights, coupled with a Ridge Regression solver to curtail weight quantization error. Experimental results attest to the effectiveness of our approach. Notably, ERQ surpasses the state-of-the-art GPTQ by 22.36% in accuracy for W3A4 ViT-S. Yunshan Zhong, You Huang, Yuxin Zhang 0002, Rongrong Ji |
ICML | 1 |
| 2023 | Bi-directional Masks for Efficient N: M Sparse TrainingabstractWe focus on addressing the dense backward propagation issue for training efficiency of N:M fine-grained sparsity that preserves at most N out of M consecutive weights and achieves practical speedups supported by the N:M sparse tensor core. Therefore, we present a novel method of Bi-directional Masks (Bi-Mask) with its two central innovations in: 1) Separate sparse masks in the two directions of forward and backward propagation to obtain training acceleration. It disentangles the forward and backward weight sparsity and overcomes the very dense gradient computation. 2) An efficient weight row permutation method to maintain performance. It picks up the permutation candidate with the most eligible N:M weight blocks in the backward to minimize the gradient gap between traditional unidirectional masks and our bi-directional masks. Compared with existing uni-directional scenario that applies a transposable mask and enables backward acceleration, our Bi-Mask is experimentally demonstrated to be more superior in performance. Also, our Bi-Mask performs on par with or even better than methods that fail to achieve backward acceleration. Project of this paper is available at https://github.com/zyxxmu/Bi-Mask. Yuxin Zhang 0002, Yiting Luo, Mingbao Lin, Yunshan Zhong, Jingjing Xie, Fei Chao 0001, Rongrong Ji |
ICML | 4 |
| 2023 | Lottery Jackpots Exist in Pre-Trained ModelsabstractNetwork pruning is an effective approach to reduce network complexity with acceptable performance compromise. Existing studies achieve the sparsity of neural networks via time-consuming weight training or complex searching on networks with expanded width, which greatly limits the applications of network pruning. In this paper, we show that high-performing and sparse sub-networks without the involvement of weight training, termed "lottery jackpots", exist in pre-trained models with unexpanded width. Our presented lottery jackpots are traceable through empirical and theoretical outcomes. For example, we obtain a lottery jackpot that has only 10% parameters and still reaches the performance of the original dense VGGNet-19 without any modifications on the pre-trained weights on CIFAR-10. Furthermore, we improve the efficiency for searching lottery jackpots from two perspectives. First, we observe that the sparse masks derived from many existing pruning criteria have a high overlap with the searched mask of our lottery jackpot, among which, the magnitude-based pruning results in the most similar mask with ours. In compliance with this insight, we initialize our sparse mask using the magnitude-based pruning, resulting in at least 3× cost reduction on the lottery jackpot searching while achieving comparable or even better performance. Second, we conduct an in-depth analysis of the searching process for lottery jackpots. Our theoretical result suggests that the decrease in training loss during weight searching can be disturbed by the dependency between weights in modern networks. To mitigate this, we propose a novel short restriction method to restrict change of masks that may have potential negative impacts on the training loss, which leads to a faster convergence and reduced oscillation for searching lottery jackpots. Consequently, our searched lottery jackpot removes 90% weights in ResNet-50, while it easily obtains more than 70% top-1 accuracy using only 5 searching epochs on ImageNet. Yuxin Zhang 0002, Mingbao Lin, Yunshan Zhong, Fei Chao 0001, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationabstractLearning to synthesize data has emerged as a promising direction in zero-shot quantization (ZSQ), which represents neural networks by low-bit integer without accessing any of the real data. In this paper, we observe an interesting phenomenon of intra-class heterogeneity in real data and show that existing methods fail to retain this property in their synthetic images, which causes a limited performance increase. To address this issue, we propose a novel zero-shot quantization method referred to as IntraQ. First, we propose a local object reinforcement that locates the target objects at different scales and positions of the synthetic images. Second, we introduce a marginal distance constraint to form class-related features distributed in a coarse area. Lastly, we devise a soft inception loss which injects a soft prior label to prevent the synthetic images from being over-fitting to a fixed object. Our IntraQ is demonstrated to well retain the intra-class heterogeneity in the synthetic images and also observed to perform state-of-the-art. For example, compared to the advanced ZSQ, our IntraQ obtains 9.17% increase of the top-1 accuracy on ImageNet when all layers of MobileNetV1 are quantized to 4-bit. Code is at https://github.com/zysxmu/IntraQ Yunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu, Baochang Zhang 0001, Yonghong Tian 0001, Rongrong Ji |
CVPR | 1 |
| 2022 | Fine-grained Data Distribution Alignment for Post-Training Quantization
Yunshan Zhong, Mingbao Lin, Mengzhao Chen, Ke Li 0015, Yunhang Shen, Fei Chao 0001, Yongjian Wu 0001, Rongrong Ji |
ECCV (11) | 1 |
| 2022 | Dynamic Dual Trainable Bounds for Ultra-low Precision Super-Resolution Networks
Yunshan Zhong, Mingbao Lin, Xunchao Li, Ke Li 0015, Yunhang Shen, Fei Chao 0001, Yongjian Wu 0001, Rongrong Ji |
ECCV (18) | 1 |
| 2019 | Re-Identification Supervised Texture GenerationabstractThe estimation of 3D human body pose and shape from a single image has been extensively studied in recent years. However, the texture generation problem has not been fully discussed. In this paper, we propose an end-to-end learning strategy to generate textures of human bodies under the supervision of person re-identification. We render the synthetic images with textures extracted from the inputs and maximize the similarity between the rendered and input images by using the re-identification network as the perceptual metrics. Experiment results on pedestrian images show that our model can generate the texture from a single image and demonstrate that our textures are of higher quality than those generated by other available methods. Furthermore, we extend the application scope to other categories and explore the possible utilization of our generated textures. Jian Wang 0042, Yunshan Zhong, Yachun Li, Chi Zhang 0026 |
CVPR | 2 |
| 2019 | Re-ID Driven Localization Refinement for Person SearchabstractPerson search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually, and the detected bounding boxes may be sub-optimal for the following re-ID task. To alleviate this issue, we propose a re-ID driven localization refinement framework for providing the refined detection boxes for person search. Specifically, we develop a differentiable ROI transform layer to effectively transform the bounding boxes from the original images. Thus, the box coordinates can be supervised by the re-ID training other than the original detection task. With this supervision, the detector can generate more reliable bounding boxes, and the downstream re-ID model can produce more discriminative embeddings based on the refined person localizations. Extensive experimental results on the widely used benchmarks demonstrate that our proposed method performs favorably against the state-of-the-art person search methods. Chuchu Han, Jiacheng Ye, Yunshan Zhong, Xin Tan 0002, Chi Zhang 0026, Changxin Gao, Nong Sang |
ICCV | 3 |