VLDB 2026 Research / reviewers in the wild / expert
Hualiang Wang
dblp:302/5416
· DBLP profile ↗
28ranked-venue papers
4as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive FocusabstractMulti-subject personalized image generation aims to synthesize customized images containing multiple specified subjects without requiring test-time optimization. However, achieving fine-grained independent control over multiple subjects remains challenging due to difficulties in preserving subject fidelity and preventing cross-subject attribute leakage. We present FocusDPO, a framework that adaptively identifies focus regions based on dynamic semantic correspondence and supervision image complexity. During training, our method progressively adjusts these focal areas across noise timesteps, implementing a weighted strategy that rewards information-rich patches while penalizing regions with low prediction confidence. The framework dynamically adjusts focus allocation during the DPO process according to the semantic complexity of reference images and establishes robust correspondence mappings between generated and reference subjects. Extensive experiments demonstrate that our method substantially enhances the performance of existing pre-trained personalized generation models, achieving state-of-the-art results on both single-subject and multi-subject personalized image synthesis benchmarks. Our method effectively mitigates attribute leakage while preserving superior subject fidelity across diverse generation scenarios, advancing the frontier of controllable multi-subject image synthesis. Qiaoqiao Jin, Siming Fu, Dong She, Weinan Jia, Hualiang Wang, Mu Liu, Jidong Jiang |
AAAI | 5 |
| 2026 | DeepSparse: A Foundation Model for Sparse-View CBCT ReconstructionabstractCone-beam computed tomography (CBCT) is a critical 3D imaging technology in the medical field, while the high radiation exposure required for high-quality imaging raises significant concerns, particularly for vulnerable populations. Sparse-view reconstruction reduces radiation by using fewer X-ray projections while maintaining image quality, yet existing methods face challenges such as high computational demands and poor generalizability to different datasets. To overcome these limitations, we propose DeepSparse, the first foundation model for sparse-view CBCT reconstruction, featuring DiCE (Dual-Dimensional Cross-Scale Embedding), a novel network that integrates multi-view 2D features and multi-scale 3D features. Additionally, we introduce the HyViP (Hybrid View Sampling Pretraining) framework, which pretrains the model on large datasets with both sparse-view and dense-view projections, and a two-step finetuning strategy to adapt and refine the model for new datasets. Extensive experiments and ablation studies demonstrate that our proposed DeepSparse achieves superior reconstruction quality compared to state-of-the-art methods, paving the way for safer and more efficient CBCT imaging. The code will be publicly available at https://github.com/xmed-lab/DeepSparse. Yiqun Lin, Jixiang Chen 0001, Hualiang Wang, Jiewen Yang, Jiarong Guo, Yi Zhang 0018, Xiaomeng Li 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image TranslationabstractOptical coherence tomography angiography (OCTA) shows its great importance in imaging microvascular networks by providing accurate 3D imaging of blood vessels, but it relies upon specialized sensors and expensive devices. For this reason, previous works show the potential to translate the readily available 3D Optical Coherence Tomography (OCT) images into 3D OCTA images. However, existing OCTA translation methods directly learn the mapping from the OCT domain to the OCTA domain in continuous and infinite space with guidance from only a single view, i.e., the OCTA project map, resulting in suboptimal results. To this end, we propose the multi-view Tri-alignment framework for OCT to OCTA 3D image translation in discrete and finite space, named MuTri. In the first stage, we pre-train two vector-quantized variational auto-encoder (VQ-VAE) by reconstructing 3D OCT and 3D OCTA data, providing semantic prior for subsequent multi-view guidances. In the second stage, our multi-view tri-alignment facilitates another VQVAE model to learn the mapping from the OCT domain to the OCTA domain in discrete and finite space. Specifically, a contrastive-inspired semantic alignment is proposed to maximize the mutual information with the pre-trained models from OCT and OCTA views, to facilitate codebook learning. Meanwhile, a vessel structure alignment is proposed to minimize the structure discrepancy with the pre-trained models from the OCTA project map view, benefiting from learning the detailed vessel structure information. We also collect the first large-scale dataset, namely, OCTA2024, which contains a pair of OCT and OCTA volumes from 846 subjects. Our codes and datasets are available at: https://github.com/xmed-lab/MuTri. Zhuangzhuang Chen, Hualiang Wang, Chubin Ou, Xiaomeng Li 0001 |
CVPR | 2 |
| 2025 | Token Activation Map to Visually Explain Multimodal LLMsabstractMultimodal large language models (MLLMs) are broadly empowering various fields. Despite their advancements, the explainability of MLLMs remains less explored, hindering deeper understanding, model credibility, and effective visualization. Unlike conventional vision models (e.g., CNNs, ViTs, CLIP) that produce a single output, MLLMs generate sequences of tokens progressively, where each generated token depends on the previous context. Therefore, earlier context tokens can introduce redundant activations that interfere with the explanation of later tokens beyond their original information. Existing studies often overlook this issue, but our observations reveal that these redundant correlations can significantly hurt the reliability of explanations. To address this, we propose an estimated causal inference method to mitigate the interference of context to achieve high-quality MLLM explanation, with a novel rank Gaussian filter to further reduce activation noises. We term this method Token Activation Map (TAM) to highlight the consideration of interactions between tokens. TAM also indicates that it excels at explaining multiple tokens of MLLM, which is different from the Class Activation Map (CAM) for a single prediction. Our TAM method significantly outperforms existing SoTA methods, showcasing high-quality visualization results that can be utilized for various scenarios, such as object localization, failure case analysis, video visualization, MLLMs visual comparison, and model understanding (e.g., color, shape, action, location, visual reasoning, multi-turn conversation, etc). The code is available at github.com/xmed-lab/TAM. Yi Li 0050, Hualiang Wang, Xinpeng Ding, Xiaomeng Li 0001 |
ICCV | 2 |
| 2025 | Exploration of Teaching Reform in the Course of Intelligent Prediction in the Information Ageabstract"Intelligent Prediction in the Information Era" is a specialized general education course designed for undergraduates. This study focuses on the teaching needs and issues in the current teaching model of this course, conducting an in-depth investigation from three aspects: teaching content, teaching methods, and evaluation system. There search develops a set of teaching content that integrates and correlates intelligent prediction knowledge across multiple disciplines for undergraduates and proposes a new syllabus. Furthermore, it explores teaching methods that inspire students' active and innovative research thinking and designs assessment standards to guide students in actively exploring intelligent prediction techniques. There search findings provide significant reference value for the reform and innovation of teaching models and methods in undergraduate specialized general education courses. Jian Ma 0006, Yujie Cheng, Hualiang Wang, Hongmei Liu 0004, Laifa Tao, Chen Lu 0001 |
INDIN | 3 |
| 2025 | Cross-View Generalized Diffusion Model for Sparse-View CT Reconstruction
Jixiang Chen 0001, Yiqun Lin, Yi Qin 0006, Hualiang Wang, Xiaomeng Li 0001 |
MICCAI (16) | 4 |
| 2025 | VAMPIRE: Uncovering Vessel Directional and Morphological Information from OCTA Images for Cardiovascular Disease Risk Factor Prediction
Lehan Wang, Hualiang Wang, Chubin Ou, Lushi Chen, Yunyi Liang, Xiaomeng Li 0001 |
MICCAI (15) | 2 |
| 2025 | MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
Yuanpeng Nie, Hualiang Wang, Wei Li 0320, Junzhi Ning, Hongqiu Wang, Jiyao Liu, Junjun He |
MICCAI (5) | 3 |
| 2025 | A closer look at the explainability of Contrastive language-image pre-training
Yi Li 0050, Hualiang Wang, Yiqun Duan, Jiheng Zhang, Xiaomeng Li 0001 |
Pattern Recognit. | 2 |
| 2025 | SemiGMMPoint: Semi-supervised point cloud segmentation based on Gaussian mixture models
Xianwei Zhuang, Hualiang Wang, Xiaoxuan He, Siming Fu, Haoji Hu |
Pattern Recognit. | 2 |
| 2024 | Robustness-Guided Image Synthesis for Data-Free QuantizationabstractQuantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks. Jianhong Bai, Huanpeng Chu, Hualiang Wang, Zuozhu Liu, Ruizhe Chen, Xiaoxuan He, Lianrui Mu, Chengfei Cai, Haoji Hu |
AAAI | 4 |
| 2024 | C2RV: Cross-Regional and Cross-View Learning for Sparse-View CBCT ReconstructionabstractCone beam computed tomography (CBCT) is an important imaging technology widely used in medical scenarios, such as diagnosis and preoperative planning. Using fewer projection views to reconstruct CT, also known as sparse-view reconstruction, can reduce ionizing radiation and further benefit interventional radiology. Compared with sparse-view reconstruction for traditional parallel/fan-beam CT, CBCT reconstruction is more challenging due to the increased dimensionality caused by the measurement process based on cone-shaped X-ray beams. As a 2D-to-3D reconstruction problem, although implicit neural representations have been introduced to enable efficient training, only local features are considered and different views are processed equally in previous works, resulting in spatial inconsistency and poor performance on complicated anatomies. To this end, we propose C2RV by leveraging explicit multi-scale volumetric representations to enable cross-regional learning in the 3D space. Additionally, the scale-view cross-attention module is introduced to adaptively aggregate multi-scale and multi-view features. Extensive experiments demonstrate that our C2RV achieves consistent and significant improvement over previous state-of-the-art methods on datasets with diverse anatomy. Code is available at https://github.com/xmed-lab/C2RV-CBCT. Yiqun Lin, Jiewen Yang, Hualiang Wang, Xinpeng Ding, Wei Zhao 0029, Xiaomeng Li 0001 |
CVPR | 3 |
| 2024 | Multimodal Survival Ensemble Network: Integrating Genomic and Histopathological Insights for Enhanced Cancer PrognosisabstractCancer’s inherent heterogeneity demands a multimodal approach to provide an accurate prognosis, taking into account histological, clinical, and genomic data. As the field of artificial intelligence evolves with advancements in multimodal learning, its role in survival analysis becomes increasingly critical. We introduce the Multimodal Survival Ensemble Network (MSEN), a novel weakly-supervised framework designed for the seamless integration of genomic data and histopathological images. Not only does our method preserve the heterogeneity among different genomic modalities during integration, but it also ensures superior retention of spatial information in histopathological images compared to traditional techniques. Rigorous evaluations across five datasets highlight MSEN’s superior performance, marking a progressive step in cancer prognosis. Chenyi Zhou, Hualiang Wang, Xiaomeng Li 0001, Wanlu Liu, Zuozhu Liu |
ICASSP | 2 |
| 2024 | HiA: Towards Chinese Multimodal LLMs for Comparative High-Resolution Joint Diagnosis
Xinpeng Ding, Yongqiang Chu, Renjie Pi, Hualiang Wang, Xiaomeng Li 0001 |
MICCAI (12) | 4 |
| 2024 | Learning 3D Gaussians for Extremely Sparse-View Cone-Beam CT Reconstruction
Yiqun Lin, Hualiang Wang, Jixiang Chen 0001, Xiaomeng Li 0001 |
MICCAI (7) | 2 |
| 2024 | Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images
Hualiang Wang, Yiqun Lin, Xinpeng Ding, Xiaomeng Li 0001 |
MICCAI (9) | 1 |
| 2024 | A Novel Design of a Unilateral Nuclear Magnetic Resonance Sensor for Soil Moisture Detection Based on a Simplified Analytical ModelabstractSoil moisture (SM) is a key state variable in terrestrial systems because it controls the exchange of water and energy between the continental surface and the atmosphere. Nuclear magnetic resonance (NMR) technology is widely used for the analysis of porous media in SM due to its unique sensitivity to hydrogen protons. Unlike traditional laboratory NMR systems, unilateral NMR (UNMR) systems allow for the placement of detection targets outside the sensor space, thereby enabling in situ detection capabilities. However, in the existing designs of UNMR sensors, the magnetic field location is typically determined after the sensor has been designed. In this study, a novel UNMR sensor design scheme is proposed based on a simplified analytical model (SAM) to solve this problem. In contrast to conventional practices, this scheme places a primary emphasis on the identification of detection positions as its initial step, followed by the computation of magnet structure parameters. Concurrently, the mechanical design of the proposed UNMR sensor offers a more adaptable approach to regulation. The scheme consists of two components: magnetic field calculation and optimization of structural parameters. Notably, the proposed model exhibits a remarkable enhancement in calculation efficiency, surpassing the baseline by more than 70 times within a single iteration, compared with the traditional analytical model (TAM). The goodness of fit between the measured magnetic field distribution and the optimized results surpasses 0.99, thereby providing additional evidence of the sensor’s effectiveness. In addition, the sensor’s performance is demonstrated through measurements conducted on samples with varying SM content. Tingting Lin 0001, Hualiang Wang, Zhengping Li, Jinbao Zhu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CLIPN for Zero-Shot OOD Detection: Teaching CLIP to Say NoabstractOut-of-distribution (OOD) detection refers to training the model on an in-distribution (ID) dataset to classify whether the input images come from unknown classes. Considerable effort has been invested in designing various OOD detection methods based on either convolutional neural networks or transformers. However, zero-shot OOD detection methods driven by CLIP, which only require class names for ID, have received less attention. This paper presents a novel method, namely CLIP saying "no" (CLIPN), which empowers the logic of saying "no" within CLIP. Our key motivation is to equip CLIP with the capability of distinguishing OOD and ID samples using positive-semantic prompts and negation-semantic prompts. Specifically, we design a novel learnable "no" prompt and a "no" text encoder to capture negation semantics within images. Subsequently, we introduce two loss functions: the image-text binary-opposite loss and the text semantic-opposite loss, which we use to teach CLIPN to associate images with "no" prompts, thereby enabling it to identify unknown samples. Furthermore, we propose two threshold-free inference algorithms to perform OOD detection by utilizing negation semantics from "no" prompts and the text encoder. Experimental results on 9 benchmark datasets (3 ID datasets and 6 OOD datasets) for the OOD detection task demonstrate that CLIPN, based on ViT-B-16, outperforms 7 well-used algorithms by at least 2.34% and 11.64% in terms of AUROC and FPR95 for zero-shot OOD detection on ImageNet-1K. Our CLIPN can serve as a solid foundation for effectively leveraging CLIP in downstream OOD tasks. The code is available on https://github.com/xmed-lab/CLIPN. Hualiang Wang, Yi Li 0050, Huifeng Yao, Xiaomeng Li 0001 |
ICCV | 1 |
| 2023 | On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail Learning
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao, Yang Feng 0011, Huanpeng Chu, Haoji Hu |
ICLR | 3 |
| 2023 | Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionabstractRecently, large-scale pre-trained vision-language models have presented benefits for alleviating class imbalance in long-tailed recognition. However, the long-tailed data distribution can corrupt the representation space, where the distance between head and tail categories is much larger than the distance between two tail categories. This uneven feature space distribution causes the model to exhibit unclear and inseparable decision boundaries on the uniformly distributed test set, which lowers its performance. To address these challenges, we propose the uniformly category prototype-guided vision-language framework to effectively mitigate feature space bias caused by data imbalance. Especially, we generate a set of category prototypes uniformly distributed on a hypersphere. Category prototype-guided mechanism for image-text matching makes the features of different classes converge to these distinct and uniformly distributed category prototypes, which maintain a uniform distribution in the feature space, and improve class boundaries. Additionally, our proposed irrelevant text filtering and attribute enhancement module allows the model to ignore irrelevant noisy text and focus more on key attribute information, thereby enhancing the robustness of our framework. In the image recognition fine-tuning stage, to address the positive bias problem of the learnable classifier, we design the class feature prototype-guided classifier, which compensates for the performance of tail classes while maintaining the performance of head classes. Our method outperforms previous vision-language methods for long-tailed learning work by a large margin and achieves state-of-the-art performance. Xiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 0005, Hualiang Wang |
ACM Multimedia | 5 |
| 2023 | Towards Distribution-Agnostic Generalized Category DiscoveryabstractData imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on classifying close-set samples and detecting open-set samples during testing, it's still essential to be able to classify unknown subjects as human beings. In this paper, we formally define a more realistic task as distribution-agnostic generalized category discovery (DA-GCD): generating fine-grained predictions for both close- and open-set classes in a long-tailed open-world setting. To tackle the challenging problem, we propose a Self-**Ba**lanced **Co**-Advice co**n**trastive framework (BaCon), which consists of a contrastive-learning branch and a pseudo-labeling branch, working collaboratively to provide interactive supervision to resolve the DA-GCD task. In particular, the contrastive-learning branch provides reliable distribution estimation to regularize the predictions of the pseudo-labeling branch, which in turn guides contrastive learning through self-balanced knowledge transfer and a proposed novel contrastive loss. We compare BaCon with state-of-the-art methods from two closely related fields: imbalanced semi-supervised learning and generalized category discovery. The effectiveness of BaCon is demonstrated with superior performance over all baselines and comprehensive analysis across various datasets. Our code is publicly available. Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li 0001, Joey Tianyi Zhou, Yang Feng 0011, Jian Wu 0001, Haoji Hu |
NeurIPS | 3 |
| 2023 | Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient BalancerabstractData privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist. Zikai Xiao, Zihan Chen 0001, Songshang Liu, Hualiang Wang, Yang Feng 0011, Jin Hao, Joey Tianyi Zhou, Jian Wu 0001, Howard H. Yang, Zuozhu Liu |
NeurIPS | 4 |
| 2023 | Class semantic enhancement network for semantic segmentation
Siming Fu, Hualiang Wang, Haoji Hu, Xiaoxuan He, Yongwen Long, Jianhong Bai, Yangtao Ou, Yuanjia Huang, Mengqiu Zhou |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Hierarchical Self-Supervised Learning for 3D Tooth Segmentation in Intra-Oral Mesh ScansabstractAccurately delineating individual teeth and the gingiva in the three-dimension (3D) intraoral scanned (IOS) mesh data plays a pivotal role in many digital dental applications, e.g., orthodontics. Recent research shows that deep learning based methods can achieve promising results for 3D tooth segmentation, however, most of them rely on high-quality labeled dataset which is usually of small scales as annotating IOS meshes requires intensive human efforts. In this paper, we propose a novel self-supervised learning framework, named STSNet, to boost the performance of 3D tooth segmentation leveraging on large-scale unlabeled IOS data. The framework follows two-stage training, i.e., pre-training and fine-tuning. In pre-training, three hierarchical-level, i.e., point-level, region-level, cross-level, contrastive losses are proposed for unsupervised representation learning on a set of predefined matched points from different augmented views. The pretrained segmentation backbone is further fine-tuned in a supervised manner with a small number of labeled IOS meshes. With the same amount of annotated samples, our method can achieve an mIoU of 89.88%, significantly outperforming the supervised counterparts. The performance gain becomes more remarkable when only a small amount of labeled samples are available. Furthermore, STSNet can achieve better performance with only 40% of the annotated samples as compared to the fully supervised baselines. To the best of our knowledge, we present the first attempt of unsupervised pre-training for 3D tooth segmentation, demonstrating its strong potential in reducing human efforts for annotation and verification. Zuozhu Liu, Xiaoxuan He, Hualiang Wang, Huimin Xiong, Yan Zhang 0004, Gaoang Wang, Jin Hao, Yang Feng 0011, Fudong Zhu, Haoji Hu |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Renovate Yourself: Calibrating Feature Representation of Misclassified Pixels for Semantic SegmentationabstractExisting image semantic segmentation methods favor learning consistent representations by extracting long-range contextual features with the attention, multi-scale, or graph aggregation strategies. These methods usually treat the misclassified and correctly classified pixels equally, hence misleading the optimization process and causing inconsistent intra-class pixel feature representations in the embedding space during learning. In this paper, we propose the auxiliary representation calibration head (RCH), which consists of the image decoupling, prototype clustering, error calibration modules and a metric loss function, to calibrate these error-prone feature representations for better intra-class consistency and segmentation performance. RCH could be incorporated into the hidden layers, trained together with the segmentation networks, and decoupled in the inference stage without additional parameters. Experimental results show that our method could significantly boost the performance of current segmentation methods on multiple datasets (e.g., we outperform the original HRNet and OCRNet by 1.1% and 0.9% mIoU on the Cityscapes test set). Codes are available at https://github.com/VipaiLab/RCH. Hualiang Wang, Huanpeng Chu, Siming Fu, Zuozhu Liu, Haoji Hu |
AAAI | 1 |
| 2022 | Meta-prototype Decoupled Training for Long-Tailed Learning
Siming Fu, Huanpeng Chu, Xiaoxuan He, Hualiang Wang, Haoji Hu |
ACCV (6) | 4 |
| 2022 | Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-Tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He, Hangxiang Fang, Zuozhu Liu, Haoji Hu |
ECCV (24) | 1 |
| 2022 | Meta-BNS FOR Adversarial Data-Free QuantizationabstractData-free quantization has recently been a promising method to perform quantization without access to the original data. However, the drawback of such approaches is the homogenization of synthetic data due to low efficiency for diverse data generation and the performance collapse of the generator. To alleviate the above issue, we propose a novel Meta-BNS for adversarial data-free quantization scheme which consists of Meta-BNS module and adversarial exploration module. Meta-BNS module automatically learns an enhancement coefficient matrix function for BN loss module to provide a suitable constrain on the generator. Adversarial exploration module leverages minimax game between the generator and quantized model via input gradient to encourage the generator to learn high-dimensional and complex real data distribution. The experimental results show that our method achieves state-of-the-art performance for various settings on data-free quantization. Siming Fu, Hualiang Wang, Yuchen Cao 0005, Haoji Hu, Wenming Tan, Tingqun Ye |
ICIP | 2 |