Huiqiong Wang

dblp:19/1682 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-9560-563XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RAIN: An embarrassingly simple approach to debiasing attribution evaluation
Jiarui Duan, Haofei Zhang, Mengqi Xue, Huiqiong Wang, Mingli Song
Comput. Vis. Image Underst.5
2025 Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?
abstract
High-quality open-source text-to-image models have lowered the threshold for obtaining photorealistic images significantly, but also face potential risks of misuse. Specifically, suspects may use synthetic data generated by these generative models to train models for specific tasks without permission, when lacking real data resources especially. Protecting these generative models is crucial for the wellbeing of their owners. In this work, we propose the first method to this important yet unresolved issue, called Training data Provenance Verification (TrainProVe). The rationale behind TrainProVe is grounded in the principle of generalization error bound, which suggests that, for two models with the same task, if the distance between their training data distributions is smaller, their generalization ability will be closer. We validate the efficacy of TrainProVe across four text-to-image models (Stable Diffusion v1.4, latent consistency model, PixArt-α, and Stable Cascade). The results show that TrainProVe achieves a verification accuracy of over 99% in determining the provenance of suspicious model training data, surpassing all previous methods. Code is available at https://github.com/xieyc99/TrainProVe.
Yuechen Xie, Jie Song 0011, Huiqiong Wang, Mingli Song
CVPR3
2025 From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition
abstract
Multilingual automatic speech recognition (ASR) models greatly facilitate recognizing low-resource languages by sharing representations across similar languages. However, the commonly adopted modeling units, e.g., character-level modeling, lack language-specific information, resulting in a susceptible word prediction to phonemes and characters. Recently, subword-level modeling has demonstrated significant effectiveness for monolingual automatic recognition systems, while it is adverse to cross-lingual feature sharing. In this paper, we propose a novel low-resource ASR method that leverages the advantages of two different modeling units. Specifically, a character-level ASR model is trained on the multilingual dataset for modeling the short-term speech and learning general speech knowledge from relevant languages. Afterwards, we convert the character-level prediction into subwords for learning contextual information of the target language. Extensive experiments on Uyghur with Kazakh and Kyrgyz as auxiliary languages have shown that our proposed method significantly reduces word error rate (WER).
Haofei Zhang, Huiqiong Wang, Mingli Song
ICASSP3
2025 Disentangled Condensation for Large-scale Graphs
abstract
Graph condensation has emerged as an intriguing technique to save the expensive training costs of Graph Neural Networks (GNNs) by substituting a condensed small graph with the original graph. Despite the promising results achieved, previous methods usually employ an entangled paradigm of redundant parameters (nodes, edges, GNNs), which incurs complex joint optimization during condensation. This paradigm has considerably impeded the scalability of graph condensation, making it challenging to condense extremely large-scale graphs and generate high-fidelity condensed graphs. Therefore, we propose to disentangle the condensation process into a two-stage GNN-free paradigm, independently condensing nodes and generating edges while eliminating the need to optimize GNNs at the same time. The node condensation module avoids the complexity of GNNs by focusing on node feature alignment with anchors of the original graph, while the edge translation module constructs the edges of the condensed nodes by transferring the original structure knowledge with neighborhood anchors. This simple yet effective approach achieves at least 10 times faster than state-of-the-art methods with comparable accuracy on medium-scale graphs. Moreover, the proposed DisCo can successfully scale up to the Ogbn-papers100M graph containing over 100 million nodes with flexible reduction rates and improves performance on the second-largest Ogbn-products dataset by over 5%. Extensive downstream tasks and ablation study on five common datasets further demonstrate the effectiveness of the proposed DisCo framework. Our code is available at https://github.com/BangHonor/DisCo.
Zhenbang Xiao, Yu Wang 0176, Shunyu Liu 0001, Bingde Hu, Huiqiong Wang, Mingli Song, Tongya Zheng
WWW5
2025 Deep feature response discriminative calibration
Linyun Zhou, Zunlei Feng, Mingli Song, Huiqiong Wang
Neurocomputing6
2024 RS-SAM: Integrating Multi-scale Information for Enhanced Remote Sensing Image Segmentation
Enkai Zhang, Anda Cao, Haofei Zhang, Huiqiong Wang, Mingli Song
ACCV (8)6
2024 Training-Free Pretrained Model Merging
abstract
Recently, model merging techniques have surfaced as a solution to combine multiple single-talent models into a single multi-talent model. However, previous endeavors in this field have either necessitated additional training or fine-tuning processes, or require that the models possess the same pre-trained initialization. In this work, we identify a common drawback in prior works w.r.t. the inconsistency of unit similarity in the weight space and the activation space. To address this inconsistency, we propose an innovative model merging framework, coined as merging under dual-space constraints (MuDSC). Specifically, instead of solely maximizing the objective of a single space, we advocate for the exploration of permutation matrices situated in a region with a unified high similarity in the dual space, achieved through the linear combination of activation and weight similarity matrices. In order to enhance usability, we have also incorporated adaptations for group structure, including Multi-Head Attention and Group Normalization. Comprehensive experimental comparisons demonstrate that MuDSC can significantly boost the performance of merged models with various task combinations and architectures. Furthermore, the visualization of the merged model within the multi-task loss landscape reveals that MuDSC enables the merged model to reside in the overlapping segment, featuring a unified lower loss for each task. Our code is publicly available at https://github.com/zju-vipa/training_free_model_merging.
Zhengqi Xu, Huiqiong Wang, Mingli Song, Jie Song 0011
CVPR3
2024 Deep Kernel Calibration
abstract
Korbinian Brodmann argued that the brain regions with different cytoarchitectures exhibited different cognitive functions, which are widely recognized as "Brodmann areas" in neuroscience. Inspired by this theory, we observe from experiments that the different well-trained convolutional kernels also hold their unique functionalities, indicated by their output activations that highly concentrate on the "different certain ranges". This discovery motivates us to devise a deep kernel-by-kernel feature calibration mechanism that "calibrates" the output distribution of a pre-trained convolutional kernel for a more concentrated activation interval, by eliminating the outlier activations. Towards this end, we develop two dedicated Brodmann calibration forms, termed as Hard Kernel Calibration (HKC) and Soft Kernel Calibration (SKC) that simply filters the overrange activations, or adaptively condense the activations in a weighted manner, respectively. As a flexible plug-and-play module, the proposed calibration demonstrates encouraging results on five benchmarks across sixteen network architectures and also triggers the functionality of significantly enhanced model robustness against adversarial attacks. The developed code is publicly available at https://github.com/tcmyxc/DKC.
Zunlei Feng, Jie Lei 0002, Huiqiong Wang, Zhongle Xie
IJCNN5
2024 GAN Doctor: Diagnosing and Treating Inherent Semantic Errors
abstract
Generative Adversarial Network (GAN), as a popular generative model in the field of Artificial Intelligence Generated Content (AIGC), has been intensively developed in previous research, with significant improvements in the quality and diversity of image generation. However, there are still many cases where the results are not satisfactory. A primary concern pertains to the chaotic and blurred local details within the generated images. In this work, through the diagnosis and analysis of high-quality and low-quality images produced by the GAN model, we identified that this issue stems from inherent semantic errors of the GAN, that is, convolutional kernels responsible for certain semantics are not properly involved in the generation process of corresponding image regions. To this end, we propose a straightforward yet effective treatment method, which constrains each image region to be generated by its corresponding semantic convolutional kernels. Experimental results demonstrate that our proposed optimization method can improve the issue of chaotic and blurred local regions in generated images and enhance the overall generation quality. Our work pioneers a novel paradigm for diagnosing and treating GANs, driving the research development and practical application of AIGC image generation technology.
Chengji Shen, Zunlei Feng, Zhongle Xie, Jie Lei 0002, Huiqiong Wang, Mingli Song
IJCNN5
2024 LG-CAV: Train Any Concept Activation Vector with Language Guidance
abstract
Concept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV.
Qihan Huang, Jie Song 0011, Mengqi Xue, Haofei Zhang, Bingde Hu, Huiqiong Wang, Hao Jiang 0014, Xingen Wang, Mingli Song
NeurIPS6
2024 Simple Graph Condensation
Zhenbang Xiao, Yu Wang 0176, Shunyu Liu 0001, Huiqiong Wang, Mingli Song, Tongya Zheng
ECML/PKDD (2)4
2022 Comparison Knowledge Translation for Generalizable Image Classification
abstract
Deep learning has recently achieved remarkable performance in image classification tasks, which depends heavily on massive annotation. However, the classification mechanism of existing deep learning models seems to contrast to humans' recognition mechanism. With only a glance at an image of the object even unknown type, humans can quickly and precisely find other same category objects from massive images, which benefits from daily recognition of various objects. In this paper, we attempt to build a generalizable framework that emulates the humans' recognition mechanism in the image classification task, hoping to improve the classification performance on unseen categories with the support of annotations of other categories. Specifically, we investigate a new task termed Comparison Knowledge Translation (CKT). Given a set of fully labeled categories, CKT aims to translate the comparison knowledge learned from the labeled categories to a set of novel categories. To this end, we put forward a Comparison Classification Translation Network (CCT-Net), which comprises a comparison classifier and a matching discriminator. The comparison classifier is devised to classify whether two images belong to the same category or not, while the matching discriminator works together in an adversarial manner to ensure whether classified results match the truth. Exhaustive experiments show that CCT-Net achieves surprising generalization ability on unseen categories and SOTA performance on target categories.
Zunlei Feng, Sai Wu, Xiaotuan Jin, Zengliang He, Mingli Song, Huiqiong Wang
IJCAI7
2021 Self-born Wiring for Neural Trees
abstract
Neural trees aim at integrating deep neural networks and decision trees so as to bring the best of the two worlds, including representation learning from the former and faster inference from the latter. In this paper, we introduce a novel approach, termed as Self-born Wiring (SeBoW), to learn neural trees from a mother deep neural network. In contrast to prior neural-tree approaches that either adopt a pre-defined structure or grow hierarchical layers in a progressive manner, task-adaptive neural trees in SeBoW evolve from a deep neural network through a construction-by-destruction process, enabling a global-level parameter optimization that further yields favorable results. Specifically, given a designated network configuration like VGG, SeBoW disconnects all the layers and derives isolated filter groups, based on which a global-level wiring process is conducted to attach a subset of filter groups, eventually bearing a lightweight neural tree. Extensive experiments demonstrate that, with a lower computational cost, SeBoW outperforms all prior neural trees by a significant margin and even achieves results on par with predominant non-tree networks like ResNets. Moreover, SeBoW proves its scalability to large-scale datasets like ImageNet, which has been barely explored by prior tree networks.
Feng Mao, Jie Song 0011, Xinchao Wang, Huiqiong Wang, Mingli Song
ICCV5
2021 Disassembling object representations without labels
Zunlei Feng, Yongming He, Yike Yuan, Huiqiong Wang, Mingli Song
Neurocomputing5
2008 Separating corneal reflections for illumination estimation
Huiqiong Wang, Stephen Lin 0001, Xiuqing Ye, Weikang Gu
Neurocomputing1
2007 A Generic Framework for Efficient 2-D and 3-D Facial Expression Analogy
abstract
Facial expression analogy provides computer animation professionals with a tool to map expressions of an arbitrary source face onto an arbitrary target face. In the recent past, several algorithms have been presented in the literature that aim at putting the expression analogy paradigm into practice. Some of these methods exclusively handle expression mapping between 3-D face models, while others enable the transfer of expressions between images of faces only. None of them, however, represents a more general framework that can be applied to either of these two face representations. In this paper, we describe a novel generic method for analogy-based facial animation that employs the same efficient framework to transfer facial expressions between arbitrary 3-D face models, as well as between images of performer's faces. We propose a novel geometry encoding for triangle meshes, vertex-tent-coordinates, that enables us to formulate expression transfer in the 2-D and the 3-D case as a solution to a simple system of linear equations. Our experiments show that our method outperforms many previous analogy-based animation approaches in terms of achieved animation quality, computation time and generality.
Mingli Song, Zhao Dong 0001, Christian Theobalt, Huiqiong Wang, Zicheng Liu 0001, Hans-Peter Seidel
IEEE Trans. Multim.4
2006 Subtle Facial Expression Modeling with Vector Field Decomposition
abstract
Facial expression means not only the feature motions but also the subtle appearance changes in the face. The details are important to facial expression synthesis. In this paper, a novel method based on vector field decomposition is presented to model subtle facial expressions. The intensity change of facial expression image is represented by Helmholtz-Hodge vector field decomposition. Based on this method, subtle facial expression mapping is carried out much faster and better. Experiment results show that our method is robust and convincible to model the subtle facial expression.
Mingli Song, Huiqiong Wang, Jiajun Bu, Chun Chen 0001, Zicheng Liu 0001
ICIP2
2005 Separating Reflections in Human Iris Images for Illumination Estimation
abstract
A method is presented for separating corneal reflections in an image of human irises to estimate illumination from the surrounding scene. Previous techniques for reflection separation have demonstrated success in only limited cases, such as for uniform colored lighting and simple object textures, so they are not applicable to irises which exhibit intricate textures and complicated reflections of the environment. To make this problem feasible, we present a method that capitalizes on physical characteristics of human irises to obtain an illumination estimate that encompasses the prominent light contributors in the scene. Results of this algorithm are presented for eyes of different colors, including light colored eyes for which reflection separation is necessary to determine a valid illumination estimate.
Huiqiong Wang, Stephen Lin 0001, Xiaopei Liu, Sing Bing Kang
ICCV1