VLDB 2026 Research / reviewers in the wild / expert
Huaidong Zhang
dblp:186/6659
· DBLP profile ↗
47ranked-venue papers
3as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-author · 29 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 27 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attribute-Centric Cross-Modal Alignment for Weakly Supervised Text-Based Person Re-IDabstractWeakly supervised text-based person re-identification (Text-ReID) confronts the challenge of matching target person images with textual descriptions, hindered by the absence of identity annotations during training. Traditional approaches, which rely solely on global features, overlook the rich, fine-grained information within both text and image modalities. Besides, merely aligning features at the semantic level is insufficient due to the significant differences in feature representation spaces between the two modalities. Existing methods also neglect the information inequality caused by person-irrelevant factors in images. In this paper, we introduce a novel framework called Attribute-Centric Cross-modal Alignment (ACCA), specifically designed to overcome these issues. Our approach concentrates on two main aspects: visual-text attribute alignment and prediction distribution alignment. To effectively capture fine-grained information without identity labels, we implement a visual-text attribute alignment method based on momentum contrastive learning to synchronize visual and textual attribute features within a unified embedding space. We also propose a unique strategy for negative sample filtering and enrichment, creating robust and comprehensive negative attribute sample spaces to support the attribute alignment. Additionally, we establish two methods of label-free prediction distribution alignment to encourage the learning of invariant feature representations across modalities. The first method, bias-reduction distribution alignment, aligns features and predictions within each text-image pair by utilizing semantic information from the text and reduces the impact of person-irrelevant factors in images. The second method, global-attribute distribution alignment, enhances the interaction between global and local prediction distributions across visual and textual modalities. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets validate our superior performances across all standard benchmarks. Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Multim. | 5 |
| 2025 | PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem EquilibriumabstractPersonalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances of facial features. In this study, we delve into the temporal dynamics of the text-to-image conditioning process, emphasizing the crucial role of stage partitioning in introducing new concepts. We present PersonaMagic, a stage-regulated generative technique designed for high-fidelity face customization. Using a simple MLP network, our method learns a series of embeddings within a specific timestep interval to capture face concepts. Additionally, we develop a Tandem Equilibrium mechanism that adjusts self-attention responses in the text encoder, balancing text description and identity preservation, improving both areas. Extensive experiments confirm the superiority of PersonaMagic over state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, its robustness and flexibility are validated in non-facial domains, and it can also serve as a valuable plug-in for enhancing the performance of pretrained personalization models. Xinzhe Li 0003, Jiahui Zhan, Shengfeng He, Yangyang Xu 0003, Junyu Dong, Huaidong Zhang, Yong Du 0003 |
AAAI | 6 |
| 2025 | FR²Seg: Continual Segmentation Across Multiple Sites via Fourier Style Replay and Adaptive Consistency RegularizationabstractIn clinical imaging, medical segmentation networks typically require continually adapting to new data from multiple sites over time, as aggregating all data for learning at once can be impractical due to storage limitations and privacy concerns. However, existing methods basically overlook domain-specific characteristics and fall short of adequately capturing domain-invariant knowledge during continual learning, leading to undesired catastrophic forgetting of previous sites and inferior generalization to new sites. To tackle this issue, this paper introduces FR2Seg, to sufficiently exploit both domain-specific and domain-invariant knowledge for efficient continual learning with the aid of low-frequency cues. For the former aspect, we propose a Fourier style replay module to synthesize pseudo images with old-site styles for data augmentation during new-site training, effectively preventing catastrophic forgetting without sacrificing data privacy. For the latter, we present a Fourier adaptive consistency regularization to identify and constrain the optimization of domain-invariant parameters with explicit awareness of knowledge transferability across sites, ensuring excellent generalizability to new sites. Experimental results on two public datasets confirm our method's superiority over existing state-of-the-art continual learning methods. Xuemiao Xu, Huaidong Zhang, Harry Qin |
AAAI | 5 |
| 2025 | MSSDA: Multi-Sub-Source Domain Adaptation for Diabetic Foot Neuropathy RecognitionabstractDiabetic foot neuropathy (DFN) is a critical factor leading to diabetic foot ulcers, which is one of the most common and severe complications of diabetes mellitus (DM) and is associated with high risks of amputation and mortality. Despite its significance, existing datasets do not directly derive from plantar data and lack continuous, long-term foot-specific information. To advance DFN research, we have collected a novel dataset comprising continuous plantar pressure data to recognize diabetic foot neuropathy. This dataset includes data from 94 DM patients with DFN and 41 DM patients without DFN. Moreover, traditional methods divide datasets by individuals, potentially leading to significant domain discrepancies in some feature spaces due to the absence of mid-domain data. In this paper, we propose an effective domain adaptation method to address this proplem. We split the dataset based on convolutional feature statistics and select appropriate sub-source domains to enhance efficiency and avoid negative transfer. We then align the distributions of each source and target domain pair in specific feature spaces to minimize the domain gap. Comprehensive results validate the effectiveness of our method on both the newly proposed dataset for DFN recognition and an existing dataset. Zhixin Yan, Shibin Wu, Huaidong Zhang, Peiru Zhou |
AAAI | 5 |
| 2025 | MODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow IntertwiningabstractMultimodal unsupervised domain adaptation leverages un-labeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel modal-affinity measurement to evaluate information quality. Additionally, we introduce a modal-affinity distillation technique to control sample-level information exchange, ensuring reliable multimodal interaction based on affinity evaluations within the feature space. Extensive experiments on three multimodal datasets demonstrate that our framework consistently outperforms state-of-the-art methods, particularly in high-noise environments. Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He |
CVPR | 4 |
| 2025 | Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head AnimationabstractSinging is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by singing holds significant research value. Our work focuses on 3D singing facial animation driven by mixed singing audio, and to the best of our knowledge, no prior studies have explored this area. Additionally, the absence of existing 3D singing datasets poses a considerable challenge. To address this, we collect a novel audiovisual dataset, ChorusHead which features synchronized mixed vocal audio and 3D motions for chorus singing. In addition, We propose a partner-aware 3D chorus head generation framework driven by mixed audio inputs. The proposed framework extracts emotional features from the background music and dependence between singers and models the head movement in a latent space from the Variational Autoencoder (VAE), enabling diverse interactive head animation generation. Extensive experimental results demonstrate that our approach effectively generates 3D facial animations of interacting singers, achieving notable improvements in realism and handling background music interference with strong robustness. The dataset will be released for research purposes at the project page: https://xxiexm.github.io/PaChorus/. Xiumei Xie, Zikai Huang, Xuemiao Xu, Huaidong Zhang |
CVPR | 6 |
| 2025 | NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian SplattingabstractNeural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, NexusGS comprises three key steps: Epipolar Depth Nexus, Flow-Resilient Depth Blending, and Flow-Filtered Depth Pruning. These steps leverage optical flow and camera poses to compute accurate depth maps, while mitigating the inaccuracies often associated with optical flow. By incorporating epipolar depth priors, NexusGS ensures reliable dense point cloud coverage and supports stable 3DGS training under sparse-view conditions. Experiments demonstrate that NexusGS significantly enhances depth accuracy and rendering quality, surpassing state-of-the-art methods by a considerable margin. Furthermore, we validate the superiority of our generated point clouds by substantially boosting the performance of competing methods. Project page: https://usmizuki.github.io/NexusGS/. Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du 0003 |
CVPR | 6 |
| 2025 | ViewSRD: 3D Visual Grounding Via Structured Multi-View Decomposition
Ronggang Huang, Haoxin Yang, Yan Cai 0021, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
ICCV | 5 |
| 2025 | SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose EstimationabstractExisting 3D Human Pose Estimation (HPE) methods achieve high accuracy but suffer from computational overhead and slow inference, while knowledge distillation methods fail to address spatial relationships between joints and temporal correlations in multi-frame inputs. In this paper, we propose Sparse Correlation and Joint Distillation (SCJD), a novel framework that balances efficiency and accuracy for 3D HPE. SCJD introduces Sparse Correlation Input Sequence Downsampling to reduce redundancy in student network inputs while preserving inter-frame correlations. For effective knowledge transfer, we propose Dynamic Joint Spatial Attention Distillation, which includes Dynamic Joint Embedding Distillation to enhance the student’s feature representation using the teacher’s multi-frame context feature, and Adjacent Joint Attention Distillation to improve the student network’s focus on adjacent joint relationships for better spatial understanding. Additionally, Temporal Consistency Distillation aligns the temporal correlations between teacher and student networks through upsampling and global supervision. Extensive experiments demonstrate that SCJD achieves state-of-the-art performance. Code is available at https://github.com/wileychan/SCJD. Xuemiao Xu, Haoxin Yang, Huaidong Zhang, Pheng-Ann Heng |
ICME | 7 |
| 2025 | L3Net: Localized and Layered Reparameterization for incremental learning
Xuandi Luo, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
Neural Networks | 2 |
| 2025 | Unambiguous granularity distillation for asymmetric image retrieval
Haoquan Zhang, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng-Ann Heng, Shengfeng He |
Neural Networks | 8 |
| 2025 | Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation LearningabstractThe vulnerability of 3D point cloud analysis to unpredictable rotations poses an open yet challenging problem: orientation-aware 3D domain generalization. Cross-domain robustness and adaptability of 3D representations are crucial but not easily achieved through rotation augmentation. Motivated by the inherent advantages of intricate orientations in enhancing generalizability, we propose an innovative rotation-adaptive domain generalization framework for 3D point cloud analysis. Our approach aims to alleviate orientational shifts by leveraging intricate samples in an iterative learning process. Specifically, we identify the most challenging rotation for each point cloud and construct an intricate orientation set by optimizing intricate orientations. Subsequently, we employ an orientation-aware contrastive learning framework that incorporates an orientation consistency loss and a margin separation loss, enabling effective learning of categorically discriminative and generalizable features with rotation consistency. Extensive experiments and ablations conducted on 3D cross-domain benchmarks firmly establish the state-of-the-art performance of our proposed approach in the context of orientation-aware 3D domain generalization. Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | SITA: Structurally Imperceptible and Transferable Adversarial Attacks for Stylized Image Generation
Jingdan Kang, Haoxin Yang, Yan Cai 0021, Huaidong Zhang, Xuemiao Xu, Yong Du 0003, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Recurrent Diffusion for 3D Point Cloud Generation From a Single ImageabstractSingle-image 3D shape reconstruction has attracted significant attention with the advance of generative models. Recent studies have utilized diffusion models to achieve unprecedented shape reconstruction quality. However, these methods, in each sampling step, perform denoising in a single forward pass, leading to cumulative errors that severely impact the geometric consistency of the generated shapes with the input targets and face difficulties in reconstructing rich details of complex 3D shapes. Moreover, the performance of current works suffers significant degradation due to limited information when only a single image is used as input during testing, further affecting the quality of 3D shape generation. In this paper, we present a recurrent diffusion framework, aiming to improve generation performance during single image-to-shape inference. Diverging from denoising in a single forward pass, we recursively refine the noise prediction in a self-rectified manner with the explicit guidance of the input target, thereby markedly suppressing cumulative errors and improving detail modeling. To enhance the geometric perception ability of the network during single-image inference, we further introduce a multi-view training scheme equipped with a view-robust conditional generation mechanism, which effectively promotes generation quality even when only a single image is available during inference. The effectiveness of our method is demonstrated through extensive evaluations on two public 3D shape datasets, where it surpasses state-of-the-art methods both qualitatively and quantitatively. Dewang Ye, Huaidong Zhang, Xuemiao Xu, Huajie Sun, Yewen Xu, Yuexia Zhou |
IEEE Trans. Image Process. | 3 |
| 2025 | Delving Into Invisible Semantics for Generalized One-Shot Neural Human RenderingabstractTraditional human neural radiance fields often overlook crucial body semantics, resulting in ambiguous reconstructions, particularly in occluded regions. To address this problem, we propose the Super-Semantic Disentangled Neural Renderer (SSD-NeRF), which employs rich regional semantic priors to enhance human rendering accuracy. This approach initiates with a Visible-Invisible Semantic Propagation module, ensuring coherent semantic assignment to occluded parts based on visible body segments. Furthermore, a Region-Wise Texture Propagation module independently extends textures from visible to occluded areas within semantic regions, thereby avoiding irrelevant texture mixtures and preserving semantic consistency. Additionally, a view-aware curricular learning approach is integrated to bolster the model's robustness and output quality across different viewpoints. Extensive evaluations confirm that SSD-NeRF surpasses leading methods, particularly in generating quality and structurally semantic reconstructions of unseen or occluded views and poses. Yihong Lin, Xuemiao Xu, Huaidong Zhang, Harry Qin, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | DreamAnime: Learning Style-Identity Textual Disentanglement for Anime and BeyondabstractText-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character or an art style, we teach our model to encapsulate its essence through novel "words" in the embedding space of a pre-existing text-to-image model. Crucially, we disentangle the concepts of style and identity into two separate "words", thus providing the ability to manipulate them independently. These distinct "words" can then be pieced together into natural language sentences, promoting an intuitive and personalized creative process. Empirical results suggest that this disentanglement into separate word embeddings successfully captures a broad range of unique and complex concepts, with each word focusing on style or identity as appropriate. Comparisons with existing methods illustrate DreamAnime's superior capacity to accurately interpret and recreate the desired concepts across various applications and tasks. Chenshu Xu, Yangyang Xu 0003, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Learning an Interpretable Stylized Subspace for 3D-Aware Animatable ArtformsabstractThroughout history, static paintings have captivated viewers within display frames, yet the possibility of making these masterpieces vividly interactive remains intriguing. This research paper introduces 3DArtmator, a novel approach that aims to represent artforms in a highly interpretable stylized space, enabling 3D-aware animatable reconstruction and editing. Our rationale is to transfer the interpretability and 3D controllability of the latent space in a 3D-aware GAN to a stylized sub-space of a customized GAN, revitalizing the original artforms. To this end, the proposed two-stage optimization framework of 3DArtmator begins with discovering an anchor in the original latent space that accurately mimics the pose and content of a given art painting. This anchor serves as a reliable indicator of the original latent space local structure, therefore sharing the same editable predefined expression vectors. In the second stage, we train a customized 3D-aware GAN specific to the input artform, while enforcing the preservation of the original latent local structure through a meticulous style-directional difference loss. This approach ensures the creation of a stylized sub-space that remains interpretable and retains 3D control. The effectiveness and versatility of 3DArtmator are validated through extensive experiments across a diverse range of art styles. With the ability to generate 3D reconstruction and editing for artforms while maintaining interpretability, 3DArtmator opens up new possibilities for artistic exploration and engagement. Chenxi Zheng, Bangzhen Liu, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | D3still: Decoupled Differential Distillation for Asymmetric Image RetrievalabstractExisting methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these one-to-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is then decomposed into three components: feature representation knowledge, inconsistent pairwise similarity differential knowledge, and consistent pairwise similarity differential knowledge. This strategic decomposition aligns the retrieval ranking of the query network with the gallery network effectively. Extensive experiments on various bench-mark datasets reveal that D3still surpasses state-of-the-art methods in asymmetric image retrieval. Code is available at https://github.com/SCY-X/D3still. Yihong Lin, Xuemiao Xu, Huaidong Zhang, Yong Du 0003, Shengfeng He |
CVPR | 5 |
| 2024 | Beyond Textual Constraints: Learning Novel Diffusion Conditions with Fewer ExamplesabstractIn this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which addresses the inconsistencies typically brought about by Classifier-free guidance in few-shot training con-texts. Our extensive experiments across a variety of condition modalities demonstrate the effectiveness and efficiency of our framework, yielding results comparable to those obtained with datasets a thousand times larger. Our codes are available at https://github.com/Yuyan9Yu/BeyondTextConstraint. Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Shengfeng He, Huaidong Zhang |
CVPR | 6 |
| 2024 | Mask4Align: Aligned Entity Prompting with Color Masks for Multi-Entity Localization ProblemsabstractIn Visual Question Answering (VQA), recognizing and localizing entities pose significant challenges. Pretrained vision-and-language models have addressed this problem by providing a text description as the answer. However, in visual scenes with multiple entities, textual descriptions struggle to distinguish the entities from the same category effectively. Consequently, the VQA dataset is limited by the limitations of text description and cannot adequately cover scenarios involving multiple entities. To address this challenge, we introduce a Mask for Align (Mask4Align) method, which can determine the entity's position in the given image that best matches the user-input question. This method incorporates colored masks into the image, enabling the VQA model to handle discrimination and localization challenges associated with multiple entities. To process an arbitrary number of similar entities, Mask4Align is designed hierarchically to discern subtle differences, achieving precise localization. Since Mask4Align directly utilizes pre-trained models, it does not introduce additional training overhead. Extensive experiments conducted on both the gaze target prediction task dataset and our proposed multi-entity localization dataset showcase the superiority of Mask4Align. Code and dataset are available at https://github.com/haoquanzhang/mask4align. Haoquan Zhang, Ronggang Huang, Huaidong Zhang |
CVPR | 4 |
| 2024 | Beat-It: Beat-Synchronized Multi-condition 3D Dance Generation
Zikai Huang, Xuemiao Xu, Huaidong Zhang, Chenxi Zheng, Harry Qin, Shengfeng He |
ECCV (19) | 4 |
| 2024 | Multi-person Pose Forecasting with Individual Interaction Perceptron and Prior Learning
Xuemiao Xu, Huaidong Zhang |
ECCV (18) | 5 |
| 2024 | rmD4-VTON: Dynamic Semantics Disentangling for Differential Diffusion Based Virtual Try-On
Zhaotong Yang, Zicheng Jiang, Xinzhe Li 0003, Huiyu Zhou 0001, Junyu Dong, Huaidong Zhang, Yong Du 0003 |
ECCV (46) | 6 |
| 2024 | Exploiting Multi-View Clues for Context-Aware Unified Lumbar MRI Identification and DiagnosisabstractLumbar disc herniation, as one of the most common spinal degeneration diseases, significantly affects the quality of people’s lives. Effective identification and diagnosis of this disease is highly demanded and crucial to improve lumbar disc health care. In this paper, we propose a unified framework for diagnosing multiple lumbar degeneration diseases in MRI. Considering the basis of diagnosis is the accurate lumbar identification of vertebrae and discs, we thus tailor an anatomical knowledge-based process to identify the index of the detected vertebrae and discs. Specifically, the main difficulty of diagnosis lies in the accurate classification of the disc degenerative level if only one view of MRI is available. To combat this problem, we introduce multi-view and multi-scale MRI clues to the model learning, and equip our framework with a context-guided multi-view feature fusion module to fully exploit spatial-correlations and semantic-correlations in multi-view MRI, leading to significant improvements of diagnosis. Extensive results on two public datasets demonstrate the superiority of our proposed framework over the existing competitives in terms of lumbar localization, identification, and diagnosis. Xuemiao Xu, Huaidong Zhang, Rongchen Zhao, Harry Qin |
IJCNN | 4 |
| 2024 | VrdONE: One-stage Video Visual Relation DetectionabstractVideo Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional methods for VidVRD, challenged by its complexity, typically split the task into two parts: one for identifying what relation categories are present and another for determining their temporal boundaries. This split overlooks the inherent connection between these elements. Addressing the need to recognize entity pairs' spatiotemporal interactions across a range of durations, we propose VrdONE, a streamlined yet efficacious one-stage model. VrdONE combines the features of subjects and objects, turning predicate detection into 1D instance segmentation on their combined representations. This setup allows for both relation category identification and binary mask generation in one go, eliminating the need for extra steps like proposal generation or post-processing. VrdONE facilitates the interaction of features across various frames, adeptly capturing both short-lived and enduring relations. Additionally, we introduce the Subject-Object Synergy (SOS) module, enhancing how subjects and objects perceive each other before combining. VrdONE achieves state-of-the-art performances on the VidOR benchmark and ImageNet-VidVRD, showcasing its superior capability in discerning relations across different temporal scales. The code is available at https://github.com/lucaspk512/vrdone. Xinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu, Weiying Zheng, Huaidong Zhang, Shengfeng He |
ACM Multimedia | 6 |
| 2024 | Adaptive multi-text union for stable text-to-image synthesis learning
Jiechang Qian, Huaidong Zhang, Xuemiao Xu, Huajie Sun, Fanzhi Zeng, Yuexia Zhou |
Pattern Recognit. | 3 |
| 2024 | GaFL: Geometric-aware Feature Learning for universal 3D models recognition
Huajie Sun, Huaidong Zhang, Xuemiao Xu, Chang'an Yi, Dewang Ye, Yuexia Zhou |
Pattern Recognit. | 3 |
| 2024 | G²Face: High-Fidelity Reversible Face Anonymization via Generative and Geometric PriorsabstractReversible face anonymization, unlike traditional face pixelization, seeks to replace sensitive identity information in facial images with synthesized alternatives, preserving privacy without sacrificing image clarity. Traditional methods, such as encoder-decoder networks, often result in significant loss of facial details due to their limited learning capacity. Additionally, relying on latent manipulation in pre-trained GANs can lead to changes in ID-irrelevant attributes, adversely affecting data utility due to GAN inversion inaccuracies. This paper introduces G2Face, which leverages both generative and geometric priors to enhance identity manipulation, achieving high-quality reversible face anonymization without compromising data utility. We utilize a 3D face model to extract geometric information from the input face, integrating it with a pre-trained GAN-based decoder. This synergy of generative and geometric priors allows the decoder to produce realistic anonymized faces with consistent geometry. Moreover, multi-scale facial features are extracted from the original face and combined with the decoder using our novel identity-aware feature fusion blocks (IFF). This integration enables precise blending of the generated facial patterns with the original ID-irrelevant features, resulting in accurate identity manipulation. Extensive experiments demonstrate that our method outperforms existing state-of-the-art techniques in face anonymization and recovery, while preserving high data utility. Code is available athttps://github.com/Harxis/G2Face. Haoxin Yang, Xuemiao Xu, Huaidong Zhang, Harry Qin, Yi Wang 0017, Pheng-Ann Heng, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Delving Into Important Samples of Semi-Supervised Old Photo Restoration: A New Dataset and MethodabstractThe degradation of printed photographs due to inadequate preservation is a major problem that can be addressed through deep learning-based restoration methods. However, these methods are often limited by their reliance on annotated data, making them less effective for new domains with limited training samples. In this paper, we propose a semi-supervised old photo restoration network that employs a continuous important sample mining strategy. Specifically, we explore the learning potential of limited data from three aspects: correcting imbalanced data distribution, assigning significant pseudo labels, and learning from unlabeled data. First, we coordinate a random mask augmented strategy with the Double-consistency Alignment method to address the unbalanced damaged category (scratched damage is more prevalent than other artifact types). Second, we develop a novel Perceptual-aware Pseudo-label Propagation method that selects initial recovered results as reliable pseudo-labels to continuously expand the sample pool. Lastly, we propose a Damage-augmented Contrastive Learning method that constructs positive, anchor, and negative samples within a semi-supervised framework to mine correlations of unlabeled data more effectively. To evaluate our approach, we introduce the Old Photo Detection Dataset (OPDD) and the Old Photo Restoration Dataset (OPRD), both of which consist of 563 (6,179 augmented) photo pairs recovered by professional artists. Our extensive experiments show that our approach significantly outperforms existing methods. Furthermore, we demonstrate the effectiveness of our approach by training an external old photographic plate restoration network using the deuterogenic old photographic film dataset and obtaining promising results. Huaidong Zhang, Xuemiao Xu, Chenshu Xu, Kun Zhang 0001, Shengfeng He |
IEEE Trans. Multim. | 2 |
| 2023 | Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image RetrievalabstractPrevious Knowledge Distillation based efficient image retrieval methods employ a lightweight network as the stu-dent model for fast inference. However, the lightweight stu-dent model lacks adequate representation capacity for effective knowledge imitation during the most critical early training period, causing final performance degeneration. To tackle this issue, we propose a Capacity Dynamic Distillation framework, which constructs a student model with editable representation capacity. Specifically, the employed student model is initially a heavy model to fruitfully learn distilled knowledge in the early training epochs, and the stu-dent model is gradually compressed during the training. To dynamically adjust the model capacity, our dynamic frame-work inserts a learnable convolutional layer within each residual block in the student model as the channel importance indicator. The indicator is optimized simultaneously by the image retrieval loss and the compression loss, and a retrieval- guided gradient resetting mechanism is proposed to release the gradient conflict. Extensive experiments show that our method has superior inference speed and accu-racy, e.g., on the VeRi-776 dataset, given the ResNet101 as a teacher, our method saves 67.13% model parameters and 65.67% FLOPs without sacrificing accuracy. Code is avail-able at https://github.com/SCY-X/Capacity_Dynamic_Distillation. Huaidong Zhang, Xuemiao Xu, Jianqing Zhu, Shengfeng He |
CVPR | 2 |
| 2023 | Where is My Spot? Few-shot Image Generation via Latent Subspace OptimizationabstractImage generation relies on massive training data that can hardly produce diverse images of an unseen category according to a few examples. In this paper, we address this dilemma by projecting sparse few-shot samples into a continuous latent space that can potentially generate infinite unseen samples. The rationale behind is that we aim to locate a centroid latent position in a conditional StyleGAN, where the corresponding output image on that centroid can maximize the similarity with the given samples. Although the given samples are unseen for the conditional StyleGAN, we assume the neighboring latent subspace around the centroid belongs to the novel category, and therefore introduce two latent subspace optimization objectives. In the first one we use few-shot samples as positive anchors of the novel class, and adjust the StyleGAN to produce the corresponding results with the new class label condition. The second objective is to govern the generation process from the other way around, by altering the centroid and its surrounding latent subspace for a more precise generation of the novel class. These reciprocal optimization objectives inject a novel class into the StyleGAN latent subspace, and therefore new unseen samples can be easily produced by sampling images from it. Extensive experiments demonstrate superior few-shot generation performances compared with state-of-the-art methods, especially in terms of diversity and generation quality. Code is available at https://github.com/chansey0529/LSO. Chenxi Zheng, Bangzhen Liu, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
CVPR | 3 |
| 2023 | Center transfer for supervised domain adaptation
Xiuyu Huang, Nan Zhou 0010, Jian Huang 0010, Huaidong Zhang, Witold Pedrycz, Kup-Sze Choi |
Appl. Intell. | 4 |
| 2023 | EFSCNN: Encoded Feature Sphere Convolution Neural Network for fast non-rigid 3D models classification and retrieval
Zhaolong Dang, Huaidong Zhang, Xuemiao Xu, Harry Qin, Fanzhi Zeng |
Comput. Vis. Image Underst. | 3 |
| 2023 | Contextual-Assisted Scratched Photo RestorationabstractPrinted photographs can be easily warped, wrinkled, and even deteriorated over time. Existing methods treat the restoration of scratches as a pure inpainting problem that neglects the underlying corrupted contextual knowledge. However, important underlying contents are hidden behind the scratches, which are essential hints for producing a semantically consistent result. Motivated by this insight, we explore how to harmonize the scratch-free features and noisy but essential scratch features to produce a visually consistent restoration. Specifically, in this paper, we propose an automatic retouching approach for scratched photographs with the aid of scratch/background context. We explicitly process scratch and background context in two stages. In the first stage, we mainly extract global scratch features, while the mask is introduced in the second stage to filter out and inpaint the scratches. Both contexts are carefully reciprocated for a faithful restoration. Particularly, we propose a Scratch Contextual Assisted Module (SCAM) to adaptively learn texture within the detected mask. This module utilizes the distance between the scratch mask-out feature and scratch encoder feature for modeling the pixel-wise correspondence, which determines the importance of the encoder feature within the scratch mask. Furthermore, to facilitate the evaluation of scratch restoration methods, we create two new scratched photo datasets which have 238 scratch/scratch-free photo pairs to promote the development in the scratch restoration field, namely Old Scratched Photo Dataset (OSPD) and Modern Scratched Photo Dataset (MSPD). Extensive experimental results on the proposed datasets demonstrate that our model outperforms existing methods. To extend the application, we also perform the proposed method on video samples and obtain visual-pleasing results. The code can be found athttps://github.com/cwyyt/Contextual-assisted-Scratched-Photo-Restoration. Huaidong Zhang, Xuemiao Xu, Shengfeng He, Kun Zhang 0001, Harry Qin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Representative Feature Alignment for Adaptive Object DetectionabstractUnsupervised domain adaptation for object detection aims to generalize the object detector trained on the label-rich source domain to the unlabeled target domain. Recently, existing works adopt the instance-level alignment or pixel-level alignment to perform domain transfer, which can effectively avoid the negative transfer due to the diverse background between domains. However, we find that they treat all the regions of an instance feature equally without suppressing background area. They do not segment the specific texture and discriminative regions of objects, which are transferable during adaptation. We call the features that combine the local structure feature and semantic discriminant features as representative features. We propose a novel Representative Feature Alignment (RFA) model to align the features extracted from representative patterns of objects, i.e. representative features, for domain adaptation. Specifically, the representative features are extracted by the Representative Feature Extraction (RFE) submodules. The RFE submodules take the features extracted from different intermediate layers of the detector as input, and filter out the representative features layer-by-layer via integrating class weighting generator, category selection and class activation mapping. Then the representative features from multi-layers are further adaptively aggregated to obtain the final representative features, which are utilized to conduct feature alignment in a class-aware manner. Our representative features are free of untransferable regions and background areas, which leads to better feature alignment. Extensive experimental results show that the proposed model outperforms state-of-the-art methods on a few benchmark datasets. Shan Xu 0006, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Yangyang Xu 0003, Liangui Dai, Kup-Sze Choi, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Panel-Page-Aware Comic Genre UnderstandingabstractUsing a sequence of discrete still images to tell a story or introduce a process has become a tradition in the field of digital visual media. With the surge in these media and the requirements in downstream tasks, acquiring their main topics or genres in a very short time is urgently needed. As a representative form of the media, comic enjoys a huge boom as it has gone digital. However, different from natural images, comic images are divided by panels, and the images are not visually consistent from page to page. Therefore, existing works tailored for natural images perform poorly in analyzing comics. Considering the identification of comic genres is tied to the overall story plotting, a long-term understanding that makes full use of the semantic interactions between multi-level comic fragments needs to be fully exploited. In this paper, we propose [Formula: see text]Comic, a Panel-Page-aware Comic genre classification model, which takes page sequences of comics as the input and produces class-wise probabilities. [Formula: see text]Comic utilizes detected panel boxes to extract panel representations and deploys self-attention to construct panel-page understanding, assisted with interdependent classifiers to model label correlation. We develop the first comic dataset for the task of comic genre classification with multi-genre labels. Our approach is proved by experiments to outperform state-of-the-art methods on related tasks. We also validate the extensibility of our network to perform in the multi-modal scenario. Finally, we show the practicability of our approach by giving effective genre prediction results for whole comic books. Chenshu Xu, Xuemiao Xu, Nanxuan Zhao, Huaidong Zhang, Chengze Li, Xueting Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | SA-DPNet: Structure-aware dual pyramid network for salient object detection
Xuemiao Xu, Huaidong Zhang, Guoqiang Han 0002 |
Pattern Recognit. | 3 |
| 2021 | Learning Semantic Context from Normal Samples for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context. This work presents a Semantic Context based Anomaly Detection Network, SCADN, for unsupervised anomaly detection by learning the semantic context from the normal samples. To achieve this, we first generate multi-scale striped masks to remove a part of regions from the normal samples, and then train a generative adversarial network to reconstruct the unseen regions. Note that the masks are designed in multiple scales and stripe directions, and various training examples are generated to obtain the rich semantic context . In testing, we obtain an error map by computing the difference between the reconstructed image and the input image for all samples, and infer the abnormal samples based on the error maps. Finally, we perform various experiments on three public benchmark datasets and a new dataset LaceAD collected by us, and show that our method clearly outperforms the current state-of-the-art methods. Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 2 |
| 2021 | Object Detection in Densely Packed Scenes via Semi-Supervised Learning with Dual ConsistencyabstractDeep neural networks have been shown to be very powerful tools for object detection in various scenes. Their remarkable performance, however, heavily depends on the availability of a large number of high quality labeled data, which are time-consuming and costly to acquire for scenes with densely packed objects. We present a novel semi-supervised approach to addressing this problem, which is designed based on a common teacher-student model, integrated with a novel intersection-over-union (IoU) aware consistency loss and a new proposal consistency loss. The IoU-aware consistency loss evaluates the IoU over the prediction pairs of the teacher model and the student model, which enforces the prediction of the student model to approach closely to that of the teacher model. The IoU-aware consistency loss also reweights the importance of different prediction pairs to suppress the low-confident pairs. The proposal consistency loss ensures proposal consistency between the two models, making it possible to involve the region proposal network in the training process with unlabeled data. We also construct a new dataset, namely RebarDSC, containing 2,125 rebar images annotated with 350,348 bounding boxes in total (164.9 annotations per image average), to evaluate the proposed method. Extensive experiments are conducted over both the RebarDSC dataset and the famous large public dataset SKU-110K. Experimental results corroborate that the proposed method is able to improve the object detection performance in densely packed scenes, consistently outperforming state-of-the-art approaches. Dataset is available in https://github.com/Armin1337/RebarDSC. Huaidong Zhang, Xuemiao Xu, Harry Qin, Kup-Sze Choi |
IJCAI | 2 |
| 2021 | Fast scene labeling via structural inference
Huaidong Zhang, Chu Han, Xiaodan Zhang 0003, Yong Du 0003, Xuemiao Xu, Guoqiang Han 0002, Harry Qin, Shengfeng He |
Neurocomputing | 1 |
| 2021 | D4Net: De-deformation defect detection network for non-rigid products with large patterns
Xuemiao Xu, Huaidong Zhang, Wing W. Y. Ng |
Inf. Sci. | 3 |
| 2021 | Perceptual-Aware Sketch Simplification Based on Integrated VGG LayersabstractDeep learning has been recently demonstrated as an effective tool for raster-based sketch simplification. Nevertheless, it remains challenging to simplify extremely rough sketches. We found that a simplification network trained with a simple loss, such as pixel loss or discriminator loss, may fail to retain the semantically meaningful details when simplifying a very sketchy and complicated drawing. In this paper, we show that, with a well-designed multi-layer perceptual loss, we are able to obtain aesthetic and neat simplification results preserving semantically important global structures as well as fine details without blurriness and excessive emphasis on local structures. To do so, we design a multi-layer discriminator by fusing all VGG feature layers to differentiate sketches and clean lines. The weights used in layer fusing are automatically learned via an intelligent adjustment mechanism. Furthermore, to evaluate our method, we compare our method to state-of-the-art methods through multiple experiments, including visual comparison and intensive user study. Xuemiao Xu, Minshan Xie, Peiqi Miao, Wenpeng Xiao, Huaidong Zhang, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Context-Aware and Scale-Insensitive Temporal Repetition CountingabstractTemporal repetition counting aims to estimate the number of cycles of a given repetitive action. Existing deep learning methods assume repetitive actions are performed in a fixed time-scale, which is invalid for the complex repetitive actions in real life. In this paper, we tailor a context-aware and scale-insensitive framework, to tackle the challenges in repetition counting caused by the unknown and diverse cycle-lengths. Our approach combines two key insights: (1) Cycle lengths from different actions are unpredictable that require large-scale searching, but, once a coarse cycle length is determined, the variety between repetitions can be overcome by regression. (2) Determining the cycle length cannot only rely on a short fragment of video but a contextual understanding. The first point is implemented by a coarse-to-fine cycle refinement method. It avoids the heavy computation of exhaustively searching all the cycle lengths in the video, and, instead, it propagates the coarse prediction for further refinement in a hierarchical manner. We secondly propose a bidirectional cycle length estimation method for a context-aware prediction. It is a regression network that takes two consecutive coarse cycles as input, and predicts the locations of the previous and next repetitive cycles. To benefit the training and evaluation of temporal repetition counting area, we construct a new and largest benchmark, which contains 526 videos with diverse repetitive actions. Extensive experiments show that the proposed network trained on a single dataset outperforms state-of-the-art methods on several benchmarks, indicating that the proposed framework is general enough to capture repetition patterns across domains. Code and data are available in https://github.com/Xiaodomgdomg/Deep-Temporal-Repetition-Counting. Huaidong Zhang, Xuemiao Xu, Guoqiang Han 0002, Shengfeng He |
CVPR | 1 |
| 2020 | Dual pyramid network for salient object detection
Xuemiao Xu, Huaidong Zhang, Guoqiang Han 0002 |
Neurocomputing | 3 |
| 2020 | Unsupervised Domain Adaptation via Importance SamplingabstractUnsupervised domain adaptation aims to generalize a model from the label-rich source domain to the unlabeled target domain. Existing works mainly focus on aligning the global distribution statistics between source and target domains. However, they neglect distractions from the unexpected noisy samples in domain distribution estimation, leading to domain misalignment or even negative transfer. In this paper, we present an importance sampling method for domain adaptation (ISDA), to measure sample contributions according to their “informative” levels. In particular, informative samples, as well as outliers, can be effectively modeled using feature-norm and prediction entropy of the network. The importance of information is further formulated as the importance sampling losses in features and label spaces. In this way, the proposed model mitigates the noisy outliers while enhancing the important samples during domain alignment. In addition, our model is easy to implement yet effective, and it does not introduce any extra parameters. Extensive experiments on several benchmark datasets show that our method outperforms state-of-the-art methods under both the standard and partial domain adaptation settings. Xuemiao Xu, Hai He, Huaidong Zhang, Yangyang Xu 0003, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Fast User-Guided Single Image Reflection Removal via Edge-Aware Cascaded NetworksabstractTaking photos through a glass window leads to glare or reflection, which might distract the viewer from the scene behind the window. In this paper, we involve user interaction to tackle the ill-posedness of the reflection removal problem. Users are allowed to draw strokes or lassos to indicate the background and reflection layers. Instead of designing hand-crafted features, we propose the edge-aware cascaded networks for reflection removal. The proposed network is a two-stage pipeline. The first stage takes the edge hints converted from user guidance and the image with reflection as input, and then separates the input image into the background and reflection layers. The second stage involves a refinement network to recover the missing details of the background layers. We simulate different types of user guidance, and the networks are trained on simulated data. The cascaded networks are end-to-end and perform with a single feed-forward pass, enabling fast editing. Extensive experimental evaluations demonstrate that the proposed used-guided reflection removal network yields better performance than the state-of-the-art methods on real-world scenarios. Furthermore, we show that novice users can easily generate reflection-free images, and large improvements in reflection removal quality can be obtained in just one minute. Huaidong Zhang, Xuemiao Xu, Hai He, Shengfeng He, Guoqiang Han 0002, Harry Qin, Dapeng Oliver Wu |
IEEE Trans. Multim. | 1 |
| 2019 | Highlight-assisted nighttime vehicle detection using a multi-level fusion network and label hierarchy
Yaoyang Mo, Guoqiang Han 0002, Huaidong Zhang, Xuemiao Xu |
Neurocomputing | 3 |