VLDB 2026 Research / reviewers in the wild / expert
Xuemiao Xu
dblp:74/6722
· DBLP profile ↗
89ranked-venue papers
10as first author
66since 2021 · last 2026
0000-0002-8006-3663ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 6 first-author · 42 since 2021Artificial intelligence and machine learning · 42 · 3 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive DiffusionabstractMonocular 3D human pose estimation remains a challenging task due to inherent depth ambiguities and occlusions. Compared to traditional methods based on Transformers or Convolutional Neural Networks (CNNs), recent diffusion-based approaches have shown superior performance, leveraging their probabilistic nature and high-fidelity generation capabilities. However, these methods often fail to account for the spatial and temporal correlations across predicted frames, resulting in limited temporal consistency and inferior accuracy in predicted 3D pose sequences. To address these shortcomings, this paper proposesStarPose, an autoregressive diffusion framework that effectively incorporates historical 3D pose predictions and spatial-temporal physical guidance to significantly enhance both the accuracy and temporal coherence of pose predictions. Unlike existing approaches,StarPosemodels the 2D-to-3D pose mapping as an autoregressive diffusion process. By synergically integrating previously predicted 3D poses with 2D pose inputs via a Historical Pose Integration Module (HPIM), the framework generates rich and informative historical pose embeddings that guide subsequent denoising steps, ensuring temporally consistent predictions. In addition, a fully plug-and-play Spatial-Temporal Physical Guidance (STPG) mechanism is tailored to refine the denoising process in an iterative manner, which further enforces spatial anatomical plausibility and temporal motion dynamics, rendering robust and realistic pose estimates. Extensive experiments on benchmark datasets demonstrate thatStarPoseoutperforms state-of-the-art methods, achieving superior accuracy and temporal consistency in 3D human pose estimation. Code is available at https://github.com/wileychan/StarPose. Haoxin Yang, Xuemiao Xu, Cuifeng Sun, Shaoyu Huang, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction IdentificationabstractDriver distraction remains a leading cause of traffic accidents, posing a critical threat to road safety globally. As intelligent transportation systems evolve, accurate and real-time identification of driver distraction has become essential. However, existing methods struggle to capture both global contextual and fine-grained local features while contending with noisy labels in training datasets. To address these challenges, we propose DSDFormer, a novel framework that integrates the strengths of Transformer and Mamba architectures through a Dual State Domain Attention (DSDA) mechanism, enabling a balance between long-range dependencies and detailed feature extraction for robust driver behavior recognition. Additionally, we introduce Temporal Reasoning Confident Learning (TRCL), an unsupervised approach that refines noisy labels by leveraging spatiotemporal correlations in video sequences. Beyond achieving state-of-the-art results on AUC-V1, AUC-V2, and 100-Driver datasets, the proposed model is deployable in real-time on embedded platforms such as NVIDIA Jetson AGX Orin and Xavier. Extensive experimental results confirm that DSDFormer and TRCL significantly improve both the accuracy and robustness of driver distraction detection, offering a scalable solution to enhance road safety. Our code has been released athttps://github.com/zhangzr23/driver-noises-learning Junzhou Chen 0001, Heqiang Huang, Xuemiao Xu, Bin Sheng 0001, Hong Yan 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Attribute-Centric Cross-Modal Alignment for Weakly Supervised Text-Based Person Re-IDabstractWeakly supervised text-based person re-identification (Text-ReID) confronts the challenge of matching target person images with textual descriptions, hindered by the absence of identity annotations during training. Traditional approaches, which rely solely on global features, overlook the rich, fine-grained information within both text and image modalities. Besides, merely aligning features at the semantic level is insufficient due to the significant differences in feature representation spaces between the two modalities. Existing methods also neglect the information inequality caused by person-irrelevant factors in images. In this paper, we introduce a novel framework called Attribute-Centric Cross-modal Alignment (ACCA), specifically designed to overcome these issues. Our approach concentrates on two main aspects: visual-text attribute alignment and prediction distribution alignment. To effectively capture fine-grained information without identity labels, we implement a visual-text attribute alignment method based on momentum contrastive learning to synchronize visual and textual attribute features within a unified embedding space. We also propose a unique strategy for negative sample filtering and enrichment, creating robust and comprehensive negative attribute sample spaces to support the attribute alignment. Additionally, we establish two methods of label-free prediction distribution alignment to encourage the learning of invariant feature representations across modalities. The first method, bias-reduction distribution alignment, aligns features and predictions within each text-image pair by utilizing semantic information from the text and reduces the impact of person-irrelevant factors in images. The second method, global-attribute distribution alignment, enhances the interaction between global and local prediction distributions across visual and textual modalities. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets validate our superior performances across all standard benchmarks. Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Multim. | 3 |
| 2026 | Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head AnimationabstractSinging-driven 3D head animation is a compelling yet underexplored task with broad applications in virtual avatars, entertainment, and education. Existing speech-driven approaches, which typically map audio directly to motion through implicit phoneme-to-viseme correspondences, often yield over-smoothed, emotionally flat, and semantically inconsistent results. These limitations render them inadequate for the unique demands of singing-driven animation. To address this challenge, we propose Think2Sing, a unified diffusion-based framework that integrates pretrained large language models to generate semantically consistent and temporally coherent 3D head animations conditioned on both lyrics and acoustics. Central to our framework is the introduction of motion subtitles, a structured, time-aligned representation generated via a Singing Chain-of-Thought process with acoustic-guided retrieval. These subtitles provide region-specific expressive cues that serve as interpretable priors for animation synthesis. We further formulate head animation as motion intensity prediction over key facial regions, enabling fine-grained control and more faithful expressive modeling. To support this paradigm, we construct the first multimodal singing dataset with synchronized 3D motion, acoustic descriptors, and aligned motion subtitles, enabling semantically grounded and expressive motion learning. Extensive experiments demonstrate that Think2Sing significantly outperforms state-of-the-art methods in realism, expressiveness, and emotional fidelity. Furthermore, our framework supports flexible subtitle-conditioned editing, enabling precise and user-controllable animation synthesis. Zikai Huang, Xuemiao Xu, Xiaofen Xing, Harry Qin, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Large model-assisted video summarization via global entity unification and robust importance scoring
Donglei Chen, Shaoyu Huang, Xuemiao Xu, Yongwei Nie, Ping Li 0016, C. L. Philip Chen |
Vis. Comput. | 3 |
| 2025 | FR²Seg: Continual Segmentation Across Multiple Sites via Fourier Style Replay and Adaptive Consistency RegularizationabstractIn clinical imaging, medical segmentation networks typically require continually adapting to new data from multiple sites over time, as aggregating all data for learning at once can be impractical due to storage limitations and privacy concerns. However, existing methods basically overlook domain-specific characteristics and fall short of adequately capturing domain-invariant knowledge during continual learning, leading to undesired catastrophic forgetting of previous sites and inferior generalization to new sites. To tackle this issue, this paper introduces FR2Seg, to sufficiently exploit both domain-specific and domain-invariant knowledge for efficient continual learning with the aid of low-frequency cues. For the former aspect, we propose a Fourier style replay module to synthesize pseudo images with old-site styles for data augmentation during new-site training, effectively preventing catastrophic forgetting without sacrificing data privacy. For the latter, we present a Fourier adaptive consistency regularization to identify and constrain the optimization of domain-invariant parameters with explicit awareness of knowledge transferability across sites, ensuring excellent generalizability to new sites. Experimental results on two public datasets confirm our method's superiority over existing state-of-the-art continual learning methods. Xuemiao Xu, Huaidong Zhang, Harry Qin |
AAAI | 4 |
| 2025 | Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head AnimationabstractSinging is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by singing holds significant research value. Our work focuses on 3D singing facial animation driven by mixed singing audio, and to the best of our knowledge, no prior studies have explored this area. Additionally, the absence of existing 3D singing datasets poses a considerable challenge. To address this, we collect a novel audiovisual dataset, ChorusHead which features synchronized mixed vocal audio and 3D motions for chorus singing. In addition, We propose a partner-aware 3D chorus head generation framework driven by mixed audio inputs. The proposed framework extracts emotional features from the background music and dependence between singers and models the head movement in a latent space from the Variational Autoencoder (VAE), enabling diverse interactive head animation generation. Extensive experimental results demonstrate that our approach effectively generates 3D facial animations of interacting singers, achieving notable improvements in realism and handling background music interference with strong robustness. The dataset will be released for research purposes at the project page: https://xxiexm.github.io/PaChorus/. Xiumei Xie, Zikai Huang, Xuemiao Xu, Huaidong Zhang |
CVPR | 5 |
| 2025 | EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual InsightsabstractTraffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints to anomaly scenarios such as crashes and honking. Our contributions are twofold. First, we compile AV-TAU, the first large-scale audio-visual dataset for TAU, providing 29,865 traffic anomaly videos and 149,325 Q&A pairs, while supporting five essential TAU tasks. Second, we develop EchoTraffic, a multimodal LLM that integrates audio and visual data for TAU, through our audio-insight frame selector and dynamic connector to effectively extract crucial audio cues for anomaly understanding with a two-phase training framework. Experimental results on AV-TAU manifest that EchoTraffic sets a new SOTA performance in TAU, outperforming the existing multimodal LLMs. Our contributions, including AV-TAU and EchoTraffic, pave a new direction for multimodal TAU. Zhenghao Xing, Hao Chen 0193, Binzhu Xie, Xuemiao Xu, Jianye Hao, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng |
CVPR | 6 |
| 2025 | Negative Learning and Dual Contrastive for Unsupervised Visible-Infrared Person Re-identificationabstractUnsupervised visible-infrared person re-identification (US-VI-ReID) aims to identify target person images from different modalities without requiring annotations. Existing works generally learn modality-invariant features by using pseudo-labels. However, the inherent noise in these labels misleads the network’s training. Besides, these methods neglect more fine-grained information. To address these issues, we proposed a Negative Learning and Dual Contrastive (NLDC) framework to efficiently learn invariant feature representations. Specifically, we introduced a cross-modal negative learning method with complementary labels to robustly handle label noise. We also designed a low-similarity label selection strategy to construct reliable complementary label sets, which support the negative learning process. Additionally, we introduce both intra-modality and inter-modality instance contrastive losses to achieve fine-grained feature alignment. Experimental results on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our proposed method, achieving competitive performance. Xuemiao Xu |
ICASSP | 2 |
| 2025 | ViewSRD: 3D Visual Grounding Via Structured Multi-View Decomposition
Ronggang Huang, Haoxin Yang, Yan Cai 0021, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
ICCV | 4 |
| 2025 | Seeing 3D Through 2D Lenses: 3D Few-Shot Class-Incremental Learning via Cross-Modal Geometric RectificationabstractThe rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP's hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically aligns 3D part structures with CLIP's intermediate spatial priors through attention-driven geometric fusion. Additionally, a Texture Amplification Module synthesizes minimal yet discriminative textures to suppress noise and reinforce cross-modal consistency. To further stabilize incremental prototypes, we employ a Base-Novel Discriminator that isolates geometric variations. Extensive experiments demonstrate that our method significantly improves 3D few-shot class-incremental learning, achieving superior geometric coherence and robustness to texture bias across cross-domain and within-domain settings. Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Shengfeng He |
ICCV | 2 |
| 2025 | RecDreamer: Consistent Text-to-3D Generation via Uniform Score DistillationabstractCurrent text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises because existing methods struggle to maintain consistency across varying poses and are biased toward a canonical pose. While recent work has improved pose control and approximation, these efforts are still limited by this inherent bias, which skews the guidance during generation.
To address this, we propose a solution called RecDreamer, which reshapes the underlying data distribution to achieve more consistent pose representation. The core idea behind our method is to rectify the prior distribution, ensuring that pose variation is uniformly distributed rather than biased toward a canonical form. By modifying the prescribed distribution through an auxiliary function, we can reconstruct the density of the distribution to ensure compliance with specific marginal constraints. In particular, we ensure that the marginal distribution of poses follows a uniform distribution, thereby eliminating the biases introduced by the prior knowledge.
We incorporate this rectified data distribution into existing score distillation algorithms, a process we refer to as uniform score distillation. To efficiently compute the posterior distribution required for the auxiliary function, RecDreamer introduces a training-free classifier that estimates pose categories in a plug-and-play manner. Additionally, we utilize various approximation techniques for noisy states, significantly improving system performance.
Our experimental results demonstrate that RecDreamer effectively mitigates the Multi-Face Janus problem, leading to more consistent 3D asset generation across different poses. Chenxi Zheng, Yihong Lin, Bangzhen Liu, Xuemiao Xu, Yongwei Nie, Shengfeng He |
ICLR | 4 |
| 2025 | SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose EstimationabstractExisting 3D Human Pose Estimation (HPE) methods achieve high accuracy but suffer from computational overhead and slow inference, while knowledge distillation methods fail to address spatial relationships between joints and temporal correlations in multi-frame inputs. In this paper, we propose Sparse Correlation and Joint Distillation (SCJD), a novel framework that balances efficiency and accuracy for 3D HPE. SCJD introduces Sparse Correlation Input Sequence Downsampling to reduce redundancy in student network inputs while preserving inter-frame correlations. For effective knowledge transfer, we propose Dynamic Joint Spatial Attention Distillation, which includes Dynamic Joint Embedding Distillation to enhance the student’s feature representation using the teacher’s multi-frame context feature, and Adjacent Joint Attention Distillation to improve the student network’s focus on adjacent joint relationships for better spatial understanding. Additionally, Temporal Consistency Distillation aligns the temporal correlations between teacher and student networks through upsampling and global supervision. Extensive experiments demonstrate that SCJD achieves state-of-the-art performance. Code is available at https://github.com/wileychan/SCJD. Xuemiao Xu, Haoxin Yang, Huaidong Zhang, Pheng-Ann Heng |
ICME | 2 |
| 2025 | StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion ModelsabstractThe advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods have made progress toward unified solutions, their reliance on post hoc processing introduces considerable application inconvenience and compromises forensic reliability. We propose StableGuard, a novel framework that seamlessly integrates a binary watermark into the diffusion generation process, ensuring copyright protection and tampering localization in Latent Diffusion Models through an end-to-end design. We develop a Multiplexing Watermark VAE (MPW-VAE) by equipping a pretrained Variational Autoencoder (VAE) with a lightweight latent residual-based adapter, enabling the generation of paired watermarked and watermark-free images. These pairs, fused via random masks, create a diverse dataset for training a tampering-agnostic forensic network. To further enhance forensic synergy, we introduce a Mixture-of-Experts Guided Forensic Network (MoE-GFN) that dynamically integrates holistic watermark patterns, local tampering traces, and frequency-domain cues for precise watermark verification and tampered region detection. The MPW-VAE and MoE-GFN are jointly optimized in a self-supervised, end-to-end manner, fostering a reciprocal training between watermark embedding and forensic accuracy. Extensive experiments demonstrate that StableGuard consistently outperforms state-of-the-art methods in image fidelity, watermark verification, and tampering localization. Haoxin Yang, Bangzhen Liu, Xuemiao Xu, Yuyang Yu, Zikai Huang, Shengfeng He |
NeurIPS | 3 |
| 2025 | Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detectionabstract3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives of point-cloud registration and memory-based anomaly detection. Our key insight is that both tasks rely on modeling local geometric structures and leveraging feature similarity across samples. By embedding feature extraction into the registration learning process, our framework jointly optimizes alignment and representation learning. This integration enables the network to acquire features that are both robust to rotations and highly effective for anomaly detection. Extensive experiments on the Anomaly-ShapeNet and Real3D-AD datasets demonstrate that our method consistently outperforms existing approaches in effectiveness and generalizability. Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang 0006, Haoxin Yang, Yongwei Nie, Shengfeng He |
NeurIPS | 3 |
| 2025 | Real-Time Smoke Detection With Split Top-K Transformer and Adaptive Dark Channel Prior in Foggy EnvironmentsabstractSmoke detection is essential for fire prevention, yet it is significantly hampered by the visual similarities between smoke and fog. To address this challenge, a split top-k attention transformer framework (STKformer) is proposed. The STKformer incorporates split top-k attention (STKA), which partitions the attention map for top-k selection to retain informative self-attention values while capturing long-range dependencies. This approach effectively filters out irrelevant attention scores, preventing information loss. Furthermore, the adaptive dark-channel-prior guidance network (ADGN) is designed to enhance smoke recognition under foggy conditions. ADGN employs pooling operations instead of minimum value filtering, allowing for efficient dark channel extraction with learnable parameters and adaptively reducing the impact of fog. The extracted prior information subsequently guides feature extraction through a priorformer block, improving model robustness. Additionally, a cross-stage fusion module (CSFM) is introduced to aggregate features from different stages efficiently, enabling flexible adaptation to smoke features at various scales and enhancing detection accuracy. Comprehensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple datasets, with an accuracy of 89.68% on dataset for smoke detection in fog, 99.76% on CCTV images of smoke, and 99.76% on UAV images of wildfire. The method maintains high speed and lightweight characteristics, validated with an inference speed of 211.46 FPS on an NVIDIA Jetson AGX Orin after TensorRT acceleration, confirming its effectiveness and efficiency for real-world applications. The source code is available athttps://github.com/Jiongze-Yu/STKformerhttps://github.com/Jiongze-Yu/STKformer. Jiongze Yu, Heqiang Huang, Yuhang Ma 0002, Yueying Wu 0001, Junzhou Chen 0001, Xuemiao Xu, Zhihan Lyu, Guodong Yin |
IEEE Internet Things J. | 7 |
| 2025 | L3Net: Localized and Layered Reparameterization for incremental learning
Xuandi Luo, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
Neural Networks | 5 |
| 2025 | Unambiguous granularity distillation for asymmetric image retrieval
Haoquan Zhang, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng-Ann Heng, Shengfeng He |
Neural Networks | 7 |
| 2025 | Tgcpn: two-level grid context propagation network for 3D small object detection
Lei Pu, Xuemiao Xu, Chang'an Yi, Yuexia Zhou, Yewen Xu |
Pattern Anal. Appl. | 3 |
| 2025 | Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation LearningabstractThe vulnerability of 3D point cloud analysis to unpredictable rotations poses an open yet challenging problem: orientation-aware 3D domain generalization. Cross-domain robustness and adaptability of 3D representations are crucial but not easily achieved through rotation augmentation. Motivated by the inherent advantages of intricate orientations in enhancing generalizability, we propose an innovative rotation-adaptive domain generalization framework for 3D point cloud analysis. Our approach aims to alleviate orientational shifts by leveraging intricate samples in an iterative learning process. Specifically, we identify the most challenging rotation for each point cloud and construct an intricate orientation set by optimizing intricate orientations. Subsequently, we employ an orientation-aware contrastive learning framework that incorporates an orientation consistency loss and a margin separation loss, enabling effective learning of categorically discriminative and generalizable features with rotation consistency. Extensive experiments and ablations conducted on 3D cross-domain benchmarks firmly establish the state-of-the-art performance of our proposed approach in the context of orientation-aware 3D domain generalization. Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Gaussian Prompter: Linking 2D Prompts for 3D Gaussian SegmentationabstractInteractive 3D segmentation in radiance fields is crucial for advanced 3D scene understanding and manipulation. However, existing methods often struggle to achieve both volumetric completeness and segmentation accuracy, primarily because they fail to consider the critical links between 2D prompt-based segmentations across multiple views. Motivated by this gap, we introduce Gaussian Prompter, a novel approach specifically designed for 3D Gaussian Splatting. The core idea behind Gaussian Prompter is to seamlessly integrate a Gaussian-centric segmentation paradigm by effectively linking various 2D prompts from multi-view segmentations to ensure consistent 3D segmentation. To realize this, we employ two tailored approaches: GaussBlend and PinPrompt. GaussBlend aggregates multi-view 2D segmentation masks into a cohesive 3D segmentation, ensuring both accuracy and completeness. PinPrompt leverages high-confidence prompts from adjacent views to enhance segmentation precision further. Additionally, to address the lack of complex datasets in 3D segmentation, we introduce the SegMip-360 dataset, which includes over 350 precisely annotated masks across seven scenes. Extensive experiments demonstrate that the Gaussian Prompter significantly outperforms state-of-the-art methods in both segmentation accuracy and completeness. Our code and video demonstrations can be found at our repository and project page. Honghan Pan, Bangzhen Liu, Xuemiao Xu, Chenxi Zheng, Yongwei Nie, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Open-Set Mixed Domain Adaptation via Visual-Linguistic Focal EvolvingabstractWe introduce a new task, Open-set Mixed Domain Adaptation (OSMDA), which considers the potential mixture of multiple distributions in the target domains, thereby better simulating real-world scenarios. To tackle the semantic ambiguity arising from multiple domains, our key idea is that the linguistic representation can serve as a universal descriptor for samples of the same category across various domains. We thus propose a more practical framework for cross-domain recognition via visual-linguistic guidance. On the other hand, the presence of multiple domains also poses a new challenge in classifying both known and unknown categories. To combat this issue, we further introduce a visual-linguistic focal evolving approach to gradually enhance the classification ability of a known/unknown binary classifier from two aspects. Specifically, we start with identifying highly confident focal samples to expand the pool of known samples by incorporating those from different domains. Then, we amplify the feature discrepancy between known and unknown samples through dynamic entropy evolving via an adaptive entropies min/max game, enabling us to accurately identify possible unknown samples in a gradual manner. Extensive experiments demonstrate our method’s superiority against the state-of-the-arts in both open-set and open-set mixed domain adaptation. Bangzhen Liu, Yangyang Xu 0003, Xuemiao Xu, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SITA: Structurally Imperceptible and Transferable Adversarial Attacks for Stylized Image Generation
Jingdan Kang, Haoxin Yang, Yan Cai 0021, Huaidong Zhang, Xuemiao Xu, Yong Du 0003, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Recurrent Diffusion for 3D Point Cloud Generation From a Single ImageabstractSingle-image 3D shape reconstruction has attracted significant attention with the advance of generative models. Recent studies have utilized diffusion models to achieve unprecedented shape reconstruction quality. However, these methods, in each sampling step, perform denoising in a single forward pass, leading to cumulative errors that severely impact the geometric consistency of the generated shapes with the input targets and face difficulties in reconstructing rich details of complex 3D shapes. Moreover, the performance of current works suffers significant degradation due to limited information when only a single image is used as input during testing, further affecting the quality of 3D shape generation. In this paper, we present a recurrent diffusion framework, aiming to improve generation performance during single image-to-shape inference. Diverging from denoising in a single forward pass, we recursively refine the noise prediction in a self-rectified manner with the explicit guidance of the input target, thereby markedly suppressing cumulative errors and improving detail modeling. To enhance the geometric perception ability of the network during single-image inference, we further introduce a multi-view training scheme equipped with a view-robust conditional generation mechanism, which effectively promotes generation quality even when only a single image is available during inference. The effectiveness of our method is demonstrated through extensive evaluations on two public 3D shape datasets, where it surpasses state-of-the-art methods both qualitatively and quantitatively. Dewang Ye, Huaidong Zhang, Xuemiao Xu, Huajie Sun, Yewen Xu, Yuexia Zhou |
IEEE Trans. Image Process. | 4 |
| 2025 | IM-Diff: Implicit Multi-Contrast Diffusion Model for Arbitrary Scale MRI Super-ResolutionabstractDiffusion models have garnered significant attention for MRI Super-Resolution (SR) and have achieved promising results. However, existing diffusion-based SR models face two formidable challenges: 1) insufficient exploitation of complementary information from multi-contrast images, which hinders the faithful reconstruction of texture details and anatomical structures; and 2) reliance on fixed magnification factors, such as 2× or 4×, which is impractical for clinical scenarios that require arbitrary scale magnification. To circumvent these issues, this paper introduces IM-Diff, an implicit multi-contrast diffusion model for arbitrary-scale MRI SR, leveraging the merits of both multi-contrast information and the continuous nature of implicit neural representation (INR). Firstly, we propose an innovative hierarchical multi-contrast fusion (HMF) module with reference-aware cross Mamba (RCM) to effectively incorporate target-relevant information from the reference image into the target image, while ensuring a substantial receptive field with computational efficiency. Secondly, we introduce multiple wavelet INR magnification (WINRM) modules into the denoising process by integrating the wavelet implicit neural non-linearity, enabling effective learning of continuous representations of MR images. The involved wavelet activation enhances space-frequency concentration, further bolstering representation accuracy and robustness in INR. Extensive experiments on three public datasets demonstrate the superiority of our method over existing state-of-the-art SR models across various magnification factors. Lanqing Liu, Kang Wang 0004, Xuemiao Xu, Zhanli Hu, Harry Qin |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Delving Into Invisible Semantics for Generalized One-Shot Neural Human RenderingabstractTraditional human neural radiance fields often overlook crucial body semantics, resulting in ambiguous reconstructions, particularly in occluded regions. To address this problem, we propose the Super-Semantic Disentangled Neural Renderer (SSD-NeRF), which employs rich regional semantic priors to enhance human rendering accuracy. This approach initiates with a Visible-Invisible Semantic Propagation module, ensuring coherent semantic assignment to occluded parts based on visible body segments. Furthermore, a Region-Wise Texture Propagation module independently extends textures from visible to occluded areas within semantic regions, thereby avoiding irrelevant texture mixtures and preserving semantic consistency. Additionally, a view-aware curricular learning approach is integrated to bolster the model's robustness and output quality across different viewpoints. Extensive evaluations confirm that SSD-NeRF surpasses leading methods, particularly in generating quality and structurally semantic reconstructions of unseen or occluded views and poses. Yihong Lin, Xuemiao Xu, Huaidong Zhang, Harry Qin, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | DreamAnime: Learning Style-Identity Textual Disentanglement for Anime and BeyondabstractText-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character or an art style, we teach our model to encapsulate its essence through novel "words" in the embedding space of a pre-existing text-to-image model. Crucially, we disentangle the concepts of style and identity into two separate "words", thus providing the ability to manipulate them independently. These distinct "words" can then be pieced together into natural language sentences, promoting an intuitive and personalized creative process. Empirical results suggest that this disentanglement into separate word embeddings successfully captures a broad range of unique and complex concepts, with each word focusing on style or identity as appropriate. Comparisons with existing methods illustrate DreamAnime's superior capacity to accurately interpret and recreate the desired concepts across various applications and tasks. Chenshu Xu, Yangyang Xu 0003, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Learning an Interpretable Stylized Subspace for 3D-Aware Animatable ArtformsabstractThroughout history, static paintings have captivated viewers within display frames, yet the possibility of making these masterpieces vividly interactive remains intriguing. This research paper introduces 3DArtmator, a novel approach that aims to represent artforms in a highly interpretable stylized space, enabling 3D-aware animatable reconstruction and editing. Our rationale is to transfer the interpretability and 3D controllability of the latent space in a 3D-aware GAN to a stylized sub-space of a customized GAN, revitalizing the original artforms. To this end, the proposed two-stage optimization framework of 3DArtmator begins with discovering an anchor in the original latent space that accurately mimics the pose and content of a given art painting. This anchor serves as a reliable indicator of the original latent space local structure, therefore sharing the same editable predefined expression vectors. In the second stage, we train a customized 3D-aware GAN specific to the input artform, while enforcing the preservation of the original latent local structure through a meticulous style-directional difference loss. This approach ensures the creation of a stylized sub-space that remains interpretable and retains 3D control. The effectiveness and versatility of 3DArtmator are validated through extensive experiments across a diverse range of art styles. With the ability to generate 3D reconstruction and editing for artforms while maintaining interpretability, 3DArtmator opens up new possibilities for artistic exploration and engagement. Chenxi Zheng, Bangzhen Liu, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Face Expression Recognition via Product-Cross Dual Attention and Neutral-Aware Anchor Loss
Yongwei Nie, Qing Zhang 0006, Xuemiao Xu, Guiqing Li, Hongmin Cai |
CVM (2) | 4 |
| 2024 | D3still: Decoupled Differential Distillation for Asymmetric Image RetrievalabstractExisting methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these one-to-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is then decomposed into three components: feature representation knowledge, inconsistent pairwise similarity differential knowledge, and consistent pairwise similarity differential knowledge. This strategic decomposition aligns the retrieval ranking of the query network with the gallery network effectively. Extensive experiments on various bench-mark datasets reveal that D3still surpasses state-of-the-art methods in asymmetric image retrieval. Code is available at https://github.com/SCY-X/D3still. Yihong Lin, Xuemiao Xu, Huaidong Zhang, Yong Du 0003, Shengfeng He |
CVPR | 4 |
| 2024 | Beyond Textual Constraints: Learning Novel Diffusion Conditions with Fewer ExamplesabstractIn this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which addresses the inconsistencies typically brought about by Classifier-free guidance in few-shot training con-texts. Our extensive experiments across a variety of condition modalities demonstrate the effectiveness and efficiency of our framework, yielding results comparable to those obtained with datasets a thousand times larger. Our codes are available at https://github.com/Yuyan9Yu/BeyondTextConstraint. Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Shengfeng He, Huaidong Zhang |
CVPR | 4 |
| 2024 | Beat-It: Beat-Synchronized Multi-condition 3D Dance Generation
Zikai Huang, Xuemiao Xu, Huaidong Zhang, Chenxi Zheng, Harry Qin, Shengfeng He |
ECCV (19) | 2 |
| 2024 | Multi-person Pose Forecasting with Individual Interaction Perceptron and Prior Learning
Xuemiao Xu, Huaidong Zhang |
ECCV (18) | 3 |
| 2024 | Exploiting Multi-View Clues for Context-Aware Unified Lumbar MRI Identification and DiagnosisabstractLumbar disc herniation, as one of the most common spinal degeneration diseases, significantly affects the quality of people’s lives. Effective identification and diagnosis of this disease is highly demanded and crucial to improve lumbar disc health care. In this paper, we propose a unified framework for diagnosing multiple lumbar degeneration diseases in MRI. Considering the basis of diagnosis is the accurate lumbar identification of vertebrae and discs, we thus tailor an anatomical knowledge-based process to identify the index of the detected vertebrae and discs. Specifically, the main difficulty of diagnosis lies in the accurate classification of the disc degenerative level if only one view of MRI is available. To combat this problem, we introduce multi-view and multi-scale MRI clues to the model learning, and equip our framework with a context-guided multi-view feature fusion module to fully exploit spatial-correlations and semantic-correlations in multi-view MRI, leading to significant improvements of diagnosis. Extensive results on two public datasets demonstrate the superiority of our proposed framework over the existing competitives in terms of lumbar localization, identification, and diagnosis. Xuemiao Xu, Huaidong Zhang, Rongchen Zhao, Harry Qin |
IJCNN | 3 |
| 2024 | VrdONE: One-stage Video Visual Relation DetectionabstractVideo Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional methods for VidVRD, challenged by its complexity, typically split the task into two parts: one for identifying what relation categories are present and another for determining their temporal boundaries. This split overlooks the inherent connection between these elements. Addressing the need to recognize entity pairs' spatiotemporal interactions across a range of durations, we propose VrdONE, a streamlined yet efficacious one-stage model. VrdONE combines the features of subjects and objects, turning predicate detection into 1D instance segmentation on their combined representations. This setup allows for both relation category identification and binary mask generation in one go, eliminating the need for extra steps like proposal generation or post-processing. VrdONE facilitates the interaction of features across various frames, adeptly capturing both short-lived and enduring relations. Additionally, we introduce the Subject-Object Synergy (SOS) module, enhancing how subjects and objects perceive each other before combining. VrdONE achieves state-of-the-art performances on the VidOR benchmark and ImageNet-VidVRD, showcasing its superior capability in discerning relations across different temporal scales. The code is available at https://github.com/lucaspk512/vrdone. Xinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu, Weiying Zheng, Huaidong Zhang, Shengfeng He |
ACM Multimedia | 3 |
| 2024 | Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryabstractHuman Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pretrained model may not provide an ideal optimization starting point for the test-time optimization. Inspired by meta-learning, we incorporate the test-time optimization into training, performing a step of test-time optimization for each sample in the training batch before really conducting the training optimization over all the training samples. In this way, we obtain a meta-model, the meta-parameter of which is friendly to the test-time optimization. At test time, after several test-time optimization steps starting from the meta-parameter, we obtain much higher HMR accuracy than the test-time optimization starting from the simply pretrained regression model. Furthermore, we find test-time HMR objectives are different from training-time objectives, which reduces the effectiveness of the learning of the meta-model. To solve this problem, we propose a dual-network architecture that unifies the training-time and test-time objectives. Our method, armed with meta-learning and the dual networks, outperforms state-of-the-art regression-based and optimization-based HMR approaches, as validated by the extensive experiments. The codes are available at https://github.com/fmx789/Meta-HMR. Yongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang 0006, Jian Zhu 0001, Xuemiao Xu |
NeurIPS | 6 |
| 2024 | Adaptive multi-text union for stable text-to-image synthesis learning
Jiechang Qian, Huaidong Zhang, Xuemiao Xu, Huajie Sun, Fanzhi Zeng, Yuexia Zhou |
Pattern Recognit. | 4 |
| 2024 | GaFL: Geometric-aware Feature Learning for universal 3D models recognition
Huajie Sun, Huaidong Zhang, Xuemiao Xu, Chang'an Yi, Dewang Ye, Yuexia Zhou |
Pattern Recognit. | 4 |
| 2024 | PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised LearningabstractRecently, Point-MAE has extended Masked Autoencoders (MAE) to point clouds for 3D self-supervised learning, which however faces two problems: (1) the shape similarity between the masked point cloud and original point cloud is high, and (2) the pretext task of reconstructing the original point cloud is straightforward which fails to compel the network to learn deep representative features. In this paper, we tackle these problems by proposing a PatchMixing strategy and a teacher-student training framework. First, with PatchMixing, we mix selected point patches of multiple point clouds and attempt to infer the object information from the resulting mixed point cloud. Due to the interference of other objects, the task is challenging but facilitates representation learning. Second, rather than directly restoring the original point cloud, we propose a novel pretext task that involves a two-branch teacher model and a student model. These models process the multiple input point clouds in different ways (no mixing, mixing + unmixing, mixing + masking), but are expected to output similar features, thereby compelling the network to extract essential features from the input. Extensive experiments show that our well-designed PatchMixing strategy and effective teacher-student learning architecture yield impressive results. Specifically, our model achieves a remarkable 92.9% classification accuracy in the Linear SVM task on the ModelNet40 dataset. Through pre-training and fine-tuning on downstream tasks, our method achieves an 89.8% classification accuracy on the most challenging split of ScanObjectNN and an outstanding 94.0% on ModelNet40. Chengxing Lin 0001, Wenju Xu, Jian Zhu 0001, Yongwei Nie, Ruichu Cai, Xuemiao Xu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | G²Face: High-Fidelity Reversible Face Anonymization via Generative and Geometric PriorsabstractReversible face anonymization, unlike traditional face pixelization, seeks to replace sensitive identity information in facial images with synthesized alternatives, preserving privacy without sacrificing image clarity. Traditional methods, such as encoder-decoder networks, often result in significant loss of facial details due to their limited learning capacity. Additionally, relying on latent manipulation in pre-trained GANs can lead to changes in ID-irrelevant attributes, adversely affecting data utility due to GAN inversion inaccuracies. This paper introduces G2Face, which leverages both generative and geometric priors to enhance identity manipulation, achieving high-quality reversible face anonymization without compromising data utility. We utilize a 3D face model to extract geometric information from the input face, integrating it with a pre-trained GAN-based decoder. This synergy of generative and geometric priors allows the decoder to produce realistic anonymized faces with consistent geometry. Moreover, multi-scale facial features are extracted from the original face and combined with the decoder using our novel identity-aware feature fusion blocks (IFF). This integration enables precise blending of the generated facial patterns with the original ID-irrelevant features, resulting in accurate identity manipulation. Extensive experiments demonstrate that our method outperforms existing state-of-the-art techniques in face anonymization and recovery, while preserving high data utility. Code is available athttps://github.com/Harxis/G2Face. Haoxin Yang, Xuemiao Xu, Huaidong Zhang, Harry Qin, Yi Wang 0017, Pheng-Ann Heng, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Delving Into Important Samples of Semi-Supervised Old Photo Restoration: A New Dataset and MethodabstractThe degradation of printed photographs due to inadequate preservation is a major problem that can be addressed through deep learning-based restoration methods. However, these methods are often limited by their reliance on annotated data, making them less effective for new domains with limited training samples. In this paper, we propose a semi-supervised old photo restoration network that employs a continuous important sample mining strategy. Specifically, we explore the learning potential of limited data from three aspects: correcting imbalanced data distribution, assigning significant pseudo labels, and learning from unlabeled data. First, we coordinate a random mask augmented strategy with the Double-consistency Alignment method to address the unbalanced damaged category (scratched damage is more prevalent than other artifact types). Second, we develop a novel Perceptual-aware Pseudo-label Propagation method that selects initial recovered results as reliable pseudo-labels to continuously expand the sample pool. Lastly, we propose a Damage-augmented Contrastive Learning method that constructs positive, anchor, and negative samples within a semi-supervised framework to mine correlations of unlabeled data more effectively. To evaluate our approach, we introduce the Old Photo Detection Dataset (OPDD) and the Old Photo Restoration Dataset (OPRD), both of which consist of 563 (6,179 augmented) photo pairs recovered by professional artists. Our extensive experiments show that our approach significantly outperforms existing methods. Furthermore, we demonstrate the effectiveness of our approach by training an external old photographic plate restoration network using the deuterogenic old photographic film dataset and obtaining promising results. Huaidong Zhang, Xuemiao Xu, Chenshu Xu, Kun Zhang 0001, Shengfeng He |
IEEE Trans. Multim. | 3 |
| 2024 | Fully Deformable Network for Multiview Face Image SynthesisabstractPhotorealistic multiview face synthesis from a single image is a challenging problem. Existing works mainly learn a texture mapping model from the source to the target faces. However, they rarely consider the geometric constraints on the internal deformation arising from pose variations, which causes a high level of uncertainty in face pose modeling, and hence, produces inferior results for large pose variations. Moreover, current methods typically suffer from undesired facial details loss due to the adoption of the de-facto standard encoder-decoder architecture without any skip connections (SCs). In this article, we directly learn and exploit geometric constraints and propose a fully deformable network to simultaneously model the deformations of both landmarks and faces for face synthesis. Specifically, our model consists of two parts: a deformable landmark learning network (DLLN) and a gated deformable face synthesis network (GDFSN). The DLLN converts an initial reference landmark to an individual-specific target landmark as delicate pose guidance for face rotation. The GDFSN adopts a dual-stream structure, with one stream estimating the deformation of two views in the form of convolution offsets according to the source pose and the converted target pose, and the other leveraging the predicted deformation offsets to create the target face. In this way, individual-aware pose changes are explicitly modeled in the face generator to cope with geometric transformation, by adaptively focusing on pertinent regions of the source face. To compensate for offset estimation errors, we introduce a soft-gating mechanism for adaptive fusion between deformable features and primitive features. Additionally, a pose-aligned SC (PASC) is tailored to propagate low-level input features to the appropriate positions in the output features for further enhancing the facial details and identity preservation. Extensive experiments on six benchmarks show that our approach performs favorably against the state-of-the-arts, especially with large pose changes. Code is available at https://github.com/cschengxu/FDFace. Xuandi Luo, Xuemiao Xu, Shengfeng He, Kun Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Appearance-Preserved Portrait-to-Anime Translation via Proxy-Guided Domain AdaptationabstractConverting a human portrait to anime style is a desirable but challenging problem. Existing methods fail to resolve this problem due to the large inherent gap between two domains that cannot be overcome by a simple direct mapping. For this reason, these methods struggle to preserve the appearance features in the original photo. In this article, we discover an intermediate domain, the coser portrait (portraits of humans costuming as anime characters), that helps bridge this gap. It alleviates the learning ambiguity and loosens the mapping difficulty in a progressive manner. Specifically, we start from learning the mapping between coser and anime portraits, and present a proxy-guided domain adaptation learning scheme with three progressive adaptation stages to shift the initial model to the human portrait domain. In this way, our model can generate visually pleasant anime portraits with well-preserved appearances given the human portrait. Our model adopts a disentangled design by breaking down the translation problem into two specific subtasks of face deformation and portrait stylization. This further elevates the generation quality. Extensive experimental results show that our model can achieve visually compelling translation with better appearance preservation and perform favorably against the existing methods both qualitatively and quantitatively. Our code and datasets are available at https://github.com/NeverGiveU/PDA-Translation. Wenpeng Xiao, Jiajie Mai, Xuemiao Xu, Chengze Li, Xueting Liu 0001, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image RetrievalabstractPrevious Knowledge Distillation based efficient image retrieval methods employ a lightweight network as the stu-dent model for fast inference. However, the lightweight stu-dent model lacks adequate representation capacity for effective knowledge imitation during the most critical early training period, causing final performance degeneration. To tackle this issue, we propose a Capacity Dynamic Distillation framework, which constructs a student model with editable representation capacity. Specifically, the employed student model is initially a heavy model to fruitfully learn distilled knowledge in the early training epochs, and the stu-dent model is gradually compressed during the training. To dynamically adjust the model capacity, our dynamic frame-work inserts a learnable convolutional layer within each residual block in the student model as the channel importance indicator. The indicator is optimized simultaneously by the image retrieval loss and the compression loss, and a retrieval- guided gradient resetting mechanism is proposed to release the gradient conflict. Extensive experiments show that our method has superior inference speed and accu-racy, e.g., on the VeRi-776 dataset, given the ResNet101 as a teacher, our method saves 67.13% model parameters and 65.67% FLOPs without sacrificing accuracy. Code is avail-able at https://github.com/SCY-X/Capacity_Dynamic_Distillation. Huaidong Zhang, Xuemiao Xu, Jianqing Zhu, Shengfeng He |
CVPR | 3 |
| 2023 | Where is My Spot? Few-shot Image Generation via Latent Subspace OptimizationabstractImage generation relies on massive training data that can hardly produce diverse images of an unseen category according to a few examples. In this paper, we address this dilemma by projecting sparse few-shot samples into a continuous latent space that can potentially generate infinite unseen samples. The rationale behind is that we aim to locate a centroid latent position in a conditional StyleGAN, where the corresponding output image on that centroid can maximize the similarity with the given samples. Although the given samples are unseen for the conditional StyleGAN, we assume the neighboring latent subspace around the centroid belongs to the novel category, and therefore introduce two latent subspace optimization objectives. In the first one we use few-shot samples as positive anchors of the novel class, and adjust the StyleGAN to produce the corresponding results with the new class label condition. The second objective is to govern the generation process from the other way around, by altering the centroid and its surrounding latent subspace for a more precise generation of the novel class. These reciprocal optimization objectives inject a novel class into the StyleGAN latent subspace, and therefore new unseen samples can be easily produced by sampling images from it. Extensive experiments demonstrate superior few-shot generation performances compared with state-of-the-art methods, especially in terms of diversity and generation quality. Code is available at https://github.com/chansey0529/LSO. Chenxi Zheng, Bangzhen Liu, Huaidong Zhang, Xuemiao Xu, Shengfeng He |
CVPR | 4 |
| 2023 | CIRI: Curricular Inactivation for Residue-aware One-shot Video InpaintingabstractVideo inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated first frame is provided. A naive solution is to propagate the initial target to the other frames with techniques like object tracking. In this context, the main obstacles are the unreliable propagation and the partially inpainted artifacts due to the inaccurate mask. For the former problem, we propose curricular inactivation to replace the hard masking mechanism for indicating the in-painting target, which is robust to erroneous predictions in long-term video inpainting. For the latter, we explore the properties of inpainting residue and present an online residue removal method in an iterative detect-and-refine manner. Extensive experiments on several real-world datasets demonstrate the quantitative and qualitative superiorities of our proposed method in one-shot video inpainting. More importantly, our method is extremely flexible that can be integrated with arbitrary traditional inpainting models, activating them to perform the reliable one-shot video inpainting task. Video demonstrations can be found in our supplement, and our code can be found at https://github.com/Arise-zwy/CIRI. Weiying Zheng, Xuemiao Xu, Wenxi Liu, Shengfeng He |
ICCV | 3 |
| 2023 | EFSCNN: Encoded Feature Sphere Convolution Neural Network for fast non-rigid 3D models classification and retrieval
Zhaolong Dang, Huaidong Zhang, Xuemiao Xu, Harry Qin, Fanzhi Zeng |
Comput. Vis. Image Underst. | 4 |
| 2023 | Contextual-Assisted Scratched Photo RestorationabstractPrinted photographs can be easily warped, wrinkled, and even deteriorated over time. Existing methods treat the restoration of scratches as a pure inpainting problem that neglects the underlying corrupted contextual knowledge. However, important underlying contents are hidden behind the scratches, which are essential hints for producing a semantically consistent result. Motivated by this insight, we explore how to harmonize the scratch-free features and noisy but essential scratch features to produce a visually consistent restoration. Specifically, in this paper, we propose an automatic retouching approach for scratched photographs with the aid of scratch/background context. We explicitly process scratch and background context in two stages. In the first stage, we mainly extract global scratch features, while the mask is introduced in the second stage to filter out and inpaint the scratches. Both contexts are carefully reciprocated for a faithful restoration. Particularly, we propose a Scratch Contextual Assisted Module (SCAM) to adaptively learn texture within the detected mask. This module utilizes the distance between the scratch mask-out feature and scratch encoder feature for modeling the pixel-wise correspondence, which determines the importance of the encoder feature within the scratch mask. Furthermore, to facilitate the evaluation of scratch restoration methods, we create two new scratched photo datasets which have 238 scratch/scratch-free photo pairs to promote the development in the scratch restoration field, namely Old Scratched Photo Dataset (OSPD) and Modern Scratched Photo Dataset (MSPD). Extensive experimental results on the proposed datasets demonstrate that our model outperforms existing methods. To extend the application, we also perform the proposed method on video samples and obtain visual-pleasing results. The code can be found athttps://github.com/cwyyt/Contextual-assisted-Scratched-Photo-Restoration. Huaidong Zhang, Xuemiao Xu, Shengfeng He, Kun Zhang 0001, Harry Qin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Deep Texture-Aware Features for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin. Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Representative Feature Alignment for Adaptive Object DetectionabstractUnsupervised domain adaptation for object detection aims to generalize the object detector trained on the label-rich source domain to the unlabeled target domain. Recently, existing works adopt the instance-level alignment or pixel-level alignment to perform domain transfer, which can effectively avoid the negative transfer due to the diverse background between domains. However, we find that they treat all the regions of an instance feature equally without suppressing background area. They do not segment the specific texture and discriminative regions of objects, which are transferable during adaptation. We call the features that combine the local structure feature and semantic discriminant features as representative features. We propose a novel Representative Feature Alignment (RFA) model to align the features extracted from representative patterns of objects, i.e. representative features, for domain adaptation. Specifically, the representative features are extracted by the Representative Feature Extraction (RFE) submodules. The RFE submodules take the features extracted from different intermediate layers of the detector as input, and filter out the representative features layer-by-layer via integrating class weighting generator, category selection and class activation mapping. Then the representative features from multi-layers are further adaptively aggregated to obtain the final representative features, which are utilized to conduct feature alignment in a class-aware manner. Our representative features are free of untransferable regions and background areas, which leads to better feature alignment. Extensive experimental results show that the proposed model outperforms state-of-the-art methods on a few benchmark datasets. Shan Xu 0006, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Yangyang Xu 0003, Liangui Dai, Kup-Sze Choi, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Panel-Page-Aware Comic Genre UnderstandingabstractUsing a sequence of discrete still images to tell a story or introduce a process has become a tradition in the field of digital visual media. With the surge in these media and the requirements in downstream tasks, acquiring their main topics or genres in a very short time is urgently needed. As a representative form of the media, comic enjoys a huge boom as it has gone digital. However, different from natural images, comic images are divided by panels, and the images are not visually consistent from page to page. Therefore, existing works tailored for natural images perform poorly in analyzing comics. Considering the identification of comic genres is tied to the overall story plotting, a long-term understanding that makes full use of the semantic interactions between multi-level comic fragments needs to be fully exploited. In this paper, we propose [Formula: see text]Comic, a Panel-Page-aware Comic genre classification model, which takes page sequences of comics as the input and produces class-wise probabilities. [Formula: see text]Comic utilizes detected panel boxes to extract panel representations and deploys self-attention to construct panel-page understanding, assisted with interdependent classifiers to model label correlation. We develop the first comic dataset for the task of comic genre classification with multi-genre labels. Our approach is proved by experiments to outperform state-of-the-art methods on related tasks. We also validate the extensibility of our network to perform in the multi-modal scenario. Finally, we show the practicability of our approach by giving effective genre prediction results for whole comic books. Chenshu Xu, Xuemiao Xu, Nanxuan Zhao, Huaidong Zhang, Chengze Li, Xueting Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Pose- and Attribute-consistent Person Image SynthesisabstractPerson Image Synthesis aims at transferring the appearance of the source person image into a target pose. Existing methods cannot handle large pose variations and therefore suffer from two critical problems: (1) synthesis distortion due to the entanglement of pose and appearance information among different body components and (2) failure in preserving original semantics (e.g., the same outfit). In this article, we explicitly address these two problems by proposing a Pose- and Attribute-consistent Person Image Synthesis Network (PAC-GAN). To reduce pose and appearance matching ambiguity, we propose a component-wise transferring model consisting of two stages. The former stage focuses only on synthesizing target poses, while the latter renders target appearances by explicitly transferring the appearance information from the source image to the target image in a component-wise manner. In this way, source-target matching ambiguity is eliminated due to the component-wise disentanglement of pose and appearance synthesis. Second, to maintain attribute consistency, we represent the input image as an attribute vector and impose a high-level semantic constraint using this vector to regularize the target synthesis. Extensive experimental results on the DeepFashion dataset demonstrate the superiority of our method over the state of the art, especially for maintaining pose and attribute consistencies under large pose variations. Zejun Chen, Jiajie Mai, Xuemiao Xu, Shengfeng He |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-Scale Flow-Based Occluding Effect and Content Separation for Cartoon AnimationsabstractOccluding effects have been frequently used to present weather conditions and environments in cartoon animations, such as raining, snowing, moving leaves, and moving petals. While these effects greatly enrich the visual appeal of the cartoon animations, they may also cause undesired occlusions on the content area, which significantly complicate the analysis and processing of the cartoon animations. In this article, we make the first attempt to separate the occluding effects and content for cartoon animations. The major challenge of this problem is that, unlike natural effects that are realistic and small-sized, the effects of cartoons are usually stylistic and large-sized. Besides, effects in cartoons are manually drawn, so their motions are more unpredictable than realistic effects. To separate occluding effects and content for cartoon animations, we propose to leverage the difference in the motion patterns of the effects and the content, and capture the locations of the effects based on a multi-scale flow-based effect prediction (MFEP) module. A dual-task learning system is designed to extract the effect video and reconstruct the effect-removed content video at the same time. We apply our method on a large number of cartoon videos of different content and effects. Experiments show that our method significantly outperforms the existing methods. We further demonstrate how the separated effects and content facilitate the analysis and processing of cartoon videos through different applications, including segmentation, inpainting, and effect migration. Xuemiao Xu, Xueting Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | DLFormer: Discrete Latent Transformer for Video InpaintingabstractVideo inpainting remains a challenging problem to fill with plausible and coherent content in unknown areas in video frames despite the prevalence of data-driven methods. Although various transformer-based architectures yield promising result for this task, they still suffer from hallucinating blurry contents and long-term spatial-temporal inconsistency. While noticing the capability of discrete representation for complex reasoning and predictive learning, we propose a novel Discrete Latent Transformer (DLFormer) to reformulate video inpainting tasks into the discrete latent space rather the previous continuous feature space. Specifically, we first learn a unique compact discrete codebook and the corresponding autoencoder to represent the target video. Built upon these representative discrete codes obtained from the entire target video, the subsequent discrete latent transformer is capable to infer proper codes for unknown areas under a self-attention mechanism, and thus produces fine-grained content with long-term spatial-temporal consistency. Moreover, we further explicitly enforce the short-term consistency to relieve temporal visual jitters via a temporal aggregation block among adjacent frames. We conduct comprehensive quantitative and qualitative evaluations to demonstrate that our method significantly outperforms other state-of-the-art approaches in reconstructing visually-plausible and spatial-temporal coherent content with fine-grained details. Code is available at https://github.com/JingjingRenabc/dlformer. Qingqing Zheng, Xuemiao Xu |
CVPR | 4 |
| 2022 | SA-DPNet: Structure-aware dual pyramid network for salient object detection
Xuemiao Xu, Huaidong Zhang, Guoqiang Han 0002 |
Pattern Recognit. | 1 |
| 2021 | Learning Semantic Context from Normal Samples for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context. This work presents a Semantic Context based Anomaly Detection Network, SCADN, for unsupervised anomaly detection by learning the semantic context from the normal samples. To achieve this, we first generate multi-scale striped masks to remove a part of regions from the normal samples, and then train a generative adversarial network to reconstruct the unseen regions. Note that the masks are designed in multiple scales and stripe directions, and various training examples are generated to obtain the rich semantic context . In testing, we obtain an error map by computing the difference between the reconstructed image and the input image for all samples, and infer the abnormal samples based on the error maps. Finally, we perform various experiments on three public benchmark datasets and a new dataset LaceAD collected by us, and show that our method clearly outperforms the current state-of-the-art methods. Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 3 |
| 2021 | From Continuity to Editability: Inverting GANs with Consecutive ImagesabstractExisting GAN inversion methods are stuck in a paradox that the inverted codes can either achieve high-fidelity reconstruction, or retain the editing capability. Having only one of them clearly cannot realize real image editing. In this paper, we resolve this paradox by introducing consecutive images (e.g., video frames or the same person with different poses) into the inversion process. The rationale behind our solution is that the continuity of consecutive images leads to inherent editable directions. This inborn property is used for two unique purposes: 1) regularizing the joint inversion process, such that each of the inverted codes is semantically accessible from one of the other and fastened in an editable domain; 2) enforcing inter-image coherence, such that the fidelity of each inverted code can be maximized with the complement of other images. Extensive experiments demonstrate that our alternative significantly outperforms state-of-the-art methods in terms of reconstruction fidelity and editability on both the real image dataset and synthesis dataset. Furthermore, our method provides the first support of video-based GAN inversion and an interesting application of unsupervised semantic transfer from consecutive images. Source code can be found at: https://github.com/cnnlstm/InvertingGANs_with_ConsecutiveImgs. Yangyang Xu 0003, Yong Du 0003, Wenpeng Xiao, Xuemiao Xu, Shengfeng He |
ICCV | 4 |
| 2021 | Object Detection in Densely Packed Scenes via Semi-Supervised Learning with Dual ConsistencyabstractDeep neural networks have been shown to be very powerful tools for object detection in various scenes. Their remarkable performance, however, heavily depends on the availability of a large number of high quality labeled data, which are time-consuming and costly to acquire for scenes with densely packed objects. We present a novel semi-supervised approach to addressing this problem, which is designed based on a common teacher-student model, integrated with a novel intersection-over-union (IoU) aware consistency loss and a new proposal consistency loss. The IoU-aware consistency loss evaluates the IoU over the prediction pairs of the teacher model and the student model, which enforces the prediction of the student model to approach closely to that of the teacher model. The IoU-aware consistency loss also reweights the importance of different prediction pairs to suppress the low-confident pairs. The proposal consistency loss ensures proposal consistency between the two models, making it possible to involve the region proposal network in the training process with unlabeled data. We also construct a new dataset, namely RebarDSC, containing 2,125 rebar images annotated with 350,348 bounding boxes in total (164.9 annotations per image average), to evaluate the proposed method. Extensive experiments are conducted over both the RebarDSC dataset and the famous large public dataset SKU-110K. Experimental results corroborate that the proposed method is able to improve the object detection performance in densely packed scenes, consistently outperforming state-of-the-art approaches. Dataset is available in https://github.com/Armin1337/RebarDSC. Huaidong Zhang, Xuemiao Xu, Harry Qin, Kup-Sze Choi |
IJCAI | 3 |
| 2021 | Fast scene labeling via structural inference
Huaidong Zhang, Chu Han, Xiaodan Zhang 0003, Yong Du 0003, Xuemiao Xu, Guoqiang Han 0002, Harry Qin, Shengfeng He |
Neurocomputing | 5 |
| 2021 | D4Net: De-deformation defect detection network for non-rigid products with large patterns
Xuemiao Xu, Huaidong Zhang, Wing W. Y. Ng |
Inf. Sci. | 1 |
| 2021 | Learning Gated Non-Local Residual for Single-Image Rain Streak RemovalabstractThis work presents a gated non-local deep residual learning framework for image deraining. It can avoid the over-deraining or under-deraining caused by the global residual learning in existing deraining networks, since the learned soft gate in our method adaptively adjusts the amount of global residual to be passed for generating the final derained result. To generate feature maps for global residual prediction, we develop a non-local guided attention module (NLAM), which first obtains non-local features by exploiting spatial inter-dependencies among all the feature positions of local features produced by convolutional neural network (CNN), and then leverages the attention mechanism to merge the local and non-local features based on their complementary relation. Moreover, we develop a channel-wise gated prediction module to learn a soft gate on the global residual by explicitly modelling channel inter-dependencies of the feature maps obtained from NLAM. Experiments on four deraining benchmark datasets and real-world rainy images show that our network has a quantitative and qualitative improvement over state-of-the-arts. Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Haoran Xie 0001, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Multi-View Face Synthesis via Progressive Face FlowabstractExisting GAN-based multi-view face synthesis methods rely heavily on "creating" faces, and thus they struggle in reproducing the faithful facial texture and fail to preserve identity when undergoing a large angle rotation. In this paper, we combat this problem by dividing the challenging large-angle face synthesis into a series of easy small-angle rotations, and each of them is guided by a face flow to maintain faithful facial details. In particular, we propose a Face Flow-guided Generative Adversarial Network (FFlowGAN) that is specifically trained for small-angle synthesis. The proposed network consists of two modules, a face flow module that aims to compute a dense correspondence between the input and target faces. It provides strong guidance to the second module, face synthesis module, for emphasizing salient facial texture. We apply FFlowGAN multiple times to progressively synthesize different views, and therefore facial features can be propagated to the target view from the very beginning. All these multiple executions are cascaded and trained end-to-end with a unified back-propagation, and thus we ensure each intermediate step contributes to the final result. Extensive experiments demonstrate the proposed divide-and-conquer strategy is effective, and our method outperforms the state-of-the-art on four benchmark datasets qualitatively and quantitatively. Yangyang Xu 0003, Xuemiao Xu, Jianbo Jiao, Shengfeng He |
IEEE Trans. Image Process. | 2 |
| 2021 | Erratum to "Multi-View Face Synthesis via Progressive Face Flow"
Yangyang Xu 0003, Xuemiao Xu, Jianbo Jiao, Shengfeng He |
IEEE Trans. Image Process. | 2 |
| 2021 | SALMNet: A Structure-Aware Lane Marking Detection NetworkabstractLane marking detection is a fundamental task, which serves as an important prerequisite for automatic driving or driver-assistance systems. However, the complex and uncontrollable driving road environment as well as the discontinuous lane marking appearance make this task challenging. In this work, a novel deep neural network architecture is presented to detect lane markings in a complex environment by analyzing their structure information. There are two contributions to the network design. Firstly, a semantic-guided channel attention (SGCA) module is developed to select the low-level features of a deep convolutional neural network by taking the high-level features as the guidance. Secondly, a pyramid deformable convolution (PDC) module is formulated to enlarge the receptive fields and to capture the complex structures of lane markings by applying deformable convolutions on multiple feature maps with different scales. Hence, our network can better reduce false detection and enhance lane marking structures simultaneously. The experimental results on three benchmark datasets for lane marking detection show that our method outperforms other methods on all the benchmark datasets. Xuemiao Xu, Tianfei Yu, Xiaowei Hu 0001, Wing W. Y. Ng, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Transductive Zero-Shot Action Recognition via Visually Connected Graph Convolutional NetworksabstractWith the explosive growth of action categories, zero-shot action recognition aims to extend a well-trained model to novel/unseen classes. To bridge the large knowledge gap between seen and unseen classes, in this brief, we visually associate unseen actions with seen categories in a visually connected graph, and the knowledge is then transferred from the visual features space to semantic space via the grouped attention graph convolutional networks (GAGCNs). In particular, we extract visual features for all the actions, and a visually connected graph is built to attach seen actions to visually similar unseen categories. Moreover, the proposed grouped attention mechanism exploits the hierarchical knowledge in the graph so that the GAGCN enables propagating the visual-semantic connections from seen actions to unseen ones. We extensively evaluate the proposed method on three data sets: HMDB51, UCF101, and NTU RGB + D. Experimental results show that the GAGCN outperforms state-of-the-art methods. Yangyang Xu 0003, Chu Han, Harry Qin, Xuemiao Xu, Guoqiang Han 0002, Shengfeng He |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Perceptual-Aware Sketch Simplification Based on Integrated VGG LayersabstractDeep learning has been recently demonstrated as an effective tool for raster-based sketch simplification. Nevertheless, it remains challenging to simplify extremely rough sketches. We found that a simplification network trained with a simple loss, such as pixel loss or discriminator loss, may fail to retain the semantically meaningful details when simplifying a very sketchy and complicated drawing. In this paper, we show that, with a well-designed multi-layer perceptual loss, we are able to obtain aesthetic and neat simplification results preserving semantically important global structures as well as fine details without blurriness and excessive emphasis on local structures. To do so, we design a multi-layer discriminator by fusing all VGG feature layers to differentiate sketches and clean lines. The weights used in layer fusing are automatically learned via an intelligent adjustment mechanism. Furthermore, to evaluate our method, we compare our method to state-of-the-art methods through multiple experiments, including visual comparison and intensive user study. Xuemiao Xu, Minshan Xie, Peiqi Miao, Wenpeng Xiao, Huaidong Zhang, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | GDFace: Gated Deformation for Multi-View Face Image SynthesisabstractPhotorealistic multi-view face synthesis from a single image is an important but challenging problem. Existing methods mainly learn a texture mapping model from the source face to the target face. However, they fail to consider the internal deformation caused by the change of poses, leading to the unsatisfactory synthesized results for large pose variations. In this paper, we propose a Gated Deformable Face Synthesis Network to model the deformation of faces that aids the synthesis of the target face image. Specifically, we propose a dual network that consists of two modules. The first module estimates the deformation of two views in the form of convolution offsets according to the input and target poses. The second one, on the other hand, leverages the predicted deformation offsets to create the target face image. In this way, pose changes are explicitly modeled in the face generator to cope with geometric transformation, by adaptively focusing on pertinent regions of the source image. To compensate offset estimation errors, we introduce a soft-gating mechanism that enables adaptive fusion between deformable features and primitive features. Extensive experimental results on five widely-used benchmarks show that our approach performs favorably against the state-of-the-arts on multi-view face synthesis, especially for large pose changes. Xuemiao Xu, Shengfeng He |
AAAI | 1 |
| 2020 | Context-Aware and Scale-Insensitive Temporal Repetition CountingabstractTemporal repetition counting aims to estimate the number of cycles of a given repetitive action. Existing deep learning methods assume repetitive actions are performed in a fixed time-scale, which is invalid for the complex repetitive actions in real life. In this paper, we tailor a context-aware and scale-insensitive framework, to tackle the challenges in repetition counting caused by the unknown and diverse cycle-lengths. Our approach combines two key insights: (1) Cycle lengths from different actions are unpredictable that require large-scale searching, but, once a coarse cycle length is determined, the variety between repetitions can be overcome by regression. (2) Determining the cycle length cannot only rely on a short fragment of video but a contextual understanding. The first point is implemented by a coarse-to-fine cycle refinement method. It avoids the heavy computation of exhaustively searching all the cycle lengths in the video, and, instead, it propagates the coarse prediction for further refinement in a hierarchical manner. We secondly propose a bidirectional cycle length estimation method for a context-aware prediction. It is a regression network that takes two consecutive coarse cycles as input, and predicts the locations of the previous and next repetitive cycles. To benefit the training and evaluation of temporal repetition counting area, we construct a new and largest benchmark, which contains 526 videos with diverse repetitive actions. Extensive experiments show that the proposed network trained on a single dataset outperforms state-of-the-art methods on several benchmarks, indicating that the proposed framework is general enough to capture repetition patterns across domains. Code and data are available in https://github.com/Xiaodomgdomg/Deep-Temporal-Repetition-Counting. Huaidong Zhang, Xuemiao Xu, Guoqiang Han 0002, Shengfeng He |
CVPR | 2 |
| 2020 | Dual pyramid network for salient object detection
Xuemiao Xu, Huaidong Zhang, Guoqiang Han 0002 |
Neurocomputing | 1 |
| 2020 | Unsupervised Domain Adaptation via Importance SamplingabstractUnsupervised domain adaptation aims to generalize a model from the label-rich source domain to the unlabeled target domain. Existing works mainly focus on aligning the global distribution statistics between source and target domains. However, they neglect distractions from the unexpected noisy samples in domain distribution estimation, leading to domain misalignment or even negative transfer. In this paper, we present an importance sampling method for domain adaptation (ISDA), to measure sample contributions according to their “informative” levels. In particular, informative samples, as well as outliers, can be effectively modeled using feature-norm and prediction entropy of the network. The importance of information is further formulated as the importance sampling losses in features and label spaces. In this way, the proposed model mitigates the noisy outliers while enhancing the important samples during domain alignment. In addition, our model is easy to implement yet effective, and it does not introduce any extra parameters. Extensive experiments on several benchmark datasets show that our method outperforms state-of-the-art methods under both the standard and partial domain adaptation settings. Xuemiao Xu, Hai He, Huaidong Zhang, Yangyang Xu 0003, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Aggregating Attentional Dilated Features for Salient Object DetectionabstractThis paper presents a novel deep learning model to aggregate the attentional dilated features for salient object detection by exploring the complementary information between the global and local context in a convolutional neural network. There are two technical contributions to our network design. First, we develop an attentional dense atrous (dilated) spatial pyramid pooling (AD-ASPP) module to selectively use the local saliency cues captured by dilated convolutions with a small rate and the global saliency cues captured by dilated convolutions with a large rate. Second, taking the feature pyramid network as the backbone, we develop an aggregation network to integrate the refined features by formulating two consecutive chains of residual learning based modules: one chain from deep to shallow layers while another chain from shallow to deep layers. We evaluate our network on seven widely-used saliency detection benchmarks by comparing it against 21 state-of-the-art methods. Experimental results show that our network outperforms others on all the seven benchmark datasets. Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Fast User-Guided Single Image Reflection Removal via Edge-Aware Cascaded NetworksabstractTaking photos through a glass window leads to glare or reflection, which might distract the viewer from the scene behind the window. In this paper, we involve user interaction to tackle the ill-posedness of the reflection removal problem. Users are allowed to draw strokes or lassos to indicate the background and reflection layers. Instead of designing hand-crafted features, we propose the edge-aware cascaded networks for reflection removal. The proposed network is a two-stage pipeline. The first stage takes the edge hints converted from user guidance and the image with reflection as input, and then separates the input image into the background and reflection layers. The second stage involves a refinement network to recover the missing details of the background layers. We simulate different types of user guidance, and the networks are trained on simulated data. The cascaded networks are end-to-end and perform with a single feed-forward pass, enabling fast editing. Extensive experimental evaluations demonstrate that the proposed used-guided reflection removal network yields better performance than the state-of-the-art methods on real-world scenarios. Furthermore, we show that novice users can easily generate reflection-free images, and large improvements in reflection removal quality can be obtained in just one minute. Huaidong Zhang, Xuemiao Xu, Hai He, Shengfeng He, Guoqiang Han 0002, Harry Qin, Dapeng Oliver Wu |
IEEE Trans. Multim. | 2 |
| 2019 | Deep Multi-Model Fusion for Single-Image DehazingabstractThis paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural network (CNN) features at different CNN layers and generate the attentional multi-level integrated features (AMLIF). Then, from the AMLIF, we further predict a haze-free result for an atmospheric scattering model, as well as for four haze-layer separation models, and then fuse the results together to produce the final haze-free image. To evaluate the effectiveness of our method, we compare our network with several state-of-the-art methods on two widely-used dehazing benchmark datasets, as well as on two sets of real-world hazy images. Experimental results demonstrate clear quantitative and qualitative improvements of our method over the state-of-the-arts. Zijun Deng, Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Qing Zhang 0006, Harry Qin, Pheng-Ann Heng |
ICCV | 5 |
| 2019 | Highlight-assisted nighttime vehicle detection using a multi-level fusion network and label hierarchy
Yaoyang Mo, Guoqiang Han 0002, Huaidong Zhang, Xuemiao Xu |
Neurocomputing | 4 |
| 2019 | A Learning-Based Multimodel Integrated Framework for Dynamic Traffic Flow Forecasting
Teng Zhou, Guoqiang Han 0002, Xuemiao Xu, Chu Han, Yuchang Huang, Harry Qin |
Neural Process. Lett. | 3 |
| 2019 | SINet: A Scale-Insensitive Convolutional Neural Network for Fast Vehicle DetectionabstractVision-based vehicle detection approaches achieve incredible success in recent years with the development of deep convolutional neural network (CNN). However, existing CNN-based algorithms suffer from the problem that the convolutional features are scale-sensitive in object detection task but it is common that traffic images and videos contain vehicles with a large variance of scales. In this paper, we delve into the source of scale sensitivity, and reveal two key issues: 1) existing RoI pooling destroys the structure of small scale objects and 2) the large intra-class distance for a large variance of scales exceeds the representation capability of a single network. Based on these findings, we present a scale-insensitive convolutional neural network (SINet) for fast detecting vehicles with a large variance of scales. First, we present a context-aware RoI pooling to maintain the contextual information and original structure of small scale objects. Second, we present a multi-branch decision network to minimize the intra-class distance of features. These lightweight techniques bring zero extra time complexity but prominent detection accuracy improvement. The proposed techniques can be equipped with any deep network architectures and keep them trained end-to-end. Our SINet achieves state-of-the-art performance in terms of accuracy and speed (up to 37 FPS) on the KITTI benchmark and a new highway dataset, which contains a large variance of scales and extremely small objects. Xiaowei Hu 0001, Xuemiao Xu, Yongjie Xiao, Hao Chen 0011, Shengfeng He, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection
Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
ECCV (6) | 5 |
| 2018 | R³Net: Recurrent Residual Refinement Network for Saliency DetectionabstractSaliency detection is a fundamental yet challenging task in computer vision, aiming at highlighting the most visually distinctive objects in an image. We propose a novel recurrent residual refinement network (R^3Net) equipped with residual refinement blocks (RRBs) to more accurately detect salient regions of an input image. Our RRBs learn the residual between the intermediate saliency prediction and the ground truth by alternatively leveraging the low-level integrated features and the high-level integrated features of a fully convolutional network (FCN). While the low-level integrated features are capable of capturing more saliency details, the high-level integrated features can reduce non-salient regions in the intermediate prediction. Furthermore, the RRBs can obtain complementary saliency information of the intermediate prediction, and add the residual into the intermediate prediction to refine the saliency maps. We evaluate the proposed R^3Net on five widely-used saliency detection benchmarks by comparing it with 16 state-of-the-art saliency detectors. Experimental results show that our network outperforms our competitors in all the benchmark datasets. Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Harry Qin, Guoqiang Han 0002, Pheng-Ann Heng |
IJCAI | 4 |
| 2018 | Deep Attentional Features for Prostate Segmentation in Ultrasound
Yi Wang 0031, Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Xuemiao Xu, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (4) | 6 |
| 2018 | Rolling normal filtering for point clouds
Yinglong Zheng, Guiqing Li, Xuemiao Xu, Yongwei Nie |
Comput. Aided Geom. Des. | 3 |
| 2018 | Towards High-Quality Visualization of Superfluid VorticesabstractSuperfluidity is a special state of matter exhibiting macroscopic quantum phenomena and acting like a fluid with zero viscosity. In such a state, superfluid vortices exist as phase singularities of the model equation with unique distributions. This paper presents novel techniques to aid the visual understanding of superfluid vortices based on the state-of-the-art non-linear Klein-Gordon equation, which evolves a complex scalar field, giving rise to special vortex lattice/ring structures with dynamic vortex formation, reconnection, and Kelvin waves, etc. By formulating a numerical model with theoretical physicists in superfluid research, we obtain high-quality superfluid flow data sets without noise-like waves, suitable for vortex visualization. By further exploring superfluid vortex properties, we develop a new vortex identification and visualization method: a novel mechanism with velocity circulation to overcome phase singularity and an orthogonal-plane strategy to avoid ambiguity. Hence, our visualizations can help reveal various superfluid vortex structures and enable domain experts for related visual analysis, such as the steady vortex lattice/ring structures, dynamic vortex string interactions with reconnections and energy radiations, where the famous Kelvin waves and decaying vortex tangle were clearly observed. These visualizations have assisted physicists to verify the superfluid model, and further explore its dynamic behavior more intuitively. Xiaopei Liu, Chi Xiong, Xuemiao Xu, Chi-Wing Fu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | Packing Vertex Data into Hardware-Decompressible TexturesabstractMost graphics hardware features memory to store textures and vertex data for rendering. However, because of the irreversible trend of increasing complexity of scenes, rendering a scene can easily reach the limit of memory resources. Thus, vertex data are preferably compressed, with a requirement that they can be decompressed during rendering. In this paper, we present a novel method to exploit existing hardware texture compression circuits to facilitate the decompression of vertex data in graphics processing unit (GPUs). This built-in hardware allows real-time, random-order decoding of data. However, vertex data must be packed into textures, and careless packing arrangements can easily disrupt data coherence. Hence, we propose an optimization approach for the best vertex data permutation that minimizes compression error. All of these result in fast and high-quality vertex data decompression for real-time rendering. To further improve the visual quality, we introduce vertex clustering to reduce the dynamic range of data during quantization. Our experiments demonstrate the effectiveness of our method for various vertex data of 3D models during rendering with the advantages of a minimized memory footprint and high frame rate. Kin Chung Kwan, Xuemiao Xu, Tien-Tsin Wong, Wai-Man Pang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | δ-agree AdaBoost stacked autoencoder for short-term traffic flow forecasting
Teng Zhou, Guoqiang Han 0002, Xuemiao Xu, Zhizhe Lin, Chu Han, Yuchang Huang, Harry Qin |
Neurocomputing | 3 |
| 2017 | A Unified Detail-Preserving Liquid Simulation by Two-Phase Lattice Boltzmann ModelingabstractTraditional methods in graphics to simulate liquid-air dynamics under different scenarios usually employ separate approaches with sophisticated interface tracking/reconstruction techniques. In this paper, we propose a novel unified approach which is easy and effective to produce a variety of liquid-air interface phenomena. These phenomena, such as complex surface splashes, bubble interactions, as well as surface tension effects, can co-exist in one single simulation, and are created within the same computational framework. Such a framework is unique in that it is free from any complicated interface tracking/reconstruction procedures. Our approach is developed from the two-phase lattice Boltzmann method with the mean field model, which provides a unified framework for interface dynamics but is numerically unstable under turbulent conditions. Considering the drawbacks of the existing approaches, we propose techniques to suppress oscillations for significant stability enhancement, as well as derive a new subgrid-scale model to further improve stability, faithfully preserving liquid-air interface details without excessive diffusion by taking into account the density variation. The whole framework is highly parallel, enabling very efficient implementation. Comparisons with the related approaches show superiority on stable simulations with detail preservation and multiphase phenomena simultaneously involved. A set of animation results demonstrate the effectiveness of our method. Xiaopei Liu, Xuemiao Xu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | ASCII Art Synthesis from Natural PhotographsabstractWhile ASCII art is a worldwide popular art form, automatic generating structure-based ASCII art from natural photographs remains challenging. The major challenge lies on extracting the perception-sensitive structure from the natural photographs so that a more concise ASCII art reproduction can be produced based on the structure. However, due to excessive amount of texture in natural photos, extracting perception-sensitive structure is not easy, especially when the structure may be weak and within the texture region. Besides, to fit different target text resolutions, the amount of the extracted structure should also be controllable. To tackle these challenges, we introduce a visual perception mechanism of non-classical receptive field modulation (non-CRF modulation) from physiological findings to this ASCII art application, and propose a new model of non-CRF modulation which can better separate the weak structure from the crowded texture, and also better control the scale of texture suppression. Thanks to our non-CRF model, more sensible ASCII art reproduction can be obtained. In addition, to produce more visually appealing ASCII arts, we propose a novel optimization scheme to obtain the optimal placement of proportional-font characters. We apply our method on a rich variety of images, and visually appealing ASCII art can be obtained in all cases. Xuemiao Xu, Linyuan Zhong, Minshan Xie, Xueting Liu 0001, Harry Qin, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Text-aware balloon extraction from manga
Xueting Liu 0001, Chengze Li, Tien-Tsin Wong, Xuemiao Xu |
Vis. Comput. | 5 |
| 2015 | Region-based structure line detection for cartoonsabstractCartoons are a worldwide popular visual entertainment medium with a long history. Nowadays, with the boom of electronic devices, there is an increasing need to digitize old classic cartoons as a basis for further editing, including deformation, colorization, etc. To perform such editing, it is essential to extract the structure lines within cartoon images. Traditional edge detection methods are mainly based on gradients. These methods perform poorly in the face of compression artifacts and spatially-varying line colors, which cause gradient values to become unreliable. This paper presents the first approach to extract structure lines in cartoons based on regions. Our method starts by segmenting an image into regions, and then classifies them as edge regions and non-edge regions. Our second main contribution comprises three measures to estimate the likelihood of a region being a non-edge region. These measure darkness, local contrast, and shape. Since the likelihoods become unreliable as regions become smaller, we further classify regions using both likelihoods and the relationships to neighboring regions via a graph-cut formulation. Our method has been evaluated on a wide variety of cartoon images, and convincing results are obtained in all cases. Xueting Liu 0001, Tien-Tsin Wong, Xuemiao Xu |
Comput. Vis. Media | 4 |
| 2010 | Structure-based ASCII artabstractThe wide availability and popularity of text-based communication channels encourage the usage of ASCII art in representing images. Existing tone-based ASCII art generation methods lead to halftone-like results and require high text resolution for display, as higher text resolution offers more tone variety. This paper presents a novel method to generate structure-based ASCII art that is currently mostly created by hand. It approximates the major line structure of the reference image content with the shape of characters. Representing the unlimited image content with the extremely limited shapes and restrictive placement of characters makes this problem challenging. Most existing shape similarity metrics either fail to address the misalignment in real-world scenarios, or are unable to account for the differences in position, orientation and scaling. Our key contribution is a novel alignment-insensitive shape similarity (AISS) metric that tolerates misalignment of shapes while accounting for the differences in position, orientation and scaling. Together with the constrained deformation approach, we formulate the ASCII art generation as an optimization that minimizes shape dissimilarity and deformation . Convincing results and user study are shown to demonstrate its effectiveness. Xuemiao Xu, Linling Zhang, Tien-Tsin Wong |
ACM Trans. Graph. | 1 |
| 2008 | Animating animal motion from stillabstractEven though the temporal information is lost, a still picture of moving animals hints at their motion. In this paper, we infer motion cycle of animals from the "motion snapshots" (snapshots of different individuals) captured in a still picture. By finding the motion path in the graph connecting motion snapshots, we can infer the order of motion snapshots with respect to time, and hence the motion cycle. Both "half-cycle" and "full-cycle" motions can be inferred in a unified manner. Therefore, we can animate a still picture of a moving animal group by morphing among the ordered snapshots. By refining the pose, morphology, and appearance consistencies, smooth and realistic animal motion can be synthesized. Our results demonstrate the applicability of the proposed method to a wide range of species, including birds, fishes, mammals, and reptiles. Xuemiao Xu, Xiaopei Liu, Tien-Tsin Wong, Andrew Chi-Sing Leung |
ACM Trans. Graph. | 1 |