EDBT 2026 Demo / reviewers in the wild / expert
Hanyu Xuan
dblp:215/9583
· DBLP profile ↗
20ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphMamba: Graph-driven spatial order-aware Mamba for medical image segmentation
Chengjin Yu, Cailing Pu, Sangyin Lv, Xiaorui Wu, Dongsheng Ruan, Hanyu Xuan, Yuan-Ting Yan |
Pattern Recognit. | 8 |
| 2025 | DSTA-Net: Deformable Spatio-Temporal Attention Network for Video InpaintingabstractVideo inpainting aims to fill in the corrupted or missing regions of target frame with plausible contents by exploiting the information drawn from surrounding reference frames. However, existing methods always treat reference frames uniformly and assign equal importance to all reference frames, which will lead to inpainting results that are too smooth to preserve necessary details and textures. For this purpose, we design a Deformable Spatio-Temporal Attention (DSTA) network, named DSTA-Net, to automatically learn to pay higher attention on specific reference frames. Concretely, our DSTA-Net begins by employing a lightweight motion estimator to predict coarse optical flows between target frame and reference frames. These coarse optical flows are employed as a basic offset rather than directly aligning reference frames. Subsequently, our proposed DSTA module calculates the sampling point offsets, attention maps, and flatten features with the aim of capturing the appropriate pixels from the regions with the highest similarity to the reference frames to generate the corrupted or missing contents. Numerous experimental results show that our method significantly outperform recent seven benchmarks on two most commonly-used datasets. Tongxing Liu, Guoxin Qiu, Hanyu Xuan |
IEEE Signal Process. Lett. | 3 |
| 2025 | X-STA: Cross-Modal Spatial-Temporal Alignment Network for Unified Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to segment sound sources from video frames using synchronized audio cues. This task requires not only localizing the sound sources within frames but also accurately delineating their shapes. Existing AVS methods often rely on assumptions of spatial-temporal consistency between audio-visual content and are typically designed for specific learning paradigms. However, this specialization limits their ability to handle multi-granularity supervision signals and and adapt to diverse task requirements. For this purpose, we propose a Cross-modal Spatial-Temporal Alignment (X-STA) network to alleviate spatial-temporal inconsistency and overcome paradigm-specific constraint. Our X-STA introduces three key components: a novel multi-stage Cross-modal Adapter (xAdapter) that transfers knowledge from a pre-trained SAM through multi-grained representation adaptation, an innovative Cross-modal Prompter (xPrompter) that provides geometry-aware constraints for AVS through dynamic prompting strategies, and a Cross-modal Self-supervised (xSelf) mechanism that refines temporal alignment and enables self-supervised AVS. These components collectively facilitate explicit reasoning about location and geometric shape of the sound source by refining the alignment of cross-modal spatial-temporal cues. Our method achieves competitive performance across several baselines on widely-used AVS datasets, demonstrating its effectiveness in addressing the complexities of AVS. Hanyu Xuan, Tongxing Liu, Wenxiang Dong, Zhongheng Li, Shuo Chen 0003 |
IEEE Signal Process. Lett. | 1 |
| 2025 | GLCONet: Learning Multisource Perception Representation for Camouflaged Object DetectionabstractRecently, the biological perception has been a powerful tool for handling the camouflaged object detection (COD) task. However, most existing methods are heavily dependent on the local spatial information of diverse scales from convolutional operations to optimize initial features. A commonly neglected point in these methods is the long-range dependencies between feature pixels from different scale spaces that can help the model build a global structure of the object, inducing a more precise image representation. In this article, we propose a novel global-local collaborative optimization network called GLCONet. Technically, we first design a collaborative optimization strategy (COS) from the perspective of multisource perception to simultaneously model the local details and global long-range relationships, which can provide features with abundant discriminative information to boost the accuracy in detecting camouflaged objects. Furthermore, we introduce an adjacent reverse decoder (ARD) that contains cross-layer aggregation and reverse optimization to integrate complementary information from different levels for generating high-quality representations. Extensive experiments demonstrate that the proposed GLCONet method with different backbones can effectively activate potentially significant pixels in an image, outperforming 20 state-of-the-art (SOTA) methods on three public COD datasets. The source code is available at: https://github.com/CSYSI/GLCONet. Hanyu Xuan, Jian Yang 0003, Lei Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingabstractVideo inpainting aims to fill in the missing regions of the video frames with plausible content. Benefiting from the outstanding long-range modeling capacity, the transformer-based models have achieved unprecedented performance regarding inpainting quality. Essentially, coherent contents from all the frames along both spatial and temporal dimensions are concerned by a patch-wise attention module, and then the missing contents are generated based on the attention-weighted summation. In this way, attention retrieval accuracy has become the main bottleneck to improve the video inpainting performance, where the factors affecting attention calculation should be explored to maximize the advantages of transformer. Towards this end, in this paper, we theoretically certificate that noise is the culprit that entangles the process of attention calculation. Meanwhile, we propose a novel wavelet transformer network with noise robustness for video inpainting, named WaveFormer. Unlike existing transformer-based methods that utilize the whole embeddings to calculate the attention, our WaveFormer first separates the noise existing in the embedding into high-frequency components by introducing the Discrete Wavelet Transform (DWT), and then adopts clean low-frequency components to calculate the attention. In this way, the impact of noise on attention computation can be greatly mitigated and the missing content regarding different frequencies can be generated by sharing the calculated attention. Extensive experiments validate the superior performance of our method over state-of-the-art baselines both qualitatively and quantitatively. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu, Yan Yan 0002 |
AAAI | 3 |
| 2024 | Frequency-Spatial Entanglement Learning for Camouflaged Object Detection
Chunyan Xu, Jian Yang 0003, Hanyu Xuan, Lei Luo 0001 |
ECCV (6) | 4 |
| 2024 | Text-Video Completion Networks With Motion Compensation And Attention AggregationabstractThe purpose of video inpainting is to fill a specified area with reasonable content. However, in the case of multiple targets and complex textures, current methods struggle to distinguish between feature information of the targets, leading to confusing or fuzzy inpainting results. In this paper, we design a new text-video completion network based on a motion compensation and temporal attention feature aggregation. Our network utilizes information from reference frames and target frames to complete the damaged region of the target frame. We first employ motion compensation to align the features of reference frames, and then use the temporal attention module to aggregate these features, resulting in accurate and reasonable content. To evaluate the effectiveness of our method, we introduce a new text video dataset with multiple text objects and complex textures, presenting a novel and challenging task for inpainting research. Through quantitative and qualitative comparison experiments, we demonstrate that our model outperforms existing baseline models in scenarios with multiple objects and complex textures. Zhiliang Wu, Hanyu Xuan, Yan Yan 0002 |
ICASSP | 3 |
| 2024 | Robust Audio-Visual Contrastive Learning for Proposal-Based Self-Supervised Sound Source Localization in VideosabstractBy observing a scene and listening to corresponding audio cues, humans can easily recognize where the sound is. To achieve such cross-modal perception on machines, existing methods take advantage of the maps obtained by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for prospective practical applications, we argue that these map-based methods only offer a coarse-grained and indirect description of the sound source. Additionally, these methods utilize a single audio-visual tuple at a time during self-supervised learning, causing the model to lose the crucial chance to reason about the data distribution of large-scale audio-visual samples. Although the introduction of Audio-Visual Contrastive Learning (AVCL) can effectively alleviate this issue, the contrastive set constructed by randomly sampling is based on the assumption that the audio and visual segments from all other videos are not semantically related. Since the resulting contrastive set contains a large number of faulty negatives, we believe that this assumption is rough. In this paper, we advocate a novel proposal-based solution that directly localizes the semantic object-level sound source, without any manual annotations. The Global Response Map (GRM) is incorporated as an unsupervised spatial constraint to filter those instances corresponding to a large number of sound-unrelated regions. As a result, our proposal-based Sound Source Localization (SSL) can be cast into a simpler Multiple Instance Learning (MIL) problem. To overcome the limitation of random sampling in AVCL, we propose a novel Active Contrastive Set Mining (ACSM) to mine the contrastive sets with informative and diverse negatives for robust AVCL. Our approaches achieve state-of-the-art (SOTA) performance when compared to several baselines on multiple SSL datasets with diverse scenarios. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Bo Jiang 0002, Lei Luo 0001, Xavier Alameda-Pineda, Yan Yan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Deep Stereo Video InpaintingabstractStereo video inpainting aims to fill the missing regions on the left and right views of the stereo video with plausible content simultaneously. Compared with the single video inpainting that has achieved promising results using deep convolutional neural networks, inpainting the missing regions of stereo video has not been thoroughly explored. In essence, apart from the spatial and temporal consistency that single video inpainting needs to achieve, another key challenge for stereo video inpainting is to maintain the stereo consistency between left and right views and hence alleviate the 3D fatigue for viewers. In this paper, we propose a novel deep stereo video inpainting network named SVINet, which is the first attempt for stereo video inpainting task utilizing deep convolutional neural networks. SVINet first utilizes a self-supervised flow-guided deformable temporal alignment module to align the features on the left and right view branches, respectively. Then, the aligned features are fed into a shared adaptive feature aggregation module to generate missing contents of their respective branches. Finally, the parallax attention module (PAM) that uses the cross-view information to consider the significant stereo correlation is introduced to fuse the completed features of left and right views. Furthermore, we develop a stereo consistency loss to regularize the trained parameters, so that our model is able to yield high-quality stereo video inpainting results with better stereo consistency. Experimental results demonstrate that our SVINet outperforms state-of-the-art single video inpainting models. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
CVPR | 3 |
| 2023 | Semi-Supervised Video Inpainting with Cycle Consistency ConstraintsabstractDeep learning-based video inpainting has yielded promising results and gained increasing attention from re-searchers. Generally, these methods assume that the cor-rupted region masks of each frame are known and easily ob-tained. However, the annotation of these masks are labor-intensive and expensive, which limits the practical application of current methods. Therefore, we expect to relax this assumption by defining a new semi-supervised inpainting setting, making the networks have the ability of completing the corrupted regions of the whole video using the anno-tated mask of only one frame. Specifically, in this work, we propose an end-to-end trainable framework consisting of completion network and mask prediction network, which are designed to generate corrupted contents of the current frame using the known mask and decide the regions to be filled of the next frame, respectively. Besides, we introduce a cycle consistency loss to regularize the training parameters of these two networks. In this way, the completion network and the mask prediction network can constrain each other, and hence the overall performance of the trained model can be maximized. Furthermore, due to the natural existence of prior knowledge (e.g., corrupted contents and clear bor-ders), current video inpainting datasets are not suitable in the context of semi-supervised video inpainting. Thus, we create a new dataset by simulating the corrupted video of real-world scenarios. Extensive experimental results are reported to demonstrate the superiority of our model in the video inpainting task. Remarkably, although our model is trained in a semi-supervised manner, it can achieve compa-rable performance as fully-supervised methods. Zhiliang Wu, Hanyu Xuan, Changchang Sun, Weili Guan, Yan Yan 0002 |
CVPR | 2 |
| 2023 | Flow-Guided Deformable Alignment Network with Self-Supervision for Video InpaintingabstractVideo inpainting aims to utilize plausible contents to fill missing regions in the video. State-of-the-art video inpainting methods typically generate the missing contents of the target frame (current frame) by aggregating the temporal information of reference frames (neighboring frames) aligned using deformable convolution. However, these deformable convolution alignment networks often suffer from offset overflow during training, resulting in unsatisfactory alignment, thereby obtaining compromised inpainting performance. In this paper, we propose a self-supervised Flow-Guided Deformable Alignment (FGDA) network for aligning reference frames at the feature level. FGDA computes the residual of the optical flow as the offsets. This design can effectively reduce the burden of offsets learning, thereby avoiding offset overflow. Furthermore, a gradient-weighted reconstruction loss for supervised completed frame reconstruction is designed, which can use the gradients in all directions of the video frame to emphasize the difficultly reconstructed texture regions, so that detail textures get more attention during training. Experiments show that FGDA-based video inpainting model trained with gradient-weighted reconstruction loss outperforms the state-of-the-art by a significant margin in terms of PSNR and SSIM with relative improvements of 6.2% and 2.1%, respectively. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
ICASSP | 4 |
| 2023 | Semantic-Guided Completion Network for Video Inpainting in Complex Urban Scene
Hanyu Xuan, Zhiliang Wu |
PRCV (11) | 2 |
| 2023 | Divide-and-conquer model based on wavelet domain for multi-focus image fusion
Zhiliang Wu, Hanyu Xuan, Xia Yuan, Chunxia Zhao |
Signal Process. Image Commun. | 3 |
| 2023 | Divide-and-Conquer Completion Network for Video InpaintingabstractVideo inpainting aims to utilize plausible contents to complete missing regions in the video. For different components, the reconstruction targets of missing regions are different,e.g.,smoothness preserving for flat regions, sharpening for edges and textures. Typically, existing methods treat the missing regions as a whole and holistically train the model by optimizing homogenous pixel-wise losses (e.g.,MSE). In this way, the trained models will be easily dominated and determined by flat regions, failing to infer realistic details (edges and textures) that are difficult to reconstruct but necessary for practical applications. In this paper, we propose a divide-and-conquer completion network for video inpainting. In particular, our network first uses discrete wavelet transform to decompose the deep features into low-frequency components containing structural information (flat regions) and high-frequency components involving detailed texture information. Thereafter, we feed these components into different branches and adopt the temporal attention feature aggregation module to generate missing contents, separately. It hence can realize flexible supervision utilizing the intermediate supervision learning strategy for each component, which has not been noticed and explored by current state-of-the-art video inpainting methods. Furthermore, we adopt a gradient-weighted reconstruction loss to supervise the completed frame reconstruction process, which can use the gradients in all directions of the video frame to emphasize the difficultly reconstructed textures regions, making the model pay more attention to the complex detailed textures. Extensive experiments validate the superior performance of our divide-and-conquer model over state-of-the-art baselines in both quantitative and qualitative evaluations. Zhiliang Wu, Changchang Sun, Hanyu Xuan, Yan Yan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosabstractHumans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for potential practical applications, we argue that these existing map-based approaches only provide a coarse-grained and indirect description of the sound source. In this pa-per, we advocate a novel proposal-based paradigm that can directly perform semantic object-level localization, without any manual annotations. We incorporate the global re-sponse map as an unsupervised spatial constraint to weight the proposals according to how well they cover the esti-mated global shape of the sound source. As a result, our proposal-based sound source localization can be cast into a simpler Multiple Instance Learning (MIL) problem by filtering those instances corresponding to large sound-unrelated regions. Our method achieves state-of-the-art (SOTA) per-formance when compared to several baselines on multiple datasets. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Yan Yan 0002, Xavier Alameda-Pineda |
CVPR | 1 |
| 2022 | Active Contrastive Set Mining for Robust Audio-Visual Instance DiscriminationabstractThe recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual Instance Discrimination (AVID) extends instance discrimination to the audio-visual realm. Existing AVID methods construct the contrastive set by random sampling based on the assumption that the audio and visual clips from all other videos are not semantically related. We argue that this assumption is rough, since the resulting contrastive sets have a large number of faulty negatives. In this paper, we overcome this limitation by proposing a novel Active Contrastive Set Mining (ACSM) that aims to mine the contrastive sets with informative and diverse negatives for robust AVID. Moreover, we also integrate a semantically-aware hard-sample mining strategy into our ACSM. The proposed ACSM is implemented into two most recent state-of-the-art AVID methods and significantly improves their performance. Extensive experiments conducted on both action and sound recognition on multiple datasets show the remarkably improved performance of our method. Hanyu Xuan, Shuo Chen 0002, Zhiliang Wu, Jian Yang 0003, Yan Yan 0002, Xavier Alameda-Pineda |
IJCAI | 1 |
| 2021 | DAPC-Net: Deformable Alignment and Pyramid Context Completion Networks for Video InpaintingabstractVideo inpainting aims to fill missing regions with plausible content in a video sequence. Deep learning-based video inpainting methods have made promising progress over the past few years. However, these methods tend to generate degraded completion content, such as missing textural details. To address this issue, we propose a novel Deformable Alignment and Pyramid-context Completion Network for video inpainting (DAPC-Net), which takes advantage of temporal redundancy information among video sequence. Specifically, we construct a deformable convolution alignment network (DANet) for aligning reference frame at the feature level. After alignment, we further devise a pyramid-context completion network (PCNet) to complete missing regions of the target frame. Particularly, the pyramid completion mechanism and cross-scale transference strategy are used to ensure the visual and semantic coherence of the completed target frame. Experimental results show that the proposed method not only achieves better quantitative and qualitative performance but also improves the inference speed by 35.4%. Zhiliang Wu, Hanyu Xuan, Jian Yang 0003 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Discriminative Cross-Modality Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationabstractIt is theoretically insufficient to construct a complete set of semantics in the real world using single-modality data. As a typical application of multi-modality perception, the audio-visual event localization task aims to match audio and visual components to identify the simultaneous events of interest. Although some recent methods have been proposed to deal with this task, they cannot handle the practical situation of temporal inconsistency that is widespread in the audio-visual scene. Inspired by the human system which automatically filters out event-unrelated information when performing multi-modality perception, we propose a discriminative cross-modality attention network to simulate such a process. Similar to human mechanism, our network can adaptively select "where" to attend, "when" to attend and "which" to attend for audio-visual event localization. In addition, to prevent our network from getting trivial solutions, a novel eigenvalue-based objective function is proposed to train the whole network to better fuse audio and visual signals, which can obtain discriminative and nonlinear multi-modality representation. In this way, even with large temporal inconsistency between audio and visual sequence, our network is able to adaptively select event-valuable information for audio-visual event localization. Furthermore, we systemically investigate three subtasks of audio-visual event localization, i.e., temporal localization, weakly-supervised spatial localization and cross-modality localization. The visualization results also help us better understand how our network works. Hanyu Xuan, Lei Luo 0001, Zhenyu Zhang 0005, Jian Yang 0003, Yan Yan 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationabstractIn human multi-modality perception systems, the benefits of integrating auditory and visual information are extensive as they provide plenty supplementary cues for understanding the events. Despite some recent methods proposed for such application, they cannot deal with practical conditions with temporal inconsistency. Inspired by human system which puts different focuses at specific locations, time segments and media while performing multi-modality perception, we provide an attention-based method to simulate such process. Similar to human mechanism, our network can adaptively select “where” to attend, “when” to attend and “which” to attend for audio-visual event localization. In this way, even with large temporal inconsistent between vision and audio, our network is able to adaptively trade information between different modalities and successfully achieve event localization. Our method achieves state-of-the-art performance on AVE (Audio-Visual Event) dataset collected in the real life. In addition, we also systemically investigate audio-visual event localization tasks. The visualization results also help us better understand how our model works. Hanyu Xuan, Zhenyu Zhang 0005, Shuo Chen 0003, Jian Yang 0003, Yan Yan 0002 |
AAAI | 1 |
| 2020 | Channel Attention Based Generative Network for Robust Visual TrackingabstractIn recent years, Siamese trackers have achieved great success in visual tracking. Siamese networks can achieve competitive performance in both accuracy and speed. However, they may suffer from the performance degradation due to the case of large pose variations, out-of-plane, etc. In this paper, we propose a novel real-time Channel Attention based Generative Network (AGSNet) for Robust Visual Tracking. AGSNet can better recognize the targets undergoing signifi-cant appearance variations and having similar distractors. The AGSNet model introduces a channel favored feature attention to the template branch to enhance the discriminative capacity and uses a simple generative network in the instance branch to capture a variety of target appearance changes. With the end-to-end offline training, our model can achieve robust visual tracking in a long temporal span.Experimental results on benchmark datasets OTB-2013 and OTB-2015, demonstrate that our proposed tracker outperforms other approaches while runs at more than 40 frames per second. Hanyu Xuan, Jian Yang 0003 |
ICASSP | 2 |