Xin Yuan 0009

dblp:78/713-9 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-3140-3243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
abstract
Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has advanced significantly, applying it directly to opera videos remains challenging. The scarcity of datasets impedes the recovery of high-frequency details, and existing STVSR methods lack global modeling capabilities—compromising visual quality when handling opera’s characteristic large motions. To address these challenges, we pioneer a large-scale Chinese Opera Video Clip (COVC) dataset and propose the Mamba-based multiscale fusion network for space-time Opera Video Super-Resolution (MambaOVSR). Specifically, MambaOVSR involves three novel components: the Global Fusion Module (GFM) for motion modeling through a multiscale alternating scanning mechanism, and the Multiscale Synergistic Mamba Module (MSMM) for alignment across different sequence lengths. Additionally, our MambaVR block resolves feature artifacts and positional information loss during alignment. Experimental results on the COVC dataset show that MambaOVSR significantly outperforms the SOTA STVSR method by an average of 1.86 dB in terms of PSNR.
Hua Chang, Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Xin Yuan 0009, Kui Jiang
AAAI5
2026 NPFML: Non-isotropic Potential Fields with Hierarchical Decay for Deep Metric Learning
Xin Yuan 0009, Minshi Chen, Xin Xu 0007
MMM (2)2
2026 FD-HDRMamba: Frequency-Decoupled Mamba for Multi-Exposure HDR Reconstruction
Zhehan Gong, Wei Wang 0170, Xiao Wang 0029, Xin Yuan 0009
IEEE Signal Process. Lett.4
2025 VAGeo: View-specific Attention for Cross-View Object Geo-Localization
abstract
Cross-view object geo-localization (CVOGL) aims to locate an object of interest in a captured ground- or drone-view image within the satellite image. However, existing works treat ground-view and drone-view query images equivalently, overlooking their inherent viewpoint discrepancies and the spatial correlation between the query image and the satellite-view reference image. To this end, this paper proposes a novel View-specific Attention Geo-localization method (VAGeo) for accurate CVOGL. Specifically, VAGeo contains two key modules: view-specific positional encoding (VSPE) module and channel-spatial hybrid attention (CSHA) module. In object-level, according to the characteristics of different viewpoints of ground and drone query images, viewpoint-specific positional codings are designed to more accurately identify the click-point object of the query image in the VSPE module. In feature-level, a hybrid attention in the CSHA module is introduced by combining channel attention and spatial attention mechanisms simultaneously for learning discriminative features. Extensive experimental results demonstrate that the proposed VAGeo gains a significant performance improvement, i.e., improving [email protected]/[email protected] on the CVOGL dataset from 45.43%/42.24% to 48.21%/45.22% for ground-view, and from 61.97%/57.66% to 66.19%/61.87% for drone-view.
Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007
ICASSP2
2025 Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
abstract
Video-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between continuous frames to improve recognition accuracy. Previous approaches have assisted by introducing event data into the video person ReID task, but they still cannot avoid the privacy leakage problem caused by RGB images. In order to avoid privacy attacks and to take advantage of the benefits of event data, we consider using only event data. To make full use of the information in the event stream, we propose a Cross-Modality and Temporal Collaboration (CMTC) network for event-based video person ReID. First, we design an event transform network to obtain corresponding auxiliary information from the input of raw events. Additionally, we propose a differential modality collaboration module to balance the roles of events and auxiliaries to achieve complementary effects. Furthermore, we introduce a temporal collaboration module to exploit motion information and appearance cues. Experimental results demonstrate that our method outperforms others in the task of event-based video person ReID.
Renkai Li, Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007
ICASSP2
2025 RPUDet: Learning Relational Prior and Uncertainty for Robust Aerial Object Detection
abstract
Aerial object detection remains a challenging task in the computer vision community. While general object detectors perform well on natural images, they struggle with aerial images due to missed detections from small, low-resolution objects and misdetections caused by classification uncertainty between semantically similar objects. To address these issues, we propose a Relational Prior and Uncertainty detector (RPUDet). RPUDet consists of two core modules: 1) Relation Aware Reasoning Module (RARM), which leverages a relational prior graph to help detect correlated objects and reduce missed detections of low-resolution objects; and 2) Uncertainty Guided Awareness Module (UGAM), which computes a uncertainty map to identify low-confidence regions, dynamically adjusts feature weights, and refines areas with high semantic ambiguity to mitigate misdetections. Additionally, to advance aerial object detection in industrial applications, we introduce the Low-Voltage Distribution Insulator Dataset (LVDID), focusing on static scenes, in contrast to existing dynamic datasets. This enables us to evaluate RPUDet's performance across diverse real-world scenarios. We evaluate RPUDet on VisDrone2019, VEDAI, and LVDID, demonstrating superior performance compared to existing methods. The code and dataset will be available at https://github.com/Godk02/RPUDet.
Wei Liu 0183, Minshi Chen, Xiao Wang 0029, Xin Yuan 0009
ICMR5
2025 Spatial Bi-Exploration for Robust Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to segment camouflaged objects hidden within their environment. Existing COD models, aside from image features, mostly focus on a single coarse-grained spatial structure, such as depth information, texture information, or edge information. However, when faced with complex scenes where the target and background textures are similar and overlapping, or when subjected to noise interference, this design often leads to insufficient detection accuracy and robustness. To address these issues, we proposed a strategy for multiple spatial explorations and designedSpatial Bi-Exploration Network (SPNet). SPNet conducts a comprehensive analysis of complex camouflage scenarios by jointly exploring depth spatial, contour spatial, and image feature information, thereby enhancing detection performance and maintaining robustness. Unlike existing methods, SPNet leverages dual exploration of depth and contour spaces to mitigate the vulnerability of coarse structures to noise. Depth spatial information aids the model in recognizing the deep relationships between objects and the background, reducing the impact of noise on object boundaries, while contour spatial information improves edge detection accuracy. This dual approach significantly enhances robustness, especially in the face of adversarial attacks. Extensive experiments on benchmark datasets demonstrate that our model not only outperforms existing methods in detection performance but also exhibits superior robustness against adversarial attacks.
Xiao Wang 0029, Xin Yuan 0009, Nan Mu, Zheng Wang 0007
IEEE Signal Process. Lett.3
2025 Mix-Modality Person Re-Identification: A New and Practical Paradigm
abstract
Current visible-infrared cross-modality person re-identification research has only focused on exploring the bi-modality mutual retrieval paradigm, and we propose a new and more practical mix-modality retrieval paradigm. Existing Visible-Infrared Person Re-Identification (VI-ReID) methods have achieved some results in the bi-modality mutual retrieval paradigm by learning the correspondence between visible and infrared modalities. However, significant performance degradation occurs due to the modality confusion problem when these methods are applied to the new mix-modality paradigm. Therefore, this article proposes a Mix-Modality Person Re-Identification (MM-ReID) task, explores the influence of modality mixing ratio on performance, and constructs mix-modality test sets for existing datasets according to the new mix-modality testing paradigm. To solve the modality confusion problem in MM-ReID, we propose a Cross-Identity Discrimination Harmonization Loss (CIDHL) adjusting the distribution of samples in the hyperspherical feature space, pulling the centers of samples with the same identity closer, and pushing away the centers of samples with different identities while aggregating samples with the same modality and the same identity. Furthermore, we propose a Modality Bridge Similarity Optimization Strategy (MBSOS) to optimize the cross-modality similarity between the query and queried samples with the help of the similar bridge sample in the gallery. Extensive experiments demonstrate that compared to the original performance of existing cross-modality methods on MM-ReID, the addition of our CIDHL and MBSOS demonstrates a general improvement.
Wei Liu 0183, Xin Xu 0007, Hua Chang, Xin Yuan 0009, Zheng Wang 0007
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Blind 3D Video Stabilization with Spatio-Temporally Varying Motion Blur
abstract
Video stabilization is a challenging task that attempts to compensate for the overall frame shake during video acquisition. Existing three-dimensional video stabilization methods aim at modeling camera perspective projection through either data-driven training or explicit motion estimation. However, the above methods are difficult to effectively solve the issue of shaky videos with abrupt object movements, resulting in local motion blur in the direction of the movement. This phenomenon is prevalent in real-world scenarios featuring foreground blind motion scenes. Unfortunately, directly combining stabilization and deblurring methods poses challenges when dealing with this situation. In the video, the intensity of motion blur undergoes continuous changes, and the direct combination method inadequately utilizes spatiotemporal information, providing insufficient clues for cross-frame compensation. To alleviate this problem, the Cross-frame-temporal Module framework is proposed to address blind motion blur induced by various conditions, which utilizes cross-frame temporal features to estimate depth maps and camera motion. In this framework, a Blur Transform Network (BTNet) is designed to adapt to spatially varying motion blur, which transforms local regions according to the impact of blur intensities to adapt to the effects of non-uniform motion blur; furthermore, our Temporal-Aware Network (TANet) further suppresses motion blur by leveraging cross-frame temporal features. In addition, the limited availability of pair-training video data containing motion blur limits the application of this approach in practice. The Cross-frame-temporal Module framework adopts an un-pretrained in-test training strategy. Extensive experimental results have demonstrated that our method outperforms state-of-the-art methods.
Hengwei Li, Wei Wang 0170, Xiao Wang 0029, Xin Yuan 0009, Xin Xu 0007
ACM Trans. Multim. Comput. Commun. Appl.4
2022 SAM: Self Attention Mechanism for Scene Text Recognition Based on Swin Transformer
Xiang Shuai, Xiao Wang 0029, Wei Wang 0170, Xin Yuan 0009, Xin Xu 0007
MMM (1)4
2022 Rank-in-Rank Loss for Person Re-identification
abstract
Person re-identification (re-ID) is commonly investigated as a ranking problem. However, the performance of existing re-ID models drops dramatically, when they encounter extreme positive-negative class imbalance (e.g., very small ratio of positive and negative samples) during training. To alleviate this problem, this article designs a rank-in-rank loss to optimize the distribution of feature embeddings. Specifically, we propose a Differentiable Retrieval-Sort Loss (DRSL) to optimize the re-ID model by ranking each positive sample ahead of the negative samples according to the distance and sorting the positive samples according to the angle (e.g., similarity score). The key idea of the proposed DRSL lies in minimizing the distance between samples of the same category along with the angle between them. Considering that the ranking and sorting operations are non-differentiable and non-convex, the DRSL also performs the optimization of automatic derivation and backpropagation. In addition, the analysis of the proposed DRSL is provided to illustrate that the DRSL not only maintains the inter-class distance distribution but also preserves the intra-class similarity structure in terms of angle constraints. Extensive experimental results indicate that the proposed DRSL can improve the performance of the state-of-the-art re-ID models, thus demonstrating its effectiveness and superiority in the re-ID task.
Xin Xu 0007, Xin Yuan 0009, Zheng Wang 0007, Kai Zhang 0002, Ruimin Hu
ACM Trans. Multim. Comput. Commun. Appl.2