Ruonan Wei

dblp:298/3155 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-2562-6021ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Depth-Synergized Mamba Meets Memory Experts for All-Day Image Reflection Separation
abstract
Image reflection separation aims to disentangle the transmission layer and the reflection layer from a blended image. Existing methods rely on limited information from a single image, tending to confuse the two layers when their contrasts are similar, a challenge more severe at night. To address this issue, we propose the Depth-Memory Decoupling Network (DMDNet). It employs the Depth-Aware Scanning (DAScan) to guide Mamba toward salient structures, promoting information flow along semantic coherence to construct stable states. Working in synergy with DAScan, the Depth-Synergized State-Space Model (DS-SSM) modulates the sensitivity of state activations by depth, suppressing the spread of ambiguous features that interfere with layer disentanglement. Furthermore, we introduce the Memory Expert Compensation Module (MECM), leveraging cross-image historical knowledge to guide experts in providing layer-specific compensation. To address the lack of datasets for nighttime reflection separation, we construct the Nighttime Image Reflection Separation (NightIRS) dataset. Extensive experiments demonstrate that DMDNet outperforms state-of-the-art methods in both daytime and nighttime.
Siyan Fang, Ruonan Wei, Yuehuan Wang
AAAI4
2026 Mamba capsule routing-enhanced heat conduction network for event stream object detection
Yuehuan Wang, Ruonan Wei
Neurocomputing4
2026 Dynamic denoising track: Towards end-to-end multiple object tracking against attention trivialization
Ruonan Wei, Yuehuan Wang
Neurocomputing1
2026 Dual-stream frequency-domain framework with contextual graph enhancer and consensus-difference fusion for cross-view geo-localization
Haitong Li, Chaoyi Ma, Yuehuan Wang, Ruonan Wei
J. Vis. Commun. Image Represent.6
2025 End-to-End Multiple Object Tracking with Dynamic Scene Perception
abstract
End-to-end Multiple Object Tracking (MOT) frameworks integrate detection and tracking into a unified model, avoiding intermediate information loss and complicated post-processing. However, existing end-to-end MOT trackers rely on track queries of the previous frame to provide prior information. Their limited short-term temporal modeling struggle to cope with high dynamic tracking scenarios, where inter-frame target variations exhibit significant heterogeneity. To address these shortcomings, we propose a scene-perception MOT framework (SP-MOT) that encodes scene context understanding into long-term embedding and adaptively complements it with short-term cues, enabling discriminative and flexible instance representations. Specifically, SP-MOT introduces: (1) a learnable scene query that globally profiles foreground and background to capture short-term scene-level features; (2) a context understanding module to uncover long-term stable relationships across dynamic scenes based on multiple historical scene features; (3) scene-adaptive augmented decoding that leverages scene information as guidance, adaptively aggregating long-term and short-term information into object embeddings, improving the model's association ability and fault-tolerance. Extensive experiments on MOT benchmarks demonstrate that SP-MOT outperforms state-of-the-art end-to-end trackers across multiple metrics, particularly in challenging scenarios with high dynamics.
Ruonan Wei, Siyan Fang, Yuehuan Wang
ACM Multimedia1
2025 Augment One With Others: Generalizing to Unforeseen Variations for Visual Tracking
abstract
Unforeseen appearance variation is a challenging factor for visual tracking. This paper provides a novel solution from semantic data augmentation, which facilitates offline training of trackers for better generalization. We utilize existing samples to obtain knowledge to augment another in terms of diversity and hardness. First, we propose that the similarity matching space in Siamese-like models has class-agnostic transferability. Based on this, we design the Latent Augmentation (LaAug) to transfer relevant variations and suppress irrelevant ones between training similarity embeddings of different classes. Thus the model can generalize across a more diverse semantic distribution. Then, we propose the Semantic Interaction Mix (SIMix), which interacts moments between different feature samples to contaminate structure and texture attributes and retain other semantic attributes. SIMix simulates the occlusion and complements the training distribution with hard cases. The mixed features with adversarial perturbations can empirically enable the model against external environmental disturbances. Experiments on six challenging benchmarks demonstrate that three representative tracking models, i.e., SiamBAN, TransT and OSTrack, can be consistently improved by incorporating the proposed methods without extra parameters and inference cost.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
IEEE Trans. Multim.3
2024 MFRGN: Multi-scale Feature Representation Generalization Network for Ground-to-Aerial Geo-localization
abstract
Cross-area evaluation poses a significant challenge for ground-to-aerial geo-localization, in which the training and testing data are captured from entirely distinct areas. However, current methods struggle in cross-area evaluation due to their emphasis solely on learning global information from single-scale features. Some efforts alleviate this problem but rely on complex and specific technologies like pre-processing and hard sample mining. To this end, we propose a pure end-to-end solution, free from task-specific techniques, termed the Multi-scale Feature Representation Generalization Network (MFRGN) to improve generalization. Specifically, we introduce multi-scale features and explicitly utilize them by an novel global-local information representation structure with two flows, to bolster feature representations. In the global flow, we present a lightweight Self and Cross Attention Module (SCAM) to efficiently learn global embeddings. In the local flow, we develop a Global-Prompt Attention Block (GPAB) to capture discriminative features under the global embeddings as prompts. As a result, our approach generates robust descriptors representing multi-scale global and local information, thereby enhancing the model's invariance to scene variations. Extensive experiments on benchmarks show our MFRGN achieves competitive performance in same-area evaluation and improves cross-area generalization by a significant margin compared to SOTA methods. Our code is available at https://github.com/ytao-wang/MFRGN.
Jinpu Zhang, Ruonan Wei, Yuehuan Wang
ACM Multimedia3
2023 Learning Mutually in Crowd Scenes for Pedestrian Detection
abstract
Pedestrian detection in crowded scenes is a challenging problem due to the diverse occlusion patterns and highly overlap. To tackle this critical problem, we propose a mutual learning detection network. First, a self-attention mechanism is proposed to achieve mutual learning between individuals by capturing similar semantics among pedestrians. Feature representation of occluded individuals is enhanced by locally fusing similar semantics. Second, mutual loss is designed to improve the consistency of regression and classification. Specifically, regression results are leveraged to make classification score aware of the quality of predicted boxes, and the classification scores help the regression head to accelerate convergence of redundant boxes. Finally, we evaluate our proposed method on MOT20 and CityPersons datasets and achieve comparable state-of-the-art performance using less data. Compared to baseline, our detector obtains 14.4% AP and 11.7% AR gains on challenging MOT20 dataset.
Ruonan Wei, Yuehuan Wang, Jinpu Zhang
ICIP1
2023 Progressive Domain-style Translation for Nighttime Tracking
abstract
Nighttime tracking is challenging due to the lack of sufficient training data and scene diversity. Unsupervised domain adaptation is a solution by transferring knowledge from day (source domain) to night (target domain). It typically involves adversarial training with a domain discriminator on the source and target data to learn domain-invariant features. However, the imbalanced source/target distribution can cause overfitting of the domain discriminator, hindering the domain adaptability. To address this issue, we propose a Progressive Domain-Style Translation (PDST) for domain adaptive nighttime tracking. PDST decomposes and recombines domain-invariant content encodings and domain-specific style encodings of different domains. Thus the rich source domain content is translated to the target domain, expanding the inter-class diversity of the target domain to alleviate overfitting. Moreover, a momentum update manner is introduced to progressively estimate the domain-style encoding from multiple features, which more accurately reflects the statistical domain attribute than an individual image-style. Finally, we incorporate two regularization terms to constrain the content and domain-style consistency in the translation process, ensuring the generated source-like target features are valid to facilitate the training of domain adaptation. Exhaustive experiments demonstrate the domain adaptability and SOTA performance of the proposed method in nighttime tracking.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
ACM Multimedia3
2023 Spatio-temporal matching for siamese visual tracking
Jinpu Zhang, Kaiheng Dai, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
Neurocomputing4