Yuehuan Wang

dblp:99/3759 · DBLP profile ↗
← Back
45ranked-venue papers
0as first author
32since 2021 · last 2026
0000-0001-7046-7587ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 21 since 2021Artificial intelligence and machine learning · 14 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Depth-Synergized Mamba Meets Memory Experts for All-Day Image Reflection Separation
abstract
Image reflection separation aims to disentangle the transmission layer and the reflection layer from a blended image. Existing methods rely on limited information from a single image, tending to confuse the two layers when their contrasts are similar, a challenge more severe at night. To address this issue, we propose the Depth-Memory Decoupling Network (DMDNet). It employs the Depth-Aware Scanning (DAScan) to guide Mamba toward salient structures, promoting information flow along semantic coherence to construct stable states. Working in synergy with DAScan, the Depth-Synergized State-Space Model (DS-SSM) modulates the sensitivity of state activations by depth, suppressing the spread of ambiguous features that interfere with layer disentanglement. Furthermore, we introduce the Memory Expert Compensation Module (MECM), leveraging cross-image historical knowledge to guide experts in providing layer-specific compensation. To address the lack of datasets for nighttime reflection separation, we construct the Nighttime Image Reflection Separation (NightIRS) dataset. Extensive experiments demonstrate that DMDNet outperforms state-of-the-art methods in both daytime and nighttime.
Siyan Fang, Ruonan Wei, Yuehuan Wang
AAAI5
2026 Mamba capsule routing-enhanced heat conduction network for event stream object detection
Yuehuan Wang, Ruonan Wei
Neurocomputing2
2026 Dynamic denoising track: Towards end-to-end multiple object tracking against attention trivialization
Ruonan Wei, Yuehuan Wang
Neurocomputing2
2026 Dual-stream frequency-domain framework with contextual graph enhancer and consensus-difference fusion for cross-view geo-localization
Haitong Li, Chaoyi Ma, Yuehuan Wang, Ruonan Wei
J. Vis. Commun. Image Represent.4
2025 Multi-view Feature Discrepancy Attack for Single Object Tracking
abstract
Adversarial attacks on single object tracking (SOT) have attracted increasing attention. However, most previous works have focused on adding small digital perturbations to tracking sequences, assuming access to the data, which makes these attacks impractical in real-world applications. Inspired by traditional military camouflage, we propose a texture pattern attack method with a similar implementation, called the Multi-View Feature Discrepancy Attack (MFDA). Unlike the consistent texture features of traditional camouflage, we iteratively optimize non-planar textures by obtaining gradients from the target model to enhance feature discrepancies across different viewpoints. Moreover, the model’s multi-scale features and heatmaps are utilized as targets for our feature-level and decision-level attacks, respectively. We apply texture patterns to controllable regions of vehicle models and conduct extensive attack experiments. The results show that our method significantly degrades the performance of state-of-the-art Siamese-based trackers, and also exhibits attack capability against ViT-based trackers.
Zhimin Weng, Yuehuan Wang
ICASSP3
2025 Retinex-Based Self-Conditioned Diffusion Model for Low-Light Image Enhancement
abstract
The conditional diffusion models have made significant progress in image synthesis, leveraging human annotations such as class labels or text descriptions to guide the generative process. However, different from image synthesis, low-light image enhancement(LLIE) lacks strictly calibrated conditional priors to guide the enhancement process, often resulting in unsatisfactory results. To address the issue, we propose Retinex-Based Self-Conditioned Diffusion Models, dubbed RSCDM, which utilizes self-conditioned illumination representation learning and representation guidance enhancement to generate high-quality image. To be specific, in the first stage, we pretrain a retinex decomposed model (RDM) to capture illumination representation and devise a illumination-representation restoration model (IRM) to accurately reconstruct the representation from noisy images. Moreover, we further design dynamic resblock (DRB) and dynamic simplified attention gated block (DSAGB) as basic units of IRM for better fine-grained restoration. In the second stage, we employ a self-conditioned diffusion model (SDM) to generate realistic results conditioned on the illumination representation. Extensive experiments demonstrates our method outperforms the existing SOTA methods both quantitatively and qualitatively. The codes will be publicly available.
Ziwen Li 0005, Jinpu Zhang, Yuehuan Wang
ICASSP4
2025 GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
abstract
In recent years, video anomaly detection has been extensively investigated in both unsupervised and weakly supervised settings to alleviate costly temporal labeling. Despite significant progress, these methods still suffer from unsatisfactory results such as numerous false alarms, primarily due to the absence of precise temporal anomaly annotation. In this paper, we present a novel labeling paradigm, termed "glance annotation", to achieve a better balance between anomaly detection accuracy and annotation cost. Specifically, glance annotation is a random frame within each abnormal event, which can be easily accessed and is cost-effective. To assess its effectiveness, we manually annotate the glance annotations for two standard video anomaly detection datasets: UCF-Crime and XD-Violence. Additionally, we propose a customized GlanceVAD method, that leverages gaussian kernels as the basic unit to compose the temporal anomaly distribution, enabling the learning of diverse and robust anomaly representations from the glance annotations. Through comprehensive analysis and experiments, we verify that the proposed labeling paradigm can achieve an excellent trade-off between annotation cost and model performance. Extensive experimental results also demonstrate the effectiveness of our GlanceVAD approach, which significantly outperforms existing advanced unsupervised and weakly supervised methods. Our annotations and code are publicly available at https://github.com/pipixin321/GlanceVAD.
Huaxin Zhang, Xiang Wang 0012, Xiaohao Xu, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Shanjun Zhang, Nong Sang
ICME6
2025 Adaptive Language-Aware Image Reflection Removal Network
abstract
Existing image reflection removal methods struggle to handle complex reflections. Accurate language descriptions can help the model understand the image content to remove complex reflections. However, due to blurred and distorted interferences in reflected images, machine-generated language descriptions of the image content are often inaccurate, which harms the performance of language-guided reflection removal. To address this, we propose the Adaptive Language-Aware Network (ALANet) to remove reflections even with inaccurate language inputs. Specifically, ALANet integrates both filtering and optimization strategies. The filtering strategy reduces the negative effects of language while preserving its benefits, whereas the optimization strategy enhances the alignment between language and visual features. ALANet also utilizes language cues to decouple specific layer content from feature maps, improving its ability to handle complex reflections. To evaluate the model's performance under complex reflections and varying levels of language accuracy, we introduce the Complex Reflection and Language Accuracy Variance (CRLAV) dataset. Experimental results demonstrate that ALANet surpasses state-of-the-art methods for image reflection removal. The code and dataset are available at https://github.com/fashyon/ALANet.
Siyan Fang, Jinpu Zhang, Ziwen Li 0005, Yuehuan Wang
IJCAI5
2025 End-to-End Multiple Object Tracking with Dynamic Scene Perception
abstract
End-to-end Multiple Object Tracking (MOT) frameworks integrate detection and tracking into a unified model, avoiding intermediate information loss and complicated post-processing. However, existing end-to-end MOT trackers rely on track queries of the previous frame to provide prior information. Their limited short-term temporal modeling struggle to cope with high dynamic tracking scenarios, where inter-frame target variations exhibit significant heterogeneity. To address these shortcomings, we propose a scene-perception MOT framework (SP-MOT) that encodes scene context understanding into long-term embedding and adaptively complements it with short-term cues, enabling discriminative and flexible instance representations. Specifically, SP-MOT introduces: (1) a learnable scene query that globally profiles foreground and background to capture short-term scene-level features; (2) a context understanding module to uncover long-term stable relationships across dynamic scenes based on multiple historical scene features; (3) scene-adaptive augmented decoding that leverages scene information as guidance, adaptively aggregating long-term and short-term information into object embeddings, improving the model's association ability and fault-tolerance. Extensive experiments on MOT benchmarks demonstrate that SP-MOT outperforms state-of-the-art end-to-end trackers across multiple metrics, particularly in challenging scenarios with high dynamics.
Ruonan Wei, Siyan Fang, Yuehuan Wang
ACM Multimedia4
2025 "Who is Trying to Access My Account?" Exploring User Perceptions and Reactions to Risk-based Authentication Notifications
Tongxin Wei, Yuehuan Wang
NDSS4
2025 SiamDTO: Mamba-Based Spatio-Temporal Attention for Satellite Video Object Tracking
Xinhong Bai, Qishuai Nie, Yuehuan Wang
PRCV (16)3
2025 TemVLT: Vision-Language Tracking via Mamba-based Temporal Information Learning
abstract
The vision-language tracking (VLT) aims to improve target localization by leveraging natural language (NL) descriptions. However, most trackers fail to account for the discrepancy between the dynamically evolving target states and the initial NL during tracking, which leads to limitations in model performance. Therefore, developing effective temporal modeling methods to resolve this vision-language ambiguity represents a critical research challenge in VLT. In this paper, we propose a video-level VLT framework called TemVLT, which aggregates language information under the guidance of dynamic template, and stores and transmits temporal information across frames via hidden states in Mamba, resulting in more robust tracking performance. Specifically, we design the Language Adaptive Aggregation (LAA) module, which aggregates the Top-k most important language tokens based on the appearance information in dynamic template to improve the stability of language information. Additionally, we introduce the Mambabased Temporal Information Capture and Storage (TCS) module, which captures temporal information in the visual features of the current frame via Mamba layers and stores it in hidden states for enhancing the context-awareness of the next frame. Extensive results on TNL2K, LaSOT, and OTB99-Lang confirm the superiority of TemVLT in VLT over existing trackers.
Qishuai Nie, Zhimin Weng, Yuehuan Wang
SMC3
2025 Augment One With Others: Generalizing to Unforeseen Variations for Visual Tracking
abstract
Unforeseen appearance variation is a challenging factor for visual tracking. This paper provides a novel solution from semantic data augmentation, which facilitates offline training of trackers for better generalization. We utilize existing samples to obtain knowledge to augment another in terms of diversity and hardness. First, we propose that the similarity matching space in Siamese-like models has class-agnostic transferability. Based on this, we design the Latent Augmentation (LaAug) to transfer relevant variations and suppress irrelevant ones between training similarity embeddings of different classes. Thus the model can generalize across a more diverse semantic distribution. Then, we propose the Semantic Interaction Mix (SIMix), which interacts moments between different feature samples to contaminate structure and texture attributes and retain other semantic attributes. SIMix simulates the occlusion and complements the training distribution with hard cases. The mixed features with adversarial perturbations can empirically enable the model against external environmental disturbances. Experiments on six challenging benchmarks demonstrate that three representative tracking models, i.e., SiamBAN, TransT and OSTrack, can be consistently improved by incorporating the proposed methods without extra parameters and inference cost.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
IEEE Trans. Multim.4
2024 Real-Time Exposure Correction via Collaborative Transformations and Adaptive Sampling
abstract
Most of the previous exposure correction methods learn dense pixel-wise transformations to achieve promising results, but consume huge computational resources. Recently, Learnable 3D lookup tables (3D LUTs) have demon-strated impressive performance and efficiency for image enhancement. However, these methods can only perform global transformations and fail to finely manipulate local regions. Moreover, they uniformly downsample the input image, which loses the rich color information and limits the learning of color transformation capabilities. In this paper, we present a collaborative transformation framework (CoTF) for real-time exposure correction, which integrates global transformation with pixel-wise transformations in an efficient manner. Specifically, the global transformation adjusts the overall appearance using image-adaptive 3D LUTs to provide decent global contrast and sharp details, while the pixel transformation compensates for local context. Then, a relation-aware modulation module is designed to combine these two components effectively. In addition, we propose an adaptive sampling strategy to preserve more color information by predicting the sampling intervals, thus providing higher quality input data for the learning of 3D LUTs. Extensive experiments demonstrate that our method can process high-resolution images in real-time on GPUs while achieving comparable performance against current state-of-the-art methods. The code is avail-able at https://github.com/HUST-IAL/CoTF.
Ziwen Li 0005, Feng Zhang 0039, Jinpu Zhang, Yuanjie Shao, Yuehuan Wang, Nong Sang
CVPR6
2024 Joint Language Prompt and Object Tracking
abstract
Recently, utilizing natural language descriptions to assist in object tracking is becoming a trend. However, The tracking performance is compromised with the absence of the Vision-Language (VL) datasets. In this paper, we develop Joint Language Prompt and Object Tracking (JLPT), the first application of prompt learning in VL tracking. JLPT preserves knowledge from large-scale pretrained vision-based models and adopts language descriptions as prompts to aid tracking with a limited amount of datasets. Specifically, the Multimodal Fusion Prompter (MFP) constructs precise language prompts by dimension compression and correlation operation on heterologous modal. It can adaptively reinforce target features and attenuate interference features. Additionally, the Language Rectification Moudle (LRM) is introduced to dynamically adjust language prompts based on target appearance variations, providing temporal adaptability of the language prompts. JLPT achieves notable performance with minimal trainable language prompt parameters and limited VL datasets. Extensive experiments on TNL2K, LaSOT, LaSOText, and OTB99-L confirm the superiority of JLPT in VL tracking over existing trackers.
Zhimin Weng, Jinpu Zhang, Yuehuan Wang
ICME3
2024 MFRGN: Multi-scale Feature Representation Generalization Network for Ground-to-Aerial Geo-localization
abstract
Cross-area evaluation poses a significant challenge for ground-to-aerial geo-localization, in which the training and testing data are captured from entirely distinct areas. However, current methods struggle in cross-area evaluation due to their emphasis solely on learning global information from single-scale features. Some efforts alleviate this problem but rely on complex and specific technologies like pre-processing and hard sample mining. To this end, we propose a pure end-to-end solution, free from task-specific techniques, termed the Multi-scale Feature Representation Generalization Network (MFRGN) to improve generalization. Specifically, we introduce multi-scale features and explicitly utilize them by an novel global-local information representation structure with two flows, to bolster feature representations. In the global flow, we present a lightweight Self and Cross Attention Module (SCAM) to efficiently learn global embeddings. In the local flow, we develop a Global-Prompt Attention Block (GPAB) to capture discriminative features under the global embeddings as prompts. As a result, our approach generates robust descriptors representing multi-scale global and local information, thereby enhancing the model's invariance to scene variations. Extensive experiments on benchmarks show our MFRGN achieves competitive performance in same-area evaluation and improves cross-area generalization by a significant margin compared to SOTA methods. Our code is available at https://github.com/ytao-wang/MFRGN.
Jinpu Zhang, Ruonan Wei, Yuehuan Wang
ACM Multimedia5
2024 Efficient Stereo Matching Using Dynamic Graph
Yuehuan Wang
PRCV (9)2
2024 Difficulty-Aware Dynamic Network for Lightweight Exposure Correction
abstract
Recently, deep learning-based methods have been successfully applied to the field of exposure correction. However, most of the existing methods treat different locations of an image in the same way, ignoring the inhomogeneous recovery difficulty and spatially-varying visual patterns in the image, which is sub-optimal and not perfectly efficient. In this paper, we propose a difficulty-aware dynamic network (DDNet) for lightweight exposure correction. Specifically, we propose a difficulty-aware strategy that determines the difficulty of feature patches according to a difficulty mask. Then, only the difficult patches are further refined instead of the whole features, which greatly reduces the overall computational complexity. Moreover, in order to achieve spatially-varying processing with a minimal computational burden, we design a spatial-aware dynamic convolution (SDConv), which is generated by predicting a set of basic kernels and a spatial-aware weight map. Benefiting from these designs, our method can strike a good trade-off between performance and complexity. Extensive experiments on several datasets demonstrate that our approach outperforms the state-of-the-art methods both qualitatively and quantitatively while requiring cheaper computational costs.
Ziwen Li 0005, Yuanjie Shao, Feng Zhang 0039, Jinpu Zhang, Yuehuan Wang, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.5
2024 MAR: Masked Autoencoders for Efficient Action Recognition
abstract
Standard approaches for video action recognition usually operate on full input videos, which is inefficient due to the widespread spatio-temporal redundancy in videos. The recent progress in masked video modelling, specifically VideoMAE, has shown the ability of vanilla Vision Transformers (ViT) to complement spatio-temporal contexts using limited visual content. Inspired by this, we propose Masked Action Recognition (MAR), which reduces redundant computation by discarding a proportion of patches and operating only on a portion of the videos. MAR includes two essential components:cell running maskingandbridging classifier. Specifically, to enable the ViT to perceive the details beyond the visible patches, cell running masking is used to preserve the spatio-temporal correlations in videos. This ensures that the patches at the same spatial location can be observed in turn for easy reconstructions. Additionally, we notice that, although the partially observed features can reconstruct semantically explicit invisible patches, they fail to achieve accurate classification. To address this issue, we propose a bridging classifier that can help fill the semantic gap between the ViT encoded features used for reconstruction and the specialized features used for classification. Our proposed MAR can reduce the computational cost of ViT by 53%. Extensive experiments have demonstrated that MAR consistently outperforms existing ViT models by a notable margin. Notably, we found that a ViT-Large model fine-tuned by MAR achieves comparable performance to a ViT-Huge model fine-tuned by standard training methods on both Kinetics-400 and Something-Something v2 datasets. Moreover, the computation overhead of our ViT-Large model is only 14.5% of that of the ViT-Huge model. Codes have been made availablehttps://github.com/alibaba-mmai-research/Masked-Action-Recognition.
Zhiwu Qing, Shiwei Zhang 0001, Ziyuan Huang 0003, Xiang Wang 0012, Yuehuan Wang, Yiliang Lv, Changxin Gao, Nong Sang
IEEE Trans. Multim.5
2023 Learning Mutually in Crowd Scenes for Pedestrian Detection
abstract
Pedestrian detection in crowded scenes is a challenging problem due to the diverse occlusion patterns and highly overlap. To tackle this critical problem, we propose a mutual learning detection network. First, a self-attention mechanism is proposed to achieve mutual learning between individuals by capturing similar semantics among pedestrians. Feature representation of occluded individuals is enhanced by locally fusing similar semantics. Second, mutual loss is designed to improve the consistency of regression and classification. Specifically, regression results are leveraged to make classification score aware of the quality of predicted boxes, and the classification scores help the regression head to accelerate convergence of redundant boxes. Finally, we evaluate our proposed method on MOT20 and CityPersons datasets and achieve comparable state-of-the-art performance using less data. Compared to baseline, our detector obtains 14.4% AP and 11.7% AR gains on challenging MOT20 dataset.
Ruonan Wei, Yuehuan Wang, Jinpu Zhang
ICIP2
2023 Progressive Domain-style Translation for Nighttime Tracking
abstract
Nighttime tracking is challenging due to the lack of sufficient training data and scene diversity. Unsupervised domain adaptation is a solution by transferring knowledge from day (source domain) to night (target domain). It typically involves adversarial training with a domain discriminator on the source and target data to learn domain-invariant features. However, the imbalanced source/target distribution can cause overfitting of the domain discriminator, hindering the domain adaptability. To address this issue, we propose a Progressive Domain-Style Translation (PDST) for domain adaptive nighttime tracking. PDST decomposes and recombines domain-invariant content encodings and domain-specific style encodings of different domains. Thus the rich source domain content is translated to the target domain, expanding the inter-class diversity of the target domain to alleviate overfitting. Moreover, a momentum update manner is introduced to progressively estimate the domain-style encoding from multiple features, which more accurately reflects the statistical domain attribute than an individual image-style. Finally, we incorporate two regularization terms to constrain the content and domain-style consistency in the translation process, ensuring the generated source-like target features are valid to facilitate the training of domain adaptation. Exhaustive experiments demonstrate the domain adaptability and SOTA performance of the proposed method in nighttime tracking.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
ACM Multimedia4
2023 Half Aggregation Transformer for Exposure Correction
Ziwen Li 0005, Jinpu Zhang, Yuehuan Wang
PRCV (10)3
2023 A Dynamic Tracking Framework Based on Scene Perception
Jinpu Zhang, Ziwen Li 0005, Yuehuan Wang
PRCV (12)3
2023 MFA: Multi-layer Feature-aware Attack for Object Detection
abstract
Physical adversarial attacks can mislead detectors in real-world scenarios and have attracted increasing attention. However, most existing works manipulate the detector’s final outputs as attack targets while ignoring the inherent characteristics of objects. This can result in attacks being trapped in model-specific local optima and reduced transferability. To address this issue, we propose a Multi-layer Feature-aware Attack (MFA) that considers the importance of multi-layer features and disrupts critical object-aware features that dominate decision-making across different models. Specifically, we leverage the location and category information of detector outputs to assign attribution scores to different feature layers. Then, we weight each feature according to their attribution results and design a pixel-level loss function in the opposite optimized direction of object detection to generate adversarial camouflages. We conduct extensive experiments in both digital and physical worlds on ten outstanding detection models and demonstrate the superior performance of MFA in terms of attacking capability and transferability. Our code is available at: \url{https://github.com/ChenWen1997/MFA}.
Yushan Zhang, Yuehuan Wang
UAI4
2023 Low-light image enhancement with knowledge distillation
Ziwen Li 0005, Yuehuan Wang, Jinpu Zhang
Neurocomputing2
2023 Corrigendum to "Low-light image enhancement with knowledge distillation" [Neurocomputing 518 (2023) 332-343]
Ziwen Li 0005, Yuehuan Wang, Jinpu Zhang
Neurocomputing2
2023 Spatio-temporal matching for siamese visual tracking
Jinpu Zhang, Kaiheng Dai, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
Neurocomputing5
2023 SMTN: Multidimensional Fusion and Time-Domain Coding for Object Tracking in Satellite Videos
abstract
Object tracking in satellite videos contains various small targets, such as cars and ships. However, since small targets always lack salient texture features, and have low contrast to the background, it is difficult to detect targets, and distinguish different target instances, which results in tracking failure. In this letter, we propose a Siamese multidimensional fusion and time domain coding network (SMTN) with an efficient attention-based multidimensional information fusion (MDF) module and a time domain information fusion (TDF) module. The MDF module fuses multi-scale template map and search map information to make the network focus on the target, thus improving the target discriminability. In the TDF module, the aggregated temporal information of previous frames is used to adjust the current frame response map through a local-sensing Metaformer module, which suppresses the response of similar interferences. Different from other temporal fusion methods in object tracking in satellite videos, the TDF enables video satellite trackers to be optimized end-to-end, and improves the tracking success rate. Experiments demonstrate that the proposed method outperforms state-of-the-art trackers.
Jinpu Zhang, Yuehuan Wang
IEEE Geosci. Remote. Sens. Lett.3
2021 From Under- to Overexposure: Single Image Contrast Enhancement via Semi-Supervised Learning
Yuehuan Wang, Ruowang Chang
BMVC2
2021 RGNET: A Two-stage Low-light Image Enhancement Network Without Paired Supervision
abstract
Deep learning-based methods have achieved remarkable success in low-light image enhancement. However, in the absence of a large number of low/normal light image pairs, it is still a challenge to train the enhancement network with good generalization ability. In this paper, we propose a highly effective unsupervised network for low-light image enhancement (named RGNET). We divide the enhancement task into two stages, complete from coarse to precise. At the first stage, we roughly amplify the input image nonlinearly using an unsupervised network. At the second stage, we build a two-path network to restore image details, one is uesed for residual restoration and the other is used for contextual attention. With the combination of reconstruction and adversarial loss, our enhancement effects are more consistant and natural than other GAN-based methods. Both quantitative and qualitative experiments on challenging datasets demonstrate the advantages of our method in comparison with state-of-the-art methods.
Ruowang Chang, Qiong Song, Yuehuan Wang
IJCNN3
2021 YOLSO: You Only Look Small Object
Jinpu Zhang, Lei Zhang 0145, Yuehuan Wang
J. Vis. Commun. Image Represent.4
2021 Object Detection in High-Resolution Remote Sensing Images Based on a Hard-Example-Mining Network
abstract
In recent years, object detection in remote sensing images (RSIs) has attracted much attention for its application value. Compared with traditional methods that are based on manually extracted features, deep learning methods have a great advantage for object detection and have been vastly promoted. However, existing deep learning methods leave much to be desired in the field of RSI object detection due to the large-scale range of the objects and the complex image backgrounds in RSIs. Algorithms need to be specially optimized for this situation. To solve this problem, we propose an effective deep learning-based RSI object detection framework called the multiscale hard-example-mining network (MSHEMN), which is composed of three parts. First, we use the existing ResNet-50 for feature extraction. Second, we propose a multiscale region proposal network (MSRPN), which improves the existing top–down pathway feature pyramid architecture of feature pyramid network (FPN) by adding lateral connection block (LCB) and adaptive feature merge (AFM) to extract features that combine high-resolution and strong semantical information. Finally, a hard-example-mining network (HEMN), which is a cascade multistage detection network integrated with a hard example mining strategy, is proposed to make the detection network focus on hard examples during the training phase by changing the input data distribution of each stage. Extensive experiments on the High-Resolution Remote Sensing Detection (HRRSD) data set have shown the effectiveness of our proposed method, which achieves an average precision (AP) of 62.6 on the testing data set.
Lei Zhang 0145, Yuehuan Wang, Yang Huo
IEEE Trans. Geosci. Remote. Sens.2
2020 End-to-end DeepNCC framework for robust visual tracking
Kaiheng Dai, Yuehuan Wang
J. Vis. Commun. Image Represent.2
2018 Fusion of Template Matching and Foreground Detection for Robust Visual Tracking
abstract
In this paper, we present an end-to-end framework for visual tracking that contains fully convolutional template matching network and fully convolutional foreground detection network. It fuses the response maps of foreground detection and template matching for robust tracking and it can inherits all the merits of them. Besides, our network don't need additional datasets to train and only object information in the first frame is needed in training stage. We conduct extensive experiments on OTB2013 and OTB2015 and our tracker achieve state-of-the-art performance in both efficiency and accuracy.
Kaiheng Dai, Yuehuan Wang, Xiaoyun Yan, Yang Huo
ICIP2
2018 Soft Mask Correlation Filter for Visual Object Tracking
abstract
Correlation filters have shown excellent performance in visual object tracking both in speed and accuracy. However, the traditional correlation filters learn from the shifted patches rather than real background patches, which may reduce the discrimination in challenging situations. In this paper, we propose a Soft Mask Correlation Filter (SMCF) which can effectively model the object by real image patches. The soft mask is conducted on the entire frame densely and crops real background patches for training. It enables the correlation filter to pay more attention to the center part of the object rather than an axis aligned rectangle which contains background pixels. Both quantitative and qualitative evaluations conducted on tracking benchmarks demonstrate the superior performance of our method compared to the state-of-the-art trackers.
Yang Huo, Yuehuan Wang, Xiaoyun Yan, Kaiheng Dai
ICIP2
2017 Long-term object tracking based on siamese network
abstract
Although the siamese network trackers achieve competitive results both on robustness and accuracy, there is still a need to improve the overall tracking capability. In this paper, we proposed a long-term tracker based on the siamese network. We address the problem of long-term tracking where the target objects undergo significant appearance variation due to heavy deformation, occlusion, abrupt motion, and out-of-view. To tackle those problems, we suggest a multi-template fusion tracking scheme. Moreover, patch template update scheme based on optical flow are proposed to boost the overall tracking performance. The extensive results on object tracking benchmark (OTB2013) show that the proposed algorithm achieve much better performance.
Kaiheng Dai, Yuehuan Wang, Xiaoyun Yan
ICIP2
2017 Salient object detection via boosting object-level distinctiveness and saliency refinement
Xiaoyun Yan, Yuehuan Wang, Qiong Song, Kaiheng Dai
J. Vis. Commun. Image Represent.2
2016 Patch similarity based edge-preserving background estimation for single frame infrared small target detection
abstract
Edges in infrared image usually cause serious false alarms in single frame infrared small target detection. So a novel edge-preserving background estimation method is proposed for small target detection to attenuate this problem. First we will introduce the patch similarity feature of infrared image. Then, patch similarity of infrared image is utilized to formulate edge-preserving infrared background estimation. At last, estimated background will be eliminated from original infrared image to suppress edges. The effective edge-preserving ability of our approach will be shown through experiments and comparisons with state-of-the-art background estimation methods.
Yuehuan Wang, Qiong Song
ICIP2
2016 Salient object detection by multi-level features learning determined sparse reconstruction
abstract
We propose a salient object detection algorithm via multilevel features learning determined sparse reconstruction. There are three stages in our method. First, the test image are successively processed by a segmentation and semantic information generation procedures. Second, three kinds of features are extracted from semantic, global, and local levels for each superpixel to train a random forest regressor, the learned regression model is then used to generate an initial saliency map. Third, the ultimate detection result is produced using sparse reconstruction determined by the initial saliency map. Compared with most approaches, the proposed method has two obvious advantages. First, the heterogeneous regions inside salient object are often allocated similar saliency values in saliency map. Second, there are much fewer false positives in our detection results. The superior performance of our method were evaluated on four datasets with 12 state-of-the-art approaches.
Xiaoyun Yan, Yuehuan Wang, Qiong Song, Kaiheng Dai
ICIP2
2016 Multi-period visual tracking via online DeepBoost learning
Yuehuan Wang
Neurocomputing2
2014 A sub-scene modeling framework for moving cast shadow detection
abstract
In this paper, we propose an adaptive and accurate online sub-scene modeling framework for moving cast shadow detection in applications of static-camera video surveillance. To describe shadow appearance more accurately, the proposed method builds adaptive online shadow models for sub-scenes with different conditions of irradiance and reflectance. Additionally, in the correction process, object inner-edges analysis and shadow region expanding are adopted to reject shadow camouflages and recycle the misclassified shadow pixels respectively. The proposed algorithm can adaptively handle the shadow appearance changes and camouflages in both outdoor and indoor scenes without prior information about illuminations and scenarios. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Yuehuan Wang, Man Jiang, Xiaoyun Yan
ICIP2
2014 Salient region detection via color spatial distribution determined global contrasts
abstract
In this paper, we propose a novel salient region detection method via color spatial distribution determined global contrasts. First, original image is preprocessed by a texture suppression approach, and segmented into superpixels. After that, the color spatial distribution of all superpixels is computed. Then, based on values of the distribution in whole image and boundaries of image, some superpixels are determined as foreground and background queries. Next, two global contrasts based on these queries are computed respectively to produce two different saliency maps. Ultimately, color spatial distribution and the two saliency maps are accumulated to generate final saliency map. Our approach is evaluated on M-SRA 1000 dataset, and the experimental results demonstrate superior performance of our method to eight state-of-the-art approaches.
Xiaoyun Yan, Yuehuan Wang, Man Jiang
ICIP2
2014 Moving cast shadow detection using online sub-scene shadow modeling and object inner-edges analysis
Yuehuan Wang, Man Jiang, Xiaoyun Yan
J. Vis. Commun. Image Represent.2
2010 A Biologically-Inspired Top-Down Learning Model Based on Visual Attention
abstract
A biologically-inspired top-down learning model based on visual attention is proposed in this paper. Low-level visual features are extracted from learning object itself and do not depend on the background information. All the features are expressed as a feature vector, which is looked as a random variable following a normal distribution. So every learning object is represented as the mean and standard deviation. All the learning objects are combined as an object class, which is represented as class's mean and class's standard deviation stored in long-term memory (LTM). Then the learned knowledge is used to find the similar location in an attended image. Experimental results indicate that: when the attended object doesn't always appear in the background similar to that in the learning objects or their combinations change hugely between learning images and attended images, our model is excellent to other two top-down visual attention models.
Nong Sang, Longsheng Wei, Yuehuan Wang
ICPR3
2010 Motion Detection Based on Biological Correlation Model
Nong Sang, Yuehuan Wang, Qingqing Zheng
ISNN (2)3