Yun Wu 0001

dblp:32/5387-1 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0002-2476-7746ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FCD-Net: Frequency and Contrastive Learning-Driven Network for Document Image Shadow Removal
Nanfeng Jiang, Dahan Wang, Yun Wu 0001
ICDAR (2)4
2025 Dual Manifold Volume-Balanced Framework for Long-Tailed Oracle Character Recognition
Tianyu Fang, Kunchi Li, Yun Wu 0001, Dahan Wang
PRCV (7)3
2025 Dual-Branch Residual Wavelet Attention Network for Colorectal Cancer Magnifying Endoscopy Image Classification
Linrui He, Yun Wu 0001, Dahan Wang, Shunzhi Zhu, Xuyao Zhang
PRCV (14)2
2025 LDH-Net: Luminance-based Deep Hybrid Network for Document Image De-shadowing
Kunchi Li, Nanfeng Jiang, Yun Wu 0001, Dahan Wang
Image Vis. Comput.4
2025 ADR-Net: Attention-oriented detail recovery network for document image shadow removal
Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Yun Wu 0001, Shunzhi Zhu
Knowl. Based Syst.5
2024 Document Image Shadow Removal via Frequency Information-Oriented Network
Xinyue Zhou, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Guantin Li, Wang Man, Yun Wu 0001
ICPR (31)8
2024 MOA-Net: Multilevel Object Aware Network for Remote Sensing Image Semantic Segmentation
abstract
Remote sensing image semantic segmentation is an essential aspect of the intelligent analysis of remote sensing, extensively applied in urban planning, economic assessment, and disaster monitoring. However, the expansive field of view and intricate backgrounds in remote sensing images cause numerous objects of varying sizes and categories to coexist. This results in incomplete object segmentation and poses challenges in restoring the spatial distribution of objects. In this paper, we present a Multilevel Object Aware Network (MOA-Net), which addresses the challenges of remote sensing image semantic segmentation. To be specific, this method is designed from three perspectives. Firstly, we establish a Progressive Multiscale Global-Local Decoder (PMGLD), integrating global-local context information of objects at varying scales through a progressive convolution strategy. Secondly, the Orientation-Aware Attention Mechanism (OAAM) provides orientation information and guides the restoration of interclass 2-D spatial relationships. Finally, obtaining fine-grained features through CNN Stem improves edge segmentation. Experimental outcomes on the Potsdam and Vaihingen datasets indicate that our method of performance and efficiency surpass those of existing methods.
Yun Wu 0001, Dahan Wang
IJCNN2
2024 RCFormer : Interactive Image Segmentation via Reconstructing Click Vision Transformers
abstract
Click-based interactive image segmentation intends to segment an object from the background under user click guidance. Recently, Vision Transformer has made significant strides in interactive image segmentation. However, the previous studies 1) overlook the importance of different clicks in terms of their contribution to the segmentation results; and 2) suffer from inconsistency across different feature scales in the multi-scale structure. In this paper, we propose a new interactive segmentation framework, named RCFormer, with two novel components: reconstruct click patch embedding (RCPE) for encoding the importance of clicks, and multi-scale adaptive fusion (MSAF) for the adaptive fusion of feature maps across different scales. RCPE enhances the effectiveness of click interactions by spatially distinguishing the importance of clicks. MSAF adaptively fuses useful spatial information and filters the redundant feature at multi-scales. The experiments on several benchmarks show that our proposed approach achieves state-of-the-art performance. Notably, our method achieves 2.31 NoC@90 on the Berkeley dataset, improving by 8.6% over the previous best results.
PanPan Chen, Dahan Wang, Yun Wu 0001, Xu-Yao Zhang, Shunzhi Zhu
IJCNN4
2024 Character Relationship Refinement Network for Handwritten Mathematical Expression Recognition
abstract
Most current Handwritten Mathematical Expression Recognition (HMER) methods employ an attention-based encoder-decoder framework, which generates LaTeX sequences from the given images, following the paradigm of predicting "one-by-one". However, this paradigm may have some challenges: 1) without considering the connectivity between characters, the prior information in the prediction process will be ignored inadvertently, especially implicit information, such as " " and " ˆ ". 2) Some characters of high similarities, such as "6/b" and "o/O", will have negative effects on prediction results. To solve these issues, we propose a simple but effective Character Relationship Refinement Network (CRRN), which consists of Joint Character Learning (JCL) and Character Refinement Mask (CRM). Specifically, JCL calculates the relationship probability between characters and uses them to improve prediction accuracy. CRM takes the character confidence coefficient in a coarse-to-fine way that can reassign the weights of all characters to improve model discriminability on easily confused characters. With the collaboration of both modules, our proposed CRRN can outperform the state-of-the-art on popular datasets.
LiWei Jiang, Nanfeng Jiang, Yun Wu 0001, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu
IJCNN3
2024 Multi-Branch Enhanced Discriminative Network for Vehicle Re-Identification
abstract
Vehicle re-identification (ReID) is the task of identifying the same vehicle across numerous cameras. This is a complex classification task, and the fine-grained information and strong discrimination features have proven to be effective in handling the re-identification classification task. However, most existing methods focuses on extracting local area features in combination with global features, while exploring subtle distinguishing features, which is a difficult task, remains an open problem and unsolved. In this paper, we propose a multi-branch enhanced discriminative network (MED) to better extract subtle distinguishing features that have high discriminative power to improve the ReID performance. In the proposed MED method, each feature map obtained by convolutional neural network (CNN) is divided into 4 spatial sub-maps, on each of which, the vertical and the horizontal branches are used to extract the subtle distinguishing features intrinsically contained in sub-areas. The vertical and the horizontal branches are combined with the global branch to perform the ReID task. Moreover, our proposed method is capable of extracting rich fine-grained features without the need of extra manual annotation while maintaining a simple design structure. We conducted extensive experiments on the vehicle ReID datasets (VehicleID and VeRi-776), showing that the proposed MED method outperforms most existing methods. Further, we directly apply the MED method to the pedestrian ReID problem on the Market-1501, DUKEMTMC, and MSMT17 datasets, achieving the state-of-the-art (SOTA) performance as well. This demonstrates that the proposed method has good generality and can be flexibly applied to the ReID tasks.
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.3
2023 Multi-Zone Transformer Based on Self-Distillation for Facial Attribute Recognition
abstract
Recently, transformers have shown great promising performance in various computer vision tasks. However, the current transformer based methods ignore the information exchanges between transformer blocks, and they have not been applied in the facial attribute recognition task. In this paper, we propose a multi-zone transformer based on self-distillation for FAR, termed MZTS, to predict the facial attributes. A multi-zone transformer encoder is firstly presented to achieve the interactions of the different transformer encoder blocks, thus avoiding forgetting the effective information between the transformer encoder block groups during the iteration process. Furthermore, we introduce a new self-distillation mechanism based on class tokens, which distills the class tokens obtained from the last transformer encoder block group to the other shallow groups by interacting with the significant information between the different transformer blocks through attention. Extensive experiments on the challenging CelebA and LFWA datasets have demonstrated the excellent performance of the proposed method for FAR.
Si Chen 0002, Xueyan Zhu, Dahan Wang, Shunzhi Zhu, Yun Wu 0001
FG5
2023 UAM-Net: An Attention-Based Multi-level Feature Fusion UNet for Remote Sensing Image Segmentation
Yiwen Cao, Nanfeng Jiang, Dahan Wang, Yun Wu 0001, Shunzhi Zhu
PRCV (4)4
2023 Pseudo Labels Refinement with Stable Cluster Reconstruction for Unsupervised Re-identification
Jiawei Lian, Dahan Wang, Yun Wu 0001, Shunzhi Zhu, Dewu Ge
PRCV (4)5
2023 HC-GCN: hierarchical contrastive graph convolutional network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Bolun Xu, Yan Yan 0001, Xia Du, Weiwei Zhuang, Yun Wu 0001
Multim. Syst.7
2022 Exploiting Robust Memory Features for Unsupervised Reidentification
Jiawei Lian, Dahan Wang, Xia Du, Yun Wu 0001, Shunzhi Zhu
PRCV (2)4