Wen Su 0004

dblp:06/7451-4 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-6787-4384ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 To-Former: semantic segmentation of transparent object with edge-enhanced transformer
Wen Su 0004, Mengjiao Ge, Jun Yu 0001
Vis. Comput.2
2024 Image Harmonization Based on Hierarchical Dynamics
abstract
Image harmonization is an essential technique in computer vision, aiming to generate visually consistent composite images by making the foreground compatible with the background. However, current methods primarily focus on applying a global transformation perspective, overlooking the fact that different regions in a real image can exhibit significant appearance variations. Yet, there is consistency within local regions. They also have limited representation ability by using fixed background statistics (e.g., mean, and standard deviation) for foreground normalization. Hence, we propose a hierarchical dynamics appearance translation strategy that adjusts the foreground appearance based on the corresponding background, adapting the model features and parameters from local to global view. To enhance the representation ability for targets, we employ a mixed attention mechanism for local dynamics, which adaptively modifies the features of different channels and positions. Additionally, we apply dynamic region-aware convolution guided by the foreground mask for global dynamics, which learns the adaptive representation of the foreground and background and correlations to global harmonization. To further improve the harmonization result, we integrate adversarial and perceptual loss into the model training. Experiments show our method significantly reduces parameters and achieves state-of-the-art performance compared with previous methods.
Liuxue Ju, Chengdao Pu, Jun Yu 0001, Wen Su 0004
ICASSP4
2024 Rotated R-CNN: A Two-Stage Object Detection Method Adapted To Oriented Bounding Boxes
abstract
Currently, oriented object detection, as an emerging subfield within object detection, has garnered significant attention. Besides encompassing directional information, datasets of oriented objects exhibit notable characteristics, including significant variations in object scales and a wide range of aspect ratios for ground-truth bounding boxes. Nevertheless, the current state-of-the-art two-stage rotating object detection models have not sufficiently addressed these characteristics, leading to inherent limitations in accuracy. In response to these challenges, we introduce the Rotated RCNN. Our model is the first to introduce trainable anchors in the field of oriented object detection to achieve anchor distributions similar to the ground truth boxes in the oriented object dataset. Furthermore, considering the distinctive traits of oriented ground truth boxes, we have devised a novel strategy for assigning labels to more effectively choose positive and negative samples specifically designed for oriented objects. In the regression phase of the RPN, we introduce shape constraints to alleviate accuracy losses stemming from mismatches between the encoding method and oriented objects. We comprehensively evaluate our model on the DOTAv1.0 and HRSC2016 datasets, demonstrating the effectiveness of our meticulously designed model.
Chengdao Pu, Jun Yu 0001, Wen Su 0004
ICIP3
2024 Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions
abstract
The continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model.
Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang
ACM Multimedia7
2024 EG-Trans: Transparent Object Segmentation with Edge Enhanced and Global Integrated Transformers
Wen Su 0004
PRCV (9)2
2024 Camouflage Object Segmentation with Multi-scale Feature Aggregation and Boundary Generation
Wen Su 0004, Jinfeng Gao 0002, Guoqiang Jia
PRCV (8)2
2024 MDEConvFormer: estimating monocular depth as soft regression based on convolutional transformer
Wen Su 0004, Haifeng Zhang 0006, Wenzhen Yang
Multim. Tools Appl.1
2022 Monocular depth estimation with spatially coherent sliced network
Wen Su 0004, Haifeng Zhang 0006, Yuan Su, Jun Yu 0001, Zengfu Wang
Image Vis. Comput.1
2022 SSR-HEF: Crowd Counting With Multiscale Semantic Refining and Hard Example Focusing
abstract
Crowd counting based on density maps is generally regarded as a regression task. Deep learning is used to learn the mapping between image content and crowd density distribution. Although great success has been achieved, some pedestrians far away from the camera are difficult to be detected. And the number of hard examples is often larger. Existing methods with simple Euclidean distance algorithm indiscriminately optimize the hard and easy examples so that the densities of hard examples are usually incorrectly predicted to be lower or even zero, which results in large counting errors. To address this problem, we are the first to propose the hard example focusing (HEF) algorithm for the regression task of crowd counting. The HEF algorithm makes our model rapidly focus on hard examples by attenuating the contribution of easy examples. Then higher importance will be given to the hard examples with wrong estimations. Moreover, the scale variations in crowd scenes are large, and the scale annotations are labor-intensive and expensive. By proposing a multiscale semantic refining strategy, lower layers of our model can break through the limitation of deep learning to capture semantic features of different scales to sufficiently deal with the scale variation. We perform extensive experiments on six benchmark datasets to verify the proposed method. Results indicate the superiority of our proposed method over the state-of-the-art methods. Moreover, our designed model is smaller and faster.
Kewei Wang 0004, Wen Su 0004, Zengfu Wang
IEEE Trans. Ind. Informatics3
2021 Monocular Depth Estimation Using Information Exchange Network
abstract
Depth estimation from single monocular image attracts increasing attention in autonomous driving and computer vision. While most existing approaches regress depth values or classify depth labels based on features extracted from limited image area, the resulting depth maps are still perceptually unsatisfying. Neither local context nor low-level semantic information is sufficient to predict depth. Learning based approaches suffer from inherent defects of supervision signals. This paper addresses monocular depth estimation with a general information exchange convolutional neural network. We maintain a high-resolution prediction throughout the network. Meanwhile, both low-resolution features capturing long-range context and fine-grained features describing local context can be refined with information exchange path stage by stage. Mutual channel attention mechanism is applied to emphasize interdependent feature maps and improve the feature representation of specific semantics. The network is trained under the supervision of improved log-cosh and gradient constraints so that the abnormal predictions have less impacts and the estimation can be consistent in high order. The results of ablation studies verify the efficiency of every proposed components. Experiments on the popular indoor and street-view datasets show competitive results compared with the recent state-of-the-art approaches.
Wen Su 0004, Haifeng Zhang 0006, Quan Zhou 0004, Wenzhen Yang, Zengfu Wang
IEEE Trans. Intell. Transp. Syst.1
2020 Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose Estimation
abstract
We rethink a well-known bottom-up approach for multi-person pose estimation and propose an improved one. The improved approach surpasses the baseline significantly thanks to (1) an intuitional yet more sensible representation, which we refer to as body parts to encode the connection information between keypoints, (2) an improved stacked hourglass network with attention mechanisms, (3) a novel focal L2 loss which is dedicated to “hard” keypoint and keypoint association (body part) mining, and (4) a robust greedy keypoint assignment algorithm for grouping the detected keypoints into individual poses. Our approach not only works straightforwardly but also outperforms the baseline by about 15% in average precision and is comparable to the state of the art on the MS-COCO test-dev dataset. The code and pre-trained models are publicly available on our project page1.
Jia Li 0026, Wen Su 0004, Zengfu Wang
AAAI2
2020 Weakly Supervised Local-Global Relation Network for Facial Expression Recognition
abstract
To extract crucial local features and enhance the complementary relation between local and global features, this paper proposes a Weakly Supervised Local-Global Relation Network (WS-LGRN), which uses the attention mechanism to deal with part location and feature fusion problems. Firstly, the Attention Map Generator quickly finds the local regions-of-interest under the supervision of image-level labels. Secondly, bilinear attention pooling is employed to generate and refine local features. Thirdly, Relational Reasoning Unit is designed to model the relation among all features before making classification. The weighted fusion mechanism in the Relational Reasoning Unit makes the model benefit from the complementary advantages between different features. In addition, contrastive losses are introduced for local and global features to increase the inter-class dispersion and intra-class compactness at different granularities. Experiments on lab-controlled and real-world facial expression dataset show that WS-LGRN achieves state-of-the-art performance, which demonstrates its superiority in FER.
Haifeng Zhang 0006, Wen Su 0004, Jun Yu 0001, Zengfu Wang
IJCAI2
2020 Crowd counting with crowd attention convolutional neural network
Wen Su 0004, Zengfu Wang
Neurocomputing2
2019 Expression-identity Fusion Network for Facial Expression Recognition
abstract
Research shows that the facial expression recognition is strongly related to the person's identity. This paper presents an expression-identity fusion network to address the great inter-subject variations in facial expression recognition. The model is designed to jointly learn identity-related features and expression-related features via two branches with the same expression image input. A bilinear module is introduced to fuse two kinds of features and learn the relationship between them. Experimental results show that identity-related features can greatly boost the performance of facial expression recognition. Our method outperforms most of the state-of-the-art. On two popular facial expression databases (CK+ and Oulu-CASIA), our method achieves 96.02% and 85.21% recognition accuracy, respectively.
Haifeng Zhang 0006, Wen Su 0004, Zengfu Wang
ICASSP2
2019 Monocular Depth Estimation as Regression of Classification using Piled Residual Networks
abstract
Predicting depth from single monocular image is a challenging task in scene understanding. Most existing work predicts depth by regression or classification with features extracted from local neighborhood area. However, neither regression nor classification achieves the final satisfying solution and local context can be insufficient to predict the depth. This paper innovatively addresses this problem as regression of class related features on a piled residual convolutional neural network. Our framework works at two stages. First, a well-designed deep convolutional neural network model is employed to classify the depths in difference-scale invariance space. The model utilizes all scales of context though piled residual paths. The deeper layers that capture high-level semantic features with long-range context can be directly refined using fine-grained features with local context from earlier convolutions. We then apply centered information gain loss to the model to produce intra-class compact and inter-class discriminative features. Second, to obtain depths instead of class labels, we infer depth regression with convolutional layers which model the mapping from class discriminative features to continuous depth values. Experiments on the popular indoor and outdoor datasets show competitive results compared with the recent state of the art methods.
Wen Su 0004, Haifeng Zhang 0006, Jia Li 0013, Wenzhen Yang, Zengfu Wang
ACM Multimedia1
2019 Widening residual refine edge reserved neural network for semantic segmentation
Wen Su 0004, Zengfu Wang
Multim. Tools Appl.1
2017 Widening residual skipped network for semantic segmentation
abstract
Over the past two years deep convolutional neural networks have pushed the performance of computer vision systems to soaring heights on semantic segmentation. In this study, the authors present a novel semantic segmentation method of using a deep fully convolutional neural network to achieve image segmentation results with more precise boundary localisation. The above segmentation engine is trainable, and consists of an encoder network with widening residual skipped connections and a decoder network with a pixel‐wise classification layer. Here the encoder network with widening residual skipped connections allows the combination of shallow layer features and deep layer semantic features, and the decoder network with classification layer maps the low‐resolution encoder features to full resolution image with pixel‐wise classification. The experimental results on PASCAL VOC 2012 semantic segmentation dataset and Cityscapes dataset show that the proposed method is effective and competitive.
Wen Su 0004, Zengfu Wang
IET Image Process.1
2016 Regularized fully convolutional networks for RGB-D semantic segmentation
abstract
The prospect of semantic segmentation using depth is alluring. In most of the previous work features are only combined by using simple classification strategies. The inter-feature and inter-label relationships have been ignored. This paper proposes a novel unified framework for RGB-D semantic segmentation. We use regularized fully convolutional networks whose inputs are depth map and hand-crafted features. Relationships between those features and their labels are learnt and utilized by rigorously imposing regularization in fully connected layers. The regularized fully convolutional networks can be efficiently launched using a GPU implementation at an affordable training cost. Experiments demonstrate that our regularized fully convolutional networks taking features as inputs obtain competitive results on the PASCAL VOC 2011 dataset and NYUDv2.
Wen Su 0004, Zengfu Wang
VCIP1