Yuan Wu 0004

dblp:41/5176-4 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 FusionFormer: A Concise Unified Feature Fusion Transformer for 3D Pose Estimation
abstract
Depth uncertainty is a core challenge in 3D human pose estimation, especially when the camera parameters are unknown. Previous methods try to reduce the impact of depth uncertainty by multi-view and/or multi-frame feature fusion to utilize more spatial and temporal information. However, they generally lead to marginal improvements and their performance still cannot match the camera-parameter-required methods. The reason is that their handcrafted fusion schemes cannot fuse the features flexibly, e.g., the multi-view and/or multi-frame features are fused separately. Moreover, the diverse and complicated fusion schemes make the principle for developing effective fusion schemes unclear and also raises an open problem that whether there exist more simple and elegant fusion schemes. To address these issues, this paper proposes an extremely concise unified feature fusion transformer (FusionFormer) with minimized handcrafted design for 3D pose estimation. FusionFormer fuses both the multi-view and multi-frame features in a unified fusion scheme, in which all the features are accessible to each other and thus can be fused flexibly. Experimental results on several mainstream datasets demonstrate that FusionFormer achieves state-of-the-art performance. To our best knowledge, this is the first camera-parameter-free method to outperform the existing camera-parameter-required methods, revealing the tremendous potential of camera-parameter-free models. These impressive experimental results together with our concise feature fusion scheme resolve the above open problem. Another appealing feature of FusionFormer we observe is that benefiting from its effective fusion scheme, we can achieve impressive performance with smaller model size and less FLOPs.
Yanlu Cai, Yuan Wu 0004, Cheng Jin 0001
AAAI3
2024 PoseIRM: Enhance 3D Human Pose Estimation on Unseen Camera Settings via Invariant Risk Minimization
abstract
Camera-parameter-free multi-view pose estimation is an emerging technique for 3D human pose estimation (HPE). They can infer the camera settings implicitly or explicitly to mitigate the depth uncertainty impact, showcasing significant potential in real applications. However, due to the limited camera setting diversity in the available datasets, the inferred camera parameters are always simply hard-coded into the model during training and not adaptable to the input in inference, making the learned models cannot generalize well under unseen camera settings. A natural solution is to artificially synthesize some samples, i.e., 2D-3D pose pairs, under massive new camera settings. Un-fortunately, to prevent over-fitting the existing camera setting, the number of synthesized samples for each new camera setting should be comparable with that for the existing one, which multiplies the scale of training and even makes it computationally prohibitive. In this paper, we propose a novel HPE approach under the invariant risk minimization (IRM) paradigm. Precisely, we first synthesize 2D poses from myriad camera settings. We then train our model under the IRM paradigm, which targets at learning a common optimal model across all camera settings and thus enforces the model to automatically learn the camera parameters based on the input data. This allows the model to accurately infer 3D poses on unseen data by training on only a hand-ful of samples from each synthesized setting and thus avoid the unbearable training cost increment. Another appealing feature of our method is that benefited from the capability of IRM in identifying the invariant features, its performance on the seen camera settings is enhanced as well. Compre-hensive experiments verify the superiority of our approach.
Yanlu Cai, Yuan Wu 0004, Cheng Jin 0001
CVPR3
2023 ETR: An Efficient Transformer for Re-ranking in Visual Place Recognition
abstract
Visual place recognition is to estimate the geographical location of a given image, which is usually addressed by recognizing its similar reference images from a database. The reference images are usually retrieved via similarity search using global descriptor, and the local descriptors are used to re-rank the initial retrieved candidates. The local descriptors re-ranking can significantly improve the accuracy of global retrieval but comes at a high computational cost. To achieve a good trade-off between accuracy and efficiency, we propose an Efficient Transformer for Re-ranking (ETR), utilizing both global and local descriptors to re-rank the top candidates in a single shot. In contrast to traditional re-ranking methods, we leverage self-attention to capture relationships between local descriptors in a single image and cross-attention to explore the similarity of the image pairs. We show that the proposed model can be regarded as a general re-ranking algorithm for significantly boosting the performance of other global-only retrieval methods. Extensive experimental results show that our method outperforms state-of-the-arts and is orders of magnitude faster in terms of computational efficiency.
Heming Jing, Yingbin Zheng, Yuan Wu 0004, Cheng Jin 0001
WACV5
2022 Attention-Based Transformation from Latent Features to Point Clouds
abstract
In point cloud generation and completion, previous methods for transforming latent features to point clouds are generally based on fully connected layers (FC-based) or folding operations (Folding-based). However, point clouds generated by FC-based methods are usually troubled by outliers and rough surfaces. For folding-based methods, their data flow is large, convergence speed is slow, and they are also hard to handle the generation of non-smooth surfaces. In this work, we propose AXform, an attention-based method to transform latent features to point clouds. AXform first generates points in an interim space, using a fully connected layer. These interim points are then aggregated to generate the target point cloud. AXform takes both parameter sharing and data flow into account, which makes it has fewer outliers, fewer network parameters, and a faster convergence speed. The points generated by AXform do not have the strong 2-manifold constraint, which improves the generation of non-smooth surfaces. When AXform is expanded to multiple branches for local generations, the centripetal constraint makes it has properties of self-clustering and space consistency, which further enables unsupervised semantic segmentation. We also adopt this scheme and design AXformNet for point cloud completion. Considerable experiments on different datasets show that our methods achieve state-of-the-art results.
Kaiyi Zhang 0002, Ximing Yang, Yuan Wu 0004, Cheng Jin 0001
AAAI3
2022 ST2PE: Spatial and Temporal Transformer for Pose Estimation
Yuan Wu 0004, Yanlu Cai, Rui Feng 0001, Cheng Jin 0001
ICANN (2)1
2022 GLTA-GCN: Global-Local Temporal Attention Graph Convolutional Network for Unsupervised Skeleton-Based Action Recognition
abstract
Unsupervised skeleton-based action recognition has attracted increasing attention. Existing methods have several limitations: (1) Many actions are highly related to local joints, which is often neglected. (2) Most methods directly employ joint coordinates as frame feature and do not utilize skeleton graph, e.g., topological information. (3) Long-range dependency is not captured well. In this work, a novel unsupervised method called Global-Local Temporal Attention Graph Convolutional Network (GLTA-GCN) is proposed to alleviate the above problems. The network consists of two branches, local and global branches. Each one utilizes graph convolution units and self-attention mechanism to better extract spatio-temporal features. Furthermore, two loss functions are designed to constrain the model to extract more essential local joint feature and maintain intrinsic structural information. Extensive experiments demonstrate that GLTA-GCN achieves state-of-the-art performance. Our code is released on https://github.com/HaoyueQiu/GLTA-GCN.
Haoyue Qiu, Yuan Wu 0004, Mengmeng Duan, Cheng Jin 0001
ICME2
2022 Memory Enhanced Spatial-Temporal Graph Convolutional Autoencoder for Human-Related Video Anomaly Detection
Sibo Luo, Shangshang Wang, Yuan Wu 0004, Cheng Jin 0001
PRCV (3)3
2021 CPCGAN: A Controllable 3D Point Cloud Generative Adversarial Network with Semantic Label Generating
abstract
Generative Adversarial Networks (GAN) are good at generating variant samples of complex data distributions. Generating a sample with certain properties is one of the major tasks in the real-world application of GANs. In this paper, we propose a novel generative adversarial network to generate 3D point clouds from random latent codes, named Controllable Point Cloud Generative Adversarial Network(CPCGAN). A two-stage GAN framework is utilized in CPCGAN and a sparse point cloud containing major structural information is extracted as the middle-level information between the two stages. With their help, CPCGAN has the ability to control the generated structure and generate 3D point clouds with semantic labels for points. Experimental results demonstrate that the proposed CPCGAN outperforms state-of-the-art point cloud GANs.
Ximing Yang, Yuan Wu 0004, Kaiyi Zhang 0002, Cheng Jin 0001
AAAI2
2021 NTU-DensePose: A New Benchmark for Dense Pose Action Recognition
abstract
Skeleton-based action recognition has recently gained a lot of attention in computer vision. The previous skeleton-based datasets used sparse poses to represent the human body, which always leads to a large loss of human body detail information. Therefore, the previous skeleton-based methods generally performed worse than the image-based methods. In this paper, we propose a dense-pose-based action recognition dataset NTU-DensePose. This dataset automatically annotates 37,060 video samples with two dense poses, IUV equidistant annotation and IUV equivalent annotation. Each dense pose annotation contains more than 240 keypoints per instance. So the dense-pose-based action recognition method can capture more subtle details and predict human action more accurately than the previous skeleton-based methods. To the best of our knowledge, NTU-DensePose is the first dense-pose-based action recognition dataset.
Mengmeng Duan, Haoyue Qiu, Zimo Zhang, Yuan Wu 0004
IEEE BigData4
2021 SOF: A Synthetic Occluded Face Dataset
abstract
In this paper, we propose an occluded face dataset named SOF (Synthetic Occluded Face) and describe in detail the construction method of SOF. We synthesize the occluded image into the face image after the thin plate spline deformation to obtain the occluded face image, and then make the image more real through guided filter. At the end of this paper, we use the SOF training set and general face datasets Ms-Celeb-1M, CASIA-WebFace to train the mainstream face recognition algorithms FaceNet, SphereFace and ArcFace, and test on a variety of test sets. A series of experiments proved that the occluded face dataset generated by this method could improve the accuracy of occluded face recognition of mainstream face recognition algorithms.
Mengmeng Duan, Lurui Jin, Yuan Wu 0004
IEEE BigData4
2020 OSD: An Occlusion Skeleton Dataset for Action Recognition
abstract
Currently available 2D skeleton datasets for action recognition mostly contain nonoccluded skeleton samples. Models trained on such datasets lack generalization ability in occlusion situations. In this paper we propose an occlusion projection method, which projects a 3D occlusion object into 2D plane to generate a 2D occluded area. Based on this method, we build a 2D occlusion skeleton dataset named OSD with 56,800 occluded skeleton samples and 60 distinct classes. Experimental results show that the model trained on OSD has better generalization ability in occlusion situations compared with the model trained on datasets with nonoccluded samples, which proves the effectiveness of OSD.
Yuan Wu 0004, Haoyue Qiu, Rui Feng 0001
IEEE BigData1