Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xuesheng Qian

dblp:192/4158 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0001-6225-7889ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 54% Image recognition and object detection · 26% 3D vision · 11%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
object tracking
1.222023
Effective Local and Global Search for Fast Long-Term Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Video Annotation for Visual Tracking via Selection and Refinement · ICCV 2021
Computer vision › 3D vision
human mesh recovery
0.712023
3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature Aggregation · IEEE Trans. Multim. 2023
Computer vision › Video understanding and tracking › object tracking › robust tracking
long-term tracking
0.712023
Effective Local and Global Search for Fast Long-Term Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Video understanding and tracking › video object detection
temporal feature aggregation
0.712023
3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature Aggregation · IEEE Trans. Multim. 2023
Computer vision › Image recognition and object detection › object detection
bounding box annotation
0.512021
Video Annotation for Visual Tracking via Selection and Refinement · ICCV 2021
Computer vision › Image recognition and object detection › object counting
crowd counting
0.512021
Spatial Uncertainty-Aware Semi-Supervised Crowd Counting · ICCV 2021
Computer vision › Image recognition and object detection › object counting › crowd counting
semi-supervised crowd counting
0.512021
Spatial Uncertainty-Aware Semi-Supervised Crowd Counting · ICCV 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Spatial Uncertainty-Aware Semi-Supervised Crowd Counting · ICCV 2021
Computer vision › Video understanding and tracking
video annotation
0.512021
Video Annotation for Visual Tracking via Selection and Refinement · ICCV 2021

Methods — techniques the papers use, named apart from their topics

template update · 0.7deep neural network · 0.7confidence estimation · 0.7bounding box regression · 0.7visual-geometry refinement network · 0.5temporal assessment network · 0.5teacher-student framework · 0.5surrogate tasks · 0.5differential transformation layer · 0.5
YearPublicationVenuePosition
2024 A Video Is Worth Three Views: Trigeminal Transformers for Video-Based Person Re-Identification
abstract
Video-based person Re-Identification (Re-ID) is a hot research topic in intelligent transportation systems, which aims to retrieve video sequences of the same person under non-overlapping surveillance cameras. Compared with static images, video sequences contain more visual information from multiple views, such as spatial and temporal views. However, previous Re-ID methods usually focus on single limited views, lacking diverse observations from different views. To capture richer perceptions and extract more comprehensive representations, we propose a novel learning framework namedTrigeminal Transformers (TMT)to tackle video-based person Re-ID. More specifically, we first design aView-wise Projector (VP)to jointly transform raw videos from spatial, temporal and spatial-temporal views. In addition, inspired by the great success of Vision Transformers (ViT), we introduce the Transformer structure for information enhancement and aggregation. In our work, threeSelf-view Transformers (ST)are proposed to exploit the relationships of local features for information enhancement in spatial, temporal and spatial-temporal. Moreover, aCross-view Transformer (CT)is proposed to aggregate the multi-view features for comprehensive representations. Experimental results indicate that our approach can obtain better performance than some other state-of-the-art approaches on four public Re-ID benchmarks.
Xuehu Liu, Chenyang Yu, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu
IEEE Trans. Intell. Transp. Syst.4
2023 Effective Local and Global Search for Fast Long-Term Tracking
abstract
Compared with short-term tracking, long-term tracking remains a challenging task that usually requires the tracking algorithm to track targets within a local region and re-detect targets over the entire image. However, few works have been done and their performances have also been limited. In this paper, we present a novel robust and real-time long-term tracking framework based on the proposed local search module and re-detection module. The local search module consists of an effective bounding box regressor to generate a series of candidate proposals and a target verifier to infer the optimal candidate with its confidence score. For local search, we design a long short-term updated scheme to improve the target verifier. The verification capability of the tracker can be improved by using several templates updated at different times. Based on the verification scores, our tracker determines whether the tracked object is present or absent and then chooses the tracking strategies of local or global search, respectively, in the next frame. For global re-detection, we develop a novel re-detection module that can estimate the target position and target size for a given base tracker. We conduct a series of experiments to demonstrate that this module can be flexibly integrated into many other tracking algorithms for long-term tracking and that it can improve long-term tracking performance effectively. Numerous experiments and discussions are conducted on several popular tracking datasets, including VOT, OxUvA, TLP, and LaSOT. The experimental results demonstrate that the proposed tracker achieves satisfactory performance with a real-time speed. Code is available at https://github.com/difhnp/ELGLT.
Haojie Zhao, Bin Yan 0004, Dong Wang 0004, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature Aggregation
abstract
Estimating 3D human body shapes and poses from videos is a challenging computer vision task. The intrinsic temporal information embedded in adjacent frames is helpful in making accurate estimations. Existing approaches learn temporal features of the target frames simply by aggregating features of their adjacent frames, using off-the-shelf deep neural networks. Consequently these approaches cannot explicitly and effectively use the correlations between adjacent frames to help infer the parameters of the target frames. In this paper, we propose a novel framework that can measure the correlations amongst adjacent frames in the form of an estimated confidence metric. The confidence value will indicate to what extent the adjacent frames can help predict the target frames’ 3D shapes and poses. Based on the estimated confidence values, temporally aggregated features are then obtained by adaptively allocating different weights to the temporal predicted features from the adjacent frames. The final 3D shapes and poses are estimated by regressing from the temporally aggregated features. Experimental results on three benchmark datasets show that the proposed method outperforms state-of-the-art approaches (even without the motion priors involved in training). In particular, the proposed method is more robust against corrupted frames.
Hongrun Zhang, Yanda Meng, Yitian Zhao, Xuesheng Qian, Yihong Qiao, Xiaoyun Yang, Yalin Zheng
IEEE Trans. Multim.4
2022 Secure olympics games with technology: Intelligent border surveillance for the 2022 Beijing winter olympics
Mengfan Chen, Xuesheng Qian
J. Syst. Archit.5
2021 BI-GCN: Boundary-Aware Input-Dependent Graph Convolution Network for Biomedical Image Segmentation
Yanda Meng, Hongrun Zhang, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng
BMVC6
2021 Video Annotation for Visual Tracking via Selection and Refinement
abstract
Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data.
Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang
ICCV7
2021 Spatial Uncertainty-Aware Semi-Supervised Crowd Counting
abstract
Semi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via regularized surrogate task (binary segmentation) for crowd counting problems. Different from existing semi-supervised learning-based crowd counting methods, to exploit the unlabeled data, our proposed spatial uncertainty-aware teacher-student framework focuses on high confident regions’ information while addressing the noisy supervision from the unlabeled data in an end-to-end manner. Specifically, we estimate the spatial uncertainty maps from the teacher model’s surrogate task to guide the feature learning of the main task (density regression) and the surrogate task of the student model at the same time. Besides, we introduce a simple yet effective differential transformation layer to enforce the inherent spatial consistency regularization between the main task and the surrogate task in the student model, which helps the surrogate task to yield more reliable predictions and generates high-quality uncertainty maps. Thus, our model can also address the task-level perturbation problems that occur spatial inconsistency between the primary and surrogate tasks in the student model. Experimental results on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised methods. Code is available at : https://github.com/smallmax00/SUA_crowd_counting
Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng
ICCV5