Yu Shi 0003

dblp:55/4736-3 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0001-9117-8282ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HMVformer++: Hierarchical multi-view fusion transformer for efficient 3D human pose estimation
Kangkang Zhou, Xiangyuan Lan, Yu Shi 0003
Pattern Recognit.7
2025 ESMformer: Error-aware self-supervised transformer for multi-view 3D human pose estimation
Kangkang Zhou, Xiaohu Shao, Yu Shi 0003
Pattern Recognit.7
2024 Deep Semantic Graph Transformer for Multi-View 3D Human Pose Estimation
abstract
Most Graph Convolutional Networks based 3D human pose estimation (HPE) methods were involved in single-view 3D HPE and utilized certain spatial graphs, existing key problems such as depth ambiguity, insufficient feature representation, or limited receptive fields. To address these issues, we propose a multi-view 3D HPE framework based on deep semantic graph transformer, which adaptively learns and fuses multi-view significant semantic features of human nodes to improve 3D HPE performance. First, we propose a deep semantic graph transformer encoder to enrich spatial feature information. It deeply mines the position, spatial structure, and skeletal edge knowledge of joints and dynamically learns their correlations. Then, we build a progressive multi-view spatial-temporal feature fusion framework to mitigate joint depth uncertainty. To enhance the pose spatial representation, deep spatial semantic feature are interacted and fused across different viewpoints during monocular feature extraction. Furthermore, long-time relevant temporal dependencies are modeled and spatial-temporal information from all viewpoints is fused to intermediately supervise the depth. Extensive experiments on three 3D HPE benchmarks show that our method achieves state-of-the-art results. It can effectively enhance pose features, mitigate depth ambiguity in single-view 3D HPE, and improve 3D HPE performance without providing camera parameters. Codes and models are available at https://github.com/z0911k/SGraFormer.
Kangkang Zhou, Yu Shi 0003
AAAI5
2024 Hierarchical Spatial-Temporal Adaptive Graph Fusion for Monocular 3D Human Pose Estimation
abstract
Single-view 3D human pose estimation (HPE) based on Graph Convolutional Networks currently suffers from problems such as insufficient feature representation and depth ambiguity. To address these issues, this letter proposes a hierarchical spatial-temporal adaptive graph fusion framework to improve monocular 3D HPE performance. Firstly, to enhance the spatial semantic feature representation of human nodes, a progressive adaptive graph feature capture strategy is developed, which adaptively constructs global-to-local attention graph features of all human joints in a coarse-to-fine manner. A spatial-temporal attention fusion module is then constructed to model long-term sequential dependencies and mitigate depth ambiguity. The temporal attention factors of related frames are captured and utilized to intermediately supervise the joint depth. The spatial semantic information among all joints in the same frame and temporal contextual knowledge of the joints across relevant frames are fused to build spatial-temporal correlations and optimize the final features. Extensive experiments on two popular benchmarks show that our method outperforms several state-of-the-art approaches and improves 3D HPE performance.
Kangkang Zhou, Yu Shi 0003
IEEE Signal Process. Lett.5
2024 Hierarchical Synergy-Enhanced Multimodal Relational Network for Video Question Answering
abstract
Video question answering (VideoQA) is challenging as it requires reasoning about natural language and multimodal interactive relations. Most existing methods apply attention mechanisms to extract interactions between the question and the video or to extract effective spatio-temporal relational representations. However, these methods neglect the implication of relations between intra- and inter-modal interactions for multimodal learning, and they fail to fully exploit the synergistic effect of multiscale semantics in answer reasoning. In this article, we propose a novel hierarchical synergy-enhanced multimodal relational network (HMRNet) to address these issues. Specifically, we devise (i) a compact and unified relation-oriented interaction module that explores the relation between intra- and inter-modal interactions to enable effective multimodal learning; and (ii) a hierarchical synergistic memory unit that leverages a memory-based interaction scheme to complement and fuse multimodal semantics at multiple scales to achieve synergistic enhancement of answer reasoning. With careful design of each component, our HMRNet has fewer parameters and is computationally efficient. Extensive experiments and qualitative analyses demonstrate that the HMRNet is superior to previous state-of-the-art methods on eight benchmark datasets. We also demonstrate the effectiveness of the different components of our method.
Min Peng 0004, Xiaohu Shao, Yu Shi 0003
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer
abstract
This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors. We achieve this with a pyramidal multimodal transformer (PMT) model, which simply incorporates a learnable word embedding layer, a few convolutional and transformer layers. We use the anisotropic pyramid to fulfill video-language interactions across different spatio-temporal scales. In addition to the canonical pyramid, which includes both bottom-up and top-down pathways with lateral connections, novel strategies are proposed to decompose the visual feature stream into spatial and temporal sub-streams at different scales and implement their interactions with the linguistic semantics while preserving the integrity of local and global semantics. We demonstrate better or on-par performances with high computational efficiency against state-of-the-art methods on five VideoQA benchmarks. Our ablation study shows the scalability of our model that achieves competitive results for text-to-video retrieval by leveraging feature extractors with reusable pre-trained weights, and also the effectiveness of the pyramid. Code available at: https://github.com/Trunpm/PMT-AAAI23.
Min Peng 0004, Yu Shi 0003
AAAI3
2023 Multi-Semantic Alignment Co-Reasoning Network for Video Question Answering
abstract
Video question answering challenges models on understanding textual questions with varying complexity and searching for clues from visual content with different hierarchical semantics. In this paper, we propose a novel Multi-Semantic Alignment Co-Reasoning Network (MACN) to accomplish an interactive inference between the question and the video input. The design of our MACN comprises two modules of Question-Centric Interaction (QCI) and Contextual Semantic Reasoning (CSR). Specifically, QCI establishes a question-centric heterogeneous graph model to align visual content at different temporal scales with questions to enable the extraction of visual representations under better textual understanding. CSR exploits self-attention mechanisms to extract the contextual dependencies of visual semantics at different hierarchies to achieve co-reasoning of answer clues. Experiments on three benchmarks demonstrate that our proposed method is superior to previous state-of-the-art performance.
Min Peng 0004, Yu Shi 0003
ICIP4
2023 Progressive Multi-View Fusion for 3D Human Pose Estimation
abstract
In multi-view 3D human pose estimation (HPE), viewpoint images have large variability due to factors like camera angles and occlusion, making feature extraction and fusion across viewpoints challenging. To address these concerns, we propose a progressive multi-view 3D HPE transformer framework, which achieves effective intra-view pose feature extraction and cross-view fusion by embedding various multi-view fusion methods in the feature extraction process. In order to fully extract spatial semantic features of human joints, we first construct a cross-view spatial fusion module performing spatial feature fusion across adjacent views while mining useful spatial knowledge. To enhance the pose features and alleviate the depth ambiguity problem, we further develop a multi-view spatial-temporal fusion module to extract effective temporal contextual information within the viewpoint and fuse spatial-temporal features across multiple viewpoints. Extensive experiments on two popular 3D HPE benchmarks validate the efficacy and superiority of our method. It outperforms several state-of-the-art methods, effectively alleviates depth ambiguity, and improves 3D pose accuracy without providing camera parameters or complex loss functions.
Kangkang Zhou, Xunyi Zhao, Yu Shi 0003
ICIP7
2023 Attention Based Network with DA-Loss for X-ray Contraband Automatic Detection
abstract
X-ray security check is one of the most effective security measures widely used in airports, high-speed trains, subways, and other important places. Due to the difference in imaging mechanism between X-ray images and visible images, X-ray contraband detection suffers from problems of large intra-class differences and small inter-class differences. The Softmax loss does not encourage compactness within class and separability between classes explicitly, resulting in insensitivity to target. We address this problem by proposing an attention mechanism based ACD-Net, which can quickly focus on the obscured objects in luggage. To optimize inter-class interval, we evaluate impacts of several cosine-based loss function on X-ray contraband detection performance, including Cosine Embedding Loss, L-Softmax Loss, CosFace and ArcFace. Then we propose a cosine loss function DA-Loss focusing on intra-class compactness and inter-class difference at the same time. Extensive experiments show the superiority of our proposed ACD-Net model with DALoss classification loss function, which has better separability of large intra-class differences and small inter-class differences, and outperform the state-of-the-arts.
Peiwen Li, Yu Shi 0003, Xiaohu Shao
ICME4
2023 TextBFA: Arbitrary Shape Text Detection with Bidirectional Feature Aggregation
Yu Shi 0003
ICONIP (8)4
2023 Two-Stream Heterogeneous Graph Network with Dynamic Interactive Learning for Video Question Answering
abstract
Video question answering (VideoQA) challenges the joint learning of visual and linguistic knowledge. Whilst the dynamic video-question interaction is not well explored in previous methods, the search for answer clues from textual semantics in the interaction is not valued. To address these issues, this paper proposes a novel Two-Stream Heterogeneous Graph Network (TSHGNet) using Dynamic Interactive Learning (DIL) to accomplish effective reasoning between videos and questions. Inspired by the way people answer questions, the two-stream architecture of TSHGNet is designed to extract question-driven visual cues and video-driven textual semantics, respectively. Therein, for each stream, DIL gradually refines the comprehension of question semantics and the extraction of visual representations through heterogeneous graph interactions. Extensive experiments and qualitative analyses demonstrate improved performances of the proposed TSHGNet on three benchmark datasets in comparison with previous state-of-the-art methods, and the effectiveness of different components of our method.
Min Peng 0004, Xiaohu Shao, Yu Shi 0003
IJCNN3
2023 Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose Estimation
abstract
In multi-view 3D human pose estimation (HPE), information from different viewpoints is highly variable due to complex factors such as background and occlusion, making cross-view feature extrac tion and fusion difficult. Most existing methods have problems of over-reliance on camera parameters or insufficient semantic feature extraction. To address these issues, this paper proposes a hierar chical multi-view fusion transformer (HMVformer) framework for 3D HPE, incorporating cross-view feature fusion methods into the spatial and temporal feature extraction process in a coarse-to-fine manner. To begin, global to local attention graph features are ex tracted and incorporated with the original pose features to better preserve the spatial structure semantic knowledge. Then, various cross-view feature fusion modules are built and embedded into the pose feature extraction for consistent and distinctive information fusion across multiple viewpoints. Furthermore, sequential tem poral information is extracted and fused with spatial knowledge for feature refinement and depth uncertainty reduction. Extensive experiments on three popular 3D HPE benchmarks show that HMV former achieves state-of-the-art results without relying on complex loss functions or providing camera parameters, simple but effective in mitigating depth ambiguity and improving 3D pose prediction accuracy. Codes and models are available1.
Kangkang Zhou, Yu Shi 0003
ACM Multimedia5
2022 Spatio-Temporal Attention Graph for Monocular 3d Human Pose Estimation
abstract
Single-view 3D human pose estimation (HPE) based on Graph Convolutional Networks (GCNs) currently suffers from problems such as insufficient spatial feature representation, difficult fusion of various information, and depth ambiguity in 2D to 3D pose mapping. This paper proposes a framework for monocular 3D human pose learning based on spatio-temporal attention graph. Firstly, we build a spatial graph feature acquisition scheme to obtain spatial semantic feature of 3D human pose with strong representativeness, by constructing a global to local attention graph through a coarse-to-fine way. And then we capture contextual information of temporal related images in the sequence as attention factors, evaluate their influence on the target image and achieve effective integration of spatio-temporal characteristics to mitigate the depth ambiguity problem. Extensive experimental results on two challenging benchmark datasets (Human3.6M and HumanEva-I) show that our method can effectively improve the accuracy of 3D HPE and outperform the state-of-the-arts.
Xiaohu Shao, Yu Shi 0003
ICIP5
2022 Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering
abstract
Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore the visual appearance-motion information at different temporal scales, it is unknown how to incorporate the multilevel processing capacity of a deep learning model with such multiscale information. Targeting these issues, this paper proposes a novel Multilevel Hierarchical Network (MHN) with multiscale sampling for VideoQA. MHN comprises two modules, namely Recurrent Multimodal Interaction (RMI) and Parallel Visual Reasoning (PVR). With a multiscale sampling, RMI iterates the interaction of appearance-motion information at each scale and the question embeddings to build the multilevel question-guided visual representations. Thereon, with a shared transformer encoder, PVR infers the visual cues at each level in parallel to fit with answering different question types that may rely on the visual information at relevant levels. Through extensive experiments on three VideoQA datasets, we demonstrate improved performances than previous state-of-the-arts and justify the effectiveness of each part of our method.
Min Peng 0004, Yu Shi 0003
IJCAI4
2022 Robust Face Alignment via Deep Progressive Reinitialization and Adaptive Error-Driven Learning
abstract
Regression-based face alignment involves learning a series of mapping functions to predict the true landmarks from an initial estimation of the alignment. Most existing approaches focus on learning efficacious mapping functions from some feature representations to improve performance. The issues related to the initial alignment estimation and the final learning objective, however, receive less attention. This work proposes a deep regression architecture with progressive reinitialization and a new error-driven learning loss function to explicitly address the above two issues. Given an image with a rough face detection result, the full face region is first mapped by a supervised spatial transformer network to a normalized form and trained to regress coarse positions of landmarks. Then, different face parts are further respectively reinitialized to their own normalized states, followed by another regression sub-network to refine the landmark positions. To deal with the inconsistent annotations in existing training datasets, we further propose an adaptive landmark-weighted loss function. It dynamically adjusts the importance of different landmarks according to their learning errors during training without depending on any hyper-parameters manually set by trial and error. A high level of robustness to annotation inconsistencies is thus achieved. The whole deep architecture permits training from end to end, and extensive experimental analyses and comparisons demonstrate its effectiveness and efficiency. The source code, trained models, and experimental results are made available at https://github.com/shaoxiaohu/Face_Alignment_DPR.git.
Xiaohu Shao, Junliang Xing, Jiangjing Lyu, Yu Shi 0003, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 EILPR: Toward End-to-End Irregular License Plate Recognition Based on Automatic Perspective Alignment
abstract
Automatic License plate recognition (ALPR) remains a challenging task in face of some difficulties such as multi-line character distribution and license plate (LP) deformation due to camera angles. Most existing ALPR methods either focus on single-line LP or perform horizontal multi-line LP detection and recognition with character-level annotations. In this paper, we propose a novel end-to-end irregular license plate recognition (EILPR) to detect and recognize the LP of multi-line text or arbitrary shooting angles, using only plate-level annotations for training. In EILPR, a coarse-to-fine strategy is adopted to extract the LP features accurately for sequence recognition. Firstly, a coarse rectangular box of the LP is located, along with the corresponding predicted LP class which is single-line or double-line. Then, considering the fact that a LP mainly generates perspective distortion in the image due to its rigid feature, we propose a new automatic perspective alignment network (APAN) to extract the fine LP features connecting the detection and recognition. For recognition, a location-aware 2D attention based recognition network is performed to recognize the multi-line and multinational LP based on the extracted features. Experiments on several datasets show that EILPR achieves the state-of-the-art performance, demonstrating the effectiveness of the proposed method.
Chaojie Li, Yu Shi 0003
IEEE Trans. Intell. Transp. Syst.6
2021 Multi-View Face Recognition Using Deep Attention-Based Face Frontalization
abstract
Face frontalization has been widely used in face recognition to alleviate distribution discrepancy between multi-view faces. Given a profile face, existing models learn to synthesize a frontal face from the whole region indistinguishably, often resulting in unsatisfactory frontalization caused by a lack of synthetic focus and disturbances of trivial backgrounds. This paper proposes a novel Deep Attention-based Face Frontalization (DAFF) method to address the above issues explicitly. We first inject the 3D spatial prior of the input face into an encoder-decoder model. This process locates the discriminative foreground for decomposing meaningful convolutional embeddings. After that, we propose a novel objective that served as the generator’s geometric guidance to pay more attention to the target’s essential regions. Therefore, we can leverage the attentional constraints to perform recovery refinement at both embedding and texture levels. Extensive experiments show that DAFF achieves satisfactory frontalization and competitive recognition performance under constrained and in-the-wild benchmarks.
Xiaohu Shao, Junliang Xing, Ruihan Pan, Yu Shi 0003
ICME6
2020 2D License Plate Recognition based on Automatic Perspective Rectification
abstract
License plate recognition (LPR) remains a challenging task in face of some difficulties such as image deformation and multi-line character distribution. Text rectification that is crucial to eliminate the effects of image deformation has attracted increasing attentions in scene text recognition. However, current text rectification methods are not designed specifically for LPR, which did not take the features of plate deformation into account. Considering the fact that a license plate (LP) can only generate perspective distortion in the image due to its rigid feature, in this paper we propose a novel perspective rectification network (PRN) to automatically estimate the perspective transformation and rectify the distorted LP accordingly. For recognition, we propose a location-aware 2D attention based recognition network that is capable of recognizing both single-line and double-line plates with perspective deformation. The rectification network and recognition network are connected for end-to-end training. Experiments on common datasets show that the proposed method achieves the state-of-the-art performance, demonstrating the effectiveness of the proposed approach.
Zhaohong Guo, Dahan Wang, Yu Shi 0003
ICPR5
2019 A Novel Apex-Time Network for Cross-Dataset Micro-Expression Recognition
abstract
The automatic recognition of micro-expression has been boosted ever since the successful introduction of deep learning approaches. As researchers working on such topics are moving to learn from the nature of micro-expression, the practice of using deep learning techniques has evolved from processing the entire video clip of micro-expression to the recognition on apex frame. Using the apex frame is able to get rid of redundant video frames, but the relevant temporal evidence of micro-expression would be thereby left out. This paper proposes a novel Apex-Time Network (ATNet)to recognize micro-expression based on spatial information from the apex frame as well as on temporal information from the respective-adjacent frames. Through extensive experiments on three benchmarks, we demonstrate the improvement achieved by learning such temporal information. Specially, the model with such temporal information is more robust in cross-dataset validations.
Min Peng 0004, Tao Bi, Yu Shi 0003, Tong Chen 0008
ACII4
2019 DFQA: Deep Face Image Quality Assessment
Xiaohu Shao, Pingling Deng, Yu Shi 0003
ICIG (2)6
2019 Improving Irregular Text Recognition by Integrating Gabor Convolutional Network
abstract
Scene text, especially irregular text, is difficult to recognize due to the arbitrary-oriented characters and irregular arrangement. Most existing methods address the irregular text by rectifying it into a regular one, which achieve good performance. However, these methods are possible to remove character information in some curved texts. To overcome this issue, we focus on extracting features that are robust to orientation changes instead of rectifying. In this work, we propose an end-to-end trainable model that combines a Gabor Convolutional Network (GCN) and a Sequence Recognition Network (SRN). The GCN is capable of extracting more robust features against the orientation, which is produced by incorporating Gabor filters of different orientations into Convolutional Neural Network (CNN). The SRN is an attention-based sequence-to-sequence model that sequentially outputs characters from the robust features. We evaluate the recognition accuracy of the proposed method on various benchmark datasets of scene text, including both regular and irregular texts. The extensive experimental results show that our proposed method achieves the state-of-the-art recognition performance on most of the irregular benchmarks as well as a regular benchmark.
Zhaohong Guo, Yu Shi 0003
ICTAI6