Xina Cheng

dblp:160/0094 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-7319-1635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Interaction-aware representation learning for action quality assessment in freestyle skiing big air
Shiyue Chen, Xina Cheng, Takeshi Ikenaga
Comput. Vis. Image Underst.4
2026 Statistic temporal checking and spatial consistency based 3D size reconstruction of multiple objects from indoor monocular videos
Xina Cheng, Takeshi Ikenaga
Image Vis. Comput.2
2025 Geometry-Aware Contextual Reasoning-Based Indoor Accessibility Detection System for Visually Impaired Wheelchair Users
Fanxiang Zhou, Xina Cheng, Takeshi Ikenaga
PRCV (6)3
2025 Temperature-Aware Dynamic Fusion Network for Few-Shot Segmentation of Infrared Images
Bo Wang 0016, Xina Cheng, Yuan Li 0058, Xiangrong Zhang, Xu Tang 0004, Dingheng Wang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 LF Tracy: A Unified Single-Pipeline Paradigm for Salient Object Detection in Light Field Cameras
Jiaming Zhang 0001, Kunyu Peng, Xina Cheng, Zhiyong Li 0001, Kailun Yang 0001
ICPR (17)5
2024 Towards Precise 3D Human Pose Estimation with Multi-Perspective Spatial-Temporal Relational Transformers
abstract
3D human pose estimation captures the human joint points in three-dimensional space while keeping the depth information and physical structure. That is essential for applications that require precise pose information, such as humancomputer interaction, scene understanding, and rehabilitation training. Due to the challenges in data collection, mainstream datasets of 3D human pose estimation are primarily composed of multi-view video data collected in laboratory environments, which contains rich spatial-temporal correlation information besides the image frame content. Given the remarkable selfattention mechanism of transformers, capable of capturing the spatial-temporal correlation from multi-view video datasets, we propose a multi-stage framework for 3D sequence-to-sequence (seq2seq) human pose detection. Firstly, the spatial module represents the human pose feature by intra-image content, while the frame-image relation module extracts temporal relationships and 3D spatial positional relationship features between the multiperspective images. Secondly, the self-attention mechanism is adopted to eliminate the interference from non-human body parts and reduce computing resources. Our method is evaluated on Human3.6M, a popular 3D human pose detection dataset. Experimental results demonstrate that our approach achieves stateof-the-art performance on this dataset. The source code will be available at https://github.com/WUJINHUAN/3D-human-pose.
Jianbin Jiao, Xina Cheng, Xiaoting Yin, Hao Shi 0004, Kailun Yang 0001
IJCNN2
2024 Bidirectional temporal and frame-segment attention for sparse action segmentation of figure skating
Xina Cheng, Yuan Li 0058, Takeshi Ikenaga
Comput. Vis. Image Underst.2
2024 Motion-aware and data-independent model based multi-view 3D pose refinement for volleyball spike analysis
abstract
Abstract In the volleyball game, estimating the 3D pose of the spiker is very valuable for training and analysis, because the spiker’s technique level determines the scoring or not of a round. The development of computer vision provides the possibility for the acquisition of the 3D pose. Most conventional pose estimation works are data-dependent methods, which mainly focus on reaching a high level on the dataset with the controllable scene, but fail to get good results in the wild real volleyball competition scene because of the lack of large labelled data, abnormal pose, occlusion and overlap. To refine the inaccurate estimated pose, this paper proposes a motion-aware and data-independent method based on a calibrated multi-camera system for a real volleyball competition scene. The proposed methods consist of three key components: 1) By utilizing the relationship of multi-views, an irrelevant projection based potential joint restore approach is proposed, which refines the wrong pose of one view with the other three views projected information to reduce the influence of occlusion and overlap. 2) Instead of training with a large amount labelled data, the proposed motion-aware method utilizes the similarity of specific motion in sports to achieve construct a spike model. Based on the spike model, joint and trajectory matching is proposed for coarse refinement. 3) To finely refine, a point distribution based posterior decision network is proposed. While expanding the receptive field, the pose estimation task is decomposed into a classification decision problem, which greatly avoids the dependence on a large amount of labelled data. The experimental dataset videos with four synchronous camera views are from a real game, the Game of 2014 Japan Inter High School of Men Volleyball. The experiment result achieves 76.25%, 81.89%, and 86.13% success rate at the 30mm, 50mm, and 70mm error range, respectively. Since the proposed refinement framework is based on a real volleyball competition, it is expected to be applied in the volleyball analysis.
Xina Cheng, Takeshi Ikenaga
Multim. Tools Appl.2
2024 Key points trajectory and multi-level depth distinction based refinement for video mirror and glass segmentation
abstract
Abstract Mirror and glass are ubiquitous materials in the 3D indoor living environment. However, the existing vision system always tends to neglect or misdiagnose them since they always perform the special visual feature of reflectivity or transparency, which causes severe consequences, i.e., a robot or drone may crash into a glass wall or be wrongly positioned by the reflections in mirrors, or wireless signals with high frequency may be influenced by these high-reflective materials. The exploration of segmenting mirrors and glass in static images has garnered notable research interest in recent years. However, accurately segmenting mirrors and glass within dynamic scenes remains a formidable challenge, primarily due to the lack of a high-quality dataset and effective methodologies. To accurately segment the mirror and glass regions in videos, this paper proposes key points trajectory and multi-level depth distinction to improve the segmentation quality of mirror and glass regions that are generated by any existing segmentation model. Firstly, key points trajectory is used to extract the special motion feature of reflection in the mirror and glass region. And the distinction in trajectory is used to remove wrong segmentation. Secondly, a multi-level depth map is generated for region and edge segmentation which contributes to the accuracy improvement. Further, an original dataset for video mirror and glass segmentation (MAGD) is constructed, which contains 9,960 images from 36 videos with corresponding manually annotated masks. Extensive experiments demonstrate that the proposed method consistently reduces the segmentation errors generated from various state-of-the-art models and reach the highest successful rate at 0.969, mIoU (mean Intersection over Union) at 0.852, and mPA (mean Pixel Accuracy) at 0.950, which is around 40% - 50% higher on average on an original video mirror and glass dataset.
Xina Cheng, Takeshi Ikenaga
Multim. Tools Appl.3
2024 Learning consensus-aware semantic knowledge for remote sensing image captioning
Yunpeng Li 0010, Xiangrong Zhang, Xina Cheng, Xu Tang 0004, Licheng Jiao
Pattern Recognit.3
2024 Intertemporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change Captioning
abstract
Remote sensing image change captioning (RSICC) is more challenging than remote sensing change detection task, which requires extracting occurred changes in similar remote sensing image (RSI) pairs while generating change caption. However, few works have been investigated on RSICC, the main challenges come from how to learn abundant change clues and face the modality gap. To handle these problems, we rethink this task from the perspective of obtaining and aligning symmetrical change features for temporal RSIs. In this work, the proposed intertemporal interaction and symmetric difference learning network are cascaded through several multitemporal integration units to model differences from coarse to fine representations. Specifically, we design a cross-temporal attention (CTA) mechanism to probe direct interaction between bi-temporal RSIs for motivating information coupling between intralevel representations and suppressing irrelevant interferences. To learn robust change features, a symmetric difference transformer (SDT) module is devised to guarantee temporal symmetry between the “before-to-after” and “after-to-before” change representations. Besides, the bi-directional triplet ranking loss is adopted to guide the network to learn strongly discriminative and temporal-symmetric change representation. Extensive experimental results on Dubai-CC and LEVIR-CC datasets demonstrate that our framework with the proposed components can achieve excellent performance and surpass recent state-of-the-art methods.https://github.com/romanticLYP/TISDNet
Yunpeng Li 0010, Xiangrong Zhang, Xina Cheng, Puhua Chen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Hierarchical Knowledge Graph for Multilabel Classification of Remote Sensing Images
abstract
Multilabel classification in remote sensing (RS) images aims to correctly predict multiple object labels in an RS image with the primary challenge of mining correlations among multiple labels. In this context, we argue that a scene can be treated as a high-level depiction of the interactions among multiple interconnected objects within the image. However, hierarchical relationships between the scene and local objects are often neglected in other state-of-the-art approaches. In this article, we consider multilabel classification as a global-to-local prediction process, whereas the scene of an image is first identified, followed by recognition of local objects in the image. To achieve this, we propose a novel hierarchical knowledge graph (HKG)-based framework for multilabel classification in RS images (ML-HKG). Specifically, we first construct a hierarchical KG to depict label correlations between scenes and objects and represent the hierarchical knowledge as interrelated scene- and object-level label embeddings. Subsequently, we generate a scene-aware enhanced feature map by recognizing scene categories in an image under the guidance of scene-level knowledge embeddings. Afterward, object-level embeddings are used to derive category-specific visual representations for final multilabel prediction. Extensive experiments on the UCM and AID datasets demonstrate the effectiveness of our framework.
Xiangrong Zhang, Xina Cheng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 High-Resolution Remote Sensing Image Segmentation With Global-Guided Normalization and Local Affinity Distillation
abstract
In recent years, high-resolution (HR) remote sensing images (RSIs) segmentation has received growing attention. The huge number of pixels poses a challenge to the semantic segmentation algorithm, which is limited by the storage of GPUs, so the current methods for processing HR RSIs are categorized into two main categories, i.e., global methods and local methods. The former downsamples the original image and loses a lot of feature details. The latter crops the original image and fails to obtain global contextual information. Both types of methods lead to limited segmentation accuracy. In this article, we propose an end-to-end framework, called global injection network (GINet), which explores two levels of feature distribution and feature relationship to achieve tradeoff between global context and local details. In concrete terms, we propose the global-guided normalization (GGN) module, which injects global context information into local branch and modulates local features using global features to enhance the global perception of local branch. In addition, to constrain the spatial consistency of two branches, inspired by the knowledge distillation technique, we propose local affinity distillation (LAD) loss, which distills the relations in local features into global features to keep the similarity of the relationships corresponding to patches in the two branches. The comprehensive experimental results on three large-scale land-cover classification datasets, DeepGlobe ($2448 \times 2448$), Inria Aerial ($5000 \times 5000$), and GID-15 ($7200 \times 6800$), confirm the effectiveness and superiority of our method in HR semantic segmentation tasks.
Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Xina Cheng, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 CTACL:Hyperspectral Image Change Detection Based on Adaptive Contrastive Learning
abstract
Hyperspectral image change detection (HSI-CD) can accurately identify changing regions by capturing subtle spectral differences and has become a research hotspot in the field of remote sensing (RS). Convolutional neural networks (CNNs) have excellent local context modeling capabilities and have been proven to be powerful feature extractors in HSI-CD. However, due to its inherent network structure limitation, CNN cannot well mine and represent the sequential properties of spectral features, especially the medium and long-term dependencies. In contrast, transformer-based network architecture shows a strong ability to model long-distance dependencies, which can fully mine and extract global features, but exhibits weak performance in extracting local information. To this end, we propose HSI-CD network based on adaptive contrastive learning (CTACL). Specifically, we first propose a parallel network of CNNs and transformers to mine local and global temporal-spatial-spectral features of HSI, respectively. Second, we propose adaptive contrastive learning to pre-train the network to learn the latent features of a large amount of unlabeled data and better mine and utilize local and global information. Experimental results on the farmland dataset show that the proposed method performs well.
Shunli Tian, Xiangrong Zhang, Guanchun Wang, Xiao Han 0012, Puhua Chen, Xina Cheng
IGARSS6
2023 A Figure Skating Jumping Dataset for Replay-Guided Action Quality Assessment
abstract
In competitive sports, judges often scrutinize replay videos from multiple views to adjudicate uncertain or contentious actions, and ultimately ascertain the definitive score. Most existing action quality assessment methods regress from a single video or a pairwise exemplar and input videos, which are limited by the viewpoint and zoom scale of videos. To end this, we construct a Replay Figure Skating Jumping dataset (RFSJ), containing additional view information provided by the post-match replay video and fine-grained annotations. We also propose a Replay-Guided approach for action quality assessment, learned by a Triple-Stream Contrastive Transformer and a Temporal Concentration Module. Specifically, besides the pairwise input and exemplar, we contrast the input and its replay by an extra contrastive module. Then the consistency of scores guides the model to learn features of the same action under different views and zoom scales. In addition, based on the fact that errors or highlight moments of athletes are crucial factors affecting scoring, these moments are concentrated in parts of the video rather than a uniform distribution. The proposed temporal concentration module encourages the model to concentrate on these features, then cooperates with the contrastive regression module to obtain an effective scoring mechanism. Extensive experiments demonstrate that our method achieves Spearman's Rank Correlation of 0.9346 on the proposed RFSJ dataset, improving over the existing state-of-the-art methods.
Xina Cheng, Takeshi Ikenaga
ACM Multimedia2
2023 Bidirectional Multiple Object Tracking Based on Trajectory Criteria in Satellite Videos
abstract
Multiple object tracking (MOT) in satellite videos requires to detect all objects belonging to specified categories and identify each object, which plays a basic and necessary role in automatic driving, traffic surveillance, and smart city. The traditional MOT methods in satellite videos mostly follow the detection–association framework. However, the detection–association framework works under a strict assumption that all objects are correctly localized by the detector. In practice, MOT in satellite videos faces challenges such as low resolution, tiny objects, and the wide field of view, which leads to the degradation of detector performance. In order to reduce the impact of detector degradation, we propose a bidirectional MOT framework based on trajectory criteria (BMTC) in satellite videos. In BMTC, the single object tracking (SOT) tracker carries out locating the objects between consecutive frames and the detector is just used for finding new objects. Therefore, it is less dependent on the detector performance. According to the characteristics of satellite videos, the trajectory criteria are designed to control the state of the tracker, which includes trajectory density, the limit of consecutive virtual motion predictions, and trajectory similarity measurement. Invalid fragment trajectory backtracking is implemented to alleviate the misalignment caused by the above subsection trajectory criteria. The method is validated on the VISO benchmark and SkySat-1 dataset. The experimental results show the improvement of completeness and accuracy, and the proposed tracker achieves the state-of-the-art performance.
Xiangrong Zhang, Zhongjian Huang, Xina Cheng, Jie Feng 0003, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Semantic Texture Complexity Model for Feature Generation and Selection in Visual SLAM
abstract
Simultaneous Localization And Mapping (SLAM) plays an irreplaceable role in autonomous driving, robotics and augmented reality for whose ability of locating and providing the map of the unknown environment. The traditional SLAM algorithms describe the environment by extracting, tracking and matching the local feature points, which often leads to the feature points gathering in certain areas and affects the performance of pose detection. In order to obtain a SLAM environment map with accurate representative feature points, this paper proposes a semantic texture complexity model for feature generation and selection. The semantic texture complexity model respectively evaluates the texture complexity of each semantic area so that even the area with few corners or edges can be represented by feature points. For different semantic regions, different feature point generation and selection schemes in the Visual-Inertial Odometry (VIO) module are adopted according to the texture complexity, which avoids the concentration of feature points and effectively improves the positioning accuracy. In this paper, we use the uHumans2 dataset to test the positioning accuracy. The test results show that the proposed algorithm reduces the error by 10%-30% compared with the related work.
Xina Cheng, Licheng Jiao
ICPR2
2017 Event state based particle filter for ball event detection in volleyball game analysis
abstract
The ball state tracking and detection technology plays a significant role in volleyball game analysis for volleyball team supporting and tactics development. This paper proposes a ball event detection method to achieve high detection rate by solving challenges including: the great variety of event length, the large intra-class difference of one event and the influence caused by ball trajectories. Proposed state vector covers both the event type and the event period length so that the system model can transits various lengths of event period and predicts event types by volleyball game rules. The curve segmental observation model avoids the tracking error influence to evaluate the event period likelihood by referring neighbouring trajectories of the ball. And according to the standard of the ball event, the feature of the distance between the ball and specific court line are extracted to evaluate the ball event type in observation. At last a two-layer estimation method estimates the posterior state which is a joint probability distribution. Experiments of the proposed method implemented on 3D trajectories tracked from multi-view volleyball game videos shows the detection rate reaches 90.43%.
Xina Cheng, Norikazu Ikoma, Masaaki Honda, Takeshi Ikenaga
FUSION1
2017 Simultaneous physical and conceptual ball state estimation in volleyball game analysis
abstract
Automatically extraction of accurate volleyball game data from game videos plays an important role in making contribution to game data analysis, TV broadcasting and performance evaluations. In this paper, a particle filter based physical and conceptual ball state estimation method is proposed to track the 3D ball trajectory and ball event simultaneously with high accuracy. The physical ball state includes 3D ball position and velocity. Besides the ball event, the conceptual state also includes flag of the external force on the ball. The system model is adaptive to this external force predicted through proposed spatial hitting points dense distribution. Observation of the external force is evaluated by hitting point likelihood, which uses not only the past tracked trajectory but also the image noise feature so that image noise is transferred into useful feature and unidirectional dependency on trajectory is avoided. Experimental results based on multi-view HDTV video sequences show the tracking success rate of ball state achieves 92.43%.
Xina Cheng, Norikazu Ikoma, Masaaki Honda, Takeshi Ikenaga
VCIP1
2016 Anti-occlusion observation model and automatic recovery for multi-view ball tracking in sports analysis
abstract
The 3D position of the ball plays a crucial role in professional sport analysis. In ball sports, tracking of ball's precise position accurately is highly required, whose performance is affected by inaccurate 3D coordinates and occlusion problem. In this paper, we propose anti-occlusion observation model and automatic recovery by 3D ball detection based on multiview videos to track the ball in 3D space. The anti-occlusion observation model evaluates each camera's image and eliminates the influence of the cameras in which the ball is occluded. The automatic recovery method detects the ball's 3D position by homography relation of the multi-video and generates a new distribution to initiate the tracker when tracking failure is detected. Experimental results based on the HDTV video sequences, which were captured by four cameras located at the corners of the court, show that the success rate of the 3D ball tracking achieves 99.14%.
Xina Cheng, Masaaki Honda, Norikazu Ikoma, Takeshi Ikenaga
ICASSP1