Kaiping Xu

dblp:167/9636 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-7388-7878ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TrackNetV6: A Unified Framework for Lightweight and Robust Fast-Moving Tiny Ball Tracking
abstract
Although vision-based tiny ball tracking has advanced in specific sports, existing methods remain heavily coupled to domain-specific distributions, severely constraining cross-domain generalization. Concurrently, lightweight designs sacrifice representational capacity, while high-performance models incur prohibitive computational costs. To address these challenges, we propose TrackNetV6, a unified fast-moving tiny ball tracking framework that reconciles efficiency with accuracy. Central to our framework is a novel and compact decoding paradigm rooted in the Linear Multistep Method (LMM), designed to supersede conventional single-step feature fusion. This paradigm orchestrates two core components: a Cross-Scale Semantic Consensus Predictor (CSCP) that distills multi-scale features into semantic-correlation location priors, and a Prior-guided Context Corrector (PCC) that injects these priors into current-scale mappings for stable refinement. By iteratively alternating between these components, the model progressively strengthens feature representation for precise tracking. Furthermore, we introduce a Direction-aware Dynamic Fusion (DDF) module as the bottleneck layer, which explicitly models the direction-sensitive feature relationships of the fast-moving ball by synergizing the dynamic interaction between wavelet-based high-frequency cues and deep semantics. Extensive experiments across badminton, table tennis, and tennis benchmarks demonstrate that TrackNetV6 not only sets a new state-of-the-art (SOTA) but also delivers exceptional real-time inference at 183 FPS. Code will be available at https://github.com/Gi-gigi/TrackNetV6.
Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu, Yifei Cao, Zhizhen Li
ICMR4
2026 SUAD: semantic understanding and attribute discrimination for visual grounding
Xiya Bu, Jizhe Yu, Yu Liu 0035, Kaiping Xu, Guolong Wang 0001, Yifei Cao
Expert Syst. Appl.4
2025 Vision-Text Interaction with Orientation-Awareness for Referring Remote Sensing Image Segmentation
Xiaoshuai Wu, Yu Liu 0035, Kaiping Xu, Hao Zhang 0218, Jizhe Yu
ICANN (2)3
2025 Visual Grounding with Feature Enhancement and Language-Aware Attribute Guidance
Xiya Bu, Jizhe Yu, Yu Liu 0035, Kaiping Xu
ICMR4
2025 PAP-SAM: Global-Local Prior Adaptive Perception SAM for Co-Salient Object Detection
Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu
ICMR4
2025 PA2Net: Pyramid Attention Aggregation Network for Saliency Detection
Jizhe Yu, Yu Liu 0035, Xiaoshuai Wu, Kaiping Xu, Jiangquan Li
MMM (3)4
2024 Towards More Accurate Tiny Object Tracking: Benchmark and Algorithm
abstract
Despite the significant progress in tiny object tracking algorithms based on convolutional neural networks, the performance falls short of the ideal level due to dataset challenges like limited scale, limited match backdrop, low match popularity, and low-resolution(720p) images. The lack of specialized large-scale tiny object tracking datasets leads data-hungry tracking algorithms to rely on small and saturated datasets from TrackNet. This paper propels the development of tiny object tracking and contributes the first large-scale dataset, Badminton100K, for tiny ball tracking. We provide over 100K annotations covering over 200 videos from 12 professional match events. Our dataset covers high-resolution(1080p) match videos in broad and diverse contexts. Moreover, considering the fast-moving tiny balls in videos often leads to challenges such as blur, afterimage, overlap, and disappearance. As another contribution, we also propose a novel Multi-stage and Multi-scale Fusion Network (MMFNet), which can extract local high-resolution details in the shallow stages and capture global semantic information in the deep stages, addressing these challenging tasks through multiple stages multi-scale feature fusion, enabling accurate identification and localization of tiny balls. Extensive experiments show that our method achieves unprecedented state-of-the-art tracking performance on Badminton100K. The dataset and the algorithm code are available at https://github.com/Gi-gigi/MMFNet.
Jizhe Yu, Yu Liu 0035, Hongkui Wei, Kaiping Xu
CSCWD4
2024 DCMFNet: Deep Cross-Modal Fusion Network for Different Modalities with Iterative Gated Fusion
abstract
Cross-modal fusion aims to establish a consistent correspondence between arbitrary modalities. Due to the inherent differences between these modalities, accurately modeling their correspondence is a challenging task. Referring image segmentation (RIS) is a fundamental cross-modal task that intends to segment a desired object from an image based on a given natural language expression. In this paper, we propose an efficient algorithm called the Deep Cross-Modal Fusion Network (DCMFNet) to address this challenge. The proposed algorithm leverages the contextual information from linguistic context to guide the modeling of the visual context, gradually highlighting the referent in the image. The network architecture employs an innovative fusion strategy known as Iterative Gated Fusion (IGF) to capture the consistency relationship between multi-modal features. IGF iteratively adjusts the relative importance of features at each level based on high-level semantics, emphasizing the shared information while suppressing the irrelevant parts. Specifically, IGF consists of cascaded fusion units and gating units. The fusion units integrate high-level semantics with the features from the previous layer to enhance the representation. The gating units perceive the discrepancy between the enhanced features and the original representation, and selectively weight and integrate the important features for further refinement. Through multi-layer iterative optimization, IGF gradually establishes a fine-grained correspondence between arbitrary modalities. Extensive experimental results on the Referring Image Segmentation task demonstrate the effectiveness and utility of the proposed method.
Mingcheng Xue, Yu Liu 0035, Kaiping Xu, Jiangquan Li, Chengyang Yu
Graphics Interface4
2024 Global-Guided Weighted Enhancement for Salient Object Detection
Jizhe Yu, Yu Liu 0035, Hongkui Wei, Kaiping Xu, Yifei Cao, Jiangquan Li
ICANN (2)4
2024 Towards Highly Effective Moving Tiny Ball Tracking via Vision Transformer
Jizhe Yu, Yu Liu 0035, Hongkui Wei, Kaiping Xu, Yifei Cao, Jiangquan Li
ICIC (3)4
2022 Structured Multimodal Fusion Network for Referring Image Segmentation
abstract
Referring image segmentation aims to segment one particular object referred by a natural language expression in the image. One major challenge of this task is how to understand and align vision and language to distinguish the referent. Another major challenge is how to refine the segmentation mask of the referent. In this paper, we focus on dissecting and enhancing the interaction between modalities to address these challenges. Specifically, we propose a Structured Multimodal Fusion Network (SMFN), which consists of a multimodal tree, a cross-modal transformer, and a mask refinement module. SMFN first exploits multimodal fusion structures to deeply integrate visual and linguistic features so that the referent can be accurately distinguished and then further utilizes a mask refinement module to aggregate multi-scale visual features to clarify boundaries. We conduct extensive experiments on the four benchmark datasets and achieve new state-of-the-art performances under different evaluation metrics.
Mingcheng Xue, Yu Liu 0035, Kaiping Xu, Chengyang Yu
ICMI3
2020 A Context-Based Network For Referring Image Segmentation
abstract
Referring image segmentation is an important task aiming at segmenting out the object referred by a natural language expression. Current works usually employ the methods of concatenating the visual and linguistic features. They underestimate the importance of language-to-vision and object-to-object relationships when the natural language expression has multiple entities. Therefore, we propose a new network named Context-Based Network(CBN) to improve the accuracy of locating the correct referent. The CBN is composed of two modules: Intra Relation Selection(Intra-RS) and Inter Relation Selection(Inter-RS). The Intra-RS can capture object-to-object relationships in an embedding visual and linguistic feature space and the Inter-RS uses the multi-scale linguistic features as a guide to match the most similar region from the image feature maps. Besides, we apply spatial pyramid pooling to get global information to solve the limited receptive field problem. Experimental results on four public datasets showed that CBN achieved comparable performance to the other state-of-art methods.
Yu Liu 0035, Kaiping Xu, Zhehuan Zhao, Sipei Liu
ICIP3
2020 DRGCN: Deep Relation GCN for Group Activity Recognition
Yiqiang Feng, Shimin Shan, Yu Liu 0035, Zhehuan Zhao, Kaiping Xu
ICONIP (4)5
2018 Bridge Video and Text with Cascade Syntactic Structure
abstract
We present a video captioning approach that encodes features by progressively completing syntactic structure (LSTM-CSS). To construct basic syntactic structure (i.e., subject, predicate, and object), we use a Conditional Random Field to label semantic representations (i.e., motions, objects). We argue that in order to improve the comprehensiveness of the description, the local features within object regions can be used to generate complementary syntactic elements (e.g., attribute, adverbial). Inspired by redundancy of human receptors, we utilize a Region Proposal Network to focus on the object regions. To model the final temporal dynamics, Recurrent Neural Network with Path Embeddings is adopted. We demonstrate the effectiveness of LSTM-CSS on generating natural sentences: 42.3% and 28.5% in terms of BLEU@4 and METEOR. Superior performance when compared to state-of-the-art methods are reported on a large video description dataset (i.e., MSR-VTT-2016).
Guolong Wang 0001, Zheng Qin 0003, Kaiping Xu, Shuxiong Ye
COLING3
2018 Collision-Free LSTM for Human Trajectory Prediction
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001, Shuxiong Ye, Huidi Zhang
MMM (1)1
2017 Recognizing Emotions Based on Human Actions in Videos
Guolong Wang 0001, Zheng Qin 0003, Kaiping Xu
MMM (2)3
2016 Human activities prediction by learning combinatorial sparse representations
abstract
Human activities prediction is to enable early recognition of unfinished activities from videos only containing the beginning parts, which is a challenge problem. Prediction of human activities is necessarily applied in particular scenes(e.g. surveillance systems, human-computer interfaces). To solve this problem, we propose a novel framework which classifies videos into activity classes by using combinatorial sparse representations (CSR). The major contributions of our work include: (1) dividing each video into multiple equal-length segments, where the local spatio-temporal features extracted; (2) concatenating combinatorial sparse activity dictionaries, formed by overcomplete dictionary of each segment; (3) computing combinatorial sparse coefficients of each segment, based on activity dictionaries above; (4) formulating probability of each activity to estimate the correct class. The results of our experiments show that the proposed method outperforms existing state-of-the-art comparison methods.
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001
ICIP1
2016 Recognize human activities from multi-part missing videos
abstract
Recognizing human activities from multi-part missing videos is a challenge problem. When the multiple missing parts are continuous, the problem is reduced to activity recognition in videos with single part missing at any position which is focused on by many researches. However, in many practical applications, some temporal gaps always appear in captured videos due to random frame loss(e.g. noise interfere). To solve this problem, we propose a novel framework: 1) dividing each video into multiple equal-length segments, where the local spatio-temporal features extracted; 2) concatenating combinatorial sparse activity dictionaries, formed by over-complete dictionary of each segment; 3) computing combinatorial sparse coefficients of each segment, based on activity dictionaries above; 4) formulating probability of each activity to estimate the correct class. Our experiments achieve superior performance not only in videos with single part missing at any position, but also in videos with multiple parts missing.
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001
ICME1
2015 From LTL Formulae to Büchi Automata: A Direct Translation Using On-the-Fly De-Generalization
abstract
In this paper, we present a conversion algorithm to translate a linear temporal logic (LTL) formula to a Büchi automaton (BA) directly. A label, acceptance degree (AD), is presented to record acceptance conditions satisfied in each state or transition of an automaton. The AD for an automaton is a set of {U, F, R, G}-subformula of the given LTL formula. According to ADs attached to states and transitions, the on-the-fly de-generalization algorithm is presented. This on-the-fly de-generalization algorithm is used to transform a generalized Büchi automaton (GBA) into a Büchi automaton. It is different from the execution of the classic de-generalization algorithm that the on-the-fly de-generalization algorithm is performed during the expansion of the given LTL formula. A direct conversion algorithm based on the on-the-fly de-generalization algorithm is conceived and implemented. We compare the conversion algorithm presented in this paper with previous works, and show that it is more efficient for a series of formulae in usual use and random formulae generated by LBTT 1.2.1 (an LTL-to-BA translator testbench).
Lai-Xiang Shan, Zheng Qin 0003, Kaiping Xu, Xu Chen 0017
APSEC3