Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bowen Chen 0004

dblp:12/7780-4 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-0042-6207ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 71% Vision and language · 22% Graph learning · 7%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action segmentation
1.622025
Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025
Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation · ECCV (54) 2024
Computer vision › Video understanding and tracking › action segmentation › human action segmentation
skeleton-based action segmentation
0.912025
Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025
Machine learning › Graph learning › graph structure learning
relation graph learning
0.312025
Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025

Methods — techniques the papers use, named apart from their topics

spatio-temporal modeling · 0.9large language model · 0.9graph neural network · 0.9contrastive learning · 0.9skeleton-based action representation · 0.8language model · 0.8
YearPublicationVenuePosition
2026 Topology-Motion Decoupling Framework With Textual Regularization for Skeleton-Based Temporal Action Segmentation
abstract
Skeleton-based temporal action segmentation aims to capture key information in long skeleton motion sequences to temporally segment and identify actions at a fine-grained level. Existing approaches have achieved promising results by improving the modeling of topological spatial relationships and long-term temporal dependencies. However, current methods often overlook the distinct nature of motion and topological information, applying a monolithic modeling paradigm to both. This approach fails to fully exploit their differential contributions to precise boundary localization and effective class discrimination. To address these limitations, we propose a novel Topology-Motion Decoupling Framework (TMD). Our framework incorporates three key designs. First, an auxiliary Differential Motion Perception Branch explicitly models the temporal gradients of skeletal trajectory to decouple boundary-sensitive motion features. Second, we introduce two effective fusion modules that integrate the complementary features from both branches for mutual enhancement. Finally, a Boundary-Aware Textual Regularization scheme leverages a dual set of semantic prompts for boundary/non-boundary to differentially guide the feature learning process. By design, TMD explicitly mitigates semantic and temporal confusion between actions, thereby enhancing inter-class discriminability and boundary awareness. Extensive experiments on five challenging public datasets demonstrate that our TMD achieves state-of-the-art performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multiscale Skeleton-Based Temporal Action Segmentation Using Hierarchical Temporal Modeling and Prediction Ensemble
abstract
Skeleton-based temporal action segmentation (TAS) decomposes untrimmed skeleton sequence into meaningful segments. The variance in temporal scale challenges the skeleton modeling network to seek a balance between over-segmentation and under-segmentation. Current methods often rely on parallel multiscale feature extractors and additional refinement modules to mitigate the multiscale issue, which brings significant computations and complexity. To address these issues, this article proposes multiscale skeleton-based TAS (MSTAS), consisting of temporal probability pyramid (TPP) and smoothed multiscale ensemble (SME). TPP represents each action as a collection of multiscale probability distributions using a U-shape hierarchical temporal pyramid. Subsequently, SME takes the average of distributions instead of deploying additional refinement stages to achieve action segmentation. Considering the over-confident issue that exists in each scale, SME incorporates a novel label smoothing phase to improve the probability distributions by dynamically calibrating the confidence of each scale. Experimental results on four public datasets show that the MSTAS achieves state-of-the-art performance with less computation overheads, such as +1.1% accuracy and +2.8% [email protected] on the challenging LARa dataset with 70% fewer parameters and 80% fewer GFLOPS. Benefiting from confidence calibration, the MSTAS efficiently utilizes more temporal scales while keeping better calibration for ambiguous action instances. Additionally, the U-shape pyramid demonstrates a strong compatibility with classical refinement module, enabling the efficient extraction of multiscale motion representations.
Bowen Chen 0004, Haoyu Ji 0001, Weihong Ren, Qiyi Tong, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.1
2025 Interaction-Aware Transformer Network for Human-Object Interaction Detection
abstract
human-object interaction (HOI) detection tackles the problem of joint localization and classification of HOIs. Recent HOI detection methods are mainly based on transformer networks, where the explicit priors at the object level (e.g., scene layout, object appearance, or category) are usually fed into the transformer to improve the object query ability. Though these methods have achieved remarkable results, they did not pay enough attention to the implicit action-level information, which is the fundamental element of HOI. In this work, we propose an interaction-aware transformer network (IATN) to obtain the interaction-aware query, by jointly utilizing implicit action-level priors and explicit object-level priors. Specifically, we design an action-aware module (AAM) to aggregate implicit action priors from the scene level and instance level, respectively. Then, we design an action-oriented graph (AOG), where human feature and object feature are graph nodes and action semantics represent graph edges, to aggregate priors jointly from action level and object level. Afterwards, the interaction-aware query is acquired and finally adopted to obtain the HOI predictions. Besides, we leverage knowledge distillation to enhance the action-level priors by transferring the final HOI predictions to the intermediate features. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed interaction-aware model.
Weibo Jiang, Weihong Ren, Jiandong Tian, Hanwei Ma, Bowen Chen 0004, Honghai Liu 0001
IEEE Trans. Cybern.5
2025 Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
abstract
Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results.
Haoyu Ji 0001, Bowen Chen 0004, Weihong Ren, Wenze Huang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Image Process.2
2025 Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action Segmentation
abstract
The skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks (GCNs) for intraframe spatial modeling and temporal convolutional networks (TCNs) for interframe temporal modeling to discern motion patterns. However, these approaches often overlook the distinctive nature of essential action elements across various actions, including engaged core body parts and key subactions. This oversight limits the ability to distinguish different actions within a given sequence. To address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along joints and sequences to identify core joints and key subactions at different scales. Specifically, in terms of the spatial domain, the intrasnippet cross-joint attention (CJA) module divides the sequence into distinct snippets and computes attention to establish intricate joint semantic relationships, emphasizing the identification of core motion joints. In terms of the temporal domain, in the encoder, the intrasnippet cross-frame attention (CFA) module segments the sequence in a blockwise expansion manner and establishes interframe relationships to highlight the most discriminative frames. In the decoder, clip-level representations at various temporal scales are initially generated through an hourglass-like sampling process, followed by the intrasnippet cross-scale attention (CSA) module to integrate the key clip information across different time scales. The performance evaluation on five public datasets demonstrates that ME-ST achieves state-of-the-art (SOTA) performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation
Haoyu Ji 0001, Bowen Chen 0004, Xinglong Xu, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
ECCV (54)2
2024 Computational Interpersonal Communication Model for Screening Autistic Toddlers: A Case Study of Response-to-Name
abstract
Interpersonal communication facilitates symptom measures of autistic sociability to enhance clinical decision-making in identifying children with autism spectrum disorder (ASD). Traditional methods are carried out by clinical practitioners with assessment scales, which are subjective to quantify. Recent studies employ engineering technologies to analyze children's behaviors with quantitative indicators, but these methods only generate specific rule-driven indicators that are not adaptable to diverse interaction scenarios. To tackle this issue, we propose a Computational Interpersonal Communication Model (CICM) based on psychological theory to represent dyadic interpersonal communication as a stochastic process, providing a scenario-independent theoretical framework for evaluating autistic sociability. We apply CICM to the response-to-name (RTN) with 48 subjects, including 30 toddlers with ASD and 18 typically developing (TD), and design a joint state transition matrix as quantitative indicators. Paired with machine learning, our proposed CICM-driven indicators achieve consistencies of 98.44% and 83.33% with RTN expert ratings and ASD diagnosis, respectively. Beyond outstanding screening results, we also reveal the interpretability between CICM-driven indicators and expert ratings based on statistical analysis.
Bingrui Zhou, Zhiyong Wang 0009, Bowen Chen 0004, Chunchun Hu, Huiping Li 0004, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics4
2023 WSCFER: Improving Facial Expression Representations by Weak Supervised Contrastive Learning
abstract
The major challenge of Facial Expression Recog-nition (FER) is to learn class discriminative representations, and the existing works mainly address it by designing various classification networks from class level. However, learning representations at class level is limited due to the inconspicuous class discrimination among different facial expressions. Thus, in this paper, we propose a Weak Supervised Contrastive learning FER (WSCFER) method to improve facial expression representations by simultaneously learning instance-level representations which are highly complementary to the general class-level representations. Specifically, our proposed WSCFER consists of three components: a major task for FER classification, an auxiliary task for Weak Supervised Contrastive (WSC) learning which pulls augmented samples of the same image together while pushing apart instance samples from different classes, and a Partial Consistency Loss (PCL) for optimizing the two embedding spaces from both the class level and the instance level. We compare WSC with some state-of-the-art contrastive methods and find that it can efficiently learn instance-level representations but avoid overemphasizing irrelevant parts, which is crucial for FER. WSCFER achieves superior performance on several in-the-wild databases, and it also shows the promising potential for learning representations under noisy annotations.
Bowen Chen 0004, Xiu Xu, Weihong Ren, Honghai Liu 0001
IROS2
2022 AutoENP: An Auto Rating Pipeline for Expressing Needs via Pointing Protocol
abstract
Early screening for ASD (Autism Spectrum Disorder) is crucial and also challenging due to the limited medical resource. Expressing Needs with Pointing (ENP) is a low-cost yet effective protocol for early screening. However, the current methods need to manually trim video for analyzing ENP protocol, which is labour-intensive. Also, they detect discriminative signs with separately high-level clues (e.g., pose, object detection), but ignore the temporal action relationships between child and clinician, which usually leads to invalid detection. In contrast to previous approaches, we propose an Auto Rating Pipeline for Expressing Needs via Pointing Protocol, named AutoENP. Specifically, we introduce action segmentation into early screening, to capture temporal interaction relationships without manually intervention. To detect fine-grained hand motions, we fuse global, local and fine-grained features to fully understand the screening scene. Besides, we integrate focal loss and center loss to improve the detection accuracy for rare actions. To evaluate the proposed pipeline, we collected 22 ENP videos containing 7 actions with above 40,000 frames. Experimental results demonstrate that our model achieves 82.1% and 84.7% action accuracy for child and clinician, respectively. Moreover, 18 in 22 children’s ENP levels are reported correctly against the clinician’s diagnoses.
Bowen Chen 0004, Weihong Ren, Honghai Liu 0001, Huiping Li 0004, Xiu Xu, Bingrui Zhou
ICPR1