Zihang Guo

dblp:366/2443 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Video understanding and tracking · 67% Language models and text generation · 33%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 77% Audio and music processing · 23%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining › misinformation detection
fake news detection
0.912025
Event Consistency-aware Robust Fake News Detection · ACM Multimedia 2025
Web and social media mining › misinformation detection › fake news detection
multimodal fake news detection
0.912025
Event Consistency-aware Robust Fake News Detection · ACM Multimedia 2025
Multimedia analysis and retrieval › video analysis
video understanding and tracking
0.912025
Event Consistency-aware Robust Fake News Detection · ACM Multimedia 2025
Computer vision › Video understanding and tracking › sign language recognition
continuous sign language recognition
0.712023
C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language Recognition · ICCV 2023
Computer vision › Video understanding and tracking
sign language recognition
0.712023
C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language Recognition · ICCV 2023

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 1.7contrastive learning · 1.7recurrent neural network · 0.7language model · 0.7connectionist temporal classification · 0.7
YearPublicationVenuePosition
2026 SNTT: A Sparsity-aware NTT Accelerator Based on FPGA for Zero-Knowledge Proof
Zihang Guo, Qiang Liu 0011, Ray C. C. Cheung, Zhaohui Guo
ISCAS1
2026 Adaptive hierarchical control of quadcopters via safe reinforcement learning from human demonstration
Junkai Tan, Shuangsi Xue, Zihang Guo, Hui Cao 0003
Eng. Appl. Artif. Intell.3
2026 Neural estimator-based finite-time formation control for manipulator end effectors with obstacle avoidance
Shuangsi Xue, Zihang Guo, Hui Cao 0003, Badong Chen
Inf. Sci.2
2026 Dynamic event-triggered finite-time actor-critic-identifier-based approximate optimal control for unknown nonlinear drifted systems
Shuangsi Xue, Junkai Tan, Zihang Guo, Qingshu Guan, Hui Cao 0003, Badong Chen
Inf. Sci.3
2026 Human-robotics hybrid shared control with guaranteed performance: A fixed-time game-theoretic learning approach
Shuangsi Xue, Junkai Tan, Zihang Guo, Tiansen Niu, Hui Cao 0003, Badong Chen
Inf. Sci.3
2026 Neural Adaptive Finite-Time Formation Tracking Control for Manipulator End Effectors Under Input Constraints
abstract
This work investigates the formation tracking issue for multirobot manipulator end-effectors under input constraints. A distributed formation control law is designed to guarantee the finite-time boundedness of tracking errors within the framework. To estimate the significant bias of dynamics discovered during practical multirobot collaborative manipulation tasks, a bias radial basis function neural network (RBFNN) is integrated, along with a designed adaptive updating law for expeditious approximation. In addition, an anti-windup compensator within a finite-time framework is specifically introduced to mitigate the input saturation issue arising from torque limitations in joint actuators. Finally, the system’s semi-global practical finite-time boundedness (SGPFTB) is rigorously established through Lyapunov theory. Five planar manipulators are employed in comparative computational experiments to validate the feasibility of the presented control strategy.
Shuangsi Xue, Zihang Guo, Junkai Tan, Kai Qu, Hui Cao 0003, Badong Chen
IEEE Trans. Syst. Man Cybern. Syst.2
2025 DDNet: Exploring Dual Dependencies for Long-Term Time Series Forecasting
abstract
Recent Transformer-based methods have advanced multivariate time series forecasting by focusing primarily on temporal dependencies (cross-time dependencies). However, these methods often overlook crucial multivariate correlations (cross-channel dependencies), leading to suboptimal performance. In this paper, we propose a novel Dual Dependencies modeling Network (DDNet) to model both cross-time and cross-channel dependencies effectively. DDNet employs aggregation tokens in the Aggregation Stage to capture cross-channel relationships and utilizes these tokens in the Broadcast Stage to incorporate cross-time dependencies. Additionally, we propose a fine-grained Patch-Wise Normalization technique to address the overfitting challenges when modeling cross-channel dependencies. Extensive experiments on benchmark datasets demonstrate that DDNet consistently surpasses state-of-the-art methods while maintaining low computational complexity.
Zihang Guo, Zhenping Mou, Jieru Guo
ICASSP1
2025 Event Consistency-aware Robust Fake News Detection
abstract
With the rapid development of short video platforms (such as Kuaishou and TikTok), these platforms have increasingly become important channels for the spread of fake news.Therefore, multi-modal fake news detection has attracted extensive attention.Existing studies mainly focus on directly integrating multi-modal information or discovering implicit clues in posts to improve detection performance.However, due to the abuse of video editing techniques, event-irrelevant segments (e.g., advertisements) are frequently mixed into videos, introducing noise information, thereby weakening models' ability to learn crucial information.Moreover, video creators often inject personal tampered information into original news content through audio modality manipulation, potentially distorting the factual. To address these challenges, we propose a novel Event Consistency-aware Robust Fake News Detection (ECR-FND) framework, comprising two key components: an Event-aware Video Denoising Learning (EVDL) and an Audio Tampering-information Capturing Module (ATCM).Specifically, the EVDL filters out the event-irrelevant segments within video modality to focus on core news events. The ATCM adaptively amplifies tampering information in audio modality, enhancing the model's capacity to detect manipulation attempts.Extensive experiments on two benchmark datasets (FakeSV and FakeTT) demonstrate ECR-FND's effectiveness.Our source code is available at https://github.com/immc-lab/ECR-FND.
Zihang Guo, Huaiwen Zhang
ACM Multimedia2
2025 Neural observer-based fixed-time formation control of multiagent systems
Zihang Guo, Shuangsi Xue, Junkai Tan, Hui Cao 0003
Neurocomputing1
2025 Data-driven optimal shared control of unmanned aerial vehicles
Junkai Tan, Shuangsi Xue, Zihang Guo, Hui Cao 0003, Badong Chen
Neurocomputing3
2025 Boundary Refinement Network for Polyp Segmentation With Deformable Attention
abstract
Early and accurate polyp segmentation is crucial for the diagnosis and treatment of colorectal cancer. However, polyp segmentation faces many challenges: different polyp sizes, complex shapes, and ambiguous intestinal wall boundaries. To solve these problems, we propose a novel polyp segmentation network named DeformSegNet. Specifically, we first introduce a polyp perception module (PPM), which combines the dynamic multi-kernel spatial selection network (DMS-Net) and a transformer encoder to effectively locate polyps of different sizes. Next, we design a deformation-aware separable module (DSM), which consists of deformable attention that adaptively adjusts the sampling position, enabling the network to adapt to complex and diverse polyp boundaries. Finally, a cross-attention aggregation module (CAAM) effectively retains low-level features, further enhancing the boundary features and suppressing false positives. DeformSegNet achieves competitive segmentation accuracy on five polyp datasets, demonstrating excellent learning and generalization capabilities.
Wangsheng Wu, Zihang Guo, Dongdong Zhu
IEEE Signal Process. Lett.4
2025 Hierarchical Safe Reinforcement Learning Control for Leader-Follower Systems With Prescribed Performance
abstract
This paper proposes a hierarchical safe reinforcement learning with prescribed performance control (HSRL-PPC) scheme to address the challenges of interconnected leader-follower systems operating in complex environments. The framework consists of two levels: at the higher level, the leader agent detects and avoids moving obstacles while planning optimal paths; at the lower level, the follower agent tracks the leader within strict prescribed performance bounds. We formulate the optimal prescribed performance safe control problem and solve it using the Hamilton-Jacobi-Bellman (HJB) equation. Due to system nonlinearity and obstacle complexity, we approximate the leader’s optimal value function using a state-following neural network that efficiently extrapolates training data to neighboring states, while employing a regular critic neural network for the follower’s value function approximation. Lyapunov stability analysis demonstrates the closed-loop system’s theoretical guarantees. Experimental results from two simulation examples and hardware tests with a quadcopter-vehicle system validate the effectiveness of the proposed approach in achieving safe navigation and precise tracking performance in dynamic environments. Note to Practitioners—Challenges exist in unpredictable obstacles and agent limitations for the interconnected leader-follower system. To provide a safe, efficient, and reliable control scheme, hierarchical safe reinforcement learning with prescribed performance control is proposed in this paper. The hierarchical structure is utilized to coordinate the leader and follower agents in the interconnected system, where the leader agent plans the optimal path and avoids obstacles, and the follower agent tracks the leader within prescribed performance bounds. Based on the proposed hierarchical structure, engineers can design efficient and safe control schemes for interconnected leader-follower systems with moving obstacles. In future work, we will address the problem of external disturbances and uncertainties in the interconnected leader-follower system.
Junkai Tan, Shuangsi Xue, Zihang Guo, Hui Cao 0003, Badong Chen
IEEE Trans Autom. Sci. Eng.4
2025 Prescribed Performance Robust Approximate Optimal Tracking Control via Stackelberg Game
abstract
Real-world applications of nonlinear systems tracking control are always challenging due to the existence of uncertainties and disturbances. To design a robust optimal tracking controller for uncertain nonlinear systems with disturbances and actuator saturation, this paper investigates the prescribed performance robust optimal tracking control problem. A prescribed performance mechanism is constructed to convert the dynamics of tracking error into transformed error dynamics, which keeps the system’s operating states within specific bounds, ensuring tracking with predefined error constraints. For the optimal tracking controller design, an optimal index is established to optimize the performance of tracking control, and a robust optimal index is established to optimize the disturbance effect on the tracking error. To achieve robust optimal tracking control that minimizes both optimal and robust optimal indexes, a Stackelberg game is constructed, which provides a hierarchical game structure for the optimal controller and the worst disturbance. The robust optimal controller is approximated online using reinforcement learning techniques. An actor-critic-identifier algorithm is designed to approximate the optimal value function, optimal controller, and drifted system parameters. Lyapunov theory is utilized to analyze the closed-loop system’s stability. To demonstrate the effectiveness of the proposed robust optimal control method, two numerical simulations and a hardware experiment on a quadcopter system are conducted. The experiment results demonstrate that our method successfully achieves prescribed performance tracking control when actuators are saturated and disturbances are present. Note to Practitioners—In this paper, the probelm of mixed$H_{2}/H_{\infty }$prescribed-performance optimal tracking control for nonlinear systems with input saturation is investigated. To constrain the operating states of the system within certain bounds, the prescribed performance transformation is designed to achieve tracking with predefined error constraints. For the optimal controller design, the$H_{2}$index is established to minimize the optimal tracking performance, and the$H_{\infty }$index is designed to minimize the disturbance effect on the tracking error. A Stackelberg-based non-zero sum game between the optimal controller and the worst disturbance is established to design the mixed$H_{2}/H_{\infty }$optimal tracking controller. The designed optimal controller is approximated online using reinforcement learning. Effectiveness of the proposed method is demonstrated by two numerical simulations and a hardware experiment on a quadcopter system. Based on the proposed high-performance controller, engineers can design a high-performance robust optimal tracking controller for uncertain nonlinear systems with extreme conditions of disturbances and actuator saturation.
Junkai Tan, Shuangsi Xue, Zihang Guo, Hui Cao 0003, Dongyu Li
IEEE Trans Autom. Sci. Eng.4
2025 Learning to Diversify for Robust Video Moment Retrieval
abstract
In this paper, we focus on diversifying the Video Moment Retrieval (VMR) model into more scenes. Most existing video moment retrieval methods focus on aligning video moments and queries by capturing the cross-modal relationship, which largely ignores the cross-instance relationship behind the representation learning. Thus, they may easily get trouble into the inaccurate cross-instance contrastive relationship in the training process: 1) Existing approaches can hardly identify similar semantic content across different scenes. They incorrectly treat such instances as negative samples (termed faulty negatives), which forces the model to learn the features from query-irrelevant scenes. 2) Existing methods perform unsatisfactorily in locating the queries with subtle differences. They neglect to mine the hard negative samples that belong to similar scenes but have different semantic content. In this paper, we propose a novel robust video moment retrieval method that prevents the model from overfitting the query-irrelevant scene features by accurately capturing both the cross-modal and cross-instance relationships. Specifically, we first develop a scene-independent cross-modal reasoning module that filters out the redundant scene contents and infers the video semantics under the guidance of query information. Then, the faulty and hard negative samples are mined from the negative ones and calibrated for their contribution to the overall loss in contrastive learning. We validate our contributions through extensive experiments on cross-scene video moment retrieval settings, where the training and test data are from different scenes. Experimental results show that the proposed robust video moment retrieval model can effectively retrieve target videos by capturing the real cross-modal and cross-instance relationships.
Huilin Ge, Zihang Guo, Zhiwen Qiu
IEEE Trans. Circuits Syst. Video Technol.3
2023 C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language Recognition
abstract
Continuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation learning, a temporal module aggregating the local and global temporal information of frame sequence, and the connectionist temporal classification (CTC) loss, which aligns video features with gloss sequence. Unfortunately, the language prior implicit in the gloss sequence is ignored throughout the modeling process. Furthermore, the contextualization of glosses is further ignored in alignment learning, as CTC makes an independence assumption between glosses. In this paper, we propose a Cross-modal Contextualized Sequence Transduction (C2ST) for CSLR, which effectively incorporates the knowledge of gloss sequence into the process of video representation learning and sequence transduction. Specifically, we introduce a cross-modal context learning framework for CSLR, in which the linguistic features of gloss sequences are extracted by a language model, and recurrently integrate with visual features for video modelling. Moreover, we introduce the contextualized sequence transduction loss that incorporates the contextual information of gloss sequences in label prediction, without making any independence assumptions between the glosses. Our method sets the new state of the art on three widely used large-scale sign language recognition datasets: Phoenix-2014, Phoenix-2014-T, and CSL-Daily. On CSL-Daily, our approach achieves an absolute gain of 4.9% WER compared to the best published results.
Huaiwen Zhang, Zihang Guo, Yang Yang 0121, De Hu
ICCV2