Zhao Guo

dblp:07/10652 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation
abstract
The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.
Longhao Li, Zhao Guo, Hongjie Chen 0001, Yuhang Dai, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Hui Bu, Jie Li 0001, Jian Kang 0006, Ruibin Yuan, Ziya Zhou, Wei Xue 0002, Lei Xie 0001
AAAI2
2025 Bird-Inspired Tendon Coupling Improves Paddling Efficiency by Shortening Phase Transition Times
abstract
Drag-based swimming using rowing appendages, fins, and webbed feet is a widely adopted mode of locomotion in aquatic animals. To develop efficient underwater and swimming vehicles, various bioinspired drag-based paddle designs have been proposed, often facing a trade-off between propulsive efficiency and versatility. Webbed feet generate effective propulsive force during the power phase, while being lightweight, robust, and partially foldable during the recovery phase. However, the time-consuming process of mechanically folding and unfolding webbed feet extends the transition periods between the recovery and power phases, which in turn increased drag, and reduces overall paddling efficiency. In this study, we draw inspiration from the coupling tendons of aquatic birds. We implement tendon coupling mechanisms to minimize the transition time between the recovery and power phases. Hardware experiments demonstrate that our proposed mechanism improves propulsive efficiency by factors of$\mathbf{2. 0}$and$\mathbf{2. 4}$compared to designs without extensor tendons and based on passive paddles, respectively. Additionally, we find that distal leg joint clutching-previously shown to enhance efficiency in terrestrial walking-plays a negligible role in swimming locomotion. In sum, we present a novel principle for efficient drag-based leg and paddle design, with implications for understanding the swimming mechanics of aquatic birds and advancing bioinspired aquatic propulsion systems.
Jianfeng Lin 0002, Zhao Guo, Alexander Badri-Spröwitz
ICRA2
2025 EEG-TFNet: Spatiotemporal and Spectral Feature Integration for EEG-Based AD Detection
An Zeng, Zhao Guo, Dan Pan 0001, Yiqun Zhang 0006, Huisi Hong
ISBRA (1)2
2024 Design and Modeling of A Compact Serial Variable Stiffness Actuator (SVSA-III) with Linear Stiffness Profile
abstract
Variable stiffness actuator (VSA) can imitate natural muscles in their compliance capbility, which can provide flexible adaptability for robots, improving the safety of robots interacting with the environment or human. This paper presents a new compact serial variable stiffness actuator ((SVSA-III)) with linear stiffness profile based on symmetrical variable lever arm mechanism. The stiffness motor is used to regulate the position of the pivot located on the Archimedean Spiral Relocation Mechanism (ASRM), so that the stiffness of the actuator can be adjusted (softening or hardening). By designing the lever length, the range of stiffness adjustment can change from 0.3Nm/degree to therotical infinity. Moreover, the continuous linear stiffness profile of the actuator can be customized by solving the transcendental equation of the relationship between the actuator stiffness and the rotation angle of the stiffness motor. SVSA-III has the advantages of compact structure, wide-range stiffness regulation, reduced control difficulty, and linear stiffness profile. Two experiments of step response and stiffness tracking have proved the high accuracy and fast response for both theoretical stiffness and position adjustment.
Shuowen Yi, Junbei Liao, Zhao Guo
ICRA4
2023 Design and Stiffness Analysis of a Bio-Inspired Soft Actuator with Bi-Direction Tunable Stiffness Property
abstract
Modulating the stiffness of soft actuators is crucial for improving the efficiency of interaction with the environment. However, current stiffness modulation mechanisms are hard to achieve high lateral stiffness and a wide range of bending stiffness simultaneously. Here, we draw inspiration from the anatomical structure of the finger and propose a bi-directional tunable stiffness actuator (BTSA). BTSA is a soft-rigid hybrid structure that combines air-tendon hybrid actuation (ATA) and bone-like structures (BLS). We develop a corresponding fabrication method and a stiffness analysis model to support the design of BLS. The results show that the influence of the BLS on bending deformation is negligible, with a distal point distance error of less than 1.5 mm. Moreover, the bi-directional tunable stiffness is proved to be functional. The bending stiffness can be tuned by ATA from 0.23 N/mm to 0.70 N/mm, with a magnification of 3 times. The addition of BLS improves lateral stiffness up to 4.2 times compared with the one without BLS, and the lateral stiffness can be tuned decoupling within 1.2 to 2.1 times (e.g. from 0.35 N/mm to 0.46 N/mm when the bending angle is 45 deg). Finally, a four-BTSA gripper is developed to conduct horizontal lifting and grasping tasks to demonstrate the advantages of BTSA.
Jianfeng Lin 0002, Ruikang Xiao, Zhao Guo
IROS3
2022 Learning ultrasound scanning skills from human demonstrations
Xutian Deng, Ziwei Lei, Zhao Guo, Chenguang Yang 0001, Miao Li 0002
Sci. China Inf. Sci.5
2019 Versatile Reactive Bipedal Locomotion Planning Through Hierarchical Optimization
abstract
When experiencing disturbances during locomotion, human beings use several strategies to maintain balance, e.g. changing posture, modulating step frequency and location. However, when it comes to the gait generation for humanoid robots, modifying step time or body posture in real time introduces nonlinearities in the walking dynamics, thus increases the complexity of the planning. In this paper, we propose a two-layer hierarchical optimization framework to address this issue and provide the humanoids with the abilities of step time and step location adjustment, Center of Mass (CoM) height variation and angular momentum adaptation. In the first layer, times and locations of consecutive two steps are modulated online based on the current CoM state using the Linear Inverted Pendulum Model. By introducing new optimization variables to substitute the hyperbolic functions of step time, the derivatives of the objective function and feasibility constraints are analytically derived, thus reduces the computational cost. Then, taking the generated horizontal CoM trajectory, step times and step locations as inputs, CoM height and angular momentum changes are optimized by the second layer nonlinear model predictive control. This whole procedure will be repeated until the termination condition is met. The improved recovery capability under external disturbances is validated in simulation studies.
Jiatao Ding, Chengxu Zhou, Zhao Guo, Xiaohui Xiao, Nikolaos G. Tsagarakis
ICRA3
2018 Design and Evaluation of a Motorized Robotic Bed Mover With Omnidirectional Mobility for Patient Transportation
abstract
Patient transportation in hospitals faces many challenges, including the limited manpower, work-related injuries, and low efficiency of current bed pushing methods. This paper presents a new motorized robotic bed mover with omnidirectional mobility to address this problem. This device is composed of an omnidirectional mobility unit, a force sensing-based human-machine interface, and control hardware with batteries and electronics. The proposed bed mover can be attached to the bottom of a manual hospital stretcher, transforming it into a powered omnidirectional bed (OmniBed) that can be used only by one person. The function of the OmniBed is compared with that of a conventional powered bed, which only provides forward assistance with a fifth powered wheel. We perform a pilot study with 14 subjects to evaluate the performance of this OmniBed and benefits for hospital application. The experimental results show that the OmniBed can half the manpower while decreasing back muscle activities, revealing the potential health benefits for older staffs. The OmniBed also shows the promising signs of high precision and handling in small spaces with its one-step "parallel-parking" ability. This device is more ergonomic, more effective, and safer than the conventional powered bed.
Zhao Guo, Xiaohui Xiao, Haoyong Yu
IEEE J. Biomed. Health Informatics1
2017 Mechanical design of a compact Serial Variable Stiffness Actuator (SVSA) based on lever mechanism
abstract
Compliant actuator is widely accepted for physical human-robot interaction due to its safety aspect, dynamic performance improvements and energy saving abilities. In this paper, based on the variable ratio lever mechanism, a new kind of Serial Variable Stiffness Actuator (SVSA) is proposed by using an Archimedean Spiral Relocation Mechanism (ASRM) to change the position of the pivot, implementing large range of adjustable stiffness. The ASRM introduced here makes the SVSA design has continuous stiffness adjustment ability and simply mechanical structure. Within the Variable Stiffness Mechanism (VSM), two linear springs are assembled antagonistically on a spring shaft. Their displacements are perpendicular to the output link to transmit the spring force more efficiently. Stiffness modeling and analysis of the SVSA are carried out to cover large deflection angle. The physical implementation of the SVSA shows that the output stiffness of the VSM is changed from 1.72 to 150.56 Nm/rad using a linear spring with stiffness 1882 N/m, working range covered from 0 to 360°. Control experiments also proved the wide range of stiffness adjustment ability of the SVSA.
Jiantao Sun, Yubing Zhang, Zhao Guo, Xiaohui Xiao
ICRA4
2017 Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
abstract
Recent progress has been made in using attention based encoder-decoder framework for video captioning. However, most existing decoders apply the attention mechanism to every generated words including both visual words (e.g., “gun” and "shooting“) and non-visual words (e.g. "the“, "a”).However, these non-visual words can be easily predicted using natural language model without considering visual signals or attention.Imposing attention mechanism on non-visual words could mislead and decrease the overall performance of video captioning.To address this issue, we propose a hierarchical LSTM with adjusted temporal attention (hLSTMat) approach for video captioning. Specifically, the proposed framework utilizes the temporal attention for selecting specific frames to predict related words, while the adjusted temporal attention is for deciding whether to depend on the visual information or the language context information. Also, a hierarchical LSTMs is designed to simultaneously consider both low-level visual information and deep semantic information to support the video caption generation. To demonstrate the effectiveness of our proposed framework, we test our method on two prevalent datasets: MSVD and MSR-VTT, and experimental results show that our approach outperforms the state-of-the-art methods on both two datasets.
Jingkuan Song, Lianli Gao, Zhao Guo, Wu Liu 0005, Dongxiang Zhang, Heng Tao Shen
IJCAI3
2017 Video Captioning With Attention-Based LSTM and Semantic Consistency
abstract
Recent progress in using long short-term memory (LSTM) for image captioning has motivated the exploration of their applications for video captioning. By taking a video as a sequence of features, an LSTM model is trained on video-sentence pairs and learns to associate a video to a sentence. However, most existing methods compress an entire video shot or frame into a static representation, without considering attention mechanism which allows for selecting salient features. Furthermore, existing approaches usually model the translating error, but ignore the correlations between sentence semantics and visual content. To tackle these issues, we propose a novel end-to-end framework named aLSTMs, an attention-based LSTM model with semantic consistency, to transfer videos to natural sentences. This framework integrates attention mechanism with LSTM to capture salient structures of video, and explores the correlation between multimodal representations (i.e., words and visual content) for generating sentences with rich semantic content. Specifically, we first propose an attention mechanism that uses the dynamic weighted sum of local two-dimensional convolutional neural network representations. Then, an LSTM decoder takes these visual features at time t and the word-embedding feature at time t-1 to generate important words. Finally, we use multimodal embedding to map the visual and sentence features into a joint space to guarantee the semantic consistence of the sentence description and the video visual content. Experiments on the benchmark datasets demonstrate that our method using single feature can achieve competitive or even better results than the state-of-the-art baselines for video captioning in both BLEU and METEOR.
Lianli Gao, Zhao Guo, Hanwang Zhang, Xing Xu 0001, Heng Tao Shen
IEEE Trans. Multim.2
2016 Attention-based LSTM with Semantic Consistency for Videos Captioning
abstract
Recent progress in using Long Short-Term Memory (LSTM) for image description has motivated the exploration of their applications for automatically describing video content with natural language sentences. By taking a video as a sequence of features, LSTM model is trained on video-sentence pairs to learn association of a video to a sentence. However, most existing methods compress an entire video shot or frame into a static representation, without considering attention which allows for salient features. Furthermore, most existing approaches model the translating error, but ignore the correlations between sentence semantics and visual content.
Zhao Guo, Lianli Gao, Jingkuan Song, Xing Xu 0001, Jie Shao 0001, Heng Tao Shen
ACM Multimedia1
2016 Spatial and temporal scoring for egocentric video summarization
Zhao Guo, Lianli Gao, Xiantong Zhen, Fuhao Zou, Fumin Shen, Kai Zheng 0001
Neurocomputing1
2015 Power analysis of a series elastic actuator for ankle joint gait rehabilitation
abstract
Series elastic actuator (SEA) has been widely used in rehabilitation robotics, where human-robot interaction is required. Due to its intrinsic compliance, SEA can improve the usage of power for its motor, which leads to a compact and lightweight SEA design. The aim of this paper is to reduce the energy consumption and the power requirements of the motor of the SEA by optimizing the stiffness of its spring. This study is inspired by the biomechanics of a human ankle joint, which stores elastic energy during the first phases of the walking process and releases the stored energy in the next gait phases to propel the human body forward. Power analysis and optimization procedure are conducted on complete SEA models with different complexity, including inertia, damping and stiffness, and with open loop and closed loop control strategies. Simulation results demonstrate that a reduction of 56.6% of the peak motor power can be achieved with the optimized spring stiffness.
Oussama Ben Farah, Zhao Guo, Chi Zhu 0001, Haoyong Yu
ICRA2
2015 Human-Robot Interaction Control of Rehabilitation Robots With Series Elastic Actuators
abstract
Rehabilitation robots, by necessity, have direct physical interaction with humans. Physical interaction affects the controlled variables and may even cause system instability. Thus, human-robot interaction control design is critical in rehabilitation robotics research. This paper presents an interaction control strategy for a gait rehabilitation robot. The robot is driven by a novel compact series elastic actuator, which provides intrinsic compliance and backdrivablility for safe human-robot interaction. The control design is based on the actuator model with consideration of interaction dynamics. It consists mainly of human interaction compensation, friction compensation, and is enhanced with a disturbance observer. Such a control scheme enables the robot to achieve low output impedance when operating in human-in-charge mode and achieve accurate force tracking when operating in force control mode. Due to the direct physical interaction with humans, the controller design must also meet the stability requirement. A theoretical proof is provided to show the guaranteed stability of the closed-loop system under the proposed controller. The proposed design is verified with an ankle robot in walking experiments. The results can be readily extended to other rehabilitation and assistive robots driven with compliant actuators without much difficulty.
Haoyong Yu, Sunan Huang 0001, Gong Chen 0001, Yongping Pan 0001, Zhao Guo
IEEE Trans. Robotics5
2013 Design of a novel compliant differential Shape Memory Alloy actuator
abstract
This paper presents a novel compliant differential (CD) Shape Memory Alloy (SMA) actuator with improved performance compared to traditional SMA actuators. This actuator is composed of two antagonistic SMA wires and a mechanical joint coupled with a torsion spring. The torsion spring is employed to reduce the total stiffness of SMA actuator and improve the range of motion. The antagonistic wires increase the response time as one wire can be heated up while the other wire is still in the cooling process. Dynamic model of this actuator was established for control design. Experimental results proved that this new actuator can provide larger output range of motion and faster response speed than traditional SMA actuators under the same conditions. Sine wave tracking with 0.05 Hz, 0.08 Hz and 0.1 Hz were performed and our results demonstrated that this compliant actuator has good tracking performance under simple PID control.
Zhao Guo, Haoyong Yu, Liang-Boon Wee
IROS1