Guorui Liao

dblp:357/0850 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CausalFall: Fall prediction wearing motion sensors from a causal perspective
Guorui Liao, Jun Liao 0001, Shu Wang 0005, Xiurong Liang, Li Liu 0001
Expert Syst. Appl.1
2025 HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement Units
abstract
This paper considers the challenging problem of 3D Human Pose Estimation (HPE) from a sparse set of Inertial Measurement Units (IMUs). Existing efforts typically reconstruct a pose sequence by either directly tackling whole-body motions or focusing on distinctive spatio-temporal features of local body parts. Unfortunately, these methods ignore existing interdependent motor synergies amongst body parts, which may lead to pose estimation with ambiguous local parts. This observation motivates us to propose a hierarchical learning-based approach, HiPoser, which utilizes a hierarchical shared structure using Mamba blocks as the backbone to focus on the following estimation tasks, involving: 1) torso pose, 2) lower limbs pose, 3) upper limbs pose, and finally 4) global translation. These tasks selectively incorporate body motion states and are to be carried out sequentially in reconstructing part-based poses, which are amalgamated to estimate the final full-body pose with the global translation that satisfies inter-part consistencies. Our hierarchical structure allows HiPoser the flexibility in prioritizing different aspects of pose estimation, to emphasize more on detail or stability. Empirical evaluations over three benchmark datasets demonstrate the superiority of HiPoser over existing state-of-the-art models, suggesting that analyzing the synergistic movement of body parts is indeed important for advancing IMU-based 3D HPE.
Guorui Liao, Chunyuan Zheng 0001, Li Cheng 0001, Shanshan Huang 0004, Jun Liao 0001, Haoxuan Li 0001, Li Liu 0001
AAAI1
2025 Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
abstract
Wearable Human Activity Recognition (WHAR) is a prominent research area within ubiquitous computing. Multi-sensor synchronous measurement has proven to be more effective for WHAR than using a single sensor. However, existing WHAR methods use shared convolutional kernels for indiscriminate temporal feature extraction across each sensor variable, which fails to effectively capture spatio-temporal relationships of intra-sensor and inter-sensor variables. We propose the DecomposeWHAR model consisting of a decomposition phase and a fusion phase to better model the relationships between modality variables. The decomposition creates high-dimensional representations of each intra-sensor variable through the improved Depth Separable Convolution to capture local temporal features while preserving their unique characteristics. The fusion phase begins by capturing relationships between intra-sensor variables and fusing their features at both the channel and variable levels. Long-range temporal dependencies are modeled using the State Space Model (SSM), and later cross-sensor interactions are dynamically captured through a self-attention mechanism, highlighting inter-sensor spatial correlations. Our model demonstrates superior performance on three widely used WHAR datasets, significantly outperforming state-of-the-art models while maintaining acceptable computational efficiency.
Haoxuan Li 0001, Chunyuan Zheng 0001, Haonan Yuan, Guorui Liao, Jun Liao 0001, Li Liu 0001
AAAI5
2025 Visual Representation Learning through Causal Intervention for Controllable Image Editing
abstract
A key challenge for controllable image editing is that visual attributes with semantic meanings are not always independent, resulting in spurious correlations in model training. However, most existing methods ignore such issues, leading to biased causal visual representation learning and unintended changes to unrelated regions or attributes in the edited images. To bridge this gap, we propose a diffusion-based causal visual representation learning framework called CIDiffuser to capture causal representations of visual attributes based on structural causal models to address the spurious correlation. Specifically, we first decompose the image representation into a high-level semantic representation for core attributes of the image and a low-level stochastic representation for other random or less structured aspects, with the former extracted by a semantic encoder and the latter derived via a stochastic encoder. We then introduce a causal effect learning module to capture the direct causal effect, that is, the difference of potential outcomes before and after intervening on the visual attributes. In addition, a diffusion-based learning strategy is designed to optimize the representation learning process. Empirical evaluations on two benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, enabling highly controllable image editing by modifying learned visual representations.
Shanshan Huang 0004, Haoxuan Li 0001, Chunyuan Zheng 0001, Lei Wang 0197, Guorui Liao, Zhili Gong 0001, Huayi Yang, Li Liu 0001
CVPR5
2025 Shimmer: a Provably Secure Steganography Based on Entropy Collecting Mechanism
Minhao Bai, Kaiyi Pang, Guorui Liao, Jinshuai Yang, Yongfeng Huang 0001
USENIX Security Symposium3
2025 A Framework for Designing Provably Secure Steganography
Guorui Liao, Jinshuai Yang, Weizhi Shao, Yongfeng Huang 0001
USENIX Security Symposium1
2025 CAP: Causal Air Quality Index Prediction Under Interference with Unmeasured Confounding
abstract
A significant challenge in air quality index (AQI) prediction is to accurately evaluate the potential outcomes after conducting interventions in pollutant factors such as industrial emissions for each enterprise. Existed methods often suffer from spurious correlations caused by unmeasured confounders and are lack of interpretability of the model, leading to sub-optimal prediction performance. This motivates us to propose a causal AQI prediction framework (CAP) that employs a structural causal model (SCM) to characterize the causal structural variability of various AQI factors for robust AQI prediction. Specifically, we employ the front-door adjustment to explicitly eliminate unmeasured confounders by intervening in industrial emissions from the target enterprise. Meanwhile, we take industrial emissions of neighboring enterprises into account when intervening in the target enterprise and simulate the dispersion of industrial emissions through a Gaussian plume model based on meteorological factors. Experiments on two real-world datasets validate the superior performance of our model on AQI prediction compared to the state-of-the-art baselines.
Huayi Yang, Chunyuan Zheng 0001, Guorui Liao, Shanshan Huang 0004, Jun Liao 0001, Zhili Gong 0001, Haoxuan Li 0001, Li Liu 0001
WWW3
2025 A real-time system for fall prediction and protection with spatio-temporal graph neural network using multiple motion sensors
Li Liu 0001, Xiaohu Li, Guorui Liao, Shu Wang 0005, Changbo Liao, Shengfa Miao, Haimiao Wu, Jun Liao 0001, Qing Tao 0002
Expert Syst. Appl.4
2024 Predicting Fall Events by a Spatio-Temporal Topological Network with Multiple Wearable Sensors
abstract
A key challenge in sensor-based fall prediction is the fact that a fall event can often occur in various configurations of fall poses together with their own spatio-temporal dependencies. This leads us to define a spatio-temporal model to explicitly characterize these internal configurations of poses. In particular, we introduce a graph neural network with spatio-temporal topological structure to encode such latent relations among poses by capturing representative patterns in fall events. Moreover, a human body orientation estimator is devised to capture human low limbs information, and as a result, separate pose dependencies are globally consistent. Empirical evaluations on two benchmark datasets and one in-house dataset suggest our approach significantly outperforms the state-of-the-art methods.
Xiaohu Li, Guorui Liao, Mingrui Yin, Shu Wang 0005, Guoxin Su, Jun Liao 0001, Li Liu 0001
ICASSP3
2024 Fall Prediction by a Spatio-Temporal Multi-Channel Causal Model from Wearable Sensors Data
abstract
Predicting human falls from wearable devices is a complex task due to the inherent diversity and causality of multivariate physical changes, where each instance exhibits a unique style of motion events and their spatio-temporal causal dependencies. Consequently, we propose a multichannel causal model that utilizes the Granger causality test to explicitly delineate these internal configurations of motion events and their causal relationships from a spatio-temporal perspective. Particularly, our model incorporates a multi-head attention mechanism with a distillation component to capture the spatio-temporal dependencies among multiple channels of motion sensors in an end-to-end fashion. Empirical evaluations conducted on two benchmark datasets, as well as one in-house dataset collected by ourselves, indicate that our model significantly surpasses state-of-the-art approaches.
Guorui Liao, Yuxuan Liang 0002, Shu Wang 0005, Li Liu 0001
ICASSP1
2024 Co-Stega: Collaborative Linguistic Steganography for the Low Capacity Challenge in Social Media
abstract
Social media platforms, with their extensive and real-time text data, are important application environments for linguistic generative steganography. However, the fragmented and context-constrained nature of social media text leads to a notably low capacity for hiding messages, making linguistic steganography impractical in real social media platforms. More frustratingly, even for high-capacity linguistic generative steganography, the upper bound of capacity is limited to a low level when required to meet a slightly strict security level under Cachin's model. To overcome the low capacity challenge, we identified an indicator of capacity that is independent of any specific steganography method to analyze the origins of this challenge, then we proposed a novel linguistic steganography framework named Collaborative Steganography (Co-Stega). Co-Stega utilizes existing texts and contextual relevance between texts in social media, collaboratively embedding secret messages in an existing text and its contextually related text via efficient retrieval and generation respectively. Additionally, we proposed an innovative and simple technique called "Entropy Enhancement Strategy", which effectively increases the entropy of generated text, thereby enhancing capacity further. Our evaluation shows that Co-Stega significantly improves capacity and maintains text quality, making it a valuable extension for linguistic generative steganography in social media platforms.
Guorui Liao, Jinshuai Yang, Kaiyi Pang, Yongfeng Huang 0001
IH&MMSec1
2024 A spatio-temporal graph neural network for fall prediction with inertial sensors
abstract
Falls are the leading cause of unintentional human injury , having become a public health event of strong social concern. The fall prediction technology based on wearable inertial sensors is a relatively reliable solution in human activity monitoring, a user scenario with mobility and high information privacy sensitivity, and has the advantages of low cost, small size, and high precision. However, a key challenge in sensor-based fall prediction is the fact that a fall event can often occur in various configurations of fall poses together with their own spatio-temporal dependencies. This leads us to define a spatio-temporal model to explicitly characterize these internal configurations of poses. In particular, we introduce a graph neural network with spatio-temporal topological structure to encode such latent relations among poses by capturing representative patterns in fall events. Moreover, a human body orientation estimator is devised to represent human low limbs information and as a result, separate pose dependencies are globally consistent. Empirical evaluations on two benchmark datasets and one in-house dataset suggest our approach significantly outperforms the state-of-the-art methods.
Shu Wang 0005, Xiaohu Li, Guorui Liao, Changbo Liao, Ming Liu 0007, Jun Liao 0001, Li Liu 0001
Knowl. Based Syst.3