Zihui Guo

dblp:224/4662 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Hypnos: A Practical Power Side-Channel Attack via CPU Idle Time
abstract
The growing demand for high-performance computing has led to various optimization techniques, but these advancements have also raised concerns about energy consumption. In response, processor vendors have implemented power management features. On x86-based CPUs, C-states allow the processor to enter idle states, reducing power consumption during low workloads. While users cannot directly control these states, C-states provide an interface to monitor CPU idle time, offering transparency without user intervention. However, it remains unclear whether this design could be exploited for power side-channel leakages. In this paper, we propose Hypno, a new type of software-based power side-channel attack on x86-based systems. Our key observation is that the unprivileged access to the CPUIDLE interface provides fine-grained observations of the time spent in various idle states. As this time is directly correlated with CPU activities, unprivileged attackers can leverage this information to establish a new power side channel. To demonstrate the viability of Hypnos, we conduct three end-to-end case studies. First, we demonstrate cross core covert channels that operate even in isolated environments, achieving higher transmission rates than channels that read cpufreq and broader applicability than methods that rely on uncore idle states. Second, we demonstrate a website fingerprinting attack on Google Chrome with high accuracy. Lastly, we successfully break KASLR within 3 minutes.
Yusi Feng, Xin Zhang 0110, Zihui Guo, Ben Liu 0007, Yinqian Zhang
IEEE Trans. Dependable Secur. Comput.3
2025 Into the Unknown: Fuzzing CPU Non-standard Instructions with MystFuzz
abstract
Modern CPU designs have become increasingly complex, making their comprehensive verification a significant challenge. Non-standard instructions, such as illegal, reserved, and hint instructions, are often overlooked during the verification process, potentially leading to critical bugs remaining undetected. Despite recent advancements in fuzzing techniques offering hope for CPU verification, the verification of nonstandard instructions remains a significant challenge. To fill the gap in the current field of CPU verification regarding non-standard instructions, we present MystFuzz, a fuzzing method specifically tailored for non-standard instructions. MystFuzz introduces an efficient instruction space constraint mechanism, supported by a lightweight instruction simulator, to generate large-scale non-standard instructions. This design enables CPU fuzzing without relying on an external golden reference model, and the constrained instruction space dynamically adjusts throughout the fuzzing process. Combined with an efficient exception handling and recovery mechanism, it supports large-scale fuzzing of non-standard instructions in CPUs. Experimental results show significant improvements in fuzzing non-standard instructions for CPUs. Compared to widely-used tools such as riscv-torture and riscv-dv, MystFuzz achieves 215.4x and 61.1x performance improvements in fuzzing, respectively. Even with the same number of non-standard instructions generated, MystFuzz achieves a more diverse range of instruction scenarios. We evaluate five RISC-V CPUs, including XiangShan, CVA6, Rocket, NutShell, Kronos, and discover 19 new bugs (with 10 CVEs assigned) caused by non-standard instructions, highlighting the security impact of non-standard instructions.
Zihui Guo, Wenhao Cui, Miaomiao Yuan, Dan Meng 0002
ACSAC1
2025 DiveFuzz: Enhancing CPU Fuzzing via Diverse Instruction Construction
abstract
Comprehensive exploration of the CPU architectural states in fuzzing is akin to generating diverse test cases, which include a reasonable distribution of opcode and diversity in instruction execution results (typically measured through write-back data). However, our analysis of state-of-the-art CPU fuzzers reveals that they exhibit high repetition in write-back data and an imbalanced distribution of opcodes during fuzzing. This paper presents DiveFuzz, which diversifies write-back data by finely controlling the operands of instructions at runtime, coupled with correlated contextual semantics, to generate instruction streams with diverse write-back data and semantic associations. Furthermore, DiveFuzz introduces a novel mutator that monitors the fuzzing process to dynamically adjust opcode distribution and accurately eliminate false positives. Our evaluations show that DiveFuzz significantly increases the diversity of instruction write-back data and achieves a more balanced opcode distribution compared to state-of-the-art fuzzers. Across five common coverage metrics, DiveFuzz achieves coverage 204× faster than DifuzzRTL and 114× faster than Cascade. We evaluated DiveFuzz on four well-known open-source RISC-V CPUs—XiangShan, CVA6, Rocket, and NutShell—uncovering 26 new bugs, 15 of which have CVE identifiers.
Zihui Guo, Miaomiao Yuan, Yanqi Yang, Dan Meng 0002
CCS1
2025 HScheduler: An execution history-based seed scheduling strategy for hardware fuzzing
Zihui Guo, Yin Lv, Ningning Cui
Comput. Secur.1
2025 Visual context learning based on cross-modal knowledge for continuous sign language recognition
Kailin Liu, Yonghong Hou, Zihui Guo
Vis. Comput.3
2024 Improved Linear Cryptanalysis of Block Cipher BORON
abstract
Abstract BORON is a lightweight substitution–permutation network cipher proposed in 2017. We reduce the number of guessed key bits by key-bridging technology and first utilize Fast Walsh Transform on BORON to minimize the time complexity. Finally, this paper gives the better key-recovery attack against block cipher BORON than previously proposed by 2 rounds: we realize a 11-round key-recovery attack on BORON-80 and 13-round key-recovery attack on BORON-128. The attacks proposed in this paper are the best attacks against BORON-80/128 to date.
Yin Lv, Danping Shi, Lei Hu 0003, Zihui Guo, Caibing Wang
Comput. J.4
2024 FTAN: Frame-to-frame temporal alignment network with contrastive learning for few-shot action recognition
Yonghong Hou, Zihui Guo, Zhiyi Gao
Image Vis. Comput.3
2024 Global and Local Contrastive Learning for Self-Supervised Skeleton-Based Action Recognition
abstract
Contrastive learning for self-supervised skeleton-based action recognition has recently received attention. It has been observed that local crops, containing partial action sequences, can predict action patterns, which is advantageous for skeleton-based action recognition. This paper proposes a Global and Local Contrastive Learning framework (skeleton-logoCLR) with two contrastive learning routes, Global-to-Global and Global-to-Local, which utilize the similarity between global and local crops of the same skeleton sequence. Specifically, in the Global-to-Global route, we design Temporal Attention Crop-Resize (TACR) to learn global semantic information by maximizing the retention of action region in the temporal dimension. In the Global-to-Local route, the proposed Skeleton-logo Augmentation is deviced to concatenate two local crops from different sequences for local semantic learning. Moreover, instead of fusing directly, the losses of two routes are combined in a cascaded manner through the Self-Adaptive Training Strategy (SATS) to achieve stronger generalization performance. Extensive experiments are conducted on the NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets. The results demonstrate that the proposed method achieves remarkable performance.
Jinhua Hu, Yonghong Hou, Zihui Guo, Jiajun Gao
IEEE Trans. Circuits Syst. Video Technol.3
2024 Spatial-Temporal Enhanced Network for Continuous Sign Language Recognition
abstract
Continuous Sign language Recognition (CSLR) aims to generate gloss sequences based on untrimmed sign videos. Since discriminative visual features are essential for CSLR, current efforts mainly focus on strengthening the feature extractor. The feature extractor can be disassembled into a spatial representation module and a short-term temporal module for spatial and visual features modeling. However, existing methods always regard it as a monoblock and rarely implement specific refinements for such two distinct modules, which is difficult to achieve effective modeling of spatial appearance information and temporal motion information. To address the above issues, we proposed a spatial temporal enhanced network which contains a spatial-visual alignment (SVA) module and a temporal feature difference (TFD) module. Specifically, the SVA module conducts an auxiliary task between the spatial features and target gloss sequences to enhance the extraction of hand and facial expressions. Meanwhile, the TFD module is constructed to exploit the underlying dynamic between consecutive frames and inject the aggregated motion information into spatial features to assist short-term temporal modeling. Extensive experimental results demonstrate the effectiveness of the proposed modules and our network achieves state-of-the-art or competitive performance on four public CSLR datasets.
Yonghong Hou, Zihui Guo, Kailin Liu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Automatic Demirci-Selçuk Meet-In-The-Middle Attack On SIMON
abstract
Abstract Demirci–Selçuk meet-in-the-middle (DS-MITM) attack is an effective method for cryptanalysis. As far as we know, the published automatic results of DS-MITM attack are all for byte-oriented ciphers. In this article, we first propose the automatic analysis method of DS-MITM attack for bit-oriented ciphers based on constraint programming, which is integrated with key-bridging technique. Based on the automatic modeling method, we propose the first result of DS-MITM attack on SIMON, which is a family of lightweight block ciphers proposed by the National Security Agency (NSA) in 2013.
Yin Lv, Danping Shi, Qiu Chen, Lei Hu 0003, Zihui Guo
Comput. J.6
2023 Motion saliency based hierarchical attention network for action recognition
Zihui Guo, Yonghong Hou, Renyi Xiao, Chuankun Li, Wanqing Li 0001
Multim. Tools Appl.1
2023 Sign language recognition via dimensional global-local shift and cross-scale aggregation
Zihui Guo, Yonghong Hou, Wanqing Li 0001
Neural Comput. Appl.1
2023 FT-HID: a large-scale RGB-D dataset for first- and third-person human interaction analysis
Zihui Guo, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
Neural Comput. Appl.1
2023 Locality-Aware Transformer for Video-Based Sign Language Translation
abstract
Recently, the application of transformer makes significant progress in sign language translation. However, several characteristics of sign videos are neglected in existing transformer-based methods that hinder translation performance. Firstly, in sign videos, multiple consecutive frames represent a single sign gloss thus the local temporal relations are crucial. Secondly, the inconsistency between video and text demands the non-local and global context modeling ability of the model. To address these issues, a locality-aware transformer is proposed for sign language translation. Concretely, the multi-stride position encoding scheme assigns the same position index to adjacent frames with various strides to enhance the local dependency. Afterward, the adaptive temporal interaction module is utilized to capture non-local and flexible local frame correlation simultaneously. Moreover, a gloss counting task is designed to facilitate the holistic understanding of sign videos. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed framework.
Zihui Guo, Yonghong Hou, Chunping Hou
IEEE Signal Process. Lett.1
2023 Learning Spatio-Temporal Semantics and Cluster Relation for Zero-Shot Action Recognition
abstract
Zero-shot Action Recognition (ZSAR) aims at bridging the video$\rightarrow $class relation with only labeled training data of seen classes while generalizing the model to alleviate the heterogeneity of unseen actions. Most existing methods have comprehensively represented videos and action classes, however, the semantic gap and the hubness problem between them remain crucial challenges that are under-explored. In this paper, we propose an effective method to tackle the above issues. Specifically, to narrow the semantic gap, we end-to-end generate a spatio-temporal semantics for each video, which provides essential textual information to refine the video representation. Furthermore, we propose a compactness-separability loss that optimizes the intra- and inter-class relations in a unified formula and quantitatively constrains cluster distribution, thus effectively diminishing the impact of the hubness problem. Extensive experiments on UCF101, HMDB51, and Olympic Sports datasets prove the effectiveness of the proposed approach and demonstrate our approach outperforms the state-of-the-art methods.
Jiajun Gao, Yonghong Hou, Zihui Guo, Haochun Zheng
IEEE Trans. Circuits Syst. Video Technol.3
2022 Learning Using Privileged Information for Zero-Shot Action Recognition
Zhiyi Gao, Yonghong Hou, Wanqing Li 0001, Zihui Guo
ACCV (4)4
2020 SAR-NAS: Skeleton-based action recognition via neural architecture searching
Yonghong Hou, Pichao Wang, Zihui Guo, Wanqing Li 0001
J. Vis. Commun. Image Represent.4
2020 A new color image encryption scheme based on DNA encoding and spatiotemporal chaotic system
Xuejing Kang, Zihui Guo
Signal Process. Image Commun.2
2019 Single Image De-Raining via Generative Adversarial Nets
abstract
In this paper, we propose a Generative Adversarial Network for Single Image De-raining(GAN-SID). We observe that batch normalization has side effects in the de-raining task. Therefore, we introduce instance normalization to replace the traditional batch normalization layers in both generator and discriminator. Motivated by the Squeeze-and-Excitation (SE) network that can learn the importance of channels, we introduce SE module in the generator to give different weights to the learned features. To preserve image details while removing rain streaks, we propose to utilize pixel-wise loss, perceptual loss, and adversarial loss to train the proposed network. Experiments on two synthetic datasets and real world images demonstrate that the proposed method outperforms state-of-the-art de-raining works in both objective and subjective measurements.
Shichao Li 0006, Yonghong Hou, Huanjing Yue, Zihui Guo
ICME4
2019 Self-Attention Guided Deep Features for Action Recognition
abstract
Skeleton based human action recognition is an important task in computer vision. However, it is very challenging due to the complex spatio-temporal variations of skeleton joints. In this work, we propose an end-to-end trainable network consisting of a Deep Convolutional Model (DCM) and a Self-Attention Model (SAM) for human action recognition from skeleton data. Specifically, skeleton sequences are encoded into color images and fed into DCM to extract deep features. In the SAM, handcrafted features representing the motion degree of joints are extracted and the attention weights are learned by a simple yet effective linear mapping. The effectiveness of proposed method has been verified on NTU RGB+D, SYSU-3D and UTD-MHAD datasets and achieved state-of-the-art results.
Renyi Xiao, Yonghong Hou, Zihui Guo, Chuankun Li, Pichao Wang, Wanqing Li 0001
ICME3