Guoyuan Lin

dblp:17/1301 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Branch-MFA-TDNN: A Parallel Branch Speaker Verification Model for Voice IoT
abstract
The security of voice control in the Voice Internet of Things (Voice IoT) heavily relies on the fast and accurate authentication of the command issuer. In this work, we focus on the critical application scenario of Voice IoT in underground coal mines, where voice commands typically last 4–10 seconds. Speech in this scenario typically consists of short, imperative utterances and faces challenges from environmental noise and device heterogeneity. The limitations of traditional speaker verification models in temporal modeling restrict their performance in such scenarios. To address this, this paper proposes a three-dimensional attention module (Branch-MFA) designed for Voice IoT. This module employs a dual-parallel branch architecture: the MFA branch is responsible for extracting attention in the frequency and channel dimensions, and its multi-scale nature enables it to effectively focus on speaker-discriminative frequency bands that remain stable under noise and different collection devices, thereby enhancing the model’s environmental robustness; the GLTA branch, through its innovative grouped variable-length attention mechanism, specifically models the temporal structure of these short voice commands, addressing the challenge of sparse temporal information in short utterances. By integrating the dual-branch outputs through a fusion module, we construct the Branch-MFA-TDNN model. Experiments on the Cn-Celeb dataset show that this model significantly outperforms baseline models in short-utterance verification tasks, particularly for the challenging 4–10 second duration relevant to mine communications, providing an identity authentication solution for Voice IoT that combines high security and real-time performance. We have also released the code1for future comparison.
Guoyuan Lin, Jinbing Deng, Wei Chen 0036, Zehua Wang 0001, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.1
2026 VoIP Call Identification via a Dual-Level 1D-CNN With Frame and Utterance Features
abstract
The increasing use of Voice over Internet Protocol (VoIP) technology in telecom fraud has become a serious global concern. Its ability to spoof caller IDs and IP addresses, and the use of overseas or anonymized servers make VoIP-based scams difficult to trace and regulate. As a result, distinguishing VoIP calls from conventional mobile phone calls based on voice signal characteristics is crucial for enhancing anti-fraud measures. However, existing forensic techniques often struggle to accurately identify speech transmitted via VoIP. To address this challenge, we propose a dual-level 1D-CNN that leverages both frame and utterance features for effective VoIP detection. After evaluating a range of acoustic features, we primarily focus on short-frame Mel-Frequency Cepstral Coefficients (MFCCs) due to their effectiveness in capturing VoIP characteristics. Given the frame-based processing and transmission nature of VoIP, we employ a 1D-CNN, rather than the more commonly used 2D-CNN that treats spectrograms as image, to extract frame-level codec features. Finally, we propose a dual-level classification strategy: the frame-level classifier captures encoding discrepancies within individual frames, while the utterance-level classifier aggregates these frame-level features to learn global encoding patterns through global covariance pooling. Experimental results on the VoIP Phone Call Identification Database (VPCID) demonstrate that the proposed method consistently outperforms existing approaches, delivering superior accuracy and robustness across a wide range of challenging scenarios. Moreover, comprehensive ablation studies validate the effectiveness and rationale behind the design of the proposed model architecture.
Guoyuan Lin, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.1
2025 A hot-repair method for the running software with zero suspends
abstract
Abstract Repairing software defects is crucial to improving the security and robustness of software. Traditional methods repair software defects by using the “stop-repair-restart” approach. Unfortunately, in some scenarios, such as cloud environments, restarting critical software is expensive. Dynamic methods enable the defect repair while the software is running, which can avoid software restart. However, dynamic repair requires capture the safe state of the running software (process). Otherwise, it will cause execution exceptions or even process crashes. To address the complex state issues in multi-threaded environments, existing methods modify the kernel or compiler, or even pre-adding custom code to the target software, which reduces their generality. To solve this problem, we propose a hot repair method HotFix, which can fix the defects without any software suspending. HotFix places probes in the process to receive state signals, which can avoid complex and time-consuming state identification. Then, it selects the safe zone and repairs the defects when the code in the safe zone is called, which can prevent the target process from hanging for a long time. Finally, HotFix completes multi-threaded automated migration online. Experiments and analysis show that HotFix can achieve hot repair in complex environments. We found that the affected function was executed no more than 1,000 times during the fix. Introducing 3us and 11us delays per request respectively when repairing Redis and Mem cached, and the requests influenced are limited.
Guoyuan Lin, Jiazhen Cai, Jiahui Zhou, Shanqing Guo
Cybersecur.1
2025 An audio watermarking method against re-recording distortions
Guoyuan Lin, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
Pattern Recognit.1
2024 114Xray: A Large-Scale X-Ray Security Detection Benchmark and Aware Enhance Network for Real-World Prohibited Item Inspection in Baggage
Hongxia Gao, Zhenming Guan, Yaobin Huang, Hongyu Liao, Hongzhen Zheng, Runze Lin, Litao Li, Haolin Tang, Guoyuan Lin, Zhanhong Chen
PRCV (11)11
2024 One-Class Neural Network With Directed Statistics Pooling for Spoofing Speech Detection
abstract
Existing deep learning models for spoofing speech detection often struggle to effectively generalize to unseen spoofing attacks that were not present during the training stage. Moreover, the presence of class imbalance further compounds this issue by biasing the learning process towards seen attack samples. To address these challenges, we present an innovative end-to-end model called One-Class Neural Network with Directed Statistics Pooling (OCNet-DSP). Our model incorporates a feature cropping operation to attenuate high-frequency components, mitigating the risk of overfitting. Additionally, leveraging the time-frequency characteristics of speech signals, we introduce a directed statistics pooling layer that extracts more effective features for distinguishing between bonafide and spoofing classes. We also propose the Threshold One-class Softmax loss, which mitigates class imbalance by reducing the optimization weight of spoofing samples during training. Extensive comparative results demonstrate that the proposed model outperforms all existing single models, achieving an equal error rate of 0.44% and a minimum detection cost function of 0.0145 for the ASVspoof 2019 logical access database. Moreover, the proposed ensemble version, which accommodates speech inputs of varying lengths in each submodel, maintains state-of-the-art performance among reproducible ensemble models. Additionally, numerous ablation experiments, along with a cross-dataset experiment, are conducted to validate the rationality and effectiveness of the proposed model.
Guoyuan Lin, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.1
2023 MagBox: Keep the risk functions running safely in a magic box
Guoyuan Lin, Yeh-Ching Chung, YaoWen Ma
Future Gener. Comput. Syst.2
2022 MProbe: Make the code probing meaningless
abstract
Modern security methods use address space layout randomization (ASLR) to defend against code reuse attacks (CRAs). However, code probing can still obtain the content and address of the code through code probing. Code probing invalidates the widely used ASLR methods, causing researchers to lose confidence in them. On the contrary, we believe the ASLR is still effective, if it has anti-probing capability. To enhance the anti-probing capability of ASLR and defense CRAs, this paper proposes an anti-probing method MProbe. First, it detects the code probing activities of attackers, including address probing and content probing. Next, the execution permission of the probed code will be de-enabled in the original address space. At the same time, the equivalent code block in a random address space will replace the probed code. Finally, new security strategies are used to prevent the probed code blocks from being used as gadgets. Experiments and analysis show that MProbe has a good defense effect against CRAs based on code probing, and only introduces less than 3% performance overhead to the operating system (OS).
Yeh-Ching Chung, Jinbiao Xing, Guoyuan Lin
ACSAC5
2022 KPointer: Keep the code pointers on the stack point to the right code
Yeh-Ching Chung, Shanqing Guo, Guoyuan Lin
Comput. Secur.6