Guixin Ye

dblp:125/3245 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-2074-4253ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SSFuzz: Synthesizing and scheduling bug-triggering code segments for history-driven compiler testing
Tianmin Hu, Zhenye Fan, Zhanbo Ye, Guixin Ye
Empir. Softw. Eng.4
2025 Scenario: User-Device Authentication on Smart IoTs Using Commodity RFID
abstract
User and device authentication are vital to the deployment of smart Internet of Things (IoT) devices. Unfortunately, achieving robust authentication on a diverse set of heterogeneous IoT devices remains an open problem. This paper presentsScenario, a generic authentication method to support user-device authentication on a wide range of IoT devices, using RFID-based wireless sensing.Scenarioonly requires attaching an RFID tag on the target device surface. It then uses the unique RFID signal characteristics introduced by the device material and user gestures to perform device and user authentication. We developed a prototype ofScenariousing commercial off-the-shelf devices and applied it to a multi-device smart environment. Experimental results show thatScenariois reliable, giving an average identification accuracy of 97.3% and 96.7% of the device and user authentication stages in diverse environments, respectively.
Weiyuan Tong, Zhanyong Tang, Huanting Wang, Guixin Ye, Shuangjiao Zhai, Zheng Wang 0001
IEEE Trans. Dependable Secur. Comput.5
2025 Trust in a Decentralized World: Data Governance From Faithful, Private, Verifiable, and Traceable Data Feeds
abstract
Blockchain technology autonomously executes smart contracts that require external data to facilitate specific applications, underscoring the necessity for Authenticated Data Feeds (ADF). Existing solutions fall short in providing genuine authentication of data, lack private and verifiable computations across multiple data sources, and overlook data traceability, rendering current systems inadequate for complex applications. We present WuKong (WK), a data governance system that offers authenticated, privately verifiable, and traceable data feeds. WK enables a server to collect faithful data through an oracle committee and to prove computation correctness in zero-knowledge proofs, and empowers legal entities to trace a leakage source conditionally. We formally define and prove the security of WK in the universal composability framework. We implement three applications that seamlessly integrate with WK. Experimental results indicate that WK effectively liberates sensitive data from distributed, untrusted, and anonymous providers, making it accessible to various services and establishing trust in a decentralized world.
Meng Li 0006, Yifei Chen 0005, Yan Qiao 0001, Guixin Ye, Zijian Zhang 0001, Liehuang Zhu, Mauro Conti
IEEE Trans. Inf. Forensics Secur.4
2024 History-driven Compiler Fuzzing via Assembling and Scheduling Bug-triggering Code Segments
abstract
History-driven testing techniques have been proven to be an effective method for detecting compiler bugs. It employs fuzzing history (e.g., historical test cases or historical execution information) to guide to generate valid test cases. However, prior methods either have an inefficient capability in synthesizing bug-triggering test cases or suffer from a plateau of code coverage, causing a low bug-exposing ability. This paper presents ASMFUZZ, another history-driven compiler testing framework by applying a multi-metric hybrid scheduling strategy. Specifically, ASMFUZZ first extracts the bug-triggering code segments from the historical test cases that triggered bugs. The extracted bug-triggering code segments are then used to assemble new test cases. To ensure the correctness of the newly synthesized test cases, ASMFUZZ always selects the segments with code context dependencies for assembly. Duration assembly, the ingredients to be assembled are determined based on multiple feedback metrics (e.g., anomalous behaviors and code coverage). To do so, we proposed a multi-metric hybrid scheduling scheme to select optimal code segments in each testing iteration. This contributes to continuously covering deep code branches of compiler duration whole testing process, avoiding getting stuck in the plateau of code coverage. We evaluated ASMFUZZ on three mainstream JVMs including OpenJ9, HotSpot, and GraalVM involving six JDK versions. Within a 72-hour concurrent test run, ASMFUZZ exposed 16 previously unknown unique bugs, of which 11 have been confirmed by the developers. We also compared ASMFUZZ to four prior state-of-the-art fuzzers. ASMFUZZ uncovers 1.6~2.2× more bugs than comparative baselines.
Zhenye Fan, Guixin Ye, Tianmin Hu, Zhanyong Tang
ISSRE2
2024 UPBEAT: Test Input Checks of Q# Quantum Libraries
abstract
High-level programming models like Q# significantly simplify the complexity of programming for quantum computing. These models are supported by a set of foundation libraries for code development. However, errors can occur in the library implementation, and one common root cause is the lack of or incomplete checks on properties like values, length, and quantum states of inputs passed to user-facing subroutines. This paper presents Upbeat, a fuzzing tool to generate random test cases for bugs related to input checking in Q# libraries. Upbeat develops an automated process to extract constraints from the API documentation and the developer implemented input-checking statements. It leverages open-source Q# code samples to synthesize test programs. It frames the test case generation as a constraint satisfaction problem for classical computing and a quantum state model for quantum computing to produce carefully generated subroutine inputs to test if the input-checking mechanism is appropriately implemented. Under 100 hours of automated test runs, Upbeat has successfully identified 16 bugs in API implementations and 4 documentation errors. Of these, 14 have been confirmed, and 12 have been fixed by the library developers.
Tianmin Hu, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Huanting Wang, Meng Li 0006, Zheng Wang 0001
ISSTA2
2024 Anonymous, Secure, Traceable, and Efficient Decentralized Digital Forensics
abstract
Digital forensics is crucial to fight crimes around the world. Decentralized Digital Forensics (DDF) promotes it to another level by channeling the power of blockchain into digital investigations. In this work, we focus on the privacy and security of DDF. Our motivations arise from (1) how to track an anonymous-and-malicious data user who leaks only a part of the previously requested data, (2) how to achieve access control while protecting data from untrusted data centers, and (3) how to enable efficient and secure search on the blockchain. To address these issues, we propose Themis: an anonymous and secure DDF scheme with traceable anonymity, private access control, and efficient search. Our framework is boosted by establishing a Trusted Execution Environment in each authority (blockchain node) for securing the uploading, requesting, and searching. To instantiate the framework, we design a secure and robust watermarking scheme in conjunction with decentralized anonymous authentication, a private and fine-grained access control scheme, and an efficient and secure search scheme based on a dynamically updated data structure. We formally define and prove the privacy and security of Themis. We build a prototype with Ethereum and Intel SGX2 to evaluate its performance, which supports processing data from a considerable number of data providers and investigators.
Meng Li 0006, Yanzhe Shen, Guixin Ye, Jialing He, Zijian Zhang 0001, Liehuang Zhu, Mauro Conti
IEEE Trans. Knowl. Data Eng.3
2023 Optimizing HPC I/O Performance with Regression Analysis and Ensemble Learning
abstract
To improve parallel I/O performance, it is imperative to optimize the adjustable parameters across the different layers of the I/O software stack. Finding an optimal configuration for different scenarios is hampered by the complex interaction dynamics between these parameters and the large parameter space. Previous research efforts have focused on tuning these parameters using independent algorithms; however, these approaches exhibit certain shortcomings such as unstable performance results and delayed convergence rates.This paper introduces OPRAEL, an auto-tuning approach on parallel I/O tasks by ensembles and performance modeling using regression analysis. To test its effectiveness, we applied this approach on the Tianhe-II supercomputer using one well-known I/O benchmark(IOR) and two I/O kernels(S3D-I/O, BT-I/O). Leveraging our experience in predictive modeling, we optimized the tuning of the I/O stack parameters. Our experimental results show a remarkable 10.2X improvement in write performance speedup for the optimization task with BT-I/O and a 500x500x500 input. We also compared the potential of using a single search algorithm versus using reinforcement learning search in the I/O parameter auto-optimization task. Our results show that OPRAEL outperforms the traditional approach, resulting in a maximum 8.4X improvement in write performance for the 128-process IOR optimization.
Zhangyu Liu, Cheng Zhang 0007, Jianbin Fang, Lin Peng 0001, Guixin Ye, Zhanyong Tang
CLUSTER6
2023 A Generative and Mutational Approach for Synthesizing Bug-Exposing Test Cases to Guide Compiler Fuzzing
abstract
Random test case generation, or fuzzing, is a viable means for uncovering compiler bugs. Unfortunately, compiler fuzzing can be time-consuming and inefficient with purely randomly generated test cases due to the complexity of modern compilers. We present COMFUZZ, a focused compiler fuzzing framework. COMFUZZ aims to improve compiler fuzzing efficiency by focusing on testing components and language features that are likely to trigger compiler bugs. Our key insight is human developers tend to make common and repeat errors across compiler implementations; hence, we can leverage the previously reported buggy-exposing test cases of a programming language to test a new compiler implementation. To this end, COMFUZZ employs deep learning to learn a test program generator from open-source projects hosted on GitHub. With the machine-generated test programs in place, COMFUZZ then leverages a set of carefully designed mutation rules to improve the coverage and bug-exposing capabilities of the test cases. We evaluate COMFUZZ on 11 compilers for JS and Java programming languages. Within 260 hours of automated testing runs, we discovered 33 unique bugs across nine compilers, of which 29 have been confirmed and 22, including an API documentation defect, have already been fixed by the developers. We also compared COMFUZZ to eight prior fuzzers on four evaluation metrics. In a 24-hour comparative test, COMFUZZ uncovers at least 1.5× more bugs than the state-of-the-art baselines.
Guixin Ye, Tianmin Hu, Zhanyong Tang, Zhenye Fan, Shin Hwei Tan, Wenxiang Qian, Zheng Wang 0001
ESEC/SIGSOFT FSE1
2022 GENDA: A Graph Embedded Network Based Detection Approach on encryption algorithm of binary program
Xiao Li 0061, Yuanhai Chang, Guixin Ye, Xiaoqing Gong, Zhanyong Tang
J. Inf. Secur. Appl.3
2022 Detecting code vulnerabilities by learning from large-scale open source repositories
Rongze Xu, Zhanyong Tang, Guixin Ye, Huanting Wang, Xin Ke, Dingyi Fang, Zheng Wang 0001
J. Inf. Secur. Appl.3
2021 Automated conformance testing for JavaScript engines via deep compiler fuzzing
abstract
JavaScript (JS) is a popular, platform-independent programming language. To ensure the interoperability of JS programs across different platforms, the implementation of a JS engine should conform to the ECMAScript standard. However, doing so is challenging as there are many subtle definitions of API behaviors, and the definitions keep evolving.
Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Lizhong Bian, Zheng Wang 0001
PLDI1
2021 Towards practical 3D ultrasound sensing on commercial-off-the-shelf mobile devices
Shuangjiao Zhai, Guixin Ye, Zhanyong Tang, Jie Ren 0007, Dingyi Fang, Baoying Liu, Zheng Wang 0001
Comput. Networks2
2021 Combining Graph-Based Learning With Automated Data Collection for Code Vulnerability Detection
abstract
This paper presents FUNDED (Flow-sensitive vUl-Nerability coDE Detection), a novel learning framework for building vulnerability detection models. Funded leverages the advances in graph neural networks (GNNs) to develop a novel graph-based learning method to capture and reason about the program's control, data, and call dependencies. Unlike prior work that treats the program as a sequential sequence or an untyped graph, Funded learns and operates on a graph representation of the program source code, in which individual statements are connected to other statements through relational edges. By capturing the program syntax, semantics and flows, Funded finds better code representation for the downstream software vulnerability detection task. To provide sufficient training data to build an effective deep learning model, we combine probabilistic learning and statistical assessments to automatically gather high-quality training samples from open-source projects. This provides many real-life vulnerable code training samples to complement the limited vulnerable code samples available in standard vulnerability databases. We apply Funded to identify software vulnerabilities at the function level from program source code. We evaluate Funded on large real-world datasets with programs written in C, Java, Swift and Php, and compare it against six state-of-the-art code vulnerability detection models. Experimental results show that Funded significantly outperforms alternative approaches across evaluation settings.
Huanting Wang, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Yansong Feng 0002, Lizhong Bian, Zheng Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2020 Deep Program Structure Modeling Through Multi-Relational Graph-based Learning
abstract
Deep learning is emerging as a promising technique for building predictive models to support code-related tasks like performance optimization and code vulnerability detection. One of the critical aspects of building a successful predictive model is having the right representation to characterize the model input for the given task. Existing approaches in the area typically treat the program structure as a sequential sequence but fail to capitalize on the rich semantics of data and control flow information, for which graphs are a proven representation structure.
Guixin Ye, Zhanyong Tang, Huanting Wang, Dingyi Fang, Jianbin Fang, Songfang Huang, Zheng Wang 0001
PACT1
2020 Compile-time code virtualization for android applications
Zhanyong Tang, Guixin Ye, Dongxu Peng, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001
Comput. Secur.3
2020 Semantics-aware obfuscation scheme prediction for binary
Zhanyong Tang, Guixin Ye, Dongxu Peng, Dingyi Fang, Xiaojiang Chen, Zheng Wang 0001
Comput. Secur.3
2020 Using Generative Adversarial Networks to Break and Protect Text Captchas
abstract
Text-based CAPTCHAs remains a popular scheme for distinguishing between a legitimate human user and an automated program. This article presents a novel genetic text captcha solver based on the generative adversarial network. As a departure from prior text captcha solvers that require a labor-intensive and time-consuming process to construct, our scheme needs significantly fewer real captchas but yields better performance in solving captchas. Our approach works by first learning a synthesizer to automatically generate synthetic captchas to construct a base solver. It then improves and fine-tunes the base solver using a small number of labeled real captchas. As a result, our attack requires only a small set of manually labeled captchas, which reduces the cost of launching an attack on a captcha scheme. We evaluate our scheme by applying it to 33 captcha schemes, of which 11 are currently used by 32 of the top-50 popular websites. Experimental results demonstrate that our scheme significantly outperforms four prior captcha solvers and can solve captcha schemes where others fail. As a countermeasure, we propose to add imperceptible perturbations onto a captcha image. We demonstrate that our countermeasure can greatly reduce the success rate of the attack.
Guixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu, Yansong Feng 0002, Pengfei Xu 0003, Xiaojiang Chen, Jungong Han, Zheng Wang 0001
ACM Trans. Priv. Secur.1
2018 Yet Another Text Captcha Solver: A Generative Adversarial Network Based Approach
abstract
Despite several attacks have been proposed, text-based CAPTCHAs are still being widely used as a security mechanism. One of the reasons for the pervasive use of text captchas is that many of the prior attacks are scheme-specific and require a labor-intensive and time-consuming process to construct. This means that a change in the captcha security features like a noisier background can simply invalid an earlier attack. This paper presents a generic, yet effective text captcha solver based on the generative adversarial network. Unlike prior machine-learning-based approaches that need a large volume of manually-labeled real captchas to learn an effective solver, our approach requires significantly fewer real captchas but yields much better performance. This is achieved by first learning a captcha synthesizer to automatically generate synthetic captchas to learn a base solver, and then fine-tuning the base solver on a small set of real captchas using transfer learning. We evaluate our approach by applying it to 33 captcha schemes, including 11 schemes that are currently being used by 32 of the top-50 popular websites including Microsoft, Wikipedia, eBay and Google. Our approach is the most capable attack on text captchas seen to date. It outperforms four state-of-the-art text-captcha solvers by not only delivering a significant higher accuracy on all testing schemes, but also successfully attacking schemes where others have zero chance. We show that our approach is highly efficient as it can solve a captcha within 0.05 second using a desktop GPU. We demonstrate that our attack is generally applicable because it can bypass the advanced security features employed by most modern text captcha schemes. We hope the results of our work can encourage the community to revisit the design and practical use of text captchas.
Guixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu, Yansong Feng 0002, Pengfei Xu 0003, Xiaojiang Chen, Zheng Wang 0001
CCS1
2018 Exploiting Code Diversity to Enhance Code Virtualization Protection
abstract
Code virtualization built upon virtual machine (VM)technologies is emerging as a viable method for implementing code obfuscation to protect programs against unauthorized analysis. State-of-the-art VM-based protection approaches use a fixed set of virtual instructions and bytecode interpreters across programs. This, however, exposes a security vulnerability where an experienced attacker can use knowledge extracted from other programs to quickly uncover the mapping between virtual instructions and native code for applications protected under the same scheme. In this paper, we propose a novel VM-based code obfuscation system to address this problem. The core idea of our approach is to obfuscate the mapping between the opcodes of bytecode instructions and their semantics. We achieve this by partitioning each protected code region into multiple segments where the mapping of opcodes and their semantics is randomized in different ways in different segments. In this way, each bytecode instruction will be translated into different native code in different sections of the obfuscated code. This significantly increases the diversity of the program behavior. As a result, the knowledge of bytecode to native code mappings obtained from other programs will be less useful when targeting a new program. We evaluate our approach on a set of real-world applications and compare it against two state-of-the-art VM-based code obfuscation approaches. Experimental results show that our approach is effective, which provides stronger protection with comparable runtime overhead and code size.
Zhanyong Tang, Guixin Ye, Xiaoqing Gong, Wei Wangg, Dingyi Fang, Zheng Wang 0001
ICPADS3
2018 A Video-based Attack for Android Pattern Lock
abstract
Pattern lock is widely used for identification and authentication on Android devices. This article presents a novel video-based side channel attack that can reconstruct Android locking patterns from video footage filmed using a smartphone. As a departure from previous attacks on pattern lock, this new attack does not require the camera to capture any content displayed on the screen. Instead, it employs a computer vision algorithm to track the fingertip movement trajectory to infer the pattern. Using the geometry information extracted from the tracked fingertip motions, the method can accurately infer a small number of (often one) candidate patterns to be tested by an attacker. We conduct extensive experiments to evaluate our approach using 120 unique patterns collected from 215 independent users. Experimental results show that the proposed attack can reconstruct over 95% of the patterns in five attempts. We discovered that, in contrast to most people’s belief, complex patterns do not offer stronger protection under our attacking scenarios. This is demonstrated by the fact that we are able to break all but one complex patterns (with a 97.5% success rate) as opposed to 60% of the simple patterns in the first attempt. We demonstrate that this video-side channel is a serious concern for not only graphical locking patterns but also PIN-based passwords, as algorithms and analysis developed from the attack can be easily adapted to target PIN-based passwords. As a countermeasure, we propose to change the way the Android locking pattern is constructed and used. We show that our proposal can successfully defeat this video-based attack. We hope the results of this article can encourage the community to revisit the design and practical use of Android pattern lock.
Guixin Ye, Zhanyong Tang, Dingyi Fang, Xiaojiang Chen, Willy Wolff, Adam J. Aviv, Zheng Wang 0001
ACM Trans. Priv. Secur.1
2017 Cracking Android Pattern Lock in Five Attempts
Guixin Ye, Zhanyong Tang, Dingyi Fang, Xiaojiang Chen, Kwang In Kim, Ben Taylor 0001, Zheng Wang 0001
NDSS1