Chijin Zhou

dblp:228/5716 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-6446-247XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 6 first-author · 13 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Fine-Tuning Data Recovery Attack on Generative Language Models via Backdooring
abstract
Generative language models (GLMs) are increasingly integrated into modern intelligent applications to power intelligent functionalities. Developers often fine-tune open-source GLMs on proprietary data and deploy them in real-world applications. In this paper, we reveal a novel model supply chain attack that exploits this workflow: by injecting backdoors into the source code of an open-source GLM, an adversary can induce the model to memorize fine-tuning data and later regenerate it via crafted prompts. We propose LURE, a new backdoor-based data recovery attack that exploits memorization capabilities of fine-tuned models. During fine-tuning, LURE stealthily injects unique and attacker-enumerable hash prompts, and incorporates a Position-Decay Weighted Aligned Cross-Entropy Loss into the original fine-tuning loss, strengthening the association between injected prompts and corresponding data samples for effective data recovery. To achieve stealthy and transparent attack injection, LURE employs a stealthy backdoor within the model’s source code, enabling automatic injection of hash prompts during fine-tuning and thus maintaining the user’s original fine-tuning workflow. LURE also proposes several optimizations to maintain minimal impact on the performance of the original task and external training state. Extensive evaluations demonstrate the remarkable efficacy of LURE, achieving a 45%-68% data recovery rate while maintaining the attack’s transparency, stealthiness, and showcasing its ability to evade existing defenses.
Zhenya Ma, Yongheng Deng, Ziqing Qiao, Quan Zhang 0003, Chijin Zhou, Fan Wu 0014, Yaoxue Zhang, Ju Ren 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Janus: Detecting Rendering Bugs in Web Browsers via Visual Delta Consistency
abstract
Rendering lies at the heart of our modern web experience. However, the correctness of browser rendering is not always guaranteed, often leading to rendering bugs. Traditional differential testing, while successful in various domains, falls short when applied to rendering bug detection because an HTML file is likely yield different rendered outcomes across different browsers. This paper introduces Visual Delta Consistency, a test oracle to detect rendering bugs in web browsers, aiming to make rendered pages across browsers comparable. Our key insight is that any modifications made to an HTML file should uniformly influence rendering outcomes across browsers. Specifically, when presented with two HTML files that differ only by minor modifications, the reaction of all browsers should be consistent, i.e., either all browsers render them identically or all render them differently. Based on this insight, We implemented it as a practical fuzzer named Janus. It constructs pairs of slightly modified HTML files and observes the change statuses of the corresponding rendered pages across browsers for bug detection. We evaluated it on three widely-used browsers, i.e., Chrome, Safari, and Firefox. In total, Janus detected 31 non-crash rendering bugs, out of which 24 confirmed with 8 fixed.
Chijin Zhou, Quan Zhang 0003, Bingzhou Qian, Yu Jiang 0001
ICSE1
2025 LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux Distributions
abstract
An open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases.
Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi
ASE2
2024 An Empirical Study of Data Disruption by Ransomware Attacks
abstract
The threat of ransomware to the software ecosystem has become increasingly alarming in recent years, raising a demand for large-scale and comprehensive ransomware analysis to help develop more effective countermeasures against unknown attacks. In this paper, we first collect a real-world dataset MarauderMap, consisting of 7,796 active ransomware samples, and analyze their behaviors of disrupting data in victim systems. All samples are executed in isolated testbeds to collect all perspectives of six categories of runtime behaviors, such as API calls, I/O accesses, and network traffic. The total logs volume is up to 1.98 TiB. By assessing collected behaviors, we present six critical findings throughout ransomware attacks' data reconnaissance, data tampering, and data exfiltration phases. Based on our findings, we propose three corresponding mitigation strategies to detect ransomware during each phase. Experimental results show that they can enhance the capability of state-of-the-art anti-ransomware tools. We report a preliminary result of a 41%-69% increase in detection rate with no additional false positives, showing that our insights are helpful.
Yiwei Hou, Lihua Guo, Chijin Zhou, Zijing Yin, Shanshan Li 0001, Chengnian Sun, Yu Jiang 0001
ICSE3
2024 Towards More Complete Constraints for Deep Learning Library Testing via Complementary Set Guided Refinement
abstract
Deep learning library is important in AI systems. Recently, many works have been proposed to ensure its reliability. They often model inputs of tensor operations as constraints to guide the generation of test cases. However, these constraints may narrow the search space, resulting in incomplete testing. This paper introduces a complementary set-guided refinement that can enhance the completeness of constraints. The basic idea is to see if the complementary set of constraints yields valid test cases. If so, the original constraint is incomplete and needs refinement. Based on this idea, we design an automatic constraint refinement tool, DeepConstr, which adopts a genetic algorithm to refine constraints for better completeness. We evaluated it on two DL libraries, PyTorch and TensorFlow. DeepConstr discovered 84 unknown bugs, out of which 72 were confirmed, with 51 fixed. Compared to state-of-the-art fuzzers, DeepConstr increased coverage for 43.44% of operators supported by NNSmith, and 59.16% of operators supported by NeuRI.
Gwihwan Go, Chijin Zhou, Quan Zhang 0003, Xiazijian Zou, Heyuan Shi, Yu Jiang 0001
ISSTA2
2024 Imperceptible Content Poisoning in LLM-Powered Applications
abstract
Large Language Models (LLMs) have shown their superior capability in natural language processing, promoting extensive LLM-powered applications to be the new portals for people to access various content on the Internet. However, LLM-powered applications do not have sufficient security considerations on untrusted content, leading to potential threats. In this paper, we reveal content poisoning, where attackers can tailor attack content that appears benign to humans but causes LLM-powered applications to generate malicious responses. To highlight the impact of content poisoning and inspire the development of effective defenses, we systematically analyze the attack, focusing on the attack modes in various content, exploitable design features of LLM application frameworks, and the generation of attack content. We carry out a comprehensive evaluation on five LLMs, where content poisoning achieves an average attack success rate of 89.60%. Additionally, we assess content poisoning on four popular LLM-powered applications, achieving the attack on 72.00% of the content. Our experimental results also show that existing defenses are ineffective against content poisoning. Finally, we discuss potential mitigations for LLM application frameworks to counter content poisoning.
Quan Zhang 0003, Chijin Zhou, Gwihwan Go, Binqi Zeng, Heyuan Shi, Zichen Xu 0001, Yu Jiang 0001
ASE2
2024 Automatic Policy Synthesis and Enforcement for Protecting Untrusted Deserialization
Quan Zhang 0003, Zijing Yin, Chijin Zhou, Yu Jiang 0001
NDSS4
2024 Data Coverage for Guided Fuzzing
Jie Liang 0006, Chijin Zhou, Zhiyong Wu 0010, Jingzhou Fu, Zhuo Su 0005, Qing Liao 0001, Bin Gu 0006, Bodong Wu, Yu Jiang 0001
USENIX Security Symposium3
2024 PolyJuice: Detecting Mis-compilation Bugs in Tensor Compilers with Equality Saturation Based Rewriting
abstract
Tensor compilers are essential for deploying deep learning applications across various hardware platforms. While powerful, they are inherently complex and present significant challenges in ensuring correctness. This paper introduces PolyJuice, an automatic detection tool for identifying mis-compilation bugs in tensor compilers. Its basic idea is to construct semantically-equivalent computation graphs to validate the correctness of tensor compilers. The main challenge is to construct equivalent graphs capable of efficiently exploring the diverse optimization logic during compilation. We approach it from two dimensions. First, we propose arithmetic and structural equivalent rewrite rules to modify the dataflow of a tensor program. Second, we design an efficient equality saturation based rewriting framework to identify the most simplified and the most complex equivalent computation graphs for an input graph. After that, the outcome computation graphs have different dataflow and will likely experience different optimization processes during compilation. We applied it to five well-tested industrial tensor compilers, namely PyTorch Inductor, OnnxRuntime, TVM, TensorRT, and XLA, as well as two well-maintained academic tensor compilers, EinNet and Hidet. In total, PolyJuice detected 84 non-crash mis-compilation bugs, out of which 49 were confirmed with 20 fixed.
Chijin Zhou, Bingzhou Qian, Gwihwan Go, Quan Zhang 0003, Shanshan Li 0001, Yu Jiang 0001
Proc. ACM Program. Lang.1
2023 Limits of I/O Based Ransomware Detection: An Imitation Based Attack
abstract
By encrypting the data of infected hosts, cryptographic ransomware has caused billions of dollars in financial losses to a wide range of victims. Many detection techniques have been proposed to counter ransomware threats over the past decade. Their common approach is to monitor I/O behaviors from user space and apply custom heuristics to discriminate ransomware. These techniques implicitly assume that ransomware behaves very differently from benign programs in terms of heuristics. However, when we investigated the behavior of benign and ransomware programs, we found that the boundary between their behaviors was blurred. A ransomware program can still achieve its goal even though it follows the behavior patterns of benign programs. In this paper, we aim to explore the limits of ransomware detection techniques that based on I/O behaviors. To this end, we present Animagus, an imitation-based ransomware attack that imitates behaviors of benign programs to disguise its encryption tasks. It first learns behavior patterns from a benign program, and then spawns and orchestrates child processes to perform encryption tasks behaving the same as the benign program. We evaluate its effectiveness against six state-of-the-art detection techniques, and the results show that it can successfully evade these defenses. We investigate in detail why they are ineffective and how Animagus is different from existing ransomware samples. In the end, we discuss potential countermeasures and the benefits that detection tools can gain from our work.
Chijin Zhou, Lihua Guo, Yiwei Hou, Zhenya Ma, Quan Zhang 0003, Zhe Liu 0001, Yu Jiang 0001
SP1
2023 Building Dynamic System Call Sandbox with Partial Order Analysis
abstract
Attack surface reduction is a security technique that secures the operating system by removing the unnecessary code or features of a program. By restricting the system calls that programs can use, the system call sandbox is able to reduce the exposed attack surface of the operating system and prevent attackers from damaging it through vulnerable programs. Ideally, programs should only retain access to system calls they require for normal execution. Many researchers focus on adopting static analysis to automatically restrict the system calls for each program. However, these methods do not adjust the restriction policy along with program execution. Thus, they need to permit all system calls required for program functionalities. We observe that some system calls, especially security-sensitive ones, are used a few times in certain stages of a program’s execution and then never used again. This motivates us to minimize the set of required system calls dynamically. In this paper, we propose , which gradually disables access to unnecessary system calls throughout the program’s execution. To accomplish this, we utilize partial order analysis to transform the program into a partially ordered graph, which enables efficient identification of the necessary system calls at any given point during program execution. Once a system call is no longer required by the program, can restrict it immediately. To evaluate , we applied it to seven widely-used programs with an average of 615 KLOC, including web servers and databases. With partial order analysis, restricts an average of 23.50, 16.86, and 15.89 more system calls than the state-of-the-art Chestnut, Temporal Specialization, and the configuration-aware sandbox, C2C, respectively. For mitigating malicious exploitations, on average, defeats 83.42% of 1726 exploitation payloads with only a 5.07% overhead.
Quan Zhang 0003, Chijin Zhou, Zijing Yin, Zhuo Su 0005, Chengnian Sun, Yu Jiang 0001, Jia-Guang Sun 0001
Proc. ACM Program. Lang.2
2023 Towards Better Semantics Exploration for Browser Fuzzing
abstract
Web browsers exhibit rich semantics that enable a plethora of web-based functionalities. However, these intricate semantics present significant challenges for the implementation and testing of browsers. For example, fuzzing, a widely adopted testing technique, typically relies on handwritten context-free grammars (CFGs) for automatically generating inputs. However, these CFGs fall short in adequately modeling the complex semantics of browsers, resulting in generated inputs that cover only a portion of the semantics and are prone to semantic errors. In this paper, we present SaGe, an automated method that enhances browser fuzzing through the use of production-context sensitive grammars (PCSGs) incorporating semantic information. Our approach begins by extracting a rudimentary CFG from W3C standards and iteratively enhancing it to create a PCSG. The resulting PCSG enables our fuzzer to generate inputs that explore a broader range of browser semantics with a higher proportion of semantically-correct inputs. To evaluate the efficacy of SaGe, we conducted 24-hour fuzzing campaigns on mainstream browsers, including Chrome, Safari, and Firefox. Our approach demonstrated better performance compared to existing browser fuzzers, with a 6.03%-277.80% improvement in edge coverage, a 3.56%-161.71% boost in semantic correctness rate, twice the number of bugs discovered. Moreover, we identified 62 bugs across the three browsers, with 40 confirmed and 10 assigned CVEs.
Chijin Zhou, Quan Zhang 0003, Lihua Guo, Yu Jiang 0001, Qing Liao 0001, Zhiyong Wu 0010, Shanshan Li 0001, Bin Gu 0006
Proc. ACM Program. Lang.1
2022 Unicorn: detect runtime errors in time-series databases with hybrid input synthesis
abstract
The ubiquitous use of time-series databases in the safety-critical Internet of Things domain demands strict security and correctness. One successful approach in database bug detection is fuzzing, where hundreds of bugs have been detected automatically in relational databases. However, it cannot be easily applied to time-series databases: the bulk of time-series logic is unreachable because of mismatched query specifications, and serious bugs are undetectable because of implicitly handled exceptions. In this paper, we propose Unicorn to secure time-series databases with automated fuzzing. First, we design hybrid input synthesis to generate high-quality queries which not only cover time-series features but also ensure grammar correctness. Then, Unicorn uses proactive exception detection to discover minuscule-symptom bugs which hide behind implicit exception handling. With the specialized design oriented to time-series databases, Unicorn outperforms the state-of-the-art database fuzzers in terms of coverage and bugs. Specifically, Unicorn outperforms SQLsmith and SQLancer on widely used time-series databases IoTDB, KairosDB, TimescaleDB, TDEngine, QuestDB, and GridDB in the number of basic blocks by 21%-199% and 34%-693%, respectively. More importantly, Unicorn has discovered 42 previously unknown bugs.
Zhiyong Wu 0010, Jie Liang 0006, Chijin Zhou, Yu Jiang 0001
ISSTA4
2022 Empirical Study of System Resources Abused by IoT Attackers
abstract
IoT devices have been under frequent attacks in recent years, causing severe impacts. Previous research has shown the evolution and features of some specific IoT malware families or stages of IoT attacks through offline sample analysis. However, we still lack a systematic observation of various system resources abused by active attackers and the malicious intentions behind these behaviors. This makes it difficult to design appropriate protection strategies to defend against existing attacks and possible future variants.
Zijing Yin, Chijin Zhou, Yu Jiang 0001
ASE3
2022 Odin: on-demand instrumentation with on-the-fly recompilation
abstract
Instrumentation is vital to fuzzing. It provides fuzzing directions and helps detect covert bugs, yet its overhead greatly reduces the fuzzing throughput. To reduce the overhead, compilers compromise instrumentation correctness for better optimization, or seek convoluted runtime support to remove unused probes during fuzzing.
Jie Liang 0006, Chijin Zhou, Zhiyong Wu 0010, Yu Jiang 0001
PLDI3
2022 Minerva: browser API fuzzing with dynamic mod-ref analysis
abstract
Browser APIs are essential to the modern web experience. Due to their large number and complexity, they vastly expand the attack surface of browsers. To detect vulnerabilities in these APIs, fuzzers generate test cases with a large amount of random API invocations. However, the massive search space formed by arbitrary API combinations hinders their effectiveness: since randomly-picked API invocations unlikely interfere with each other (i.e., compute on partially shared data), few interesting API interactions are explored. Consequently, reducing the search space by revealing inter-API relations is a major challenge in browser fuzzing.
Chijin Zhou, Quan Zhang 0003, Lihua Guo, Jie Liang 0006, Zhe Liu 0001, Mathias Payer, Yu Jiang 0001
ESEC/SIGSOFT FSE1
2022 PATA: Fuzzing with Path Aware Taint Analysis
abstract
Taint analysis assists fuzzers in solving complex fuzzing constraints by inferring the influencing input bytes. Execution paths in real-world programs often reach loops, where constraints in these loops can be visited and recorded multiple times. Conventional taint analysis techniques experience difficulties when distinguishing between multiple occurrences of the same constraint. In this paper, we propose PATA, a fuzzer that implements path-aware taint analysis, i.e. one that distinguishes between multiple occurrences of the same variable based on the execution path information. PATA does so using the following steps. First, PATA identifies variables used in constraints and constructs the Representative Variable Sequence (RVS), consisting of occurrences of all representative constraint variables and their values. Next, PATA perturbs the input, matches its RVS with that of the original input, and looks for value changes to identify the influencing input bytes for each entry in the RVS. Finally, PATA mutates the corresponding input bytes to solve constraints in the given path. To demonstrate the effectiveness of PATA over conventional taint analysis methods, we evaluated its performance on the benchmarks Google’s fuzzer-test-suite and LAVA-M against AFL, MOPT, TortoriseFuzz, VUzzer, Angora, Redqueen, and Greyone. On Google’s fuzzer-test-suite, PATA outperformed these state-of-the-art fuzzers by 29%–1830% and 7%–87% in the number of unique paths found and basic blocks covered, respectively. More importantly, it found more bugs than the comparison fuzzers, including 17 unlisted ones. On LAVA-M, PATA performed the best out of all evaluated fuzzers and found 2602 bugs. On open-source projects, PATA found 40 previously unknown bugs, with 12 of them confirmed as CVEs.
Jie Liang 0006, Chijin Zhou, Zhiyong Wu 0010, Yu Jiang 0001, Jianzhong Liu, Zhe Liu 0001, Jia-Guang Sun 0001
SP3
2021 Industrial Oriented Evaluation of Fuzzing Techniques
abstract
Fuzzing is a promising method for discovering vulnerabilities. Recently, various techniques are developed to improve the efficiency of fuzzing, and impressive gains are observed in evaluation results. However, evaluation is complex, as many factors affect the results, for example, test suites, baseline and metrics. Even more, most experiment setups are lab-oriented, lacking industrial settings such as large code-base and parallel runs. The correlation between the academic evaluation results and the bug-finding ability in real industrial settings has not been sufficiently studied. In this paper, we test representative fuzzing techniques to reveal their efficiency in industrial settings. First, we apply typical fuzzers on academic widely used small projects from LAVAM suite. We also apply the same fuzzers on large practical projects from Google's fuzzer-test-suite, which is rarely used in academic settings. Both experiments are performed in both single and parallel run. By analyzing the results, we found that most optimizations working well on LAVA-M suite fail to achieve satisfying results on Google's fuzzer-test-suite (e.g. compared to AFL, QSYM detects 82x more synthesized bugs in LAVA-M, but only detects 26% real bugs in Google's fuzzer-test-suite), and the original AFL even outperforms most academic optimization variants in industry widely used parallel runs (e.g. AFL covers 13% more paths than AFLFast). Then, we summarize common pitfalls of those optimizations, analyze the corresponding root causes, and propose potential directions such as orchestrations and synchronization to overcome the problems. For example, when running in parallel on those large practical projects, the proposed horizontal orchestration could cover 36%-82% more paths, and discover 46%-150% more unique crashes or bugs, compared to fuzzers such as AFL, FairFuzz and QSYM.
Jie Liang 0006, Chijin Zhou, Yuanliang Chen, Zhiyong Wu 0010, Yu Jiang 0001
ICST3
2021 RIFF: Reduced Instruction Footprint for Coverage-Guided Fuzzing
Jie Liang 0006, Chijin Zhou, Yu Jiang 0001, Rui Wang 0024, Chengnian Sun, Jia-Guang Sun 0001
USENIX ATC3
2020 Zeror: Speed Up Fuzzing with Coverage-sensitive Tracing and Scheduling
abstract
Coverage-guided fuzzing is one of the most popular software testing techniques for vulnerability detection. While effective, current fuzzing methods suffer from significant performance penalty due to instrumentation overhead, which limits its practical use. Existing solutions improve the fuzzing speed by decreasing instrumentation overheads but sacrificing coverage accuracy, which results in unstable performance of vulnerability detection.
Chijin Zhou, Jie Liang 0006, Zhe Liu 0001, Yu Jiang 0001
ASE1
2019 VisFuzz: Understanding and Intervening Fuzzing with Interactive Visualization
abstract
Fuzzing is widely used for vulnerability detection. One of the challenges for an efficient fuzzing is covering code guarded by constraints such as the magic number and nested conditions. Recently, academia has partially addressed the challenge via whitebox methods. However, high-level constraints such as array sorts, virtual function invocations, and tree set queries are yet to be handled. To meet this end, we present VisFuzz, an interactive tool for better understanding and intervening fuzzing process via real-time visualization. It extracts call graph and control flow graph from source code, maps each function and basic block to the line of source code and tracks real-time execution statistics with detail constraint contexts. With VisFuzz, test engineers first locate blocking constraints and then learn its semantic context, which helps to craft targeted inputs or update test drivers. Preliminary evaluations are conducted on four real-world programs in Google fuzzer-test-suite. Given additional 15 minutes to understand and intervene the state of fuzzing, the intervened fuzzing outperform the original pure AFL fuzzing, and the path coverage improvements range from 10.84% to 150.58%, equally fuzzed by for 12 hours.
Chijin Zhou, Jie Liang 0006, Zhe Liu 0001, Chengnian Sun, Yu Jiang 0001
ASE1
2019 EnFuzz: Ensemble Fuzzing with Seed Synchronization among Diverse Fuzzers
Yuanliang Chen, Yu Jiang 0001, Fuchen Ma, Jie Liang 0006, Chijin Zhou, Xun Jiao 0002, Zhuo Su 0005
USENIX Security Symposium6
2018 PAFL: extend fuzzing optimizations of single mode to industrial parallel mode
abstract
Researchers have proposed many optimizations to improve the efficiency of fuzzing, and most optimized strategies work very well on their targets when running in single mode with instantiating one fuzzer instance. However, in real industrial practice, most fuzzers run in parallel mode with instantiating multiple fuzzer instances, and those optimizations unfortunately fail to maintain the efficiency improvements.
Jie Liang 0006, Yu Jiang 0001, Yuanliang Chen, Chijin Zhou, Jia-Guang Sun 0001
ESEC/SIGSOFT FSE5