Kaikai Zhang

dblp:321/9151 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RandSet: Randomized Corpus Reduction for Fuzzing Seed Scheduling
abstract
Seed explosion is a fundamental problem in fuzzing seed scheduling. It occurs when a fuzzer maintains a corpus with a huge number of seeds and fails to choose a promising one. Existing seed scheduling works focus on seed prioritization but suffer from the seed explosion since the seed corpus size is still huge. We tackle seed explosion from a new perspective, corpus reduction, i.e., compute a seed corpus subset. Corpus reduction can eliminate redundant seeds in the corpus and significantly reduce corpus size. However, this could lead to poor diversity in seed selection and severely impact the fuzzing performance. Meanwhile, effective corpus reduction incurs large runtime overhead. In practice, it’s challenging to adopt corpus reduction in fuzzing seed scheduling. Prior techniques like cull_queue , AFL-Cmin and MinSet all suffer from poor seed diversity. AFL-Cmin and MinSet incur prohibitive runtime overhead and are hence only applicable to one-time task initial seed selection rather than high-frequency seed scheduling. We propose a novel randomized corpus reduction technique, RandSet, that can reduce the corpus size and yield diverse seed selection simultaneously. Meanwhile, the runtime overhead of RandSet is minimal, suiting a high-frequency seed scheduling task. Our key insight is to introduce randomness into corpus reduction so as to enjoy the two benefits of a randomized algorithm: randomized output (i.e., diverse seed selection) and low runtime overhead. Specifically, we formulate the corpus reduction in seed scheduling as a classic set cover problem and compute a randomized subset of seed corpus as a set cover to cover all features of the entire corpus. We then develop a novel seed scheduling approach using the randomized corpus subset. Our technique can effectively mitigate seed explosion by scheduling a small and randomized subset of the corpus rather than the entire corpus. We implement RandSet on three popular fuzzers: AFL++, LibAFL and Centipede to showcase its general algorithmic design. We perform a comprehensive evaluation of RandSet on three benchmarks: standalone programs, FuzzBench and Magma. Our evaluation results show that RandSet can achieve significantly more diverse seed selection compared with other corpus reduction techniques. RandSet also yields high reduction ratio, achieving an average subset ratio of 4.03% and 5.99% after corpus reduction in terms of standalone programs and FuzzBench programs. In terms of fuzzing performance gain from our randomized corpus reduction, RandSet achieves a 16.58% gain on standalone programs and up to 3.57% gain on FuzzBench programs in AFL++. RandSet triggers up to 7 more ground-truth bugs than the state-of-the-art fuzzer on Magma, while introducing only 3.93% overhead on standalone programs and as low as 1.17% overhead on FuzzBench.
Yuchong Xie, Kaikai Zhang, Rundong Yang, Dongdong She
Proc. ACM Program. Lang.2
2024 DDGF: Dynamic Directed Greybox Fuzzing with Path Profiling
abstract
Coverage-Guided Fuzzing (CGF) has become the most popular and effective method for vulnerability detection. It is usually designed as an automated “black-box” tool. Security auditors start it and then just wait for the results. However, after a period of testing, CGF struggles to find new coverage gradually, thus making it inefficient. It is difficult for users to explain reasons that prevent fuzzing from making further progress and to determine whether the existing coverage is sufficient. In addition, there is no way to interact and direct the fuzzing process. In this paper, we design the dynamic directed greybox fuzzing (DDGF) to facilitate collaboration between the user and fuzzer. By leveraging Ball-Larus path profiling algorithm, we propose two new techniques: dynamic introspection and dynamic direction. Dynamic introspection reveals the significant imbalance in the distribution of path frequency through encoding and decoding. Based on the insight from introspection, users can dynamically direct the fuzzer to focus testing on the selected paths in real time. We implement DDGF based on AFL++. Experiments on Magma show that DDGF is effective in helping the fuzzer to reproduce vulnerabilities faster, with up to 100x speedup and only 13% performance overhead. DDGF shows the great potential of human-in-the-loop for fuzzing.
Haoran Fang, Kaikai Zhang, Donghui Yu, Yuanyuan Zhang 0002
ISSTA2
2024 Palette-based colour normalization for histopathology images
abstract
Abstract With the emergence of computer‐aided diagnostic (CAD) systems, the speed of histopathological image analysis and the accuracy of cancer detection have considerably improved. However, the appearance of histopathological slides can vary depending on the consistency of tissue thickness, stain concentrations, and equipment, thus affecting the judgment of CAD systems. This study proposes a palette‐based colour normalization method for histopathology images is proposed to solve the problem of colour differences between histopathology images. The method builds a graphical user interface based on an improved palette generation algorithm and colour transfer definitions, through which the user can complete the corresponding colour modification to achieve colour normalization. The evaluation of the metrics on histopathological images shows that our method is able to achieve higher image quality and better preservation of structural information of the source image compared with four other colour normalization algorithms. The peak signal‐to‐noise ratio values obtained by the proposed method on two publicly available datasets were 24.1914 and 21.3666, and the structural similarity index matrix values were 0.9871 and 0.9760. The proposed method provides new ideas for the development and design of CAD systems.
Shengzhe Shi, Kaikai Zhang, Sheng Liu 0005
IET Image Process.3
2022 A digital twin-based multidisciplinary collaborative design approach for complex engineering product development
Youde Wu, Linzhen Zhou, Pai Zheng, Yanqing Sun, Kaikai Zhang
Adv. Eng. Informatics5