Yongheng Chen

dblp:44/9048 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Towards Generic Database Management System Fuzzing
Yupeng Yang, Yongheng Chen, Jizhou Chen, Wenke Lee
USENIX Security Symposium2
2023 µFUZZ: Redesign of Parallel Fuzzing using Microservice Architecture
Yongheng Chen, Yupeng Yang, Hong Hu 0004, Dinghao Wu, Wenke Lee
USENIX Security Symposium1
2023 Prediction of drug side effects with transductive matrix co-completion
abstract
MOTIVATION: Side effects of drugs could cause severe health problems and the failure of drug development. Drug-target interactions are the basis for side effect production and are important for side effect prediction. However, the information on the known targets of drugs is incomplete. Furthermore, there could be also some missing data in the existing side effect profile of drugs. As a result, new methods are needed to deal with the missing features and missing labels in the problem of side effect prediction. RESULTS: We propose a novel computational method based on transductive matrix co-completion and leverage the low-rank structure in the side effects and drug-target data. Positive-unlabelled learning is incorporated into the model to handle the impact of unobserved data. We also introduce graph regularization to integrate the drug chemical information for side effect prediction. We collect the data on side effects, drug targets, drug-associated proteins and drug chemical structures to train our model and test its performance for side effect prediction. The experiment results show that our method outperforms several other state-of-the-art methods under different scenarios. The case study and additional analysis illustrate that the proposed method could not only predict the side effects of drugs but also could infer the missing targets of drugs. AVAILABILITY AND IMPLEMENTATION: The data and the code for the proposed method are available at https://github.com/LiangXujun/GTMCC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xujun Liang, Lingzhi Qu, Yongheng Chen
Bioinform.5
2021 Identifying Behavior Dispatchers for Malware Analysis
abstract
Malware is a major threat to modern computer systems. Malicious behaviors are hidden by a variety of techniques: code obfuscation, message encoding and encryption, etc. Countermeasures have been developed to thwart these techniques in order to expose malicious behaviors. However, these countermeasures rely heavily on identifying specific API calls, which has significant limitations as these calls can be misleading or hidden from the analyst. In this paper, we show that malicious programs share a key component which we call a behavior dispatcher, a code structure which is intercepted between various condition checks and malicious actions. By identifying these behavior dispatchers, a malware analysis can be guided into behavior dispatchers and activate hidden malicious actions more easily. We propose BDHunter, a system that automatically identifies behavior dispatchers to assist triggering malicious behaviors. BDHunter takes advantage of the observation that a dispatcher compares an input with a set of expected values to determine which malicious behaviors to execute next. We evaluate BDHunter on recent malware samples to identify behavior dispatchers and show that these dispatchers can help trigger more malicious behaviors (otherwise hidden). Our experimental results show that BDHunter identifies 77.4% of dispatchers within the top 20 candidates discovered. Furthermore, BDHunter-guided concolic execution successfully triggers 13.0x and 2.6x more malicious behaviors, compared to unguided symbolic and concolic execution, respectively. These demonstrate that BDHunter effectively identifies behavior dispatchers, which are useful for exposing malicious behaviors.
Kyuhong Park, Burak Sahin, Yongheng Chen, Jisheng Zhao, Evan Downing, Hong Hu 0004, Wenke Lee
AsiaCCS3
2021 One Engine to Fuzz 'em All: Generic Language Processor Testing with Semantic Validation
abstract
Language processors, such as compilers and interpreters, are indispensable in building modern software. Errors in language processors can lead to severe consequences, like incorrect functionalities or even malicious attacks. However, it is not trivial to automatically test language processors to find bugs. Existing testing methods (or fuzzers) either fail to generate high-quality (i.e., semantically correct) test cases, or only support limited programming languages.In this paper, we propose POLYGLOT, a generic fuzzing framework that generates high-quality test cases for exploring processors of different programming languages. To achieve the generic applicability, POLYGLOT neutralizes the difference in syntax and semantics of programming languages with a uniform intermediate representation (IR). To improve the language validity, POLYGLOT performs constrained mutation and semantic validation to preserve syntactic correctness and fix semantic errors. We have applied POLYGLOT on 21 popular language processors of 9 programming languages, and identified 173 new bugs, 113 of which are fixed with 18 CVEs assigned. Our experiments show that POLYGLOT can support a wide range of programming languages, and outperforms existing fuzzers with up to 30× improvement in code coverage.
Yongheng Chen, Hong Hu 0004, Hangfan Zhang, Yupeng Yang, Dinghao Wu, Wenke Lee
SP1
2020 SQUIRREL: Testing Database Management Systems with Language Validity and Coverage Feedback
abstract
Fuzzing is an increasingly popular technique for verifying software functionalities and finding security vulnerabilities. However, current mutation-based fuzzers cannot effectively test database management systems (DBMSs), which strictly check inputs for valid syntax and semantics. Generation-based testing can guarantee the syntax correctness of the inputs, but it does not utilize any feedback, like code coverage, to guide the path exploration.
Yongheng Chen, Hong Hu 0004, Hangfan Zhang, Wenke Lee, Dinghao Wu
CCS2
2018 Multi-modal multi-layered topic classification model for social event analysis
Yongheng Chen, Chunyan Yin, Yaojin Lin, Wanli Zuo
Multim. Tools Appl.1