EDBT 2026 Demo / reviewers in the wild / expert
Yizheng Chen 0001
dblp:09/7487-1
· DBLP profile ↗
16ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-2019-5955ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 14 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vulnerability Detection with Code Language Models: How Far are We?abstractIn the context of the rising interest in code language models (code LMs) and vulnerability detection, we study the effectiveness of code LMs for detecting vulnerabilities. Our analysis reveals significant shortcomings in existing vulnerability datasets, including poor data quality, low label accuracy, and high duplication rates, leading to unreliable model performance in realistic vulnerability detection scenarios. Additionally, the evaluation methods used with these datasets are not representative of real-world vulnerability detection. To address these challenges, we introduce Primevul, a new dataset for training and evaluating code LMs for vulnerability detection. Primevul incorporates a novel set of data labeling techniques that achieve comparable label accuracy to human-verified benchmarks while significantly expanding the dataset. It also implements a rigorous data de-duplication and chronological data splitting strategy to mitigate data leakage issues, alongside introducing more realistic evaluation metrics and settings. This comprehensive approach aims to provide a more accurate assessment of code LMs' performance in real-world conditions. Evaluating code LMs on Primevul reveals that existing benchmarks significantly overestimate the performance of these models. For instance, a state-of-the-art 7B model scored 68.26% Fl on BigVul but only 3.09% Fl on Primevul. Attempts to improve performance through advanced training techniques and larger models like GPT-3.5 and GPT-4 were unsuccessful, with results akin to random guessing in the most stringent settings. These findings underscore the considerable gap between current capabilities and the practical requirements for deploying code LMs in security roles, highlighting the need for more innovative research in this domain. Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin, Basel Alomair, David A. Wagner 0001, Baishakhi Ray, Yizheng Chen 0001 |
ICSE | 9 |
| 2023 | Part-Based Models Improve Adversarial Robustness
Chawin Sitawarin, Kornrapat Pongmala, Yizheng Chen 0001, Nicholas Carlini, David A. Wagner 0001 |
ICLR | 3 |
| 2023 | DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability DetectionabstractWe propose and release a new vulnerable source code dataset. We curate the dataset by crawling security issue websites, extracting vulnerability-fixing commits and source codes from the corresponding projects. Our new dataset contains 18,945 vulnerable functions spanning 150 CWEs and 330,492 non-vulnerable functions extracted from 7,514 commits. Our dataset covers 295 more projects than all previous datasets combined. Yizheng Chen 0001, Zhoujie Ding, Lamya Alowain, David A. Wagner 0001 |
RAID | 1 |
| 2023 | Continuous Learning for Android Malware Detection
Yizheng Chen 0001, Zhoujie Ding, David A. Wagner 0001 |
USENIX Security Symposium | 1 |
| 2021 | Learning Security Classifiers with Verified Global Robustness PropertiesabstractMany recent works have proposed methods to train classifiers with local robustness properties, which can provably eliminate classes of evasion attacks for most inputs, but not all inputs. Since data distribution shift is very common in security applications, e.g., often observed for malware detection, local robustness cannot guarantee that the property holds for unseen inputs at the time of deploying the classifier. Therefore, it is more desirable to enforce global robustness properties that hold for all inputs, which is strictly stronger than local robustness. In this paper, we present a framework and tools for training classifiers that satisfy global robustness properties. We define new notions of global robustness that are more suitable for security classifiers. We design a novel booster-fixer training framework to enforce global robustness properties. We structure our classifier as an ensemble of logic rules and design a new verifier to verify the properties. In our training algorithm, the booster increases the classifier's capacity, and the fixer enforces verified global robustness properties following counterexample guided inductive synthesis. Yizheng Chen 0001, Shiqi Wang 0002, Xiaojing Liao, Suman Jana, David A. Wagner 0001 |
CCS | 1 |
| 2021 | Cost-Aware Robust Tree Ensembles for Security Applications
Yizheng Chen 0001, Shiqi Wang 0002, Weifan Jiang, Asaf Cidon, Suman Jana |
USENIX Security Symposium | 1 |
| 2020 | Neutaint: Efficient Dynamic Taint Analysis with Neural NetworksabstractDynamic taint analysis (DTA) is widely used by various applications to track information flow during runtime execution. Existing DTA techniques use rule-based taint-propagation, which is neither accurate (i.e., high false positive rate) nor efficient (i.e., large runtime overhead). It is hard to specify taint rules for each operation while covering all corner cases correctly. Moreover, the overtaint and undertaint errors can accumulate during the propagation of taint information across multiple operations. Finally, rule-based propagation requires each operation to be inspected before applying the appropriate rules resulting in prohibitive performance overhead on large real-world applications.In this work, we propose Neutaint, a novel end-to-end approach to track information flow using neural program embeddings. The neural program embeddings model the target's programs computations taking place between taint sources and sinks, which automatically learns the information flow by observing a diverse set of execution traces. To perform lightweight and precise information flow analysis, we utilize saliency maps to reason about most influential sources for different sinks. Neutaint constructs two saliency maps, a popular machine learning approach to influence analysis, to summarize both coarse-grained and fine-grained information flow in the neural program embeddings.We compare Neutaint with 3 state-of-the-art dynamic taint analysis tools. The evaluation results show that Neutaint can achieve 68% accuracy, on average, which is 10% improvement while reducing 40× runtime overhead over the second-best taint tool Libdft on 6 real world programs. Neutaint also achieves 61% more edge coverage when used for taint-guided fuzzing indicating the effectiveness of the identified influential bytes. We also evaluate Neutaint's ability to detect real world software attacks. The results show that Neutaint can successfully detect different types of vulnerabilities including buffer/heap/integer overflows, division by zero, etc. Lastly, Neutaint can detect 98.7% of total flows, the highest among all taint analysis tools. Dongdong She, Yizheng Chen 0001, Abhishek Shah, Baishakhi Ray, Suman Jana |
SP | 2 |
| 2020 | On Training Robust PDF Malware Classifiers
Yizheng Chen 0001, Shiqi Wang 0002, Dongdong She, Suman Jana |
USENIX Security Symposium | 1 |
| 2017 | Practical Attacks Against Graph-based ClusteringabstractGraph modeling allows numerous security problems to be tackled in a general way, however, little work has been done to understand their ability to withstand adversarial attacks. We design and evaluate two novel graph attacks against a state-of-the-art network-level, graph-based detection system. Our work highlights areas in adversarial machine learning that have not yet been addressed, specifically: graph-based clustering techniques, and a global feature space where realistic attackers without perfect knowledge must be accounted for (by the defenders) in order to be practical. Even though less informed attackers can evade graph clustering with low cost, we show that some practical defenses are possible. Yizheng Chen 0001, Yacin Nadji, Athanasios Kountouras, Fabian Monrose, Roberto Perdisci, Manos Antonakakis, Nikolaos Vasiloglou |
CCS | 1 |
| 2017 | Hiding in Plain Sight: A Longitudinal Study of Combosquatting AbuseabstractDomain squatting is a common adversarial practice where attackers register domain names that are purposefully similar to popular domains. In this work, we study a specific type of domain squatting called "combosquatting," in which attackers register domains that combine a popular trademark with one or more phrases (e.g., betterfacebook[.]com, youtube-live[.]com). We perform the first large-scale, empirical study of combosquatting by analyzing more than 468 billion DNS records - collected from passive and active DNS data sources over almost six years. We find that almost 60% of abusive combosquatting domains live for more than 1,000 days, and even worse, we observe increased activity associated with combosquatting year over year. Moreover, we show that combosquatting is used to perform a spectrum of different types of abuse including phishing, social engineering, affiliate abuse, trademark abuse, and even advanced persistent threats. Our results suggest that combosquatting is a real problem that requires increased scrutiny by the security community. Panagiotis Kintis, Najmehalsadat Miramirkhani, Charles Lever, Yizheng Chen 0001, Rosa Romero Gómez, Nikolaos Pitropakis, Nick Nikiforakis, Manos Antonakakis |
CCS | 4 |
| 2017 | Measuring Network Reputation in the Ad-Bidding Process
Yizheng Chen 0001, Yacin Nadji, Rosa Romero Gómez, Manos Antonakakis, David Dagon |
DIMVA | 1 |
| 2017 | Measuring lower bounds of the financial abuse to online advertisers: A four year case study of the TDSS/TDL4 Botnet
Yizheng Chen 0001, Panagiotis Kintis, Manos Antonakakis, Yacin Nadji, David Dagon, Michael Farrell |
Comput. Secur. | 1 |
| 2016 | Financial Lower Bounds of Online Advertising Abuse - A Four Year Case Study of the TDSS/TDL4 Botnet
Yizheng Chen 0001, Panagiotis Kintis, Manos Antonakakis, Yacin Nadji, David Dagon, Wenke Lee, Michael Farrell |
DIMVA | 1 |
| 2016 | Enabling Network Security Through Active DNS Datasets
Athanasios Kountouras, Panagiotis Kintis, Charles Lever, Yizheng Chen 0001, Yacin Nadji, David Dagon, Manos Antonakakis, Rodney Joffe |
RAID | 4 |
| 2014 | DNS Noise: Measuring the Pervasiveness of Disposable Domains in Modern DNS TrafficabstractIn this paper, we present an analysis of a new class of domain names: disposable domains. We observe that popular web applications, along with other Internet services, systematically use this new class of domain names. Disposable domains are likely generated automatically, characterized by a "one-time use" pattern, and appear to be used as a way of "signaling" via DNS queries. To shed light on the pervasiveness of disposable domains, we study 24 days of live DNS traffic spanning a year observed at a large Internet Service Provider. We find that disposable domains increased from 23.1% to 27.6% of all queried domains, and from 27.6% to 37.2% of all resolved domains observed daily. While this creative use of DNS may enable new applications, it may also have unanticipated negative consequences on the DNS caching infrastructure, DNSSEC validating resolvers, and passive DNS data collection systems. Yizheng Chen 0001, Manos Antonakakis, Roberto Perdisci, Yacin Nadji, David Dagon, Wenke Lee |
DSN | 1 |
| 2014 | On the Feasibility of Large-Scale Infections of iOS Devices
Tielei Wang, Yeongjin Jang, Yizheng Chen 0001, Simon P. Chung, Billy Lau, Wenke Lee |
USENIX Security Symposium | 3 |