VLDB 2026 Research / reviewers in the wild / expert
David Tao
dblp:79/6478
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0005-2440-4389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
3 papers |
Systems and software security · 46% Blockchain and cryptocurrency security · 40% Web and mobile security · 14% | |
| Artificial intelligence
2 papers |
Deep learning architectures and training · 50% Language models and text generation · 50% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% |
Topics — the 2 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and mobile security
harmful content detection |
0.3 | 1 | 2025 | Supporting Human Raters with the Detection of Harmful Content Using Large Language Models · SP 2025 |
Web and social media mining › social media analysis
social media measurement |
0.2 | 1 | 2024 | Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates · IMC 2024 |
Methods — techniques the papers use, named apart from their topics
prompt engineering · 2.6large language model · 2.6blockchain analysis · 2.3deep learning · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MAGIKA: AI-Powered Content-Type DetectionabstractThe task of content-type detection—which entails identifying the data encoded in an arbitrary byte sequence—is critical for operating systems, development, reverse engineering environments, and a variety of security applications. In this paper, we introduce Magika, a novel AI-powered content-type detection tool. Under the hood, Magika employs a deep learning model that can execute on a single CPU with just 1MB of memory to store the model's weights. We show that Magika achieves an average F1 score of 99% across over a hundred content types and a test set of more than 1M files, outperforming all existing content-type detection tools today. To foster adoption and improvements, we open source Magika under an Apache 2 license on GitHub and we make our model and training pipeline publicly available. Our tool has already seen adoption by Gmail and Google Drive for attachment scanning, by VirusTotal to aid with malware analysis, and by prominent open-source projects such as Apache Tika. While this paper focuses on the initial version, Magika continues to evolve with support for over 200 content types now available. The latest developments can be found at https://github.com/google/magika. Yanick Fratantonio, Luca Invernizzi, Loua Farah, Kurt Thomas, Marina Zhang, Ange Albertini, Francois Galilee, Giancarlo Metitieri, Julien Cretin, Alex Petit-Bianco, David Tao, Elie Bursztein |
ICSE | 11 |
| 2025 | Supporting Human Raters with the Detection of Harmful Content Using Large Language ModelsabstractIn this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election misinformation. Using a dataset of 50,000 user comments, we demonstrate that LLMs can achieve 90 % accuracy when compared to human verdicts. We explore how to best leverage these capabilities, proposing five design patterns that integrate LLMs with human rating, such as pre-filtering non-violative content, detecting potential errors in human rating, or surfacing critical context to support human rating. We outline how to support all of these design patterns using a single, optimized prompt. Beyond these synthetic experiments, we share how piloting our proposed techniques in a real-world review queue yielded a 41.5% improvement in optimizing available human rater capacity, and a 9–11 % increase (absolute) in precision and recall for detecting violative content. Kurt Thomas, Patrick Gage Kelley, David Tao, Sarah Meiklejohn, Owen Vallis, Shunwen Tan, Blaz Bratanic, Felipe Tiengo Ferreira, Vijay Eranti, Elie Bursztein |
SP | 3 |
| 2024 | Give and Take: An End-To-End Investigation of Giveaway Scam Conversion RatesabstractThe Internet's combination of low communication cost, global reach, and functional anonymity has allowed fraudulent scam volumes to reach new heights. Designing effective interventions requires first understanding the context: how scammers reach potential victims, the earnings they make, and any potential bottlenecks for durable interventions. In this short paper, we focus on these questions in the context of cryptocurrency giveaway scams, where victims are tricked into irreversibly transferring funds to scammers under the pretense of even greater returns. Combining data from Twitter (also known as X), YouTube and Twitch livestreams, landing pages, and cryptocurrency blockchains, we measure how giveaway scams operate at scale. We find that 1 in 1000 scam tweets, and 4 in 100,000 livestream views, net a victim, and that scammers managed to extract nearly $4.62 million from just hundreds of victims during our measurement window. Enze Liu 0001, George Kappos, Eric Mugnier, Luca Invernizzi, Stefan Savage, David Tao, Kurt Thomas, Geoffrey M. Voelker, Sarah Meiklejohn |
IMC | 6 |
| 2020 | Spotlight: Malware Lead Generation at ScaleabstractMalware is one of the key threats to online security today, with applications ranging from phishing mailers to ransomware and trojans. Due to the sheer size and variety of the malware threat, it is impractical to combat it as a whole. Instead, governments and companies have instituted teams dedicated to identifying, prioritizing, and removing specific malware families that directly affect their population or business model. The identification and prioritization of the most disconcerting malware families (known as malware hunting) is a time-consuming activity, accounting for more than 20% of the work hours of a typical threat intelligence researcher, according to our survey. To save this precious resource and amplify the team’s impact on users’ online safety we present Spotlight, a large-scale malware lead-generation framework. Spotlight first sifts through a large malware data set to remove known malware families, based on first and third-party threat intelligence. It then clusters the remaining malware into potentially-undiscovered families, and prioritizes them for further investigation using a score based on their potential business impact. Fabian Kaczmarczyck, Bernhard Grill, Luca Invernizzi, Jennifer Pullman, Cecilia M. Procopiuc, David Tao, Borbala Benko, Elie Bursztein |
ACSAC | 6 |
| 2009 | Challenges in Exchanging Medication Information: Identifying Gaps in Clinical Document Exchange and Terminology Standards
Shobha Phansalkar, George A. Robinson, George Getty, James Shalaby, David Tao, Carol A. Broverman |
AMIA | 5 |