Eric Xu

dblp:87/3921 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Leveraging Large Language Models for Generating Training Datasets for Text Extraction from Thumbnails
abstract
Obtaining datasets for training AI models can often be an expensive endeavor in the domain of digital forensics. This research leverages large language models (LLMs) to automate the creation of a training dataset aimed at the extraction of text in Word thumbnails. Our method unfolds in three stages: initially, an LLM generates the text for a Word document based on a randomly chosen title. The generated text acts as the training label for a training instance. Subsequently, we apply a predefined Word format template, which organizes the document into sections with specified fonts and sizes. This content is integrated with the template to produce a formatted Word document. From this document, we extract the thumbnail and also capture a high-resolution screenshot. Therefore, each training instance comprises two types of labels (i.e., a high-resolution image and text label of a Word document) and one blurry thumbnail of the Word document. Through this method, we successfully generated 10 datasets. Each dataset contains the same font with 30K training instances. The 30k instances are further divided into three groups in terms of three different thumbnail resolution sizes: medium, large, and extra large. Each group has 10k instances.
Eric Xu, Chimezie Onwuegbuchulem, Sarfraz Shaikh, Lin Deng 0001
COMPSAC1
2025 Enhancing Digital Forensics Evidence Analysis with Large Language Models
abstract
In an era where justice and accountability increasingly depend on digital evidence, Large Language Models (LLMs) offer transformative potential for digital forensics. This three-hour Hands-on tutorial explores how LLMs can automate investigations, reveal hidden insights, and enhance evidence analysis. Through real-world case studies, interactive exercises, and hands-on labs, participants will learn to leverage LLMs for tasks such as entity identification, evidence processing, and knowledge graph reconstruction. Designed for professionals, researchers, and students, this collaborative learning experience equips attendees with practical skills to innovate in digital forensics. As LLMs reshape the field, this tutorial underscores their role in improving justice outcomes, strengthening accountability, and advancing the future of digital investigations.
Eric Xu, Lin Deng 0001
KDD (2)1
2024 Transforming Digital Forensics with Large Language Models: Unlocking Automation, Insights, and Justice
abstract
In the pursuit of justice and accountability in the digital age, the integration of Large Language Models (LLMs) with digital forensics holds immense promise. This half-day tutorial provides a comprehensive exploration of the transformative potential of LLMs in automating digital investigations and uncovering hidden insights. Through a combination of real-world case studies, interactive exercises, and hands-on labs, participants will gain a deep understanding of how to harness LLMs for evidence analysis, entity identification, and knowledge graph reconstruction. By fostering a collaborative learning environment, this tutorial aims to empower professionals, researchers, and students with the skills and knowledge needed to drive innovation in digital forensics. As LLMs continue to revolutionize the field, this tutorial will have far-reaching implications for enhancing justice outcomes, promoting accountability, and shaping the future of digital investigations.
Eric Xu, Wenbin Zhang 0002
CIKM1
2023 A Hands-on Digital Forensic Lab to Investigate Morris Worm Attack
abstract
We have developed a hands-on digital forensic lab to investigate the Morris Worm attack. In the poster, after the attack, we demonstrate a systematic approach to reconstructing the attack scenario by analyzing the worm's running processes, the networking communication used by running processes, and metadata of the files left on victims' machines.
Eric Xu, Alex S. Xu, Danny Ferreira, Lin Deng 0001
SIGCSE (2)1
2023 Derivative Based Nonbacktracking Real-World Regex Matching with Backtracking Semantics
abstract
We develop a new derivative based theory and algorithm for nonbacktracking regex matching that supports anchors and counting, preserves backtracking semantics, and can be extended with lookarounds. The algorithm has been implemented as a new regex backend in .NET and was extensively tested as part of the formal release process of .NET7. We present a formal proof of the correctness of the algorithm, which we believe to be the first of its kind concerning industrial implementations of regex matchers. The paper describes the complete foundation, the matching algorithm, and key aspects of the implementation involving a regex rewrite system, as well as a comprehensive evaluation over industrial case studies and other regex engines.
Dan Moseley, Mario Nishio, Jose Perez Rodriguez, Olli Saarikivi, Stephen Toub, Margus Veanes, Tiki Wan, Eric Xu
Proc. ACM Program. Lang.8
2021 DIFF: a relational interface for large-scale data explanation
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Ananthanarayan, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
VLDB J.5
2019 Symbolic Regex Matcher
abstract
Symbolic regex matcher is a new open source .NET regular expression matching tool and match generator in the Microsoft Automata framework. It is based on the .NET regex parser in combination with a set based representation of character classes. The main feature of the tool is that the core matching algorithms are based on symbolic derivatives that support extended regular expression operations such as intersection and complement and also support a large set of commonly used features such as bounded loop quantifiers. The particularly useful features of the tool are that it supports full UTF16 encoded strings, the match generation is backtracking free, thread safe, and parallelizes with low overhead in multithreaded applications. We discuss the main design decisions behind the tool, explain the core algorithmic ideas and how the tool works, discuss some practical usage scenarios, and compare it to existing state of the art.
Olli Saarikivi, Margus Veanes, Tiki Wan, Eric Xu
TACAS (1)4
2018 DIFF: A Relational Interface for Large-Scale Data Explanation
abstract
A range of explanation engines assist data analysts by performing feature selection over increasingly high-volume and high-dimensional data, grouping and highlighting commonalities among data points. While useful in diverse tasks such as user behavior analytics, operational event processing, and root cause analysis, today's explanation engines are designed as standalone data processing tools that do not interoperate with traditional, SQL-based analytics workflows; this limits the applicability and extensibility of these engines. In response, we propose the DIFF operator, a relational aggregation operator that unifies the core functionality of these engines with declarative relational query processing. We implement both single-node and distributed versions of the DIFF operator in MB SQL, an extension of MacroBase, and demonstrate how DIFF can provide the same semantics as existing explanation engines while capturing a broad set of production use cases in industry, including at Microsoft and Facebook. Additionally, we illustrate how this declarative approach to data explanation enables new logical and physical query optimizations. We evaluate these optimizations on several real-world production applications, and find that DIFF in MB SQL can outperform state-of-the-art engines by up to an order of magnitude.
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Anathanaraya, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
Proc. VLDB Endow.5
2002 A vacation model for the non-saturated Readers and Writers system with a threshold policy
Eric Xu, Attahiru Sule Alfa
Perform. Evaluation1