Lulu Xie

dblp:133/3878 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 CactusDB: Unlock Co-Optimization Opportunities for SQL Queries and AI/ML Model Inferences
Lixi Zhou, Kanchan Chowdhury, Lulu Xie, Jaykumar Tandel, Xinwei Fu, Jia Zou 0001
ICDE3
2025 ExBoost: Out-of-Box Co-optimization of Machine Learning and Join Queries
Kanchan Chowdhury, Lulu Xie, Lixi Zhou, Jia Zou 0001
DASFAA (1)2
2025 Privacy and Accuracy-Aware AI/ML Model Deduplication
abstract
With the growing adoption of privacy-preserving machine learning algorithms, such as Differentially Private Stochastic Gradient Descent (DP-SGD), training or fine-tuning models on private datasets has become increasingly prevalent. This shift has led to the need for models offering varying privacy guarantees and utility levels to satisfy diverse user requirements. Managing numerous versions of large models introduces significant operational challenges, including increased inference latency, higher resource consumption, and elevated costs. Model deduplication is a technique widely used by many model serving and database systems to support high-performance and low-cost inference queries and model diagnosis queries. However, none of the existing model deduplication works has considered privacy, leading to unbounded aggregation of privacy costs for certain deduplicated models and inefficiencies when applied to deduplicate DP-trained models. We formalize the problem of deduplicating DP-trained models for the first time and propose a novel privacy- and accuracy-aware deduplication mechanism to address the problem. We developed a greedy strategy to select and assign base models to target models to minimize storage and privacy costs. When deduplicating a target model, we dynamically schedule accuracy validations and apply the Sparse Vector Technique to reduce the privacy costs associated with private validation data. Compared to baselines, our approach improved the compression ratio by up to 35× for individual models (including large language models and vision transformers). We also observed up to 43× inference speedup due to the reduction of I/O operations.
Lei Yu 0002, Lixi Zhou, Li Xiong 0001, Kanchan Chowdhury, Lulu Xie, Xusheng Xiao, Jia Zou 0001
Proc. ACM Manag. Data6
2024 IDNet: A Novel Identity Document Dataset via Few-Shot and Quality-Driven Synthetic Data Generation
abstract
Effective fraud detection and analysis of government-issued identity documents, such as passports, driver’s licenses, and identity cards, are essential in thwarting identity theft and bolstering security on online platforms. The accuracy of training fraud detection and analysis tools depends on the availability of extensive and diverse identity document datasets. However, current publicly available benchmark datasets for identity document analysis, including MIDV-500, MIDV-2020, and FMIDV, fall short in several aspects: they offer a limited number of samples of ten European country document types, cover insufficient varieties of fraud patterns, and seldom include alterations in critical personal identifying fields such as portrait images, limiting their utility in training models capable of detecting realistic frauds while preserving privacy. In response to these shortcomings, our research introduces a new benchmark dataset, IDNet, designed to advance privacy-preserving fraud detection efforts, synthesized by integrating the generative models and a Bayesian optimization approach. The IDNet dataset comprises 837, 060 images of synthetically generated identity documents, totaling approximately 490 gigabytes, categorized into 20 types from 10 U.S. states and 10 European countries, which is the largest identity document dataset publicly available today. We evaluated the fidelity and utility of IDNet to demonstrate the effectiveness of our unique synthetic data generation method. We also presented two use cases of the dataset, illustrating how it can aid in training privacy-preserving fraud detection methods, and facilitating the generation of camera and video capturing of identity documents.
Lulu Xie, Yancheng Wang 0001, Soham Nag, Rajeev Goel, Niranjan Erappa Narayana Swamy, Yingzhen Yang, Chaowei Xiao, Jonathan Prisby, Ross Maciejewski, Jia Zou 0001
IEEE Big Data1
2023 Integrative Drug Discovery Platform: A Modular Approach for Efficient and Automated Virtual Screening
abstract
This paper presents a drug development platform based on virtual screening technology. The platform integrates key components such as pocket prediction, molecular docking, molecular dynamics simulation, and ADMET evaluation to achieve an efficient and automated drug virtual screening process. The platform utilizes Docker for modular encapsulation, ensuring environment isolation and convenient deployment. It also provides standardized input-output formats and a task allocation system, enabling users to quickly deploy and customize the workflow. Experimental results demonstrate the effectiveness of the platform in identifying real drugs and evaluating virtual screening results, providing an efficient and reliable solution for drug development. The platform features easy deployment and migration, independent module execution, automated workflow implementation, personalized customization and replacement, task allocation for computationally intensive steps, and complex operations in molecular dynamics simulation.
Lulu Xie, Zhonghai Zhang, Bo Duan, Gang Niu 0008, Shiwei Sun, Fa Zhang 0001, Runting Zhang, Guangming Tan
BIBM2
2013 Towards accurate acoustic localization on a smartphone
abstract
Since our daily activities are dominantly indoor, as smart phones emerge as the most popular personal computing companions, major IT companies recently launched aggressive investment on mobile indoor location services and positioning systems, e.g., on iOS or Android mobile devices. However, one major hurdle has not been conquered yet: smart phone-based high-resolution indoor localization. In this paper, we propose a practical solution for accurate ranging and localization based on acoustic communication between anchor nodes with speakers and the microphone on a smartphone. To identify different anchor nodes and enable time-of-arrival (TOA) ranging, we propose approaches for signal modulation, symbol detection and demodulation, synchronization and ranging. Experimental results show that the communication bit-error-rate and ranging accuracy is sufficient for our target applications. The preliminary results of localization demonstrate that our algorithm could achieve highaccuracy of 23cm in the offline mode with a promising potential for realtime smartphone-based indoor localization.
Xinxin Liu 0006, Lulu Xie, Xiaolin Li 0001
INFOCOM3