Chang Ge 0002

dblp:122/5424-2 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-8788-4379ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Demonstration of PKGem: Secure Enrichment of Personal Knowledge Graphs
abstract
We present PKGem, a system that provides an end-to-end secure solution to enrich personal knowledge graphs in mobile environments. This task faces two core challenges: First, the proprietary, user-centric, and locally stored nature of personal knowledge graphs makes collaborative enrichment with socially connected peers a privacy concern. Moreover, the mobile environment has strict constraints on resource and computation cost, requiring lightweight and efficient design. PKGem addresses both challenges by leveraging cryptographic techniques to enable secure data enrichment across personal knowledge graphs, while remaining practical under mobile constraints. The system is implemented as an Android application and supports a variety of real-world usage scenarios. The code of PKGem is available at https://github.com/golden-eggs-lab/pkgem, with a demonstration video link included in the repository.
Junzhou Su, Sriram Nutulapati, Chang Ge 0002
CIKM3
2025 Position-Enhanced Gradient Attack (PEGA) on Medical Language Models
abstract
Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.
Nuo Xu 0013, Christopher Stanley, John Gounley, Heidi A. Hanson, Chang Ge 0002, Caiwen Ding
MMAsia5
2024 SoK: Privacy-Preserving Data Synthesis
abstract
As the prevalence of data analysis grows, safeguarding data privacy has become a paramount concern. Consequently, there has been an upsurge in the development of mechanisms aimed at privacy-preserving data analyses. However, these approaches are task-specific; designing algorithms for new tasks is a cumbersome process. As an alternative, one can create synthetic data that is (ideally) devoid of private information. This paper focuses on privacy-preserving data synthesis (PPDS) by providing a comprehensive overview, analysis, and discussion of the field. Specifically, we put forth a master recipe that unifies two prominent strands of research in PPDS: statistical methods and deep learning (DL)-based methods. Under the master recipe, we further dissect the statistical methods into choices of modeling and representation, and investigate the DL-based methods by different generative modeling principles. To consolidate our findings, we provide comprehensive reference tables, distill key takeaways, and identify open problems in the existing literature. In doing so, we aim to answer the following questions: What are the design principles behind different PPDS methods? How can we categorize these methods, and what are the advantages and disadvantages associated with each category? Can we provide guidelines for method selection in different real-world scenarios? We proceed to benchmark several prominent DL-based methods on the task of private image synthesis and conclude that DP-MERF is an all-purpose approach. Finally, upon systematizing the work over the past decade, we identify future directions and call for actions from researchers.
Yuzheng Hu, Fan Wu 0011, Qinbin Li, Yunhui Long, Gonzalo Munilla Garrido, Chang Ge 0002, Bolin Ding, David A. Forsyth, Bo Li 0026, Dawn Song
SP6
2021 Kamino: Constraint-Aware Differentially Private Data Synthesis
abstract
Organizations are increasingly relying on data to support decisions. When data contains private and sensitive information, the data owner often desires to publish a synthetic database instance that is similarly useful as the true data, while ensuring the privacy of individual data records. Existing differentially private data synthesis methods aim to generate useful data based on applications, but they fail in keeping one of the most fundamental data properties of the structured data --- the underlying correlations and dependencies among tuples and attributes (i.e., the structure of the data). This structure is often expressed as integrity and schema constraints, or with a probabilistic generative process. As a result, the synthesized data is not useful for any downstream tasks that require this structure to be preserved. This work presents KAMINO, a data synthesis system to ensure differential privacy and to preserve the structure and correlations present in the original dataset. KAMINO takes as input of a database instance, along with its schema (including integrity constraints), and produces a synthetic database instance with differential privacy and structure preservation guarantees. We empirically show that while preserving the structure of the data, KAMINO achieves comparable and even better usefulness in applications of training classification models and answering marginal queries than the state-of-the-art methods of differentially private data synthesis.
Chang Ge 0002, Shubhankar Mohapatra, Xi He 0001, Ihab F. Ilyas
Proc. VLDB Endow.1
2019 APEx: Accuracy-Aware Differentially Private Data Exploration
abstract
Organizations are increasingly interested in allowing external data scientists to explore their sensitive datasets. Due to the popularity of differential privacy, data owners want the data exploration to ensure provable privacy guarantees. However, current systems for answering queries with differential privacy place an inordinate burden on the data analysts to understand differential privacy, manage their privacy budget, and even implement new algorithms for noisy query answering. Moreover, current systems do not provide any guarantees to the data analyst on the quality they care about, namely accuracy of query answers. We present APEx, a novel system that allows data analysts to pose adaptively chosen sequences of queries along with required accuracy bounds. By translating queries and accuracy bounds into differentially private algorithms with the least privacy loss, APEx returns query answers to the data analyst that meet the accuracy bounds, and proves to the data owner that the entire data exploration process is differentially private. Our comprehensive experimental study on real datasets demonstrates that APEx can answer a variety of queries accurately with moderate to small privacy loss, and can support data exploration for entity resolution with high accuracy under reasonable privacy settings.
Chang Ge 0002, Xi He 0001, Ihab F. Ilyas, Ashwin Machanavajjhala
SIGMOD Conference1
2019 Speculative Distributed CSV Data Parsing for Big Data Analytics
abstract
There has been a recent flurry of interest in providing query capability on raw data in today's big data systems. These raw data must be parsed before processing or use in analytics. Thus, a fundamental challenge in distributed big data systems is that of efficient parallel parsing of raw data. The difficulties come from the inherent ambiguity while independently parsing chunks of raw data without knowing the context of these chunks. Specifically, it can be difficult to find the beginnings and ends of fields and records in these chunks of raw data. To parallelize parsing, this paper proposes a speculation-based approach for the CSV format, arguably the most commonly used raw data format. Due to the syntactic and statistical properties of the format, speculative parsing rarely fails and therefore parsing is efficiently parallelized in a distributed setting. Our speculative approach is also robust, meaning that it can reliably detect syntax errors in CSV data. We experimentally evaluate the speculative, distributed parsing approach in Apache Spark using more than 11,000 real-world datasets, and show that our parser produces significant performance benefits over existing methods.
Chang Ge 0002, Yinan Li 0009, Eric Eilebrecht, Badrish Chandramouli, Donald Kossmann
SIGMOD Conference1
2019 Secure Multi-Party Functional Dependency Discovery
abstract
Data profiling is an important task to understand data semantics and is an essential pre-processing step in many tools. Due to privacy constraints, data is often partitioned into silos, with different access control. Discovering functional dependencies (FDs) usually requires access to all data partitions to find constraints that hold on the whole dataset. Simply applying general secure multi-party computation protocols incurs high computation and communication cost. This paper formulates the FD discovery problem in the secure multi-party scenario. We propose secure constructions for validating candidate FDs, and present efficient cryptographic protocols to discover FDs over distributed partitions. Experimental results show that solution is practically efficient over non-secure distributed FD discovery, and can significantly outperform general purpose multi-party computation frameworks. To the best of our knowledge, our work is the first one to tackle this problem.
Chang Ge 0002, Ihab F. Ilyas, Florian Kerschbaum
Proc. VLDB Endow.1
2016 Towards a Hybrid Design for Fast Query Processing in DB2 with BLU Acceleration Using Graphical Processing Units: A Technology Demonstration
abstract
In this paper, we show how we use Nvidia GPUs and host CPU cores for faster query processing in a DB2 database using BLU Acceleration (DB2's column store technology). Moreover, we show the benefits and problems of using hardware accelerators (more specifically GPUs) in a real commercial Relational Database Management System(RDBMS).We investigate the effect of off-loading specific database operations to a GPU, and show how doing so results in a significant performance improvement. We then demonstrate that for some queries, using just CPU to perform the entire operation is more beneficial. While we use some of Nvidia's fast kernels for operations like sort, we have also developed our own high performance kernels for operations such as group by and aggregation. Finally, we show how we use a dynamic design that can make use of optimizer metadata to intelligently choose a GPU kernel to run. For the first time in the literature, we use benchmarks representative of customer environments to gauge the performance of our prototype, the results of which show that we can get a speed increase upwards of 2x, using a realistic set of queries.
Sina Meraji, Berni Schiefer, Lan Pham, Lee Chu, Peter Kokosielis, Adam J. Storm, Wayne Young, Chang Ge 0002, Geoffrey Ng, Kajan Kanagaratnam
SIGMOD Conference8
2015 Bi-temporal Timeline Index: A data structure for Processing Queries on bi-temporal data
abstract
Following the adoption of basic temporal features in the SQL:2011 standard, there has been a tremendous interest within the database industry in supporting bi-temporal features, as a significant number of real-life workloads would greatly benefit from efficient temporal operations. However, current implementations of bi-temporal storage systems and operators are far from optimal. In this paper, we present the Bi-temporal Timeline Index, which supports a broad range of temporal operators and exploits the special properties of an in-memory column store database system. Comprehensive performance experiments with the TPC-BiH benchmark show that algorithms based on the Bi-temporal Timeline Index outperform significantly both existing commercial database systems and state-of-the-art data structures from research.
Martin Kaufmann, Peter M. Fischer 0001, Norman May, Chang Ge 0002, Anil K. Goel, Donald Kossmann
ICDE4
2015 Indexing bi-temporal windows
abstract
Bi-temporal databases support system (transaction) and application time, enabling users to query the history as recorded today and as it was known in the past. In this paper, we study windows over both system and application time, i.e., bi-temporal windows. We propose a two-dimensional index that supports one-time and continuous queries over fixed and sliding bi-temporal windows, covering static and streaming data. We demonstrate the advantages of the proposed index compared to the state-of-the-art in terms of query performance, index update overhead and space footprint.
Chang Ge 0002, Martin Kaufmann, Lukasz Golab, Peter M. Fischer 0001, Anil K. Goel
SSDBM1
2013 Lazy data structure maintenance for main-memory analytics over sliding windows
abstract
We address the problem of maintaining data structures used by memory-resident data warehouses that store sliding windows. We propose a framework that eagerly expires data from the sliding window to save space and/or satisfy data retention policies, but lazily maintains the associated data structures to reduce maintenance overhead. Using a dictionary as an example, we show that our framework enables maintenance algorithms that outperform existing approaches in terms of space overhead, maintenance overhead, and dictionary lookup overhead during query execution.
Chang Ge 0002, Lukasz Golab
DOLAP1