Siyuan Xia

dblp:292/6093 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Programmable Dataflows: Abstraction and Programming Model for Data Sharing
abstract
Abstract Data sharing is central to various applications such as fraud detection, ad matching, and improving patient care. However, each solution to data sharing is bespoke and cost-intensive, hampering value generation. We identify the lack of abstractions to control data release as the culprit of the problem. For example, it is common to have constraints on whether to share data that depend on the result of sharing, and evaluating these constraints requires sharing in the first place, leading to a standstill. To help people build solutions to a wide variety of data sharing applications, we propose programmable dataflows , which consist of two components. The first component is an abstraction, the contract , which agents use to communicate the intent of a data sharing action and evaluate its consequences before the dataflow takes place. This helps agents control the release of their data. The second component is a contract programming model (CPM), which allows agents to program data sharing applications catered to each problem’s needs with the contract abstraction. We describe how to deploy those applications on a data escrow to ensure data remains protected from unintended data releases. Our evaluation shows 1) the contract abstraction permits representing a wide range of sharing problems, 2) CPM permits writing programs for complex data sharing problems and 3) quantitatively, our improvements to CPM make sharing programs run efficiently.
Siyuan Xia, Chris Zhu, Tapan Srivastava, Bridget Fahey, Raul Castro Fernandez
VLDB J.1
2025 Suna: Scalable Causal Confounder Discovery over Relational Data
abstract
Understanding the causal relationships between treatments and outcomes is fundamental in various areas. Causal inference aims to estimate the effect of one variable on another, and critically relies on access to those variables as well as the key confounders. Unfortunately, data analysts often start with datasets lacking these columns, leading to incorrect estimations. Relational data repositories hold significant potential to augment such datasets with an admissible set of confounders necessary for causal analysis. While recent work has advocated for this potential, these approaches face notable limitations. They either assume the existence of a complete causal diagram over all datasets in the repository, which is impractical; rely on computationally infeasible techniques that do not scale to large data repositories with many features; or can only detect confounders in the absence of causal relations, and are thus ineffective when a causal effect exists. We observe that the asymmetry between causes and effects used in causal discovery can be exploited to directly identify confounders for causal queries. In this paper, we establish a connection between the existence of confounders and the presence of unconfounded ancestors of the treatment variable in the underlying causal diagram—without requiring access to the diagram. This makes it feasible to iteratively discover confounders until an admissible set is constructed. We propose Suna, a highly optimized, GPU-compatible system that implements a novel end-to-end algorithm for discovering confounders within large relational data repositories. Experiments on both real-world and synthetic datasets demonstrate that our system effectively discovers high-quality confounders. Furthermore, Suna employs algorithmic optimizations to accelerate confounder discovery without materializing joins. Our experiments show that Suna finds high-quality confounders while running >100x faster than existing confounder discovery systems.
Siyuan Xia, Daniel Alabi, Eugene Wu 0002
Proc. VLDB Endow.2
2022 Data Station: Delegated, Trustworthy, and Auditable Computation to Enable Data-Sharing Consortia with a Data Escrow
abstract
Pooling and sharing data increases and distributes its value. But since data cannot be revoked once shared, scenarios that require controlled release of data for regulatory, privacy, and legal reasons default to not sharing. Because selectively controlling what data to release is difficult, the few data-sharing consortia that exist are often built around data-sharing agreements resulting from long and tedious one-off negotiations. We introduce Data Station, a data escrow designed to enable the formation of data-sharing consortia. Data owners share data with the escrow knowing it will not be released without their consent. Data users delegate their computation to the escrow. The data escrow relies on delegated computation to execute queries without releasing the data first. Data Station leverages hardware enclaves to generate trust among participants, and exploits the centralization of data and computation to generate an audit log. We evaluate Data Station on machine learning and data-sharing applications while running on an untrusted intermediary. In addition to important qualitative advantages, we show that Data Station: i) outperforms federated learning baselines in accuracy and runtime for the machine learning application; ii) is orders of magnitude faster than alternative secure data-sharing frameworks; and iii) introduces small overhead on the critical path.
Siyuan Xia, Zhiru Zhu, Chris Zhu, Kyle Chard, Aaron J. Elmore, Ian T. Foster, Michael J. Franklin, Sanjay Krishnan, Raul Castro Fernandez
Proc. VLDB Endow.1
2021 CFR-GAN: A Generative Model for Craniofacial Reconstruction
abstract
Craniofacial reconstruction is to reconstruct the face from the skull based on the relationship between the skull and the face to help recognition. This paper proposes a deep generative model for craniofacial reconstruction: CFR-GAN, which avoids the disadvantages of traditional methods of insufficient deep information learning ability of craniofacial data and insufficient ability to express specific features of the dataset. The model is divided into two steps: rough reconstruction and refinement reconstruction. Rough reconstruction rebuilds the overall structural content of the corresponding human head through the skull, and refinement reconstruction restorate facial feature contours. This paper constructs a dataset of 2210 two-dimensional images with craniofacial depth information, which is used to train a CFRGAN model to realize facial reconstruction of skull images. Experiments are conducted from the perspectives of qualitative analysis and quantitative analysis. The results show that CFRGAN generated image retains more identity information, and the similarity between the reconstructed face image and the real face image reaches 94%, which is better than the existing methods. In summary, CFR-GAN proposed in this paper has ability to generate high-fidelity images and is efficient at craniofacial reconstructing.
Pengyue Lin, Wen Yang 0003, Siyuan Xia, Xiaoning Liu 0001, Guohua Geng
BIBM3
2021 KTabulator: Interactive Ad hoc Table Creation using Knowledge Graphs
abstract
The need to find or construct tables arises routinely to accomplish many tasks in everyday life, as a table is a common format for organizing data. However, when relevant data is found on the web, it is often scattered across multiple tables on different web pages, requiring tedious manual searching and copy-pasting to collect data. We propose KTabulator, an interactive system to effectively extract, build, or extend ad hoc tables from large corpora, by leveraging their computerized structures in the form of knowledge graphs. We developed and evaluated KTabulator using Wikipedia and its knowledge graph DBpedia as our testbed. Starting from an entity or an existing table, KTabulator allows users to extend their tables by finding relevant entities, their properties, and other relevant tables, while providing meaningful suggestions and guidance. The results of a user study indicate the usefulness and efficiency of KTabulator in ad hoc table creation.
Siyuan Xia, Nafisa Anzum, Semih Salihoglu, Jian Zhao 0010
CHI1
2021 DPGraph: A Benchmark Platform for Differentially Private Graph Analysis
abstract
Differential privacy has become an appealing choice for analyzing sensitive data while offering strong privacy protection, even for complex data types like graphs. Despite a decade of academic efforts in designing differentially private algorithms for graph analysis, few works have been used in practice. This is due to their complexity in the choice of privacy guarantees and parameter/environmental configurations, or due to their scalability issues for large datasets.
Siyuan Xia, Beizhen Chang, Karl Knopf, Yihan He, Yuchao Tao, Xi He 0001
SIGMOD Conference1
2021 ANINet: a deep neural network for skull ancestry estimation
abstract
BACKGROUND: Ancestry estimation of skulls is under a wide range of applications in forensic science, anthropology, and facial reconstruction. This study aims to avoid defects in traditional skull ancestry estimation methods, such as time-consuming and labor-intensive manual calibration of feature points, and subjective results. RESULTS: This paper uses the skull depth image as input, based on AlexNet, introduces the Wide module and SE-block to improve the network, designs and proposes ANINet, and realizes the ancestry classification. Such a unified model architecture of ANINet overcomes the subjectivity of manually calibrating feature points, of which the accuracy and efficiency are improved. We use depth projection to obtain the local depth image and the global depth image of the skull, take the skull depth image as the object, use global, local, and local + global methods respectively to experiment on the 95 cases of Han skull and 110 cases of Uyghur skull data sets, and perform cross-validation. The experimental results show that the accuracies of the three methods for skull ancestry estimation reached 98.21%, 98.04% and 99.03%, respectively. Compared with the classic networks AlexNet, Vgg-16, GoogLenet, ResNet-50, DenseNet-121, and SqueezeNet, the network proposed in this paper has the advantages of high accuracy and small parameters; compared with state-of-the-art methods, the method in this paper has a higher learning rate and better ability to estimate. CONCLUSIONS: In summary, skull depth images have an excellent performance in estimation, and ANINet is an effective approach for skull ancestry estimation.
Pengyue Lin, Siyuan Xia, Jiang Yi, Wen Yang 0003, Xiaoning Liu 0001, Guohua Geng
BMC Bioinform.2