Jie Song 0013

dblp:09/4756-13 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
3since 2021 · last 2022
0000-0002-3433-4522ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 6 first-author · 3 since 2021
YearPublicationVenuePosition
2022 Structured data transformation algebra (SDTA) and its applications
Jie Song 0013, George Alter, H. V. Jagadish
Distributed Parallel Databases1
2021 Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes
abstract
Complex data pipelines are increasingly common in diverse applications such as BI reporting and ML modeling. These pipelines often recur regularly (e.g., daily or weekly), as BI reports need to be refreshed, and ML models need to be retrained. However, it is widely reported that in complex production pipelines, upstream data feeds can change in unexpected ways, causing downstream applications to break silently that are expensive to resolve. Data validation has thus become an important topic, as evidenced by notable recent efforts from Google and Amazon, where the objective is to catch data quality issues early as they arise in the pipelines. Our experience on production data suggests, however, that on string-valued data, these existing approaches yield high false-positive rates and frequently require human intervention. In this work, we develop a corpus-driven approach to auto-validate machine-generated data by inferring suitable data-validation "patterns'' that accurately describe the underlying data-domain, which minimizes false-positives while maximizing data quality issues caught. Evaluations using production data from real data lakes suggest that \sj is substantially more effective than existing methods. Part of this technology ships as an Auto-Tag feature in Microsoft Azure Purview.
Jie Song 0013, Yeye He
SIGMOD Conference1
2021 SDTA: An Algebra for Statistical Data Transformation
abstract
Statistical data manipulation is a crucial component of many data science analytic pipelines, particularly as part of data ingestion. This task is generally accomplished by writing transformation scripts in languages such as SPSS, Stata, SAS, R, Python (Pandas) and etc. The disparate data models, language representations and transformation operations supported by these tools make it hard for end users to understand and document the transformations performed, and for developers to port transformation code across languages.
Jie Song 0013, H. V. Jagadish, George Alter
SSDBM1
2019 C2Metadata: Automating the Capture of Data Transformations from Statistical Scripts in Data Documentation
abstract
Datasets are often derived by manipulating raw data with statistical software packages. The derivation of a dataset must be recorded in terms of both the raw input and the manipulations applied to it. Statistics packages typically provide limited help in documenting provenance for the resulting derived data. At best, the operations performed by the statistical package are described in a script. Disparate representations make these scripts hard to understand for users. To address these challenges, we created Continuous Capture of Metadata (C2Metadata), a system to capture data transformations in scripts for statistical packages and represent it as metadata in a standard format that is easy to understand. We do so by devising a Structured Data Transformation Algebra (SDTA), which uses a small set of algebraic operators to express a large fraction of data manipulation performed in practce. We then implement SDTA, inspired by relational algebra, in a data transformation specification language we call SDTL. In this demonstration, we showcase C2metadata's capture of data transformations from a pool of sample transformation scripts in at least two languages: SPSS and Stata (SAS and R are under development), for social science data in a large academic repository. We will allow the audience to explore C2Metadata using a web-based interface, visualize the intermediate steps and trace the provenance and changes of data at different levels for better understanding of the process.
Jie Song 0013, George Alter, H. V. Jagadish
SIGMOD Conference1
2018 GeoAlign: Interpolating Aggregates over Unaligned Partitions
Jie Song 0013, Danai Koutra, Murali Mani, H. V. Jagadish
EDBT1
2018 GeoFlux: Hands-Off Data Integration Leveraging Join Key Knowledge
abstract
Data integration is frequently required to obtain the full value of data from multiple sources. In spite of extensive research on tools to assist users, data integration remains hard, particularly for users with limited technical proficiency. To address this barrier, we study how much we can do with no user guidance. Our vision is that the user should merely specify two input datasets to be joined and get a meaningful integrated result. It turns out that our vision can be realized if the system can correctly determine the join key, for example based on domain knowledge.
Jie Song 0013, Danai Koutra, Murali Mani, H. V. Jagadish
SIGMOD Conference1