VLDB 2026 Research / reviewers in the wild / expert
Tongyu Liu
dblp:227/9554
· DBLP profile ↗
9ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing GNN-Based Cloth Simulation with Macro-spatial and Local Dynamic Priors
Tao Peng 0006, Chuang Yin, Li Li 0094, Junping Liu, Tongyu Liu |
ICIC (6) | 7 |
| 2026 | Fusion of microstructural images and constituent properties for elastic property prediction in unidirectional composites: a hybrid ResNet34-MLP approach
Tao Peng 0006, Tongyu Liu, Junping Liu, Xinrong Hu, Li Li 0094 |
Neural Comput. Appl. | 2 |
| 2024 | Controllable Tabular Data Synthesis Using Diffusion ModelsabstractControllable tabular data synthesis plays a crucial role in numerous applications by allowing users to generate synthetic data with specific conditions. These conditions can include synthesizing tuples with predefined attribute values or creating tuples that exhibit a particular correlation with an external table. However, existing approaches lack the flexibility to support new conditions and can be time-consuming when dealing with multiple conditions. To overcome these limitations, we propose a novel approach that leverages diffusion models to first learn an unconditional generative model. Subsequently, we introduce lightweight controllers to guide the unconditional generative model in generating synthetic data that satisfies different conditions. The primary research challenge lies in effectively supporting controllability using lightweight solutions while ensuring the realism of the synthetic data. To address this challenge, we design an unconditional diffusion model tailored specifically for tabular data. Additionally, we propose a new sampling method that enables correlation-aware controls throughout the data generation process. We conducted extensive experiments across various applications for controllable tabular data synthesis, which show that our approach outperforms the state-of-the-art methods. Tongyu Liu, Ju Fan, Nan Tang 0001, Guoliang Li 0001, Xiaoyong Du 0001 |
Proc. ACM Manag. Data | 1 |
| 2024 | Tabular data synthesis with generative adversarial networks: design space and optimizations
Tongyu Liu, Ju Fan, Guoliang Li 0001, Nan Tang 0001, Xiaoyong Du 0001 |
VLDB J. | 1 |
| 2021 | Adaptive Data Augmentation for Supervised Learning over Missing DataabstractReal-world data is dirty, which causes serious problems in (supervised) machine learning (ML). The widely used practice in such scenario is to first repair the labeled source (a.k.a. train) data using rule-, statistical- or ML-based methods and then use the "repaired" source to train an ML model. During production, unlabeled target (a.k.a. test) data will also be repaired, and is then fed in the trained ML model for prediction. However, this process often causes a performance degradation when the source and target datasets are dirty with different noise patterns , which is common in practice. In this paper, we propose an adaptive data augmentation approach, for handling missing data in supervised ML. The approach extracts noise patterns from target data, and adapts the source data with the extracted target noise patterns while still preserving supervision signals in the source. Then, it patches the ML model by retraining it on the adapted data, in order to better serve the target. To effectively support adaptive data augmentation, we propose a novel generative adversarial network (GAN) based framework, called DAGAN, which works in an unsupervised fashion. DAGAN consists of two connected GAN networks. The first GAN learns the noise pattern from the target, for target mask generation. The second GAN uses the learned target mask to augment the source data, for source data adaptation. The augmented source data is used to retrain the ML model. Extensive experiments show that our method significantly improves the ML model performance and is more robust than the state-of-the-art missing data imputation solutions for handling datasets with different missing value patterns. Tongyu Liu, Ju Fan, Yinqing Luo, Nan Tang 0001, Guoliang Li 0001, Xiaoyong Du 0001 |
Proc. VLDB Endow. | 1 |
| 2020 | Relational Data Synthesis using Generative Adversarial Networks: A Design Space Exploration
Ju Fan, Tongyu Liu, Guoliang Li 0001, Junyou Chen, Yuwei Shen, Xiaoyong Du 0001 |
Proc. VLDB Endow. | 2 |
| 2020 | A game-based framework for crowdsourced data labeling
Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Tongyu Liu, Xiaoyong Du 0001 |
VLDB J. | 5 |
| 2019 | CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data LabelingabstractLarge-scale data labeling has become a major bottleneck for many applications, such as machine learning and data integration. This paper presents CrowdGame, a crowdsourcing system that harnesses the crowd to gather data labels in a cost-effective way. CrowdGame focuses on generating high-quality labeling rules to largely reduce the labeling cost while preserving quality. It first generates candidate rules, and then devises a game-based crowdsourcing approach to select rules with high coverage and accuracy. CrowdGame applies the generated rules for effective data labeling. We have implemented CrowdGame and provided a user-friendly interface for users to deploy their labeling applications. We will demonstrate CrowdGame in two representative data labeling scenarios, entity matching and relation extraction. Tongyu Liu, Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Xiaoyong Du 0001 |
SIGMOD Conference | 1 |
| 2018 | Cost-Effective Data Annotation using Game-Based CrowdsourcingabstractLarge-scale data annotation is indispensable for many applications, such as machine learning and data integration. However, existing annotation solutions either incur expensive cost for large datasets or produce noisy results. This paper introduces a cost-effective annotation approach, and focuses on the labeling rule generation problem that aims to generate high-quality rules to largely reduce the labeling cost while preserving quality. To address the problem, we first generate candidate rules, and then devise a game-based crowdsourcing approach C ROWD G AME to select high-quality rules by considering coverage and precision. C ROWD G AME employs two groups of crowd workers: one group answers rule validation tasks (whether a rule is valid) to play a role of rule generator, while the other group answers tuple checking tasks (whether the annotated label of a data tuple is correct) to play a role of rule refuter. We let the two groups play a two-player game: rule generator identifies high-quality rules with large coverage and precision, while rule refuter tries to refute its opponent rule generator by checking some tuples that provide enough evidence to reject rules covering the tuples. This paper studies the challenges in C ROWD G AME . The first is to balance the trade-off between coverage and precision. We define the loss of a rule by considering the two factors. The second is rule precision estimation. We utilize Bayesian estimation to combine both rule validation and tuple checking tasks. The third is to select crowdsourcing tasks to fulfill the game-based framework for minimizing the loss. We introduce a minimax strategy and develop efficient task selection algorithms. We conduct experiments on entity matching and relation extraction, and the results show that our method outperforms state-of-the-art solutions. Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Tongyu Liu, Xiaoyong Du 0001 |
Proc. VLDB Endow. | 5 |