VLDB 2026 Research / reviewers in the wild / expert
Tianwa Chen
dblp:225/7403
· DBLP profile ↗
6ranked-venue papers in the field
2as first author
4since 2021 · last 2025
0000-0002-5135-0313ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 2Information Retrieval & Web Search · 2Business Process & Enterprise Data · 2 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Do Experts Make Sense of Integrated Process Models?
Tianwa Chen, Barbara Weber, Graeme G. Shanks, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
CAiSE (2) | 1 |
| 2023 | DataOps-4G: On Supporting Generalists in Data Quality DiscoveryabstractData preparation has become a necessary but labor and resource intensive step to perform data analytics. To date, such activities still require considerable manual effort from experts. In this paper, we focus on a specific data preparation activity, namely data quality discovery. We explore different settings in which data workers undertake data quality discovery tasks and the implications of those settings for the efficiency and effectiveness of data workers. To this end, we propose DataOps-4G, a data curation platform for generalists, that allows users to interact with data without the need to write code. We wrap up pre-defined code snippets that implement useful functionalities to explore data quality and bundle the code into so-called DataOps. Then, we conduct a lab-based user study to evaluate our DataOps-4G platform from two perspectives: (i) effectiveness, the accuracy of the outcomes achieved by participants; and (ii) efficiency, their effort and strategies in task completion. Our experimental results uncover how effectiveness and efficiency can be affected by their task completion patterns and strategies. This opens up the possibility of popularizing data curation processes by employing non-experts (e.g., from crowdsourcing platforms) and consequently allowing experts to focus on more complex activities (e.g., building machine learning models). Shaochen Yu, Tianwa Chen, Lei Han 0003, Gianluca Demartini, Shazia Sadiq |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | A Data-Driven Analysis of Behaviors in Data Curation ProcessesabstractUnderstanding how data workers interact with data, and various pieces of information related to data preparation, is key to designing systems that can better support them in exploring datasets. To date, however, there is a paucity of research studying the strategies adopted by data workers as they carry out data preparation activities. In this work, we investigate a specific data preparation activity, namely data quality discovery , and aim to (i) understand the behaviors of data workers in discovering data quality issues, (ii) explore what factors (e.g., prior experience) can affect their behaviors, as well as (iii) understand how these behavioral observations relate to their performance. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The experiment results reveal that: (i) ‘copy–paste–modify’ is a typical strategy for writing code to complete tasks; (ii) proficiency in writing code has a significant impact on the quality of task performance, while perceived difficulty and efficacy can influence task completion patterns; and (iii) searching in external resources is a prevalent action that can be leveraged to achieve better performance. Furthermore, our experiment indicates that providing sample code within the system can help data workers get started with their task, and surfacing underlying data is an effective way to support exploration. By investigating data worker behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Based on our experiment results, we showcase a systematic approach to select from the top best code snippets created by data workers and assemble them to achieve better performance than the best individual performer in the dataset. By doing so, our findings not only provide insights into patterns of interactions with various system components and information resources when performing data curation tasks, but also build effective and efficient data curation processes through data workers’ collective intelligence. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
ACM Trans. Inf. Syst. | 2 |
| 2022 | Business process and rule integration approaches - An empirical analysis of model understanding
Wei Wang 0186, Tianwa Chen, Marta Indulska, Shazia Sadiq, Barbara Weber |
Inf. Syst. | 2 |
| 2020 | Sensemaking in Dual Artefact Tasks - The Case of Business Process Models and Business Rules
Tianwa Chen, Shazia Sadiq, Marta Indulska |
ER | 1 |
| 2020 | On Understanding Data Worker Interaction BehaviorsabstractUnderstanding how data workers interact with data and various pieces of information (e.g., code snippet examples) is key to design systems that can better support them in exploring a given dataset. To date, however, there is a paucity of research studying information seeking patterns and the strategies adopted by data workers as they carry out data curation activities. In this work, we aim at understanding the behaviors of data workers in discovering data quality issues, and how these behavioral observations relate to their performance. Specifically, we investigate how data workers use information resources and tools to support their task completion. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The collected data reveals that: (i) searching in external resources is a prevalent action that can be leveraged to achieve better performance; (ii) 'copy-paste-modify' is a typical strategy for writing code to complete tasks; (iii) providing sample code within the system could help data workers to get started with their task; and (iv) surfacing underlying data is an effective way to support exploration. By investigating the behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Our findings provide insights into patterns of interactions with various system components and information resources to perform data curation tasks. This bears implications on the design of domain-specific IR systems for data workers like code-base search. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
SIGIR | 2 |