VLDB 2026 Research / reviewers in the wild / expert
Yihao Hu 0001
dblp:234/7986-1
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0003-3048-3867ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LakeVisage: Towards Scalable, Flexible and Interactive Visualization Recommendation for Data Discovery over Data LakesabstractData discovery from data lakes is an essential application in modern data science. While many previous studies focused on improving the efficiency and effectiveness of data discovery, little attention has been paid to the usability of such applications. In particular, exploring data discovery results can be cumbersome due to the cognitive load involved in understanding raw tabular results and identifying insights to draw conclusions. To address this challenge, we introduce a new problem: visualization recommendation for data discovery over data lakes, which aims to automatically identify visualizations that highlight relevant or desired trends in the results returned by data discovery engines. We propose LakeVisage, an end-to-end framework as the first solution to this problem. Given a data lake, a data discovery engine, and a user-specified query table, LakeVisage intelligently explores the space of visualizations and recommends the most useful and "interesting" visualization plans. To this end, we developed (i) approaches to smartly construct the candidate visualization plans from the results of the data discovery engine and (ii) effective pruning strategies to filter out less interesting plans so as to accelerate the visual analysis. Experimental results on real data lakes demonstrate that our proposed techniques can achieve an order-of-magnitude speedup in visualization recommendation. We also conduct a comprehensive user study to demonstrate that LakeVisage offers convenience to users in real data analysis applications by enabling them seamlessly get started with the tasks and performing explorations flexibly. Yihao Hu 0001, Jin Wang 0007, Sajjadur Rahman |
Proc. VLDB Endow. | 1 |
| 2024 | How Database Theory Helps Teach Relational Queries in Database Education (Invited Talk)
Sudeepa Roy 0001, Amir Gilad, Yihao Hu 0001, Hanze Meng, Zhengjie Miao, Kristin Stephens-Martinez, Jun Yang 0001 |
ICDT | 3 |
| 2024 | Qr-Hint: Actionable Hints Towards Correcting Wrong SQL QueriesabstractWe describe a system called Qr-Hint that, given a (correct) target query Q* and a (wrong) working query Q, both expressed in SQL, provides actionable hints for the user to fix the working query so that it becomes semantically equivalent to the target. It is particularly useful in an educational setting, where novices can receive help from Qr-Hint without requiring extensive personal tutoring. Since there are many different ways to write a correct query, we do not want to base our hints completely on how Q* is written; instead, starting with the user's own working query, Qr-Hint purposefully guides the user through a sequence of steps that provably lead to a correct query, which will be equivalent to Q* but may still "look" quite different from it. Ideally, we would like Qr-Hint's hints to lead to the "smallest" possible corrections to Q. However, optimality is not always achievable in this case due to some foundational hurdles such as the undecidability of SQL query equivalence and the complexity of logic minimization. Nonetheless, by carefully decomposing and formulating the problems and developing principled solutions, we are able to provide provably correct and locally optimal hints through Qr-Hint. We show the effectiveness of Qr-Hint through quality and performance experiments as well as a user study in an educational setting. Yihao Hu 0001, Amir Gilad, Kristin Stephens-Martinez, Sudeepa Roy 0001, Jun Yang 0001 |
Proc. ACM Manag. Data | 1 |
| 2022 | I-Rex: An Interactive Relational Query Debugger for SQLabstractDespite the enduring popularity of SQL (Structured Query Language), it is challenging to learn and debug, even for people with considerable programming experience. There is also a lack of SQL tools with advanced debugger features like breakpoints, stepped execution, and variable watching. We present I-Rex, an interactive SQL debugger that enables users to trace the evaluation of a query by its constituent blocks, visualizing how each block computes results from its inputs, and exploring the dependencies among these blocks. I-Rex can be integrated into an autograder, which typically works by comparing the results of submitted queries against reference queries over test database instances. Instead of showing full test instances, which often overwhelm students, I-Rex automatically generates small, illustrative instances for debugging. In this demo, we show how I-Rex helps a student trace complex SQL query execution, learn the semantics of various query constructs, and understand why a query produces (or does not produce) certain results. We also show how a teacher can customize I-Rex for a set of SQL exercises over a database. Overall, we demonstrate how I-Rex supports SQL learning and debugging, thereby increasing students' self-reliance and reducing the burden on the teaching staff. Yihao Hu 0001, Zhengjie Miao, Zhiming Leong, Haechan Lim, Zachary Zheng, Sudeepa Roy 0001, Kristin Stephens-Martinez, Jun Yang 0001 |
SIGCSE (2) | 1 |
| 2020 | Feature-based Distant Domain Transfer LearningabstractIn this paper, we study a not well-investigated but important transfer learning problem termed Distant Domain Transfer Learning (DDTL). This topic is closely related to negative transfer. Unlike conventional transfer learning problems which assume that the source domain and the target domain are more or less similar to each other, DDTL aims to make efficient transfers even when the domains or the tasks are completely different. As an extreme example in image classification, there are only a sufficient amount of unlabeled images of watches, airplanes, and horses in the source domain, and the target domain only has a small set of labeled human face images. Previously, a few instance-based distant domain transfer algorithms were proposed to deal with this type of binary distant domain image classification problems. Yet most existing algorithms are very task-specific and they are only good at binary classification tasks. In this study, we propose a novel feature-based distant domain transfer learning algorithm, which requires only a tiny set of labeled target data and unlabeled source data from completely different domains. Instead of selecting intermediate instances, we introduced Distant Feature Fusion (DFF), a novel feature selection method, to discover general features cross distant domains and tasks by using convolutional autoencoder with a domain distance measurement as a feature extractor. As the novelty of this study, it can effectively handle both distant domain mutil-class image classification and binary image classification problems. More importantly, it has achieved up to 19% higher classification accuracy than "non-transfer" algorithms, and up to 9% higher than existing distant transfer algorithms. Shuteng Niu, Yihao Hu 0001, Jian Wang 0061, Yongxin Liu 0001, Houbing Song |
IEEE BigData | 2 |
| 2020 | MuSe: Multiple Deletion Semantics for Data RepairabstractWe propose to demonstrate MuSe, a system for Database repairs where constraints are expressed as Declarative Rules and can be interpreted in different ways by using four different semantics. Our framework may capture common, cross-relation, repair semantics such as that of SQL deletion triggers, causal rules, and denial constraints. Our demonstration will show the usefulness of the system in easing specification of database repair policies, for different use cases. Amir Gilad, Yihao Hu 0001, Daniel Deutch, Sudeepa Roy 0001 |
Proc. VLDB Endow. | 2 |