EDBT 2026 Demo / reviewers in the wild / expert
Hang Cui 0001
dblp:93/2906-1
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-0987-3743ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly-supervised entity matching via LLM-guided data augmentation and knowledge transfer
Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
Knowl. Based Syst. | 6 |
| 2025 | GARF+: self-supervised and interpretable data cleaning with sequence generative adversarial networks
Jinfeng Peng, Hanghai Cui, Derong Shen, Nan Tang 0001, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
VLDB J. | 7 |
| 2024 | Node Generation for Node Classification in Sparsely-Labeled Graphs
Hang Cui 0001, Tarek F. Abdelzaher |
ASONAM (1) | 1 |
| 2024 | Enhancing Deep Entity Resolution with Integrated Blocker-Matcher Training: Balancing Consensus and DiscrepancyabstractDeep entity resolution (ER) identifies matching entities across data sources using techniques based on deep learning. It involves two steps: a blocker for identifying the potential matches to generate the candidate pairs, and a matcher for accurately distinguishing the matches and non-matches among these candidate pairs. Recent deep ER approaches utilize pretrained language models (PLMs) to extract similarity features for blocking and matching, achieving state-of-the-art performance. However, they often fail to balance the consensus and discrepancy between the blocker and matcher, emphasizing the consensus while neglecting the discrepancy. This paper proposes MutualER, a deep entity resolution framework that integrates and jointly trains the blocker and matcher, balancing both the consensus and discrepancy between them. Specifically, we firstly introduce a lightweight PLM in siamese structure for the blocker and a heavier PLM in cross structure or an autoregressive large language model (LLM) for the matcher. Two optimization techniques named Mutual Sample Selection (MSS) and Similarity Knowledge Transferring (SKT) are designed to jointly train the blocker and matcher. MSS enables the blocker and matcher to mutually select the customized training samples for each other to maintain the discrepancy, while SKT allows them to share the similarity knowledge for improving their blocking and matching capabilities respectively to maintain the consensus. Extensive experiments on five datasets demonstrate that MutualER significantly outperforms existing PLM-based and LLM-based approaches, achieving leading performance in both effectiveness and efficiency. Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
CIKM | 7 |
| 2024 | Unsupervised Node Clustering via Contrastive Hard Sampling
Hang Cui 0001, Tarek F. Abdelzaher |
DASFAA (6) | 1 |
| 2023 | Soft Target-Enhanced Matching Framework for Deep Entity MatchingabstractDeep Entity Matching (EM) is one of the core research topics in data integration. Typical existing works construct EM models by training deep neural networks (DNNs) based on the training samples with onehot labels. However, these sharp supervision signals of onehot labels harm the generalization of EM models, causing them to overfit the training samples and perform badly in unseen datasets. To solve this problem, we first propose that the challenge of training a well-generalized EM model lies in achieving the compromise between fitting the training samples and imposing regularization, i.e., the bias-variance tradeoff. Then, we propose a novel Soft Target-EnhAnced Matching (Steam) framework, which exploits the automatically generated soft targets as label-wise regularizers to constrain the model training. Specifically, Steam regards the EM model trained in previous iteration as a virtual teacher and takes its softened output as the extra regularizer to train the EM model in the current iteration. As such, Steam effectively calibrates the obtained EM model, achieving the bias-variance tradeoff without any additional computational cost. We conduct extensive experiments over open datasets and the results show that our proposed Steam outperforms the state-of-the-art EM approaches in terms of effectiveness and label efficiency. Wenzhou Dou, Derong Shen, Xiangmin Zhou, Tiezheng Nie, Yue Kou, Hang Cui 0001, Ge Yu 0001 |
AAAI | 6 |
| 2022 | Empowering Transformer with Hybrid Matching Knowledge for Entity Matching
Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou, Chenchen Sun, Hang Cui 0001, Ge Yu 0001 |
DASFAA (3) | 6 |
| 2022 | Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksabstractWe study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision. Jinfeng Peng, Derong Shen, Nan Tang 0001, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
Proc. VLDB Endow. | 7 |
| 2022 | SenseLens: An Efficient Social Signal Conditioning System for True Event DetectionabstractThis article narrows the gap between physical sensing systems that measure physical signals and social sensing systems that measure information signals by (i) defining a novel algorithm for extracting information signals (building on results from text embedding) and (ii) showing that it increases the accuracy of truth discovery—the separation of true information from false/manipulated one. The work is applied in the context of separating true and false facts on social media, such as Twitter and Reddit, where users post predominantly short microblogs. The new algorithm decides how to aggregate the signal across words in the microblog for purposes of clustering the miscroblogs in the latent information signal space, where it is easier to separate true and false posts. Although previous literature extensively studied the problem of short text embedding/representation, this article improves previous work in three important respects: (1) Our work constitutes unsupervised truth discovery, requiring no labeled input or prior training. (2) We propose a new distance metric for efficient short text similarity estimation, we call Semantic Subset Matching , that improves our ability to meaningfully cluster microblog posts in the latent information signal space. (3) We introduce an iterative framework that jointly improves miscroblog clustering and truth discovery. The evaluation shows that the approach improves the accuracy of truth-discovery by 6.3%, 2.5%, and 3.8% (constituting a 38.9%, 14.2%, and 18.7% reduction in error, respectively) in three real Twitter data traces. Hang Cui 0001, Tarek F. Abdelzaher |
ACM Trans. Sens. Networks | 1 |
| 2021 | The voice of silence: interpreting silence in truth discovery on social mediaabstractThis paper enhances the interpretation of silence for purposes of truth discovery on social media. Most solutions to fact-finding problems from social media data focus on what users explicitly post. Absence of a post, however, also plays a key role in interpreting veracity of information. In this paper, we focus on (absent links in) the retweet graph. A user might abstain from propagating content for many potential reasons. For example, they might not be aware of the original post; they might find the content uninteresting; or they might doubt content veracity and refrain from propagation (among other reasons). This paper formulates a joint fact-finding and silence interpretation problem, and shows that the joint formulation significantly improves our ability to distinguish true and false claims. An unsupervised algorithm, Joint Network Embedding and Maximum Likelihood (JNEML) framework, is developed to solve this problem. We show that the joint algorithm outperforms other unsupervised baselines significantly on truth discovery tasks on three empirical data sets collected using the Twitter API. Hang Cui 0001, Tarek F. Abdelzaher |
ASONAM | 1 |
| 2019 | A Semi-Supervised Active-learning Truth Estimator for Social NetworksabstractThis paper introduces an active-learning-based truth estimator for social networks, such as Twitter, that enhances estimation accuracy significantly by requesting a well-selected (small) fraction of data to be labeled. Data assessment and truth discovery from arbitrary open online sources are a hard problem due to uncertainty regarding source reliability. Multiple truth finding systems were developed to solve this problem. Their accuracy is limited by the noisy nature of the data, where distortions, fabrications, omissions, and duplication are introduced. This paper presents a semi-supervised truth estimator for social networks, in which a portion of inputs are carefully selected to be reliably verified. The challenge is to find the subset of observations to verify that would maximally enhance the overall fact-finding accuracy. This work extends previous passive approaches to recursive truth estimation, as well as semi-supervised approaches where the estimator has no control over the choice of data to be labeled. Results show that by optimally selecting claims to be verified, we improve estimated accuracy by 12% over unsupervised baseline, and by 5% over previous semi-supervised approaches. Hang Cui 0001, Tarek F. Abdelzaher, Lance M. Kaplan |
WWW | 1 |
| 2018 | Recursive Truth Estimation of Time-Varying Sensing Data from Online Open SourcesabstractThis paper is motivated by prospective Internet of Things (IoT) applications that exploit inputs from online open sources whose reliability may be uncertain. Unlike physical signal fusion (that can leverage solid analytic foundations derived from physical properties of fused signals), data reliability assessment from arbitrary online open sources is a harder problem. At least two difficulties arise. First, source reliability is harder to estimate from first principles due to lack of visibility into the sensing and subsequent processing stages for published data. Second, by virtue of being open, some sources can be copied by others, leading to correlated errors at a large scale. This paper presents a recursive truth estimator for online public data streams that addresses the above two problems. We focus on categorical data. Many truth-finding systems were developed to cope with unreliable categorical data. Most of them are designed for batch analysis of bulk datasets. This work extends previous efforts by developing an online recursive estimator. Unlike previous recursive fact-finders, ours is the first that can jointly handle (i) changes in the population of sources over time, (ii) changes in the ground-truth state of the physical phenomenon being observed (that result in the appearance of conflicting claims), and (iii) correlated errors due to potential copying among sources. Results show that our algorithm not only outperforms other recursive fact-finders in the case of changing ground-truth state, but also improves estimation accuracy of static state. Hang Cui 0001, Tarek F. Abdelzaher, Lance M. Kaplan |
DCOSS | 1 |