Wenzhou Dou

dblp:319/3742 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0005-5823-1965ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Cost of Thinking: Increased Jailbreak Risk in Large Language Models in Education
Ke Wang 0068, Wenzhou Dou, Yifan Shuai, Zitao Liu 0001, Weiqi Luo 0002
AIED3
2026 COMA: A Collaborative Multi-Role Agent Framework for Automated Lesson Plan Generation
Xiaoli Zeng, Ying Zheng 0010, Shuyan Huang, Zitao Liu 0001, Mi Tian 0008, Mingliang Hou, Jiaqi Zheng 0012, Wenzhou Dou
WWW8
2026 Weakly-supervised entity matching via LLM-guided data augmentation and knowledge transfer
Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
Knowl. Based Syst.1
2024 Enhancing Deep Entity Resolution with Integrated Blocker-Matcher Training: Balancing Consensus and Discrepancy
abstract
Deep entity resolution (ER) identifies matching entities across data sources using techniques based on deep learning. It involves two steps: a blocker for identifying the potential matches to generate the candidate pairs, and a matcher for accurately distinguishing the matches and non-matches among these candidate pairs. Recent deep ER approaches utilize pretrained language models (PLMs) to extract similarity features for blocking and matching, achieving state-of-the-art performance. However, they often fail to balance the consensus and discrepancy between the blocker and matcher, emphasizing the consensus while neglecting the discrepancy. This paper proposes MutualER, a deep entity resolution framework that integrates and jointly trains the blocker and matcher, balancing both the consensus and discrepancy between them. Specifically, we firstly introduce a lightweight PLM in siamese structure for the blocker and a heavier PLM in cross structure or an autoregressive large language model (LLM) for the matcher. Two optimization techniques named Mutual Sample Selection (MSS) and Similarity Knowledge Transferring (SKT) are designed to jointly train the blocker and matcher. MSS enables the blocker and matcher to mutually select the customized training samples for each other to maintain the discrepancy, while SKT allows them to share the similarity knowledge for improving their blocking and matching capabilities respectively to maintain the consensus. Extensive experiments on five datasets demonstrate that MutualER significantly outperforms existing PLM-based and LLM-based approaches, achieving leading performance in both effectiveness and efficiency.
Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
CIKM1
2024 Towards Long-Text Entity Resolution with Chain-of-Thought Knowledge Augmentation from Large Language Models
Jiakai Tang, Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou
DASFAA (5)2
2023 Soft Target-Enhanced Matching Framework for Deep Entity Matching
abstract
Deep Entity Matching (EM) is one of the core research topics in data integration. Typical existing works construct EM models by training deep neural networks (DNNs) based on the training samples with onehot labels. However, these sharp supervision signals of onehot labels harm the generalization of EM models, causing them to overfit the training samples and perform badly in unseen datasets. To solve this problem, we first propose that the challenge of training a well-generalized EM model lies in achieving the compromise between fitting the training samples and imposing regularization, i.e., the bias-variance tradeoff. Then, we propose a novel Soft Target-EnhAnced Matching (Steam) framework, which exploits the automatically generated soft targets as label-wise regularizers to constrain the model training. Specifically, Steam regards the EM model trained in previous iteration as a virtual teacher and takes its softened output as the extra regularizer to train the EM model in the current iteration. As such, Steam effectively calibrates the obtained EM model, achieving the bias-variance tradeoff without any additional computational cost. We conduct extensive experiments over open datasets and the results show that our proposed Steam outperforms the state-of-the-art EM approaches in terms of effectiveness and label efficiency.
Wenzhou Dou, Derong Shen, Xiangmin Zhou, Tiezheng Nie, Yue Kou, Hang Cui 0001, Ge Yu 0001
AAAI1
2023 Domain-Generic Pre-Training for Low-Cost Entity Matching via Domain Alignment and Domain Antagonism
abstract
Entity matching (EM) is a core problem of data mining and data integration. Existing EM solutions achieve great successes by designing and training the deep learning model for the specific domain, e.g., watches and shoes. However, these methods require a high training cost (e.g., model engineering and data preprocessing) in realistic EM applications. In this paper, we develop a deep learning-based solution in the manner of the domain-generic pre-training that targets low training cost for EM through a novel combination of the domain alignment and domain antagonism. In domain alignment, we design a novel contrastive learning method that align the representation distribution of different domains. And in domain antagonism, we conduct the domain adversarial training to force the encoder to focus the domain-generic knowledge. These two optimizations ensure that the pre-trained EM model can capture the general matching knowledge and be fine-tuned into specific domains at a fairly low cost. Empirical evaluation demonstrates that this combination achieves state-of-the-art performance in both in-domain and out-of-domain settings.
Derong Shen, Wenzhou Dou, Tiezheng Nie, Yue Kou
IJCNN3
2022 SAREM: Semi-supervised Active Heterogeneous Entity Matching Framework
Jinxiu Du, Tiezheng Nie, Wenzhou Dou, Derong Shen, Yue Kou
WISA3
2022 Empowering Transformer with Hybrid Matching Knowledge for Entity Matching
Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou, Chenchen Sun, Hang Cui 0001, Ge Yu 0001
DASFAA (3)1