Wenzhe Hou

dblp:360/6087 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0001-7227-7899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity Matching
abstract
Entity Matching (EM) is a fundamental task in data integration, traditionally studied over structured data such as tables and knowledge graphs. In modern repositories, real-world objects are often represented across multiple modalities, including structured entities with symbolic attributes and visual entities in image-centric collections. This motivates cross-modal entity matching, which aims to identify visual-structured entity pairs that refer to the same object. Existing methods typically rely on pretrained vision-language models to compute entity pair similarity and derive correspondence via local ranking, which, however, can be unreliable given noisy and ambiguous cross-modal data and may produce globally inconsistent correspondences across related entities. Multimodal large language models (MLLMs) offer richer cross-modal cues for matching, but exhaustive MLLM reasoning over large candidate spaces is prohibitively expensive. To address these limitations, in this work, we propose øurq, a blackboard-based multi-agent framework for zero-shot cross-modal entity matching that performs iterative, self-correcting refinement by fusing multiple matching signals, explicitly regulating global consistency, and selectively invoking an MLLM only for hard cases. We further construct two new benchmarks from real-world visual and structured data. Extensive experiments show that øurq consistently outperforms competitive baselines, providing an effective solution for zero-shot cross-modal entity matching. We release our data and code at https://github.com/Q-17/CoMCo.
Shiqi Zhang 0011, Weixin Zeng, Wenzhe Hou, Weidong Xiao 0003, Xiang Zhao 0002
SIGIR4
2025 OptMatch: An Efficient and Generic Neural Network-Assisted Subgraph Matching Approach
abstract
The graph has been widely used to model the entities and the relationships among them in real-world applications. Subgraph matching is a core operation in graph data analysis. However, existing exact matching methods may incur high cost as their searched branches are always unpromising. In recent years, several approximate matching solutions have been proposed by exploiting neural networks. Nevertheless, the accuracy of the returned approximate results could be improved significantly. Motivated by these observations, we proposed OptMatch, an efficient and generic neural network-assisted subgraph matching approach, in this work. In particular, OptMatch proposes a novel subgraph partial embedding network and implements carefully designed search strategies to optimize search processes during the subgraph matching process. First, it can be used to accelerate existing exact matching methods. Moreover, it is also an approximate matching solution, which offers better accuracy compared to existing approximate solutions. We conduct exten-sive experiments on seven real-world data graphs to demonstrate the superiority of OptMatch in both exact and approximate subgraph matching.
Wenzhe Hou, Xiang Zhao 0002
ICDE1
2024 LearnSC: An Efficient and Unified Learning-Based Framework for Subgraph Counting Problem
abstract
Graphs are valuable data structures used to represent complex relationships between entities in a wide range of applications, such as social networks and chemical reactions. Subgraph counting problem is a well-known hard problem, as its core subroutine, the subgraph matching, is NP-complete. In this work, we propose an efficient and unified deep learning-based solution framework LearnSC, which solves the subgraph counting problem approximately. This framework offers two key advantages: (i) it is a generic solution that is orthogonal to the existing techniques of learning-based solutions; and (ii) it is equipped with a suite of optimizations to significantly improve the accuracy of the estimated results. Our experimental results on 7 datasets demonstrate that our proposal is highly accurate, robust, and scalable, making it an excellent solution for subgraph counting problem among all statistics-based and learning-based competitors.
Wenzhe Hou, Xiang Zhao 0002
ICDE1
2023 Cardinality Estimation of Subgraph Search Queries with Direction Learner
Wenzhe Hou, Xiang Zhao 0002, Wei Wang 0011
ADMA (5)1