VLDB 2026 Research / reviewers in the wild / expert
Huobin Tan
dblp:225/6585
· DBLP profile ↗
11ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-3113-6552ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 40% Representation and self-supervised learning · 20% Graph learning · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 77% Knowledge graphs · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › contrastive learning
graph contrastive learning |
1.0 | 1 | 2026 | GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
graph diffusion model |
1.0 | 1 | 2026 | MoE-Guided Graph Diffusion for Oriented Molecule Design · AAAI 2026 |
Machine learning › Graph learning › graph neural network
heterophily |
1.0 | 1 | 2026 | GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs · AAAI 2026 |
Machine learning › Generative modeling › molecular generation
molecular design |
1.0 | 1 | 2026 | MoE-Guided Graph Diffusion for Oriented Molecule Design · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
optimal transport alignment |
1.0 | 1 | 2026 | GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs · AAAI 2026 |
Recommender systems › knowledge-aware recommendation
knowledge graph-based recommendation |
0.4 | 1 | 2020 | CKAN: Collaborative Knowledge-aware Attentive Network for Recommender Systems · SIGIR 2020 |
Bioinformatics and computational biology › drug discovery
molecular optimization |
0.3 | 1 | 2026 | MoE-Guided Graph Diffusion for Oriented Molecule Design · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.0mixture of experts · 2.0optimal transport · 1.0language model · 1.0graph neural network · 1.0collaborative filtering · 0.4attentive network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoE-Guided Graph Diffusion for Oriented Molecule DesignabstractDesigning molecules with desired properties, aka the oRiented molEcule Design (RED), is a fundamental task in chemistry and materials science. While graph diffusion models (GDMs) and reinforcement learning techniques (RL) show promise in molecule structure generation and property optimization stages individually, their integration in the unified RED task often suffers from poor compatibility. The large variance among candidate molecular structures generated by GDMs can be amplified in the iterative optimization process of RL, leading to slow and unstable convergence. In this work, motivated by the adaptive and divide-and-conquer characteristics of Mixture of Experts (MoE) architecture, we propose a novel framework called MoE-Guided Graph Diffusion Model (MEGD) that incorporates the MoE architecture to guide the orchestration of GDM and RL, promoting faster and more stable convergence in the design process. MEGD is evaluated on benchmark datasets optimizing the physical and chemical properties of AI-generated molecular structures. On all three datasets, our method outperforms the best of 9 alternative models by 7.73% on the target structural properties, while not penalizing other important application-level quality metrics of the generated molecules. A real-world case study on an emerging class of material, i.e., metal-organic framework, is also conducted, which further demonstrates the effectiveness of our method in accomplishing the RED task. Shuochen Li, Xiangqi Guo, Huobin Tan |
AAAI | 3 |
| 2026 | GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed GraphsabstractRecently, structure–text contrastive learning has shown promising performance on text-attributed graphs by leveraging the complementary strengths of graph neural networks and language models. However, existing methods typically rely on homophily assumptions in similarity estimation and hard optimization objectives, which limit their applicability to heterophilic graphs. Although existing methods can mitigate heterophily through structural adjustments or neighbor aggregation, they usually treat textual embeddings as static targets, leading to suboptimal alignment. In this work, we identify the multi-granular heterophily in text-attributed graphs, including complete heterophily, partial heterophily, and latent homophily, which makes structure–text alignment particularly challenging due to mixed, noisy, and missing semantic correlations. To achieve flexible and bidirectional alignment, we propose GCL-OT, a novel graph contrastive learning framework with optimal transport, equipped with tailored mechanisms for each type of heterophily. Specifically, for partial heterophily, we design a RealSoftMax-based similarity estimator to emphasize key neighbor-word interactions while easing background noise. For complete heterophily, we introduce a prompt-based filter that adaptively excludes irrelevant noise during optimal transport alignment. Furthermore, we incorporate OT-guided soft supervision to uncover potential neighbors with similar semantics, enhancing the learning of latent homophily. Theoretical analysis shows that GCL-OT can improve the mutual information bound and Bayes error guarantees. Extensive experiments on nine benchmarks show that GCL-OT outperforms state-of-the-art methods, demonstrating its effectiveness and robustness. Yating Ren, Yikun Ban, Huobin Tan |
AAAI | 3 |
| 2025 | NuPreX: Times Series Forecasting for the Nuclear Steam Supply System
Huobin Tan, Biao Dong |
ICIC (7) | 1 |
| 2025 | Multimodal Object Detection by Adaptive Channel Enhancement and Attention Fusionabstract—Cross-modal feature fusion is a critical research area in multimodal object detection, focusing on integrating features extracted from different modalities to retain richer semantic information. While several advanced fusion strategies have been proposed, most fail to effectively address the interaction of complementary information between modalities, resulting in suboptimal information exchange and fusion that do not fully leverage the intrinsic characteristics of each modality. To address these challenges, this paper introduces an Adaptive Channel Enhancement and Attention Fusion (ACAF) module, which bridges the feature gaps across modalities, enabling smooth information interaction and attention-based multimodal feature integration. Specifically, the module adaptively reweights weaker channels in each modality using features from other modalities, thereby enhancing the expressiveness of single-modal feature representations. Additionally, an attention mechanism is employed to capture multidimensional cross-modal relationships, facilitating efficient feature fusion. Experimental results on the DroneVehicle and VEDAI datasets demonstrate that our method significantly outperforms the baseline models, achieving improvements of 3.0% and 2.4% in recall, 2.1% and 1.7% in mAP@50, and 3.0% and 4.6% in mAP@50:95, respectively. This shows that channel reweighting enhances cross-modal information interaction and fusion, leading to superior performance in multimodal object detection. Yaqi Mei, Tianyuan Zhang 0004, Huobin Tan |
IJCNN | 3 |
| 2025 | Graph neural network-based long method and blob code smell detectionabstract• We propose a graph neural network-based model for long method and blob code smell detection. • The best strategies for the class imbalance of graph data and graph pooling are determined through experiments in our method. • During model design for abstract syntax tree of code, Euclidean space and non-Euclidean space are combined. • The experiments show that our proposed method outperforms machine learning methods and deep learning methods. The concept of code smell was first proposed in the late nineties, to refer to signals that code may need refactoring. While not necessarily affecting functionality, code smell can hinder understandability and future scalability of the program. As a result, the precise detection of code smell has become an important topic in coding research. However, current detection methods are limited by imbalanced and industrial-irrelevant datasets, a lack of sufficient structural and logical information on the code, and simple model architecture. Given these limitations, this paper utilized an industry-relevant and sufficient dataset and then developed a graph neural network to better detect code smell. First, we identified Long Method and Blob as our research subjects due to their frequent occurrence and impacts on the maintainability of software. We then designed modified fuzzy sampling with focalloss to address the issue of data imbalance. Second, to deal with the large volume of data, we proposed a global and local attention scoring mechanism to extract the key information from the code. Third, in order to design a graph neural network specifically for the abstract syntax tree of code, we combined Euclidean space and non-Euclidean space. Finally, we compared our method with other machine learning methods and deep learning methods. The results demonstrate that our method outperforms the other methods on Long Method and Blob, which indicates the effectiveness of our proposed method. Minnan Zhang, Jingdong Jia, Luiz Fernando Capretz, Huobin Tan |
Sci. Comput. Program. | 5 |
| 2025 | Visual analysis of LLM-based entity resolution from scientific papersabstractThis paper focuses on the visual analytics support for extracting domain-specific entities from extensive scientific literature, a task with inherent limitations using traditional named entity resolution methods. With the advent of large language models (LLMs) such as GPT-4, significant improvements over conventional machine learning approaches have been achieved due to LLM’s capability on entity resolution integrate abilities such as understanding multiple types of text. This research introduces a new visual analysis pipeline that integrates these advanced LLMs with versatile visualization and interaction designs to support batch entity resolution. Specifically, we focus on a specific material science field of Metal-Organic Frameworks (MOFs) and a large data collection namely CSD-MOFs. Through collaboration with domain experts in material science, we obtain well-labeled synthesis paragraphs. We propose human-in-the-loop refinement over the entity resolution process using visual analytics techniques, which allows domain experts to interactively integrate insights into LLM intelligence, including error analysis and interpretation of the retrieval-augmented generation (RAG) algorithm. Our evaluation through the case study of example selection for RAG demonstrates that this visual analysis approach effectively improves the accuracy of single-document entity resolution. Weize Wu, Ruiming Li, Huobin Tan, Zipeng Liu |
Vis. Informatics | 7 |
| 2024 | A Transformer-based Knowledge Graph Embedding Model Combining Graph Paths and Local NeighborhoodabstractMany existing knowledge graph embedding methods achieve outstanding performance by exploiting the graph structure, among which graph neural network-based methods that utilize the local neighborhood are the most representative. However, the shallow network structure of graph neural networks limits the model’s expressiveness, and the existing problem of over-smoothing also prevents the model from capturing long-distance information. To address these issues, we propose a Transformer-based knowledge graph embedding method that combines graph paths and local neighborhood (TKGE-PN). First, it samples multiple graph paths by using a biased random walk algorithm starting from the central entity. Then these sampled paths are transformed into vector representations by a Transformer-based graph path encoding module. Finally, the local neighborhood encoding module aggregates all graph path vector representations to score triples. During the graph path encoding process, a masked entity relation prediction task is used to enhance the model’s ability to learn long-distance information. Experimental results show that on two standard datasets, FB15k-237 and WN18RR, the performance of TKGE-PN surpasses that of most existing models, demonstrating the effectiveness of our approach. Huobin Tan, Yating Ren |
IJCNN | 2 |
| 2022 | Heterogeneous Graph Neural Network with Hypernetworks for Knowledge Graph Embedding
Xiyang Liu 0001, Huobin Tan, Richong Zhang |
ISWC | 3 |
| 2021 | Curved SDE-Net Leads to Better Generalization for Uncertainty Estimates of DNNs
Yongguang Wang, Huobin Tan, Shuzhen Yao |
ICANN (2) | 2 |
| 2020 | CKAN: Collaborative Knowledge-aware Attentive Network for Recommender SystemsabstractSince it can effectively address the problem of sparsity and cold start of collaborative filtering, knowledge graph (KG) is widely studied and employed as side information in the field of recommender systems. However, most of existing KG-based recommendation methods mainly focus on how to effectively encode the knowledge associations in KG, without highlighting the crucial collaborative signals which are latent in user-item interactions. As such, the learned embeddings underutilize the two kinds of pivotal information and are insufficient to effectively represent the latent semantics of users and items in vector space. Guangyan Lin, Huobin Tan, Qinghong Chen, Xiyang Liu 0001 |
SIGIR | 3 |
| 2018 | Neural Collaborative Filtering: Hybrid Recommendation Algorithm with Content Information and Implicit Feedback
Guangyan Lin, Huobin Tan |
IDEAL (1) | 3 |