EDBT 2026 Demo / reviewers in the wild / expert
Zhenyu Zeng
dblp:273/0191
· DBLP profile ↗
10ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-0329-6893ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Knowledge graphs · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 63% Smart cities and intelligent transportation · 37% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs › domain-specific knowledge graph
urban knowledge graph |
1.2 | 2 | 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal Prediction · NeurIPS 2023 A Knowledge-Enhanced Framework for Imitative Transportation Trajectory Generation · ICDM 2022 |
Bioinformatics and computational biology › functional genomics
functional enrichment analysis |
0.7 | 1 | 2023 | RNAenrich: a web server for non-coding RNA enrichment · Bioinform. 2023 |
Bioinformatics and computational biology › transcriptomics
non-coding RNA analysis |
0.7 | 1 | 2023 | RNAenrich: a web server for non-coding RNA enrichment · Bioinform. 2023 |
Knowledge graphs
knowledge graph embedding |
0.7 | 1 | 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal Prediction · NeurIPS 2023 |
Knowledge graphs
link prediction |
0.7 | 1 | 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal Prediction · NeurIPS 2023 |
Smart cities and intelligent transportation
trajectory generation |
0.6 | 1 | 2022 | A Knowledge-Enhanced Framework for Imitative Transportation Trajectory Generation · ICDM 2022 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.4 | 1 | 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential Learning · KDD 2020 |
Machine learning › Deep learning architectures and training
memory-augmented neural networks |
0.4 | 1 | 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential Learning · KDD 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
multi-view sequential learning |
0.4 | 1 | 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential Learning · KDD 2020 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention and cross-attention |
0.4 | 1 | 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential Learning · KDD 2020 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.4 | 1 | 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential Learning · KDD 2020 |
Smart cities and intelligent transportation › spatio-temporal prediction
urban spatio-temporal prediction |
0.2 | 1 | 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal Prediction · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
knowledge graph embedding · 1.3knowledge graph representation learning · 1.1imitation learning · 1.1adversarial training · 1.1enrichment analysis · 0.7recurrent neural network · 0.4external memory · 0.4attention · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Match, Compare, or Select? An Investigation of Large Language Models for Entity MatchingabstractEntity matching (EM) is a critical step in entity resolution (ER). Recently, entity matching based on large language models (LLMs) has shown great promise. However, current LLM-based entity matching approaches typically follow a binary matching paradigm that ignores the global consistency among record relationships. In this paper, we investigate various methodologies for LLM-based entity matching that incorporate record interactions from different perspectives. Specifically, we comprehensively compare three representative strategies: matching, comparing, and selecting, and analyze their respective advantages and challenges in diverse scenarios. Based on our findings, we further design a compound entity matching framework (ComEM) that leverages the composition of multiple strategies and LLMs. ComEM benefits from the advantages of different sides and achieves improvements in both effectiveness and efficiency. Experimental results on 8 ER datasets and 10 LLMs verify the superiority of incorporating record interactions through the selecting strategy, as well as the further cost-effectiveness brought by ComEM. Tianshu Wang 0002, Xiaoyang Chen 0001, Xuanang Chen, Xianpei Han, Le Sun 0001, Zhenyu Zeng |
COLING | 8 |
| 2025 | DBCopilot: Natural Language Querying over Massive Databases via Schema Routing
Tianshu Wang 0002, Xiaoyang Chen 0001, Xianpei Han, Le Sun 0001, Zhenyu Zeng |
EDBT | 7 |
| 2025 | Enhancing online industrial quality index prediction with a general deep temporal feature extraction and incremental ensemble modeling frameworkabstractData-driven modeling methods for industrial quality index prediction often face the challenge of limited data representation. Using process variable snapshots from a single time step is insufficient for building high-performance soft sensors. Moreover, during online prediction, the performance of soft sensors is affected by diverse operating conditions and concept drift in industrial data streams, leading to performance degradation . To address these challenges, this paper proposes a Temporal Feature Extraction and Incremental Variational Bayesian Regression Ensemble (TFE-IVBRE) framework, which provides a general solution for various online industrial quality index prediction tasks. The TFE-IVBRE framework combines the temporal feature extraction capabilities of deep neural networks , the ability of ensemble learning to handle diverse operating conditions, and the online learning features of variational Bayesian models . An incremental update strategy is also incorporated to maintain consistent prediction performance. Experiments in the Debutanizer Column and Sulfur Recovery Unit scenarios show that the prediction performance of TFE-IVBRE significantly outperforms other static and online comparison models, with the effectiveness of its components validated through ablation studies. Finally, the overall model also demonstrates good robustness. These results offer valuable insights for the advancement of industrial soft sensor development. Zhenyu Zeng |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | MoDAFold: a strategy for predicting the structure of missense mutant protein based on AlphaFold2 and molecular dynamicsabstractProtein structure prediction is a longstanding issue crucial for identifying new drug targets and providing a mechanistic understanding of protein functions. To enhance the progress in this field, a spectrum of computational methodologies has been cultivated. AlphaFold2 has exhibited exceptional precision in predicting wild-type protein structures, with performance exceeding that of other methods. However, predicting the structures of missense mutant proteins using AlphaFold2 remains challenging due to the intricate and substantial structural alterations caused by minor sequence variations in the mutant proteins. Molecular dynamics (MD) has been validated for precisely capturing changes in amino acid interactions attributed to protein mutations. Therefore, for the first time, a strategy entitled 'MoDAFold' was proposed to improve the accuracy and reliability of missense mutant protein structure prediction by combining AlphaFold2 with MD. Multiple case studies have confirmed the superior performance of MoDAFold compared to other methods, particularly AlphaFold2. Lingyan Zheng, Shuiyang Shi, Xiuna Sun, Mingkun Lu, Yang Liao, Sisi Zhu, Hongning Zhang, Pan Fang, Zhenyu Zeng, Honglin Li 0003, Zhaorong Li, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 10 |
| 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal PredictionabstractAccurate Urban SpatioTemporal Prediction (USTP) is of great importance to the development and operation of the smart city. As an emerging building block, multi-sourced urban data are usually integrated as urban knowledge graphs (UrbanKGs) to provide critical knowledge for urban spatiotemporal prediction models. However, existing UrbanKGs are often tailored for specific downstream prediction tasks and are not publicly available, which limits the potential advancement. This paper presents UUKG, the unified urban knowledge graph dataset for knowledge-enhanced urban spatiotemporal predictions. Specifically, we first construct UrbanKGs consisting of millions of triplets for two metropolises by connecting heterogeneous urban entities such as administrative boroughs, POIs, and road segments. Moreover, we conduct qualitative and quantitative analysis on constructed UrbanKGs and uncover diverse high-order structural patterns, such as hierarchies and cycles, that can be leveraged to benefit downstream USTP tasks. To validate and facilitate the use of UrbanKGs, we implement and evaluate 15 KG embedding methods on the KG completion task and integrate the learned KG embeddings into 9 spatiotemporal models for five different USTP tasks. The extensive experimental results not only provide benchmarks of knowledge-enhanced USTP models under different task settings but also highlight the potential of state-of-the-art high-order structure-aware UrbanKG embedding methods. We hope the proposed UUKG fosters research on urban knowledge graphs and broad smart city applications. The dataset and source code are available at https://github.com/usail-hkust/UUKG/. Yansong Ning, Hao Liu 0026, Hao Wang 0073, Zhenyu Zeng, Hui Xiong 0001 |
NeurIPS | 4 |
| 2023 | RNAenrich: a web server for non-coding RNA enrichmentabstractMOTIVATION: With the rapid advances of RNA sequencing and microarray technologies in non-coding RNA (ncRNA) research, functional tools that perform enrichment analysis for ncRNAs are needed. On the one hand, because of the rapidly growing interest in circRNAs, snoRNAs, and piRNAs, it is essential to develop tools for enrichment analysis for these newly emerged ncRNAs. On the other hand, due to the key role of ncRNAs' interacting target in the determination of their function, the interactions between ncRNA and its corresponding target should be fully considered in functional enrichment. Based on the ncRNA-mRNA/protein-function strategy, some tools have been developed to functionally analyze a single type of ncRNA (the majority focuses on miRNA); in addition, some tools adopt predicted target data and lead to only low-confidence results. RESULTS: Herein, an online tool named RNAenrich was developed to enable the comprehensive and accurate enrichment analysis of ncRNAs. It is unique in (i) realizing the enrichment analysis for various RNA types in humans and mice, such as miRNA, lncRNA, circRNA, snoRNA, piRNA, and mRNA; (ii) extending the analysis by introducing millions of experimentally validated data of RNA-target interactions as a built-in database; and (iii) providing a comprehensive interacting network among various ncRNAs and targets to facilitate the mechanistic study of ncRNA function. Importantly, RNAenrich led to a more comprehensive and accurate enrichment analysis in a COVID-19-related miRNA case, which was largely attributed to its coverage of comprehensive ncRNA-target interactions. AVAILABILITY AND IMPLEMENTATION: RNAenrich is now freely accessible at https://idrblab.org/rnaenr/. Kuerbannisha Amahong, Yintao Zhang, Mingkun Lu, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Haibin Dai, Jianqing Gao, Feng Zhu 0004 |
Bioinform. | 7 |
| 2022 | A Knowledge-Enhanced Framework for Imitative Transportation Trajectory GenerationabstractDesigning efficient tools for modeling and analyzing transportation trajectories help improve various smart city services, e.g., location prediction, business recommendation and traffic forecasting. Yet with limited historical travel data, conventional tools are faced with challenges in representing the sophisticated contextual and spatiotemporal relationships inherent in transportation trajectories, while such trajectories have strong functional, geographical and time-specific characteristics. In this work, we propose an imitative transportation trajectory generation framework with multi-source urban knowledge enhancement. Specifically, we first construct Know-ST, an urban knowledge fusion framework that captures cross-domain knowledge (i.e., road network, point of interest, etc.) via tailored urban knowledge graph representation learning. The Know-ST preserves the temporal characteristics of trajectories and can continuously adapt to newly arriving trajectories. Moreover, we formulate the trajectory generation problem as a sequential imitation learning process, where an adversarially trained generator network is equipped to autonomously capture both the spatiotemporal and contextual information within individual transportation trajectories. Extensive experiments on real-world dataset show Know-ST outperforms multiple baselines under a variety of benchmarks, indicating the effectiveness of the proposed framework. Qingyan Zhu, Yize Chen, Hao Wang 0026, Zhenyu Zeng |
ICDM | 4 |
| 2022 | ConSIG: consistent discovery of molecular signature from OMIC dataabstractThe discovery of proper molecular signature from OMIC data is indispensable for determining biological state, physiological condition, disease etiology, and therapeutic response. However, the identified signature is reported to be highly inconsistent, and there is little overlap among the signatures identified from different biological datasets. Such inconsistency raises doubts about the reliability of reported signatures and significantly hampers its biological and clinical applications. Herein, an online tool, ConSIG, was constructed to realize consistent discovery of gene/protein signature from any uploaded transcriptomic/proteomic data. This tool is unique in a) integrating a novel strategy capable of significantly enhancing the consistency of signature discovery, b) determining the optimal signature by collective assessment, and c) confirming the biological relevance by enriching the disease/gene ontology. With the increasingly accumulated concerns about signature consistency and biological relevance, this online tool is expected to be used as an essential complement to other existing tools for OMIC-based signature discovery. ConSIG is freely accessible to all users without login requirement at https://idrblab.org/consig/. Feng Cheng Li, Jiayi Yin, Mingkun Lu, Qingxia Yang, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Haibin Dai, Yuzong Chen 0002, Feng Zhu 0004 |
Briefings Bioinform. | 5 |
| 2022 | Biological activities of drug inactive ingredientsabstractIn a drug formulation (DFM), the major components by mass are not Active Pharmaceutical Ingredient (API) but rather Drug Inactive Ingredients (DIGs). DIGs can reach much higher concentrations than that achieved by API, which raises great concerns about their clinical toxicities. Therefore, the biological activities of DIG on physiologically relevant target are widely demanded by both clinical investigation and pharmaceutical industry. However, such activity data are not available in any existing pharmaceutical knowledge base, and their potentials in predicting the DIG-target interaction have not been evaluated yet. In this study, the comprehensive assessment and analysis on the biological activities of DIGs were therefore conducted. First, the largest number of DIGs and DFMs were systematically curated and confirmed based on all drugs approved by US Food and Drug Administration. Second, comprehensive activities for both DIGs and DFMs were provided for the first time to pharmaceutical community. Third, the biological targets of each DIG and formulation were fully referenced to available databases that described their pharmaceutical/biological characteristics. Finally, a variety of popular artificial intelligence techniques were used to assess the predictive potential of DIGs' activity data, which was the first evaluation on the possibility to predict DIG's activity. As the activities of DIGs are critical for current pharmaceutical studies, this work is expected to have significant implications for the future practice of drug discovery and precision medicine. Minjie Mou, Wei Zhang 0218, Xichen Lian, Shuiyang Shi, Mingkun Lu, Huaicheng Sun, Feng Cheng Li, Zhenyu Zeng, Zhaorong Li, Yunqing Qiu, Feng Zhu 0004, Jianqing Gao |
Briefings Bioinform. | 11 |
| 2020 | Attention and Memory-Augmented Networks for Dual-View Sequential LearningabstractIn recent years, sequential learning has been of great interest due to the advance of deep learning with applications in time-series forecasting, natural language processing, and speech recognition. Recurrent neural networks (RNNs) have achieved superior performance in single-view and synchronous multi-view sequential learning comparing to traditional machine learning models. However, the method remains less explored in asynchronous multi-view sequential learning, and the unalignment nature of multiple sequences poses a great challenge to learn the inter-view interactions. We develop an AMANet (Attention and Memory-Augmented Networks) architecture by integrating both attention and memory to solve asynchronous multi-view learning problem in general, and we focus on experiments in dual-view sequences in this paper. Self-attention and inter-attention are employed to capture intra-view interaction and inter-view interaction, respectively. History attention memory is designed to store the historical information of a specific object, which serves as local knowledge storage. Dynamic external memory is used to store global knowledge for each view. We evaluate our model in three tasks: medication recommendation from a patient's medical records, diagnosis-related group (DRG) classification from a hospital record, and invoice fraud detection through a company's taxation behaviors. The results demonstrate that our model outperforms all baselines and other state-of-the-art models in all tasks. Moreover, the ablation study of our model indicates that the inter-attention mechanism plays a key role in the model and it can boost the predictive power by effectively capturing the inter-view interactions from asynchronous views. Yong He 0008, Nan Li 0027, Zhenyu Zeng |
KDD | 4 |