EDBT 2026 Demo / reviewers in the wild / expert
Yu Zhao 0019
dblp:57/2056-19
· DBLP profile ↗
11ranked-venue papers in the field
4as first author
10since 2021 · last 2026
0000-0002-8454-0025ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 4 (3 first)Data Mining & Knowledge Discovery · 4 (1 first)Information Retrieval & Web Search · 2Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Comprehensive Survey on Enterprise Financial Risk Analysis from Big Data and LLMs Perspective
Huaming Du, Cancan Feng, Yuqian Lei, Guisong Liu, Gang Kou, Carl Yang 0001, Yu Zhao 0019 |
PAKDD (4) | 8 |
| 2026 | Traceable Latent Variable Discovery Based on Multi-Agent CollaborationabstractRevealing the underlying causal mechanisms in the real world is crucial for scientific and technological progress. Despite notable advances in recent decades, the lack of high-quality data and the reliance of traditional causal discovery algorithms (TCDA) on the assumption of no latent confounders, as well as their tendency to overlook the precise semantics of latent variables, have long been major obstacles to the broader application of causal discovery. To address this issue, we propose a novel causal modeling framework, TLVD, which integrates the metadata-based reasoning capabilities of large language models (LLMs) with the data-driven modeling capabilities of TCDA for inferring latent variables and their semantics. Specifically, we first employ a data-driven approach to construct a causal graph that incorporates latent variables. Then, we employ multi-LLM collaboration for latent variable inference, modeling this process as a game with incomplete information and seeking its Bayesian Nash Equilibrium (BNE) to infer the possible specific latent variables. Finally, to validate the inferred latent variables across multiple real-world web-based data sources, we leverage LLMs for evidence exploration to ensure traceability. We comprehensively evaluate TLVD on three de-identified real patient datasets provided by a hospital and two benchmark datasets. Extensive experimental results confirm the effectiveness and reliability of TLVD, with average improvements of 32.67% in Acc, 62.21% in CAcc, and 26.72% in ECit across the five datasets. Huaming Du, Yu Zhao 0019, Guisong Liu, Gang Kou, Carl Yang 0001 |
WWW | 4 |
| 2025 | Causal Discovery through Synergizing Large Language Model and Data-Driven ReasoningabstractRevealing the underlying causal mechanisms in the real world is critical for scientific and technical progress. Despite advancements over the past decades, the lack of high-quality data and the inability of traditional causal discovery algorithms (TCDA) to fully comprehend the exact semantics of variables have long been major obstacles to the broader application of causal discovery. To address this issue, this paper proposes a novel causal modeling framework, LLM-CD, which integrates the metadata-based reasoning capabilities of large language models (LLMs) with the data-driven modeling abilities of TCDA for causal discovery. LLM-CD deeply couples the reasoning abilities of LLMs at various stages of TCDA, and enhances causal discovery through an iterative process. Due to the issues of overconfidence and hallucination in LLMs, LLM-CD quantifies and analyzes its uncertainty by incorporating evidence-based deep learning theory with the assumptions of TCDA. We utilize a large-scale de-identified real patient dataset provided by a hospital, a new dataset extracted from MIMIC-IV about the same disease (lung cancer), and two benchmark datasets to comprehensively evaluate LLM-CD. Extensive experimental results confirm the effectiveness and reliability of LLM-CD, with the highest improvement of 403.93% in the Recall and 25.77% in the Ratio metric across four datasets. Huaming Du, Yujia Zheng 0001, Baoyu Jing, Yu Zhao 0019, Gang Kou, Guisong Liu, Weimin Li 0003, Carl Yang 0001 |
KDD (2) | 4 |
| 2025 | MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsabstractLarge language models (LLMs) are expected to offer structured Markdown responses for the sake of readability in web chatbots (e.g., ChatGPT). Although there are a myriad of metrics to evaluate LLMs, they fail to evaluate the readability from the view of output content structure. To this end, we focus on an overlooked yet important metric --- Markdown Awareness, which directly impacts the readability and structure of the content generated by these language models. In this paper, we introduce MDEval, a comprehensive benchmark to assess Markdown Awareness for LLMs, by constructing a dataset with 20K instances covering 10 subjects in English and Chinese. Unlike traditional model-based evaluations, MDEval provides excellent interpretability by combining model-based generation tasks and statistical methods. Our results demonstrate that MDEval achieves a Spearman correlation of 0.791 and an accuracy of 84.1% with human, outperforming existing methods by a large margin. Extensive experimental results also show that through fine-tuning over our proposed dataset, less performant open-source models are able to achieve comparable performance to GPT-4o in terms of Markdown Awareness. To ensure reproducibility and transparency, MDEval is open sourced at https://github.com/SWUFE-DB-Group/MDEval-Benchmark. Zhongpu Chen, Yinfeng Liu, Long Shi 0002, Zhi-Jie Wang 0009, Xingyan Chen, Yu Zhao 0019, Fuji Ren |
WWW | 6 |
| 2024 | Representation Learning of Temporal Graphs with Structural RolesabstractTemporal graph representation learning has drawn considerable attention in recent years. Most existing works mainly focus on modeling local structural dependencies of temporal graphs. However, underestimating the inherent global structural role information in many real-world temporal graphs inevitably leads to sub-optimal graph representations. To overcome this shortcoming, we propose a novel Role-based Temporal Graph Convolution Network (RTGCN) that fully leverages the global structural role information in temporal graphs. Specifically, RTGCN can effectively capture the static global structural roles by using hypergraph convolution neural networks. To capture the evolution of nodes' structural roles, we further design structural role-based gated recurrent units. Finally, we integrate structural role proximity in our objective function to preserve global structural similarity, further promoting temporal graph representation learning. Experimental results on multiple real-world datasets demonstrate that RTGCN consistently outperforms state-of-the-art temporal graph representation learning methods by significant margins in various temporal link prediction and node classification tasks. Specifically, RTGCN achieves AUC improvement of up to 5.1% for link prediction and F1 improvement of up to 6.2% for new link prediction. In addition, RTGCN achieves AUC improvement up to 4.6% for node classification and 2.7% for structural role classification. Huaming Du, Long Shi 0002, Xingyan Chen, Yu Zhao 0019, Hegui Zhang, Carl Yang 0001, Fuzhen Zhuang, Gang Kou |
KDD | 4 |
| 2024 | Combining intra-risk and contagion risk for enterprise bankruptcy prediction using graph neural networks
Shaopeng Wei 0002, Jia Lv, Yu Guo 0009, Xingyan Chen, Yu Zhao 0019, Qing Li 0005, Fuzhen Zhuang, Gang Kou |
Inf. Sci. | 6 |
| 2024 | Temporal Knowledge Graph Reasoning With Dynamic Memory EnhancementabstractTemporal Knowledge Graph (TKG) reasoning involves predicting future facts based on historical information by learning correlations between entities and relations. Recently, many models have been proposed for the TKG reasoning task. However, most existing models cannot efficiently utilize historical information, which can be summarized in two aspects: 1) Many models only consider the historical information in a fixed time range, resulting in a lack of useful information; 2) some models use all the historical facts, thus some noise or invalid facts are introduced during reasoning. In this regard, we propose a novel TKG reasoning model with dynamic memory enhancement (DyMemR). Inspired by human memory, we introduce memory capacity, memory loss, and repetition stimulation to design a human-like memory pool that could remember potentially useful historical facts. To fully leverage the memory pool, we utilize a two-stage training strategy.The first stage is guided by the memory-based encoding module which learns embeddings from memory-based subgraphs generated through the memory pool. The second stage is the memory-based scoring module that emphasizes the historical facts in the memory pool. Finally, we extensively validate the superiority of DyMemR against various state-of-the-art baselines. Zhao Zhang 0011, Fuzhen Zhuang, Yu Zhao 0019, Deqing Wang 0001, Hongwei Zheng 0003 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Stock Movement Prediction Based on Bi-Typed Hybrid-Relational Market Knowledge Graph via Dual Attention NetworksabstractStock Movement Prediction (SMP) aims at predicting listed companies' stock future price trend, which is a challenging task due to the volatile nature of financial markets. Recent financial studies show that the momentum spillover effect plays a significant role in stock fluctuation. However, previous studies typically only learn the simple connection information among related companies, which inevitably fail to model complex relations of listed companies in real financial market. To address this issue, we first construct a more comprehensive Market Knowledge Graph (MKG) which contains bi-typed entities including listed companies and their associated executives, and hybrid-relations including the explicit relations and implicit relations. Afterward, we proposeDanSmp, a novel Dual Attention Networks to learn the momentum spillover signals based upon the constructed MKG for stock prediction. The empirical experiments on our constructed datasets against nine SOTA baselines demonstrate that the proposedDanSmpis capable of improving stock prediction with the constructed MKG. Yu Zhao 0019, Huaming Du, Shaopeng Wei 0002, Xingyan Chen, Fuzhen Zhuang, Qing Li 0005, Gang Kou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Learning Bi-Typed Multi-Relational Heterogeneous Graph Via Dual Hierarchical Attention NetworksabstractBi-typed multi-relational heterogeneous graph (BMHG) is one of the most common graphs in practice, for example, academic networks, e-commerce user behavior graph and enterprise knowledge graph. It is a critical and challenge problem on how to learn the numerical representation for each node to characterize subtle structures. However, most previous studies treat all node relations in BMHG as the same class of relation without distinguishing the different characteristics between the intra-type relations and inter-type relations of the bi-typed nodes, causing the loss of significant structure information. To address this issue, we propose a novelDualHierarchicalAttentionNetworks (DHAN) based on the bi-typed multi-relational heterogeneous graphs to learn comprehensive node representations with the intra-type and inter-type attention-based encoder under a hierarchical mechanism. Specifically, the former encoder aggregates information from the same type of nodes, while the latter aggregates node representations from its different types of neighbors. Moreover, to sufficiently model node multi-relational information in BMHG, we adopt a newly proposed hierarchical mechanism. By doing so, the proposed dual hierarchical attention operations enable our model to fully capture the complex structures of the bi-typed multi-relational heterogeneous graphs. Experimental results on various tasks against the state-of-the-arts sufficiently confirm the capability of DHAN in learning node representations on the BMHGs. Yu Zhao 0019, Shaopeng Wei 0002, Huaming Du, Xingyan Chen, Qing Li 0005, Fuzhen Zhuang, Ji Liu 0002, Gang Kou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Connecting Embeddings Based on Multiplex Relational Graph Attention Networks for Knowledge Graph Entity TypingabstractKnowledge graph entity typing (KGET) aims to infer missing entity typing instances in KGs, which is a significant subtask of KG completion. Despite of its progress, however, it still faces two non-trivial challenges: (i) most existing KGET methods extract features by encoding the existing entity typing tuples, while ignoring rich relational knowledge. (ii) they typically treat each entity typing tuple in KGs independently, and thus inevitably fail to take account of the inherent and valuable neighborhood information surrounding a tuple. To address these challenges, we build a novel Heterogeneous Relational Graph (HRG), and propose a Multiplex Relational Graph Attention Networks (MRGAT) to learn on HRG, and then utilize a Connecting Embeddings model (ConnectE) to make entity type inference. Specifically, the framework contains three components. Firstly, to effectively integrate the entity typing tuples and entity relation triples in KGs, we construct a HRG that consists of three semantic subgraphs. Secondly, we employ MRGAT to learn embeddings on HRG. In MRGAT, each subgraph of HRG is fed to its corresponding model that is capable of capturing neighborhood information. Finally, given the learned embeddings, we make entity type prediction by ConnectE. Experimental results validate the superiority of our model against various state-of-the-art baselines. Yu Zhao 0019, Han Zhou 0008, Anxiang Zhang, Ruobing Xie, Qing Li 0005, Fuzhen Zhuang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Knowledge base completion by learning pairwise-interaction differentiated embeddings
Yu Zhao 0019, Sheng Gao 0001, Patrick Gallinari, Jun Guo 0002 |
Data Min. Knowl. Discov. | 1 |