EDBT 2026 Demo / reviewers in the wild / expert
Hang Zhang 0032
dblp:49/6156-32
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0009-5918-3183ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Text2GQL: Integrating Structural Twig Linking and Evolutionary In-Context LearningabstractWhile large language models have revolutionized Text-to-SQL tasks, translating natural language into Graph Query Languages (Text2GQL) remains underexplored due to the topological heterogeneity and syntactic diversity of graph query languages (e.g., Cypher, Gremlin, SPARQL).Existing approaches often struggle with structural hallucinations and lack adaptability in cold-start scenarios.In this paper, we present a unified, trainingfree Text2GQL framework.First, Structural Twig Linking elevates schema grounding to the identification of semantic substructures ("twigs"), providing robust topological priors.Second, addressing data scarcity, Evolutionary In-Context Learning operates in a Tabula Rasa setting to implicitly construct a selfgrowing repository of verified examples driven by syntactic utility.Finally, our Adversarial Execution-Guided Correction agent enforces fidelity through synergistic static critique and dynamic verification.Experiments demonstrate significant improvements over baselines in both accuracy and executability across diverse GQLs.The code is available at https: //github.com/nf202/Text2Graph. Fang Niu, Chaokun Wang, Hang Zhang 0032, Songyao Wang |
ACL (1) | 3 |
| 2026 | What Should I Cite? A RAG Benchmark for Academic Citation PredictionabstractWith the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation prediction aims to automatically suggest appropriate references, helping scholars navigate the expanding scientific literature. Here we present CiteRAG, the first comprehensive retrieval-augmented generation (RAG)-integrated benchmark for evaluating large language models on academic citation prediction, featuring a multi-level retrieval strategy, specialized retrievers, and generators. Our benchmark makes four core contributions: (1) We establish two instances of the citation prediction task with different granularity. Task 1 focuses on coarse-grained list-specific citation prediction, while Task 2 targets fine-grained position-specific citation prediction. To enhance these two tasks, we build a dataset containing 7,267 instances for Task 1 and 8,541 instances for Task 2, enabling comprehensive evaluation of both retrieval and generation. (2) We construct a three-level large-scale corpus with 554k papers spanning many major subfields, using an incremental pipeline. (3) We propose a multi-level hybrid RAG approach to citation prediction, fine-tuning embedding models with contrastive learning to capture complex citation relationships, paired with specialized generation models. (4) We conduct extensive experiments across state-of-the-art language models, including closed-source APIs, open-source models, and our fine-tuned generators, demonstrating the effectiveness of our framework. Our open-source toolkit enables reproducible evaluation and focuses on academic literature, providing the first comprehensive evaluation framework for citation prediction and serving as a methodological template for other scientific domains. Our source code and data are released at https://github.com/LQgdwind/CiteRAG. Leqi Zheng, Jiajun Zhang 0012, Canzhi Chen, Chaokun Wang, Hongwei Li 0032, Yuying Li 0006, Yaoxin Mao, Shannan Yan, Zixin Song, Zhiyuan Feng, Zhaolu Kang, Zirong Chen, Hang Zhang 0032, Qiang Liu 0006, Liang Wang 0001, Ziyang Liu 0004 |
WWW | 13 |
| 2026 | Training-Free and Unbiased Graph Collaborative Filtering for Personalized RecommendationsabstractWith the widespread adoption of collaborative filtering techniques for personalized recommendations, exposure bias has become a significant challenge.Exposure biasrefers to the tendency of recommendation models to disproportionately favor items with high exposure over those with low exposure. In graph collaborative filtering that uses graph neural networks (GNNs) for recommendations, exposure bias can be exacerbated due to 1) the reliance on positive feedback during graph construction and 2) the effects of the neighbor aggregation step in GNNs. To tackle this challenge, we propose a novel and efficient framework called FUGCF (training-Free andUnbiasedGraphCollaborativeFiltering) to improve both the accuracy and bias mitigation of graph-based personalized recommendations. FUGCF employs a two-stage calculation strategy: it estimates exposure probabilities in the first stage and then leverages them to help derive debiased node embeddings in the second stage. Furthermore, we design a training-free estimation method for FUGCF based on closed-form solutions to enhance its computational efficiency. The extensive experiments on a synthetic dataset and three real-world datasets demonstrate the effectiveness of FUGCF in reducing exposure bias, improving recommendation accuracy, and optimizing computational efficiency. Ziyang Liu 0004, Chaokun Wang, Cheng Wu 0004, Leqi Zheng, Hao Feng 0007, Hang Zhang 0032 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | PLForge: Enhancing Language Models for Natural Language to Procedural Extensions of SQLabstractProcedural Language extensions of SQL (abbr. PL/SQL) enhance database programming by integrating procedural constructs with SQL's declarative syntax, thereby improving the reusability, modularity, and maintainability of SQL. Besides, PL/SQL in database systems presents significant challenges in real-world development, primarily due to the inherent complexity of programming. To reduce the development difficulty of PL/SQL, this paper studies the novel task of translating natural language (NL) to PL/SQL (i.e., NL-to-PL/SQL), aimed at simplifying PL/SQL development. Recent advancements in language models have shown promise in translating natural language questions into SQL queries (i.e., Text-to-SQL). However, the state-of-the-art Text-to-SQL methods focus only on single SQL queries, neglecting the procedural extensions of SQL, which limits their effectiveness for the NL-to-PL/SQL task. In this paper, we propose PLForge, a suite of pre-trained language models with parameter configurations of 3B, 7B, and 15B, tailored for NL-to-PL/SQL tasks. To enhance the PL/SQL generation capabilities of PLForge, we leverage a curated PL/SQL-centric data corpus and employ an incremental pre-training approach. Furthermore, to fully exploit the potential of PLForge, we propose a comprehensive prompt construction strategy tailored specifically for PL/SQL. Given the scarcity of NL-to-PL/SQL datasets, we develop a template-based method for generating NL-to-PL/SQL data. We conduct a series of experiments on PLForge and several baseline models. Based on execution match and exact match metrics that are designed specifically for the NL-to-PL/SQL task, the experimental results demonstrate that PLForge outperforms existing models in both in-context learning and supervised fine-tuning settings. Hang Zhang 0032, Chaokun Wang, Hongwei Li 0032, Cheng Wu 0004, Songyao Wang, Yabin Liu, Gengyuan Shi, Ziyang Liu 0004 |
Proc. ACM Manag. Data | 1 |
| 2024 | CIVET: Exploring Compact Index for Variable-Length Subsequence Matching on Time SeriesabstractNowadays the demands for managing and analyzing substantially increasing collections of time series are becoming more challenging. Subsequence matching, as a core subroutine in time series analysis, has drawn significant research attention. Most of the previous works only focus on matching the subsequences with equal length to the query. However, many scenarios require support for efficient variable-length subsequence matching. In this paper, we propose a new representation, Uniform Piecewise Aggregate Approximation (UPAA) with the capability of aligning features for variable-length time series while remaining the lower bounding property. Based on UPAA, we present a compact index structure by grouping adjacent subsequences and similar subsequences respectively. Moreover, we propose an index pruning algorithm and a data filtering strategy to efficiently support variable-length subsequence matching without false dismissals. The experiments conducted on both real and synthetic datasets demonstrate that our approach achieves considerably better efficiency, scalability, and effectiveness than existing approaches. Haoran Xiong, Hang Zhang 0032, Zeyu Wang 0007, Zhenying He, Peng Wang 0027, Xiaoyang Sean Wang |
Proc. VLDB Endow. | 2 |