EDBT 2026 Demo / reviewers in the wild / expert
Zifeng Ding
dblp:283/5849
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0009-0000-9713-2701ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningabstractSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schuetze, Volker Tresp, Yunpu Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma 0001, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, Yunpu Ma |
ACL (1) | 5 |
| 2025 | BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context LearningabstractThis paper introduces BMIKE-53, a comprehensive benchmark for cross-lingual in-context knowledge editing (IKE) across 53 languages, unifying three knowledge editing (KE) datasets: zsRE, CounterFact, and WikiFactDiff.Crosslingual KE, which requires knowledge edited in one language to generalize across others while preserving unrelated knowledge, remains underexplored.To address this gap, we systematically evaluate IKE under zero-shot, oneshot, and few-shot setups, incorporating tailored metric-specific demonstrations.Our findings reveal that model scale and demonstration alignment critically govern cross-lingual IKE efficacy, with larger models and tailored demonstrations significantly improving performance.Linguistic properties, particularly script type, strongly influence performance variation across languages, with non-Latin languages underperforming due to issues like language confusion. Ercong Nie, Mingyang Wang 0003, Zifeng Ding, Helmut Schmid, Hinrich Schütze |
ACL (1) | 4 |
| 2025 | Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question AnsweringabstractRecent works integrating Knowledge Graphs (KGs) have shown promising improvements in enhancing the reasoning capabilities of Large Language Models (LLMs).However, existing benchmarks primarily focus on closedended tasks, leaving a gap in evaluating performance on more complex, real-world scenarios.This limitation also hinders a thorough assessment of KGs' potential to reduce hallucinations in LLMs.To address this, we introduce OKGQA 1 , a new benchmark specifically designed to evaluate LLMs augmented with KGs in open-ended, real-world question answering settings.OKGQA reflects practical complexities through diverse question types and incorporates metrics to quantify both hallucination rates and reasoning improvements in LLM+KG models.To consider the scenarios in which KGs may contain varying levels of errors, we propose a benchmark variant, OKGQA-P, to assess model performance when the semantics and structure of KGs are deliberately perturbed and contaminated.In this paper, we aims to (1) explore whether KGs can make LLMs more trustworthy in an open-ended setting, and (2) conduct a comparative analysis to shed light on method design.We believe this study can facilitate a more complete performance comparison and encourages continuous improvement in integrating KGs with LLMs to mitigate hallucination, and make LLMs more trustworthy. Yuan Sui 0001, Zifeng Ding, Bryan Hooi |
ACL (1) | 3 |
| 2025 | TCP: a Benchmark for Temporal Constraint-Based PlanningabstractTemporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity.To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark, that jointly assesses both capabilities.Each instance in TCP features a naturalistic dialogue around a collaborative project, where diverse and interdependent temporal constraints are explicitly or implicitly expressed, and models must infer an optimal schedule that satisfies all constraints.To construct TCP, we generate abstract problem prototypes that are then paired with realistic scenarios from various domains and enriched into dialogues using an LLM.A human quality check is performed on a sampled subset to confirm the reliability of our benchmark.We evaluate state-of-the-art LLMs and find that even the strongest models may struggle with TCP, highlighting its difficulty and revealing limitations in LLMs' temporal constraint-based planning abilities.We analyze underlying failure cases, open source our benchmark 1 , and hope our findings can inspire future research. Zifeng Ding, Sikuan Yan, Moy Yuan, Xianglong Hu, Fangru Lin, Andreas Vlachos 0001 |
EMNLP | 1 |
| 2025 | ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar ArgumentationabstractRetrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, yet suffers from critical limitations in high-stakes domains—namely, sensitivity to noisy or contradictory evidence and opaque, stochastic decision-making. We propose \textsc{ArgRAG}, an explainable, and contestable alternative that replaces black-box reasoning with structured inference using a Quantitative Bipolar Argumentation Framework (QBAF). \textsc{ArgRAG} constructs a QBAF from retrieved documents and performs deterministic reasoning under gradual semantics. This allows faithfully explanaining and contesting decisions. Evaluated on two fact verification benchmarks, PubHealth and RAGuard, \textsc{ArgRAG} achieves strong accuracy while significantly improving transparency. Yuqicheng Zhu, Nico Potyka, Daniel Hernández 0002, Yuan He 0008, Zifeng Ding, Bo Xiong 0001, Dongzhuoran Zhou, Evgeny Kharlamov, Steffen Staab |
NeSy | 5 |
| 2025 | AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the WebabstractTextual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation.Existing datasets for automated verification of image-text claims remain limited, as they often consist of synthetic claims and lack evidence annotations to capture the reasoning behind the verdict.In this work, we introduce AVerImaTeC, a dataset consisting of 1,297 real-world image-text claims. Each claim is annotated with question-answer (QA) pairs containing evidence from the web, reflecting a decomposed reasoning regarding the verdict.We mitigate common challenges in fact-checking datasets such as contextual dependence, temporal leakage, and evidence insufficiency, via claim normalization, temporally constrained evidence annotation, and a two-stage sufficiency check. We assess the consistency of the annotation in AVerImaTeC via inter-annotator studies, achieving a $\kappa=0.742$ on verdicts and $74.7\%$ consistency on QA pairs. We also propose a novel evaluation method for evidence retrieval and conduct extensive experiments to establish baselines for verifying image-text claims using open-web evidence. Zifeng Ding, Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
NeurIPS | 2 |
| 2025 | Image Token Matters: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent EditingabstractLarge Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiveness, we find that these models still hallucinate non-existent objects. We hypothesize that one reason is due to visual priors induced during training: when certain image tokens frequently co-occur in the same spatial regions and represent shared objects, they become strongly associated with the verbalizations of those objects. As a result, the model may hallucinate by evoking visually absent tokens that often co-occur with present ones. To test this assumption, we construct a co-occurrence graph of image tokens using a segmentation dataset and employ a Graph Neural Network (GNN) with contrastive learning followed by a clustering method to group tokens that frequently co-occur in similar visual contexts. We find that hallucinations predominantly correspond to clusters whose tokens dominate the input, and more specifically, that the visually absent tokens in those clusters show much higher correlation with hallucinated objects compared to tokens present in the image. Based on this observation, we propose a hallucination mitigation method that suppresses the influence of visually absent tokens by modifying latent image embeddings during generation. Experiments show our method reduces hallucinations while preserving expressivity. Weixing Wang 0005, Zifeng Ding, Jindong Gu, Christoph Meinel, Gerard de Melo, Haojin Yang 0001 |
NeurIPS | 2 |
| 2025 | Introducing FOReCAst: The Future Outcome Reasoning and Confidence Assessment BenchmarkabstractForecasting is an important task in many domains. However, existing forecasting benchmarks lack comprehensive confidence assessment, focusing on limited question types, and often consist of artificial questions that do not reflect real-world needs. To address these gaps, we introduce FOReCAst (Future Outcome Reasoning and Confidence Assessment), a benchmark that evaluates models' ability to make predictions and their confidence in them. FOReCAst spans diverse forecasting scenarios involving Boolean questions, timeframe prediction, and quantity estimation, enabling a comprehensive evaluation of both prediction accuracy and confidence calibration for real-world applications. Zhangdie Yuan, Zifeng Ding, Andreas Vlachos 0001 |
NeurIPS | 2 |
| 2024 | Text2Loc: 3D Point Cloud Localization from Natural LanguageabstractWe tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2 × over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/. Yan Xia 0003, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers |
CVPR | 3 |
| 2024 | zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language ModelsabstractZifeng Ding, Heling Cai, Jingpei Wu, Yunpu Ma, Ruotong Liao, Bo Xiong, Volker Tresp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zifeng Ding, Heling Cai, Jingpei Wu, Yunpu Ma, Ruotong Liao, Bo Xiong 0001, Volker Tresp |
NAACL-HLT | 1 |
| 2024 | A Unified Data Augmentation Framework for Low-Resource Multi-domain Dialogue Generation
Yongkang Liu 0002, Ercong Nie, Shi Feng 0001, Zifeng Ding, Daling Wang, Yifei Zhang 0003, Hinrich Schütze |
ECML/PKDD (2) | 5 |
| 2023 | Learning Meta-Representations of One-shot Relations for Temporal Knowledge Graph Link PredictionabstractFew-shot relational learning for static knowledge graphs (KGs) has drawn greater interest in recent years, while few-shot learning for temporal knowledge graphs (TKGs) has hardly been studied. Compared to KGs, TKGs contain rich temporal information, thus requiring temporal reasoning techniques for modeling. This poses a greater challenge in learning few-shot relations in the temporal context. In this paper, we follow the previous work that focuses on few-shot relational learning on static KGs and extend two fundamental TKG reasoning tasks, i.e., interpolated and extrapolated link prediction, to the one-shot setting. We propose four new large-scale benchmark datasets and develop a TKG reasoning model for learning one-shot relations in TKGs. Experimental results show that our model can achieve superior performance on all datasets in both TKG link prediction tasks. Zifeng Ding, Bailan He, Jingpei Wu, Yunpu Ma, Zhen Han 0003, Volker Tresp |
IJCNN | 1 |
| 2023 | Improving Few-Shot Inductive Learning on Temporal Knowledge Graphs Using Confidence-Augmented Reinforcement Learning
Zifeng Ding, Jingpei Wu, Zongyue Li, Yunpu Ma, Volker Tresp |
ECML/PKDD (3) | 1 |
| 2023 | ForecastTKGQuestions: A Benchmark for Temporal Question Answering and Forecasting over Temporal Knowledge Graphs
Zifeng Ding, Zongyue Li, Ruoxia Qi, Jingpei Wu, Bailan He, Yunpu Ma, Shuo Chen 0014, Ruotong Liao, Zhen Han 0003, Volker Tresp |
ISWC | 1 |
| 2021 | Learning Neural Ordinary Equations for Forecasting Future Links on Temporal Knowledge GraphsabstractThere has been an increasing interest in inferring future links on temporal knowledge graphs (KG).While links on temporal KGs vary continuously over time, the existing approaches model the temporal KGs in discrete state spaces.To this end, we propose a novel continuum model by extending the idea of neural ordinary differential equations (ODEs) to multi-relational graph convolutional networks.The proposed model preserves the continuous nature of dynamic multi-relational graph data and encodes both temporal and structural information into continuous-time dynamic embeddings.In addition, a novel graph transition layer is applied to capture the transitions on the dynamic graph, i.e., edge formation and dissolution.We perform extensive experiments on five benchmark datasets for temporal KG reasoning, showing our model's superior performance on the future link forecasting task. Zhen Han 0003, Zifeng Ding, Yunpu Ma, Yujia Gu, Volker Tresp |
EMNLP (1) | 2 |