EDBT 2026 Demo / reviewers in the wild / expert
Zhen Tan 0001
dblp:13/10345-1
· DBLP profile ↗
16ranked-venue papers in the field
5as first author
16since 2021 · last 2026
0009-0006-9548-2330ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 12 (5 first)Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explaining the 'Unexplainable' Large Language ModelsabstractThe integration of Large Language Models (LLMs) into critical societal and scientific functions has intensified the urgent demand for transparency, reliability, and trust. While post-hoc attribution methods and Chain-of-Thought reasoning currently serve as the dominant approaches to explainability, growing evidence shows that they are often unreliable, producing brittle, misleading, or illusory explanations that fail to reflect true model behavior. This tutorial aims to unpack why these limitations arise. We first establish the theoretical intractability of complete mechanistic explanations for modern LLMs and clarify the intrinsic barriers to achieving full transparency in overparameterized models. We then pivot to a principled alternative: user-centric explainability, with a focus on concept-based interpretability and controlled data attribution. We review the theoretical foundations of these methods and survey their modern extensions that enable comprehensive explanation, inference-time intervention, and model editability. Finally, we demonstrate how such approaches support effective human--AI collaboration in high-stakes scientific and decision-critical applications. By synthesizing foundational theory, critical analysis of existing methods, and emerging techniques, this tutorial offers a coherent framework for developing the next generation of explainable and trustworthy AI systems. Zhen Tan 0001, Song Wang 0013, Tianlong Chen 0001, Jing Ma 0002, Jundong Li, Huan Liu 0001 |
WSDM | 1 |
| 2025 | Can LLMs Improve Multimodal Fact-Checking by Asking Relevant Questions?
Alimohammad Beigi, Bohan Jiang, Dawei Li 0008, Zhen Tan 0001, Pouya Shaeri, Tharindu Kumarage, Amrita Bhattacharjee, Huan Liu 0001 |
IEEE Big Data | 4 |
| 2025 | Ontology-Aware RAG for Improved Question-Answering in Cybersecurity Education
Chengshuai Zhao, Garima Agrawal, Tharindu Kumarage, Zhen Tan 0001, Yuli Deng, Ying-Chih Chen, Huan Liu 0001 |
IEEE Big Data | 5 |
| 2025 | Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety BenchmarkabstractConstruction safety research is a critical field in civil engineering, aiming to mitigate risks and prevent injuries through the analysis of site conditions and human factors. However, the limited volume and lack of diversity in existing construction safety datasets pose significant challenges to conducting in-depth analyses. To address this research gap, this paper introduces the Construction Safety Dataset (CSDataset), a well-organized comprehensive multi-level dataset that encompasses incidents, inspections, and violations recorded sourced from the Occupational Safety and Health Administration (OSHA). This dataset uniquely integrates structured attributes with unstructured narratives, facilitating a wide range of approaches driven by machine learning and large language models. We also conduct a preliminary approach benchmarking and various cross-level analyses using our dataset, offering insights to inform and enhance future efforts in construction safety. For example, we found that complaint-driven inspections were associated with a 17.3% reduction in the likelihood of subsequent incidents. Our dataset and code are released at https://github.com/zhenhuiou/Construction-Safety-Dataset-CSDataset. Zhenhui Ou, Dawei Li 0008, Zhen Tan 0001, Huan Liu 0001, Siyuan Song |
CIKM | 3 |
| 2025 | GraphRCG: Self-Conditioned Graph GenerationabstractGraph generation aims to create new graphs that closely align with a target graph distribution. Existing works often implicitly capture this distribution by aligning the output of a generator with each training sample. As such, the overview of the entire distribution is not explicitly captured and used for graph generation. In contrast, in this work, we propose a novel self-conditioned graph generation framework designed to explicitly model graph distributions and employ these distributions to guide the generation process. We first perform self-conditioned modeling to capture the graph distributions by transforming each graph sample into a low-dimensional representation and optimizing a representation generator to create new representations reflective of the learned distribution. Subsequently, we leverage these bootstrapped representations as self-conditioned guidance for the generation process, thereby facilitating the generation of graphs that more accurately reflect the learned distributions. We conduct extensive experiments on generic and molecular graph datasets. Our framework, GraphRCG, demonstrates superior performance over existing state-of-the-art graph generation methods in terms of graph quality and fidelity to training data. Song Wang 0013, Zhen Tan 0001, Tianlong Chen 0001, Huan Liu 0001, Jundong Li |
CIKM | 2 |
| 2025 | MerRec: A Large-scale Multipurpose Mercari Dataset for Consumer-to-Consumer Recommendation Systems
Lichi Li, Zain ul Abi Din, Zhen Tan 0001, Sam London, Tianlong Chen 0001, Ajay H. Daptardar |
KDD (1) | 3 |
| 2025 | SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents
Dawei Li 0008, Zhen Tan 0001, Peijia Qian, Kumar Satvik Chaudhary, Lijie Hu |
PAKDD (3) | 2 |
| 2024 | Model Attribution in LLM-Generated Disinformation: A Domain Generalization Approach with Supervised Contrastive LearningabstractModel attribution for LLM-generated disinformation poses a significant challenge in understanding its origins and mitigating its spread. This task is especially challenging because modern large language models (LLMs) produce disinformation with human-like quality. Additionally, the diversity in prompting methods used to generate disinformation complicates accurate source attribution. These methods introduce domain-specific features that can mask the fundamental characteristics of the models. In this paper, we introduce the concept of model attribution as a domain generalization problem, where each prompting method represents a unique domain. We argue that an effective attribution model must be invariant to these domain-specific features. It should also be proficient in identifying the originating models across all scenarios, reflecting real-world detection challenges. To address this, we introduce a novel approach based on Supervised Contrastive Learning. This method is designed to enhance the model's robustness to variations in prompts and focuses on distinguishing between different source LLMs. We evaluate our model through rigorous experiments involving three common prompting methods: “open-ended”, “rewriting”, and “paraphrasing”, and three advanced LLMs: “llama 2”, “chatgpt”, and “vicuna”. Our results demonstrate the effectiveness of our approach in model attribution tasks, achieving state-of-the-art performance across diverse and unseen datasets. Alimohammad Beigi, Zhen Tan 0001, Nivedh Mudiam, Canyu Chen, Kai Shu, Huan Liu 0001 |
DSAA | 2 |
| 2024 | Media Bias Matters: Understanding the Impact of Politically Biased News on Vaccine Attitudes in Social MediaabstractNews media has been frequently utilized as a political tool to stray from facts, making biased statements and claims without evidence. During the COVID-19 vaccine campaign, politically biased news (PBN) has significantly undermined public trust in vaccines. Despite medical evidence showing the benefits of these vaccines, the misperceptions of the vaccine's safety, risks, and efficacy have led to a non-negligible fraction of the population resistant to receiving the vaccine. In this paper, we analyze: (i) how inherent vaccine stances subtly influence individuals' selection of news sources and participation in social media discussions; and (ii) the impact of exposure to PBN on users' attitudes toward vaccines. In doing so, we first curate a comprehensive dataset that connects PBN with related social media discourse. Utilizing advanced deep learning and causal inference techniques, we reveal distinct user behaviors between social media groups with various vaccine stances. Moreover, we observe that individuals with moderate stances, particularly the vaccine-hesitant majority, are more vulnerable to the influence of PBN compared to those with extreme views. Our findings provide critical insights to foster this line of research. Bohan Jiang, Lu Cheng 0001, Zhen Tan 0001, Ruocheng Guo, Huan Liu 0001 |
DSAA | 3 |
| 2024 | Interpreting Pretrained Language Models via Concept Bottlenecks
Zhen Tan 0001, Lu Cheng 0001, Song Wang 0013, Bo Yuan 0017, Jundong Li, Huan Liu 0001 |
PAKDD (3) | 1 |
| 2024 | Disinformation Detection: An Evolving Challenge in the Age of LLMsabstractThe advent of generative Large Language Models (LLMs) such as ChatGPT has catalyzed transformative advancements across multiple domains. However, alongside these advancements, they have also introduced potential threats. One critical concern is the misuse of LLMs by disinformation spreaders, leveraging these models to generate highly persuasive yet misleading content that challenges the disinformation detection system. This work aims to address this issue by answering three research questions: (1) To what extent can the current disinformation detection technique reliably detect LLM-generated disinformation? (2) If traditional techniques prove less effective, can LLMs themself be exploited to serve as a robust defense against advanced disinformation? and, (3) Should both these strategies falter, what novel approaches can be proposed to counter this burgeoning threat effectively? A holistic exploration for the formation and detection of disinformation is conducted to foster this line of research. Bohan Jiang, Zhen Tan 0001, Ayushi Nirmal, Huan Liu 0001 |
SDM | 2 |
| 2024 | Label Distribution Learning-Enhanced Dual-KNN for Text ClassificationabstractMany text classification methods usually introduce external information (e.g., label descriptions and knowledge bases) to improve the classification performance. Compared to external information, some internal information generated by the model itself during training, like text embeddings and predicted label probability distributions, are exploited poorly when predicting the outcomes of some texts. In this paper, we focus on leveraging this internal information, proposing a dual k nearest neighbor (DkNN) framework with two kNN modules, to retrieve several neighbors from the training set and augment the distribution of labels. For the kNN module, it is easily confused and may cause incorrect predictions when retrieving some nearest neighbors from noisy datasets (datasets with labeling errors) or similar datasets (datasets with similar labels). To address this issue, we also introduce a label distribution learning module that can learn label similarity, and generate a better label distribution to help models distinguish texts more effectively. This module eases model overfitting and improves final classification performance, hence enhancing the quality of the retrieved neighbors by kNN modules during inference. Extensive experiments on the benchmark datasets verify the effectiveness of our method. Bo Yuan 0017, Zhen Tan 0001, Huan Liu 0001 |
SDM | 3 |
| 2023 | Virtual Node Tuning for Few-shot Node Classification
Zhen Tan 0001, Ruocheng Guo, Kaize Ding, Huan Liu 0001 |
KDD | 1 |
| 2023 | Contrastive Meta-Learning for Few-shot Node ClassificationabstractFew-shot node classification, which aims to predict labels for nodes on graphs with only limited labeled nodes as references, is of great significance in real-world graph mining tasks. To tackle such a label shortage issue, existing works generally leverage the meta-learning framework, which utilizes a number of episodes to extract transferable knowledge from classes with abundant labeled nodes and generalizes the knowledge to other classes with limited labeled nodes. In essence, the primary aim of few-shot node classification is to learn node embeddings that are generalizable across different classes. To accomplish this, the GNN encoder must be able to distinguish node embeddings between different classes, while also aligning embeddings for nodes in the same class. Thus, in this work, we propose to consider both the intra-class and inter-class generalizability of the model. We create a novel contrastive meta-learning framework on graphs, named COSMIC, with two key designs. First, we propose to enhance the intra-class generalizability by involving a contrastive two-step optimization in each episode to explicitly align node embeddings in the same classes. Second, we strengthen the inter-class generalizability by generating hard node classes for classification via a novel similarity-sensitive mix-up strategy. Extensive experiments on prevalent few-shot node classification datasets verify the effectiveness of our framework and demonstrate its superiority over other state-of-the-art baselines. Song Wang 0013, Zhen Tan 0001, Huan Liu 0001, Jundong Li |
KDD | 2 |
| 2022 | Supervised Graph Contrastive Learning for Few-Shot Node Classification
Zhen Tan 0001, Kaize Ding, Ruocheng Guo, Huan Liu 0001 |
ECML/PKDD (2) | 1 |
| 2022 | Graph Few-shot Class-incremental LearningabstractThe ability to incrementally learn new classes is vital to all real-world artificial intelligence systems. A large portion of high-impact applications like social media, recommendation systems, E-commerce platforms, etc. can be represented by graph models. In this paper, we investigate the challenging yet practical problem,Graph Few-shot Class-incremental (Graph FCL) problem, where the graph model is tasked to classify both newly encountered classes and previously learned classes. Towards that purpose, we put forward a Graph Pseudo Incremental Learning paradigm by sampling tasks recurrently from the base classes, so as to produce an arbitrary number of training episodes for our model to practice the incremental learning skill. Furthermore, we design a Hierarchical-Attention-based Graph Meta-learning framework, HAG-Meta from an optimization perspective. We present a task-sensitive regularizer calculated from task-level attention and node class prototypes to mitigate overfitting onto either novel or base classes. To employ the topological knowledge, we add a node-level attention module to adjust the prototype representation. Our model not only achieves greater stability of old knowledge consolidation, but also acquires advantageous adaptability to new knowledge with very limited data samples. Extensive experiments on three real-world datasets, including Amazon-clothing, Reddit, and DBLP, show that our framework demonstrates remarkable advantages in comparison with the baseline and other related state-of-the-art methods. Zhen Tan 0001, Kaize Ding, Ruocheng Guo, Huan Liu 0001 |
WSDM | 1 |