Delvin Ce Zhang

dblp:97/919-4 · also Ce Zhang 0004 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0001-5571-9766ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 9 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms
abstract
Audio platforms have evolved beyond entertainment.They have become central to public discourse, from podcasts and radio to What-sApp voice notes and live streams.With millions of shows and hundreds of millions of listeners, audio platforms are now a major channel for misinformation.Yet existing factchecking pipelines are mostly designed for written claims, overlooking the unique properties of spoken media.We argue that audio misinformation is not merely textual content with transcripts: it is structurally different because it is both spoken-carrying persuasive force through prosody, pacing, and emotion-and conversational-unfolding across turns, speakers, and episodes.These dual properties introduce verification difficulties that traditional methods rarely face.This position paper synthesizes evidence across modalities and platforms, examines datasets and methods, and highlights why existing pipelines fail on audio.We argue that advancing fact-checking requires rethinking verification pipelines around the spoken and conversational realities of audio.
Chaewan Chun, Delvin Ce Zhang, Dongwon Lee 0001
ACL (1)2
2026 PromptHG: Prompt-Enhanced Heterogeneous Graph for Personalized News Recommendation
Hai-Dang Kieu, Delvin Ce Zhang, Qiang Wu 0001, Min Xu 0001, Dung D. Le
ECIR (1)2
2025 Unmasking Fake Careers: Detecting Machine-Generated Career Trajectories via Multi-layer Heterogeneous Graphs
abstract
The rapid advancement of Large Language Models (LLMs) has enabled the generation of highly realistic synthetic data.We identify a new vulnerability, LLMs generating convincing career trajectories in fake resumes and explore effective detection methods.To address this challenge, we construct a dataset of machinegenerated career trajectories using LLMs and various methods, and demonstrate that conventional text-based detectors perform poorly on structured career data.We propose Career-Scape, a novel heterogeneous, hierarchical multi-layer graph framework that models career entities and their relations in a unified global graph built from genuine resumes.Unlike conventional classifiers that treat each instance independently, CareerScape employs a structure-aware framework that augments user-specific subgraphs with trusted neighborhood information from a global graph, enabling the model to capture both global structural patterns and local inconsistencies indicative of synthetic career paths.Experimental results show that CareerScape outperforms state-of-the-art baselines by 5.8-85.0%relatively, highlighting the importance of structureaware detection for machine-generated content.Our codebase is available at https://github. com/mickeymst/careerscape.
Michiharu Yamashita, Thanh Tran 0005, Delvin Ce Zhang, Dongwon Lee 0001
EMNLP3
2025 SUA: Stealthy Multimodal Large Language Model Unlearning Attack
abstract
Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing serious privacy risks.To mitigate this, MLLM unlearning methods are proposed, which finetune MLLMs to forget sensitive information.However, it remains unclear whether the knowledge has been truly forgotten or just hidden in the model.Therefore, we propose to study a novel problem of MLLM unlearning attack, which aims to recover the unlearned knowledge of an unlearned MLLM.To achieve the goal, we propose a novel framework-Stealthy Unlearning Attack (SUA)-that learns a universal noise pattern.When applied to input images, this noise can trigger the model to reveal unlearned content.While pixel-level perturbations may be visually subtle, they can be detected in the semantic embedding space, making such attacks vulnerable to potential defenses.To improve stealthiness, we introduce an embedding alignment loss that minimizes the difference between the perturbed and denoised image embeddings, ensuring that the attack remains semantically unnoticeable.Experimental results show that SUA can effectively recover unlearned information from MLLMs.Furthermore, the learned noise generalizes well-i.e., a single perturbation trained on a few samples can reveal forgotten contents in unseen images.
Xianren Zhang, Hui Liu 0033, Delvin Ce Zhang, Xianfeng Tang, Qi He 0002, Dongwon Lee 0001, Suhang Wang
EMNLP3
2025 Retrieval-Augmented Language Model for Knowledge-aware Protein Encoding
abstract
Protein language models often struggle to capture biological functions due to their lack of factual knowledge (e.g., gene descriptions). Existing solutions leverage protein knowledge graphs (PKGs) as auxiliary pre-training objectives, but lack explicit integration of task-oriented knowledge, making them suffer from limited knowledge exploitation and catastrophic forgetting. The root cause is that they fail to align PKGs with task-specific data, forcing their knowledge modeling to adapt to the knowledge-isolated nature of downstream tasks. In this paper, we propose Knowledge-aware retrieval augmented protein language model (Kara), achieving the first task-oriented and explicit integration of PKGs and protein language models. With a knowledge retriever learning to predict linkages between PKG and task proteins, Kara unifies the knowledge integration of the pre-training and fine-tuning stages with a structure-based regularization, mitigating catastrophic forgetting. To ensure task-oriented integration, Kara uses contextualized virtual tokens to extract graph context as task-specific knowledge for new proteins. Experiments show that Kara outperforms existing knowledge-enhanced models in 6 representative tasks, achieving on average 5.1% improvements.
Delvin Ce Zhang, Shuang Liang 0002, Zhengpin Li, Rex Ying, Jie Shao 0001
ICML2
2025 CORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
abstract
Delvin Ce Zhang, Dongwon Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Delvin Ce Zhang, Dongwon Lee 0001
NAACL (Long Papers)1
2024 Topic Modeling on Document Networks with Dirichlet Optimal Transport Barycenter
abstract
Texts are often interconnected in a network structure, e.g., academic papers via citations. On the one hand, though Graph Neural Networks (GNNs) have shown promising ability to derive effective embeddings for networked documents, they do not assume latent topics, resulting in uninterpretahle embeddings. On the other hand, topic models can infer interpretable document representations. However, most topic models focus on plain text and fail to leverage network structure across documents. In this paper, we propose a GNN-based topic model that both captures network connection and derives semantically interpretable text representations. For network modeling, we build our model with Optimal Transport Barycenter. For semantic interpretability, we extend optimal transport with pre-trained word embeddings.
Delvin Ce Zhang, Hady Wirawan Lauw
ICDE1
2024 Hypformer: Exploring Efficient Transformer Fully in Hyperbolic Space
abstract
Hyperbolic geometry have shown significant potential in modeling complex structured data, particularly those with underlying tree-like and hierarchical structures. Despite the impressive performance of various hyperbolic neural networks across numerous domains, research on adapting the Transformer to hyperbolic space remains limited. Previous attempts have mainly focused on modifying self-attention modules in the Transformer. However, these efforts have fallen short of developing a complete hyperbolic Transformer. This stems primarily from: (i) the absence of well-defined modules in hyperbolic space, including linear transformation layers, LayerNorm layers, activation functions, dropout operations, etc. (ii) the quadratic time complexity of the existing hyperbolic self-attention module w.r.t the number of input tokens, which hinders its scalability. To address these challenges, we propose, Hypformer, a novel hyperbolic Transformer based on the Lorentz model of hyperbolic geometry. In Hypformer, we introduce two foundational blocks that define the essential modules of the Transformer in hyperbolic space. Furthermore, we develop a linear self-attention mechanism in hyperbolic space, enabling hyperbolic Transformer to process billion-scale graph data and long-sequence inputs for the first time. Our experimental results confirm the effectiveness and efficiency of \method across various datasets, demonstrating its potential as an effective and scalable solution for large-scale data representation and large models.
Menglin Yang 0001, Harshit Verma, Delvin Ce Zhang, Jiahong Liu 0001, Irwin King, Rex Ying
KDD3
2024 Topic Modeling on Document Networks With Dirichlet Optimal Transport Barycenter
abstract
Text documents are often interconnected in a network structure, e.g., academic papers via citations, Web pages via hyperlinks. On the one hand, though Graph Neural Networks (GNNs) have shown promising ability to derive effective embeddings for such networked documents, they do not assume a latent topic structure and result inuninterpretableembeddings. On the other hand, topic models can infer semantically interpretable topic distributions for documents by associating each topic with a group of understandable key words. However, most topic models mainly focus on plain text within documents and fail to leveragenetwork structureacross documents. Network connectivity reveals topic similarity between linked documents, and modeling it could uncover meaningful semantics. Motivated by above two challenges, in this paper, we propose a GNN-based neural topic model that both captures network connectivity and derives semantically interpretable topic distributions for networked documents. For network modeling, we build the model based on the theory of Optimal Transport Barycenter, which captures network structure by allowing the topic distribution of a document to generate the content of its linked neighbors. For semantic interpretability, we extend optimal transport by incorporating semantically related words in the embedding space. Since Dirichlet prior in Latent Dirichlet Allocation successfully improves topic quality, we also analyze Dirichlet as an optimal transport prior distribution to improve topic interpretability. We design rejection sampling to simulate Dirichlet distribution. Extensive experiments on document classification, clustering, link prediction, and topic analysis verify the effectiveness of our model.
Delvin Ce Zhang, Hady Wirawan Lauw
IEEE Trans. Knowl. Data Eng.1
2023 Hyperbolic Graph Topic Modeling Network with Continuously Updated Topic Tree
abstract
Connectivity across documents often exhibits a hierarchical network structure. Hyperbolic Graph Neural Networks (HGNNs) have shown promise in preserving network hierarchy. However, they do not model the notion of topics, thus document representations lack semantic interpretability. On the other hand, a corpus of documents usually has high variability in degrees of topic specificity. For example, some documents contain general content (e.g., sports), while others focus on specific themes (e.g., basketball and swimming). Topic models indeed model latent topics for semantic interpretability, but most assume a flat topic structure and ignore such semantic hierarchy. Given these two challenges, we propose a Hyperbolic Graph Topic Modeling Network to integrate both network hierarchy across linked documents and semantic hierarchy within texts into a unified HGNN framework. Specifically, we construct a two-layer document graph. Intra- and cross-layer encoding captures network hierarchy. We design a topic tree for text decoding to preserve semantic hierarchy and learn interpretable topics. Supervised and unsupervised experiments verify the effectiveness of our model.
Delvin Ce Zhang, Rex Ying, Hady Wirawan Lauw
KDD1
2022 Dynamic Topic Models for Temporal Document Networks
abstract
Dynamic topic models explore the time evolution of topics in temporally accumulative corpora. While existing topic models focus on the dynamics of individual documents, we propose two neural topic models aimed at learning unified topic distributions that incorporate both document dynamics and network structure. For the first model, by adding a time dimension, we propose Time-Aware Optimal Transport, which measures the probability of a link between two differently timestamped documents using their semantic distance. Since the gradually evolving topological structure of network may also influence the establishment of a new link, for the second model, we further design a Temporal Point Process to capture the impact of historical neighbors on the current link formation at the network level. Experiments on four dynamic document networks demonstrate the advantage of our models in jointly modeling document dynamics and network adjacency.
Delvin Ce Zhang, Hady Wirawan Lauw
ICML1
2022 Variational Graph Author Topic Modeling
abstract
While Variational Graph Auto-Encoder (VGAE) has presented promising ability to learn representations for documents, most existing VGAE methods do not model a latent topic structure and therefore lack semantic interpretability. Exploring hidden topics within documents and discovering key words associated with each topic allow us to develop a semantic interpretation of the corpus. Moreover, documents are usually associated with authors. For example, news reports have journalists specializing in writing certain type of events, academic papers have authors with expertise in certain research topics, etc. Modeling authorship information could benefit topic modeling, since documents by the same authors tend to reveal similar semantics. This observation also holds for documents published on the same venues. However, most topic models ignore the auxiliary authorship and publication venues. Given above two challenges, we propose a Variational Graph Author Topic Model for documents to integrate both semantic interpretability and authorship and venue modeling into a unified VGAE framework. For authorship and venue modeling, we construct a hierarchical multi-layered document graph with both intra- and cross-layer topic propagation. For semantic interpretability, three word relations (contextual, syntactic, semantic) are modeled and constitute three word sub-layers in the document graph. We further propose three alternatives for variational divergence. Experiments verify the effectiveness of our model on supervised and unsupervised tasks.
Delvin Ce Zhang, Hady Wirawan Lauw
KDD1
2022 Meta-Complementing the Semantics of Short Texts in Neural Topic Models
abstract
Topic models infer latent topic distributions based on observed word co-occurrences in a text corpus. While typically a corpus contains documents of variable lengths, most previous topic models treat documents of different lengths uniformly, assuming that each document is sufficiently informative. However, shorter documents may have only a few word co-occurrences, resulting in inferior topic quality. Some other previous works assume that all documents are short, and leverage external auxiliary data, e.g., pretrained word embeddings and document connectivity. Orthogonal to existing works, we remedy this problem within the corpus itself by proposing a Meta-Complement Topic Model, which improves topic quality of short texts by transferring the semantic knowledge learned on long documents to complement semantically limited short texts. As a self-contained module, our framework is agnostic to auxiliary data and can be further improved by flexibly integrating them into our framework. Specifically, when incorporating document connectivity, we further extend our framework to complement documents with limited edges. Experiments demonstrate the advantage of our framework.
Delvin Ce Zhang, Hady Wirawan Lauw
NeurIPS1
2021 Topic Modeling for Multi-Aspect Listwise Comparisons
abstract
As a well-established probabilistic method, topic models seek to uncover latent semantics from plain text. In addition to having textual content, we observe that documents are usually compared in listwise rankings based on their content. For instance, world-wide countries are compared in an international ranking in terms of electricity production based on their national reports. Such document comparisons constitute additional information that reveal documents' relative similarities. Incorporating them into topic modeling could yield comparative topics that help to differentiate and rank documents. Furthermore, based on different comparison criteria, the observed document comparisons usually cover multiple aspects, each expressing a distinct ranked list. For example, a country may be ranked higher in terms of electricity production, but fall behind others in terms of life expectancy or government budget. Each comparison criterion, or aspect, observes a distinct ranking. Considering such multiple aspects of comparisons based on different ranking criteria allows us to derive one set of topics that inform heterogeneous document similarities. We propose a generative topic model aimed at learning topics that are well aligned to multi-aspect listwise comparisons. Experiments on public datasets demonstrate the advantage of the proposed method in jointly modeling topics and ranked lists against baselines comprehensively.
Delvin Ce Zhang, Hady Wirawan Lauw
CIKM1
2021 Representation Learning on Multi-layered Heterogeneous Network
Delvin Ce Zhang, Hady Wirawan Lauw
ECML/PKDD (2)1
2021 Semi-supervised Semantic Visualization for Networked Documents
Delvin Ce Zhang, Hady Wirawan Lauw
ECML/PKDD (3)1
2020 Topic Modeling on Document Networks with Adjacent-Encoder
abstract
Oftentimes documents are linked to one another in a network structure,e.g., academic papers cite other papers, Web pages link to other pages. In this paper we propose a holistic topic model to learn meaningful and unified low-dimensional representations for networked documents that seek to preserve both textual content and network structure. On the basis of reconstructing not only the input document but also its adjacent neighbors, we develop two neural encoder architectures. Adjacent-Encoder, or AdjEnc, induces competition among documents for topic propagation, and reconstruction among neighbors for semantic capture. Adjacent-Encoder-X, or AdjEnc-X, extends this to also encode the network structure in addition to document content. We evaluate our models on real-world document networks quantitatively and qualitatively, outperforming comparable baselines comprehensively.
Delvin Ce Zhang, Hady Wirawan Lauw
AAAI1