EDBT 2026 Demo / reviewers in the wild / expert
Kewei Cheng
dblp:175/1247
· DBLP profile ↗
15ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-2139-3451ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Framework for Rule Learning: Integrating Commonsense Knowledge from LLMs with Structured Knowledge from Knowledge Graphs
Qirui Hao, Kewei Cheng, Tongze Zhang, Hongyuan Liu 0006, Junming Shao, Carl Yang 0001 |
WWW | 2 |
| 2025 | Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-TrainingabstractYuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014 |
NAACL (Long Papers) | 5 |
| 2024 | Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the TextabstractAlthough Large Language Models (LLMs) excel at addressing straightforward reasoning tasks, they frequently struggle with difficulties when confronted by more complex multi-step reasoning due to a range of factors.Firstly, natural language often encompasses complex relationships among entities, making it challenging to maintain a clear reasoning chain over longer spans.Secondly, the abundance of linguistic diversity means that the same entities and relationships can be expressed using different terminologies and structures, complicating the task of identifying and establishing connections between multiple pieces of information.Graphs provide an effective solution to represent data rich in relational information and capture long-term dependencies among entities.To harness the potential of graphs, our paper introduces Structure Guided Prompt, an innovative three-stage task-agnostic prompting framework designed to improve the multi-step reasoning capabilities of LLMs in a zero-shot setting.This framework explicitly converts unstructured text into a graph via LLMs and instructs them to navigate this graph using taskspecific strategies to formulate responses.By effectively organizing information and guiding navigation, it enables LLMs to provide more accurate and context-aware responses.Our experiments show that this framework significantly enhances the reasoning capabilities of LLMs, enabling them to excel in a broader spectrum of natural language scenarios. Kewei Cheng, Nesreen K. Ahmed, Theodore L. Willke, Yizhou Sun |
EMNLP | 1 |
| 2024 | Inductive Meta-Path Learning for Schema-Complex Heterogeneous Information NetworksabstractHeterogeneous Information Networks (HINs) are information networks with multiple types of nodes and edges. The concept of meta-path, i.e., a sequence of entity types and relation types connecting two entities, is proposed to provide the meta-level explainable semantics for various HIN tasks. Traditionally, meta-paths are primarily used for schema-simple HINs, e.g., bibliographic networks with only a few entity types, where meta-paths are often enumerated with domain knowledge. However, the adoption of meta-paths for schema-complex HINs, such as knowledge bases (KBs) with hundreds of entity and relation types, has been limited due to the computational complexity associated with meta-path enumeration. Additionally, effectively assessing meta-paths requires enumerating relevant path instances, which adds further complexity to the meta-path learning process. To address these challenges, we propose SchemaWalk, an inductive meta-path learning framework for schema-complex HINs. We represent meta-paths with schema-level representations to support the learning of the scores of meta-paths for varying relations, mitigating the need of exhaustive path instance enumeration for each relation. Further, we design a reinforcement-learning based path-finding agent, which directly navigates the network schema (i.e., schema graph) to learn policies for establishing meta-paths with high coverage and confidence for multiple relations. Extensive experiments on real data sets demonstrate the effectiveness of our proposed paradigm. Changjun Fan, Kewei Cheng, Peng Cui 0001, Yizhou Sun, Zhong Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Neural-Symbolic Methods for Knowledge Graph Reasoning: A SurveyabstractNeural symbolic knowledge graph (KG) reasoning offers a promising approach that combines the expressive power of symbolic reasoning with the learning capabilities inherent in neural networks. This survey provides a comprehensive overview of advancements, techniques, and challenges in the field of neural symbolic KG reasoning. The survey introduces the fundamental concepts of KGs and symbolic logic, followed by an exploration of three significant KG reasoning tasks: KG completion, complex query answering, and logical rule learning. For each task, we thoroughly discuss three distinct categories of methods: pure symbolic methods, pure neural approaches, and the integration of neural networks and symbolic reasoning methods known as neural-symbolic. We carefully analyze and compare the strengths and limitations of each category of methods to provide a comprehensive understanding. By synthesizing recent research contributions and identifying open research directions, this survey aims to equip researchers and practitioners with a comprehensive understanding of the state-of-the-art in neural symbolic KG reasoning, fostering future advancements in this interdisciplinary domain. Kewei Cheng, Nesreen K. Ahmed, Ryan Rossi, Theodore L. Willke, Yizhou Sun |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Neural Compositional Rule Learning for Knowledge Graph Reasoning
Kewei Cheng, Nesreen K. Ahmed, Yizhou Sun |
ICLR | 1 |
| 2023 | Multi-Instrument Flood Monitoring With a Distributed, Decentralized, Dynamic and Context-Aware Satellite Sensor WebabstractThis work explores a new concept of operations for observation of Earth events by a satellite sensor web that is able to reason about its capabilities and plan observations based on detected or requested events on Earth’s surface. An intelligent agile satellite sensor web is shown to produce more than twice the number of observations of a nadir-looking sensor web. Ben Gorr 0001, Alan Aguilar Jaramillo, Zida Wu, Wooyeong Cho, Kewei Cheng, Molly K. Stroud, Vinay Ravindra, Cédric H. David, Huilin Gao, Yizhou Sun, Ankur Mehta, George H. Allen, Daniel Selva |
IGARSS | 5 |
| 2023 | Decentralized Market-Based Observation Assignment Strategy for Dynamic Networks in Sensor Web Mission ConceptsabstractMonitoring of short-lived and highly-dynamic processes and events such as floods or forest fires has gained an increasing interest in Earth Observation, particularly as global climate change is affecting these processes. The observation of such dynamic events is often limited by the response time of human operation of Earth-Observing satellites or UAVs.To address this bottleneck, this paper presents a Modified Asynchronous Consensus Constraint-Based Bundle Algorithm (MACCBBA) for observation task allocation in Sensor Web mission concepts for Earth Observation. This algorithm allows for the decentralized allocation of observation tasks amongst a network of Satellites and UAVs based on recently measured data processed on board, or on messages received from other sensors or from the ground. This algorithm is also capable of reaching a feasible plan in a dynamic communications network such as the ones present in some Sensor Web mission concepts and allows for complex temporal constraints and dependencies between tasks to model the value of near-simultaneous co-observations by complementary or synergistic sensors. Alan Aguilar Jaramillo, Ben Gorr 0001, Vinay Ravindra, Cédric H. David, Molly K. Stroud, Ankur Mehta, George H. Allen, Wooyeong Cho, Kewei Cheng, Huilin Gao, Yizhou Sun, Zida Wu, Daniel Selva |
IGARSS | 9 |
| 2022 | RLogic: Recursive Logical Rule Learning from Knowledge GraphsabstractLogical rules are widely used to represent domain knowledge and hypothesis, which is fundamental to symbolic reasoning-based human intelligence. Very recently, it has been demonstrated that integrating logical rules into regular learning tasks can further enhance learning performance in a label-efficient manner. Many attempts have been made to learn logical rules automatically from knowledge graphs (KGs). However, a majority of existing methods entirely rely on observed rule instances to define the score function for rule evaluation and thus lack generalization ability and suffer from severe computational inefficiency. Instead of completely relying on rule instances for rule evaluation, RLogic defines a predicate representation learning-based scoring model, which is trained by sampled rule instances. In addition, RLogic incorporates one of the most significant properties of logical rules, the deductive nature, into rule learning, which is critical especially when a rule lacks supporting evidence. To push deductive reasoning deeper into rule learning, RLogic breaks a big sequential model into small atomic models in a recursive way. Extensive experiments have demonstrated that RLogic is superior to existing state-of-the-art algorithms in terms of both efficiency and effectiveness. Kewei Cheng, Wei Wang 0010, Yizhou Sun |
KDD | 1 |
| 2022 | PGE: Robust Product Graph Embedding Learning for Error DetectionabstractAlthough product graphs (PGs) have gained increasing attentions in recent years for their successful applications in product search and recommendations, the extensive power of PGs can be limited by the inevitable involvement of various kinds of errors. Thus, it is critical to validate the correctness of triples in PGs to improve their reliability. Knowledge graph (KG) embedding methods have strong error detection abilities. Yet, existing KG embedding methods may not be directly applicable to a PG due to its distinct characteristics: (1) PG contains rich textual signals, which necessitates a joint exploration of both text information and graph structure; (2) PG contains a large number of attribute triples, in which attribute values are represented by free texts. Since free texts are too flexible to define entities in KGs, traditional way to map entities to their embeddings using ids is no longer appropriate for attribute value representation; (3) Noisy triples in a PG mislead the embedding learning and significantly hurt the performance of error detection. To address the aforementioned challenges, we propose an end-to-end noise-tolerant embedding learning framework, PGE, to jointly leverage both text information and graph structure in PG to learn embeddings for error detection. Experimental results on real-world product graph demonstrate the effectiveness of the proposed framework comparing with the state-of-the-art approaches. Kewei Cheng, Yifan Ethan Xu, Xin Dong 0001, Yizhou Sun |
Proc. VLDB Endow. | 1 |
| 2021 | UniKER: A Unified Framework for Combining Embedding and Definite Horn Rule Reasoning for Knowledge Graph InferenceabstractKnowledge graph inference has been studied extensively due to its wide applications.It has been addressed by two lines of research, i.e., the more traditional logical rule reasoning and the more recent knowledge graph embedding (KGE).Several attempts have been made to combine KGE and logical rules for better knowledge graph inference.Unfortunately, they either simply treat logical rules as additional constraints into KGE loss or use probabilistic models to approximate the exact logical inference (i.e., MAX-SAT).Even worse, both approaches need to sample ground rules to tackle the scalability issue, as the total number of ground rules is intractable in practice, making them less effective in handling logical rules.In this paper, we propose a novel framework UniKER to address these challenges by restricting logical rules to be definite Horn rules, which can fully exploit the knowledge in logical rules and enable the mutual enhancement of logical rule-based reasoning and KGE in an extremely efficient way.Extensive experiments have demonstrated that our approach is superior to existing state-of-the-art algorithms in terms of both efficiency and effectiveness. Kewei Cheng, Ziqing Yang 0002, Ming Zhang 0004, Yizhou Sun |
EMNLP (1) | 1 |
| 2018 | Streaming Link Prediction on Dynamic Attributed NetworksabstractLink prediction targets to predict the future node interactions mainly based on the current network snapshot. It is a key step in understanding the formation and evolution of the underlying networks; and has practical implications in many real-world applications, ranging from friendship recommendation, click through prediction to targeted advertising. Most existing efforts are devoted to plain networks and assume the availability of network structure in memory before link prediction takes place. However, this assumption is untenable as many real-world networks are affiliated with rich node attributes, and often, the network structure and node attributes are both dynamically evolving at an unprecedented rate. Even though recent studies show that node attributes have an added value to network structure for accurate link prediction, it still remains a daunting task to support link prediction in an online fashion on such dynamic attributed networks. As changes in the dynamic attributed networks are often transient and can be endless, link prediction algorithms need to be efficient by making only one pass of the data with limited memory overhead. To tackle these challenges, we study a novel problem of streaming link prediction on dynamic attributed networks and present a novel framework - SLIDE. Methodologically, SLIDE maintains and updates a low-rank sketching matrix to summarize all observed data, and we further leverage the sketching matrix to infer missing links on the fly. The whole procedure is theoretically guaranteed, and empirical experiments on real-world dynamic attributed networks validate the effectiveness and efficiency of the proposed framework. Jundong Li, Kewei Cheng, Liang Wu 0006, Huan Liu 0001 |
WSDM | 2 |
| 2017 | Unsupervised Sentiment Analysis with Signed Social NetworksabstractHuge volumes of opinion-rich data is user-generated in social media at an unprecedented rate, easing the analysis of individual and public sentiments. Sentiment analysis has shown to be useful in probing and understanding emotions, expressions and attitudes in the text. However, the distinct characteristics of social media data present challenges to traditional sentiment analysis. First, social media data is often noisy, incomplete and fast-evolved which necessitates the design of a sophisticated learning model. Second, sentiment labels are hard to collect which further exacerbates the problem by not being able to discriminate sentiment polarities. Meanwhile, opportunities are also unequivocally presented. Social media contains rich sources of sentiment signals in textual terms and user interactions, which could be helpful in sentiment analysis. While there are some attempts to leverage implicit sentiment signals in positive user interactions, little attention is paid on signed social networks with both positive and negative links. The availability of signed social networks motivates us to investigate if negative links also contain useful sentiment signals. In this paper, we study a novel problem of unsupervised sentiment analysis with signed social networks. In particular, we incorporate explicit sentiment signals in textual terms and implicit sentiment signals from signed social networks into a coherent model SignedSenti for unsupervised sentiment analysis. Empirical experiments on two real-world datasets corroborate its effectiveness. Kewei Cheng, Jundong Li, Jiliang Tang, Huan Liu 0001 |
AAAI | 1 |
| 2017 | Unsupervised Feature Selection in Signed Social NetworksabstractThe rapid growth of social media services brings a large amount of high-dimensional social media data at an unprecedented rate. Feature selection is powerful to prepare high-dimensional data by finding a subset of relevant features. A vast majority of existing feature selection algorithms for social media data exclusively focus on positive interactions among linked instances such as friendships and user following relations. However, in many real-world social networks, instances may also be negatively interconnected. Recent work shows that negative links have an added value over positive links in advancing many learning tasks. In this paper, we study a novel problem of unsupervised feature selection in signed social networks and propose a novel framework SignedFS. In particular, we provide a principled way to model positive and negative links for user latent representation learning. Then we embed the user latent representations into feature selection when label information is not available. Also, we revisit the principle of homophily and balance theory in signed social networks and incorporate the signed graph regularization into the feature selection framework to capture the first-order and the second-order proximity among users in signed social networks. Experiments on two real-world signed social networks demonstrate the effectiveness of our proposed framework. Further experiments are conducted to understand the impacts of different components of SignedFS. Kewei Cheng, Jundong Li, Huan Liu 0001 |
KDD | 1 |
| 2016 | FeatureMiner: A Tool for Interactive Feature SelectionabstractThe recent popularity of big data has brought immense quantities of high-dimensional data, which presents challenges to traditional data mining tasks due to curse of dimensionality. Feature selection has shown to be effective to prepare these high dimensional data for a variety of learning tasks. To provide easy access to feature selection algorithms, we provide an interactive feature selection tool FeatureMiner based on our recently released feature selection repository scikit-feature. FeatureMiner eases the process of performing feature selection for practitioners by providing an interactive user interface. Meanwhile, it also gives users some practical guidance in finding a suitable feature selection algorithm among many given a specific dataset. In this demonstration, we show (1) How to conduct data preprocessing after loading a dataset; (2) How to apply feature selection algorithms; (3) How to choose a suitable algorithm by visualized performance evaluation. Kewei Cheng, Jundong Li, Huan Liu 0001 |
CIKM | 1 |