VLDB 2026 Research / reviewers in the wild / expert
Xiyao Cheng
dblp:210/2052
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Securing Inverter-Based Resources via Knowledge-Driven Threat Modeling, Analysis, and MitigationabstractInverter-based Resources (IBRs) present unique cybersecurity challenges due to their digital control systems and connection to the electric grid. They rely on digital communications and control systems, making them vulnerable to cyberattacks. These attacks can disrupt grid operations and stability, compromise data, or cause physical damage to equipment. To address these challenges, it is essential to establish robust cybersecurity measures that meet and exceed existing industry standards. In this paper, we describe a comprehensive strategy to bolster the cybersecurity of IBRs through cutting-edge applications and technologies via a cybersecurity framework called “CIBR-Fort”, a knowledge-driven, interoperable, scalable, and manageable framework for modeling, analysis, and mitigation of cyber threats disrupting different components of IBR systems. Our knowledge-driven analysis consists of a fusion of knowledge graphs (KGs) in cybersecurity and the electric grid, achieved through link prediction leveraging Large Language Models (LLMs) and cosine similarity, attributed towards informed decision-making for threat mitigation. The evaluation results show how we can automate LLM-driven link prediction based on the fusion of two distantly separated ontologies, generating a dataset that can be used for scaling via graph learning that can be utilized for further security analyses of IBR systems. In addition, we show our knowledge-driven threat analysis can predict different attacks with 91.88% maximum accuracy. Lastly, we show how we can achieve real-time end-to-end threat mitigation with an average of 40 ms per traffic flow. Roshan Neupane, Vamsi Pusapati, Lakshmi Srinivas Edara, Xiyao Cheng, Kiran Neupane, Harshavardhan Chintapatla, Reshmi Mitra, Mert Korkali, Hyeong Suk Na, Sharan Srinivas, Prasad Calyam |
NOMS | 4 |
| 2025 | ScholarFinder: Knowledge Embedding Based Recommendations Using a Deep Embedded Clustering ModelabstractBold scientific research tasks need multi-disciplinary knowledge and collaborations that require finding scholars from particular domains with relevant knowledge. Given the variety of scholars and diversity, finding the appropriate scholar is an important and challenging problem for scientific communities. In this paper, we propose a “ScholarFinder” framework that uses contextual information (abstracts or publications) for embedding a scholar's knowledge in an unsupervised learning manner. Specifically, we implement an unsupervised embedding technique viz.,Variational AutoEncoder (VAE). For better feature representation learning, we also implement aVariational Deep Embedded Clustering (VDEC)method that further enhances downstream tasks (e.g., clustering, classification) accuracy, scalability, and performance. In addition, we incorporate a multi-task learning scheme into our VDEC model for improving the effectiveness of simultaneously learning both embedding and clustering. Subsequently, the downstream tasks can be built based on pre-trained scholars' knowledge embeddings to predict suitability of a scholar for a research task. Using a dataset involving a 20-year collection of federal grant awards, we have demonstrated how our pre-trained model improved the performance for downstream tasks. We have also investigated how our pre-trained model can be integrated into a knowledge graph to achieve better performance. Lastly, we show that our ScholarFinder model variants outperform state-of-the-art baseline models (i.e., XGBoost, GBDT, AdaBoost, DNN, GraphSAGE, DEC, VaDE) and recent LLM based models (i.e., Bert4Rec, OpenP5) by atleast 18%. Yuanxun Zhang, Xiyao Cheng, Roland Oruche, Sai Swathi Sivarathri, Prasad Calyam |
IEEE Trans. Big Data | 2 |
| 2025 | Chatbot Dialog Design for Improved Human Performance in Domain Knowledge DiscoveryabstractThe advent of machine learning (ML) has led to the widespread adoption of developing task-oriented dialog systems for scientific applications (e.g., science gateways) where voluminous information sources are retrieved and curated for domain users. Yet, there still exists a challenge in designing chatbot dialog systems that achieve widespread diffusion among scientific communities. In this article, we propose a novel Vidura advisor design framework (VADF) to develop dialog system designs for information retrieval (IR) and question-answering (QA) tasks, while enabling the quantification of system utility based on human performance in diverse application environments. We adopt a socio-technical approach in our framework for designing dialog systems by utilizing domain expert feedback, which features a sparse retriever for enabling accurate responses in QA settings using linear interpolation smoothing. We apply our VADF for an exemplar science gateway, viz. KnowCOVID-19, to conduct experiments that demonstrate the utility of dialog systems based on IR and QA performance, application utility, and perceived adoption. Experimental results show our VADF approach significantly improves IR performance against retriever baselines (up to 5% increase) and QA performance against large language models (LLMs) such as ChatGPT (up to 43% increase) on scientific literature datasets. In addition, through a usability survey, we observe that measuring application utility and human performance when applying VADF to KnowCOVID-19 translates to an increase in perceived community adoption. Roland Oruche, Xiyao Cheng, Zian Zeng, Audrey Vazzana, MD Ashraful Goni, Bruce Wang Shibo, Sai Keerthana Goruganthu, Kerk F. Kee, Prasad Calyam |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2024 | Influence Role Recognition and LLM-Based Scholar Recommendation in Academic Social NetworksabstractIdentifying scholars and their relevant publications in interdisciplinary collaborations within an academic social network (ASN) can help drive new scientific knowledge discovery. This involves a challenging and time-consuming process, which requires scholar's influence role recognition in a scholar team for a given research task. In this paper, we propose a novel “ScholarInfluencer” recommendation system that: (a) uses a classification model combined with network analysis on a heterogeneous knowledge graph to recognize the scholar influencers within interdisciplinary teams of collaborators, and (b) features a large language model (LLM) to use influence role recognition results to support user queries to produce pertinent scholar and their publication recommendations. Our novel approach involves building a heterogeneous knowledge graph using diverse ASN datasets involving entities such as scholars, publications, research grants, and the relationship among these entities. We perform an evaluation of ScholarInfluencer using four widely-used ASN datasets (i.e., NSF, DBLP, Cora and CA-HepTh). Our experiment results show that our influence role recognition model outperforms the state-of-the-art models across the different datasets; especially in the case of the NSF dataset, our model outperforms by up to 13.6%. Further, we show how our recommendation model with role recognition outperforms the model without role recognition across the different datasets; especially in the case of the NSF dataset, our model outperforms by 7%. Xiyao Cheng, Lakshmi Srinivas Edara, Yuanxun Zhang, Mayank Kejriwal, Prasad Calyam |
DSAA | 1 |
| 2023 | Knowledge Graph-based Embedding for Connecting Scholars in Academic Social NetworksabstractIn recent years, research tasks have increasingly involved using multi-disciplinary knowledge through collaborations of scholars from multiple fields. However, identifying a team of suitable collaborators from diverse fields for a given research task is a challenging and time-consuming process. In this paper, we propose a novel “ScholarTeamFinder” model that uses knowledge graph based link prediction to identify collaborators within an academic social network (ASN) to form a research team to address a multi-disciplinary research problem. Our approach involves building a heterogeneous knowledge graph within an ASN using entities such as scholars, publications, research grants, and the relationship among these entities. Following this, we use graph-based deep learning to learn the node embedding from the knowledge graph that can be used for scholar team recommendation. More specifically, we used the classical meth-path2vec as our base graph learning algorithm and improved its performance by considering semantic meaning of entities and encoding edge embeddings in the graph. Finally, we propose a beam-search algorithm for scholar team prediction based on our model embeddings. Our evaluation of ScholarTeamFinder is performed using large ASN datasets including a unique dataset (i.e., NSF award dataset) of federal grant awards collected over the last ten years and the scholars’ publication data, as well as three other widely used datasets (i.e., APS, SCHOLAT and Gowalla). Experiment results show that our model outperforms the state-of-the-art models across the different datasets. Xiyao Cheng, Yuanxun Zhang, Harsh Joshi, Mayank Kejriwal, Prasad Calyam |
DSAA | 1 |
| 2023 | Science gateway adoption using plug-in middleware for evidence-based healthcare data managementabstractSummary There is a growing need for next‐generation science gateways to increase the accessibility of emerging large‐scale datasets for data consumers (e.g., clinicians, researchers) who aim to combat COVID‐19‐related challenges. Such science gateways that enable access to distributed computing resources for large‐scale data management need to be made more programmable, extensible, and scalable. In this article, we propose a novel socio‐technical approach for developing a next‐generation healthcare science gateway, namely, OnTimeEvidence that addresses data consumer challenges surrounding the COVID‐19 pandemic related data analytics. OnTimeEvidence implements an intelligent agent, namely, Vidura Advisor that integrates an evidence‐based filtering method to transform manual practices and improve scalability of data analytics. It also features a plug‐in management middleware that improves the programmability and extensibility of the science gateway capabilities using microservices. Lastly, we present a usability study that shows the important factors from data consumers' perspective to adopt OnTimeEvidence with chatbot‐assisted middleware support to increase their productivity and collaborations to access vast publication archives for rapid knowledge discovery tasks. Roland Oruche, Eric D. Milman, Mauro Lemus, Xiyao Cheng, Songjie Wang, Prasad Calyam, Kerk F. Kee |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Joint Learning for Emotion Classification and Emotion Cause DetectionabstractWe present a neural network-based joint approach for emotion classification and emotion cause detection, which attempts to capture mutual benefits across the two sub-tasks of emotion analysis.Considering that emotion classification and emotion cause detection need different kinds of features (affective and event-based separately), we propose a joint encoder which uses a unified framework to extract features for both sub-tasks and a joint model trainer which simultaneously learns two models for the two sub-tasks separately.Our experiments on Chinese microblogs show that the joint approach is very promising. Ying Chen 0012, Xiyao Cheng, Shoushan Li |
EMNLP | 3 |
| 2018 | Hierarchical Convolution Neural Network for Emotion Cause Detection on Microblogs
Ying Chen 0012, Xiyao Cheng |
ICANN (1) | 3 |
| 2017 | An Emotion Cause Corpus for Chinese Microblogs with Multiple-User StructuresabstractA notably challenging problem in emotion analysis is recognizing the cause of an emotion. Although there have been a few studies on emotion cause detection, most of them work on news reports or a few of them focus on microblogs using a single-user structure (i.e., all texts in a microblog are written by the same user). In this article, we focus on emotion cause detection for Chinese microblogs using a multiple-user structure (i.e., texts in a microblog are successively written by several users). First, based on the fact that the causes of an emotion of a focused user may be provided by other users in a microblog with the multiple-user structure, we design an emotion cause annotation scheme which can deal with such a complicated case, and then provide an emotion cause corpus using the annotation scheme. Second, based on the analysis of the emotion cause corpus, we formalize two emotion cause detection tasks for microblogs (current-subtweet-based emotion cause detection and original-subtweet-based emotion cause detection). Furthermore, in order to examine the difficulty of the two emotion cause detection tasks and the contributions of texts written by different users in a microblog with the multiple-user structure, we choose two popular classification methods (SVM and LSTM) to do emotion cause detection. Our experiments show that the current-subtweet-based emotion cause detection is much more difficult than the original-subtweet-based emotion cause detection, and texts written by different users are very helpful for both emotion cause detection tasks. This study presents a pilot study of emotion cause detection which deals with Chinese microblogs using a complicated structure. Xiyao Cheng, Ying Chen 0012, Bixiao Cheng, Shoushan Li, Guodong Zhou 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |