Prasad Calyam

dblp:06/1313 · DBLP profile ↗
← Back
9ranked-venue papers in the field
0as first author
7since 2021 · last 2025
0000-0002-7666-5389ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4Data Mining & Knowledge Discovery · 3Database Systems & Data Management · 1Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Disentangling Complex Questions in LLMs via Multi-Hop Dependency Graphs
abstract
While Large language models (LLMs) have shown to exhibit remarkable performance in a wide range of NLP tasks, they often struggle to interpret and reason over multi-hop questions in open-domain question answering (ODQA) settings. While popular prompt approaches such as Chain-of-Thought and Plan-and-Solve facilitate more manageable questions for OQDA via task decomposition, these approaches are prone to generating erroneous and redundant intermediate steps in multi-hop queries due to limited capacity for modeling complex entity relationships. In this paper, we introduce a novel prompt approach for multi-hop QA viz., MoDeGraph (Multi-Hop Dependency Graphs), that is designed to steer LLMs to extract and model entity relationships in complex questions. MoDeGraph constructs a dependency graph from LLM-generated entity-relation triples to enable more coherent and human-like multi-step reasoning. Experimental results in knowledge-intensive tasks for multi-hop QA demonstrate our approach produces more coherent and faithful reasoning chains as well as consistent increase in QA performance across several benchmark datasets.
Roland Oruche, Alphaeus Dmonte, Vani Seth, Zian Zeng, Yuanxun Zhang, Marcos Zampieri, Prasad Calyam
CIKM7
2024 Influence Role Recognition and LLM-Based Scholar Recommendation in Academic Social Networks
abstract
Identifying scholars and their relevant publications in interdisciplinary collaborations within an academic social network (ASN) can help drive new scientific knowledge discovery. This involves a challenging and time-consuming process, which requires scholar's influence role recognition in a scholar team for a given research task. In this paper, we propose a novel “ScholarInfluencer” recommendation system that: (a) uses a classification model combined with network analysis on a heterogeneous knowledge graph to recognize the scholar influencers within interdisciplinary teams of collaborators, and (b) features a large language model (LLM) to use influence role recognition results to support user queries to produce pertinent scholar and their publication recommendations. Our novel approach involves building a heterogeneous knowledge graph using diverse ASN datasets involving entities such as scholars, publications, research grants, and the relationship among these entities. We perform an evaluation of ScholarInfluencer using four widely-used ASN datasets (i.e., NSF, DBLP, Cora and CA-HepTh). Our experiment results show that our influence role recognition model outperforms the state-of-the-art models across the different datasets; especially in the case of the NSF dataset, our model outperforms by up to 13.6%. Further, we show how our recommendation model with role recognition outperforms the model without role recognition across the different datasets; especially in the case of the NSF dataset, our model outperforms by 7%.
Xiyao Cheng, Lakshmi Srinivas Edara, Yuanxun Zhang, Mayank Kejriwal, Prasad Calyam
DSAA5
2024 Deep Contrastive Active Learning for Out-of-domain Filtering in Dialog Systems
abstract
Task-oriented dialog systems have shown to foster effective human-chatbot collaborations for accomplishing goal-specific tasks through intent classification. In a real-world setting, collecting and training over user intents incurs a labeling-cost challenge for human annotators. While existing human-AI collaborative approaches such as active learning (AL) can properly resolve such labeling-cost challenges, most existing AL algorithms assume the unlabeled pool has similar distributions as the in domain (IND) training set. To address the conflict between AL and out-of-domain (OOD) data samples, we present Deep Contrastive Active Learning (DeCAL), a deep novel AL framework that uses contrastive learning techniques for query intent classification in task-oriented dialogs. DeCAL features an acquisition function that filters OOD samples by computing a distance-based confidence score over unlabeled samples using their neighboring features. To validate DeCAL, we compare against deep AL baselines via the performance of acquired IND/OOD samples and using the classification accuracy metric. Experimental results on benchmark datasets demonstrate De-CAL outperforms deep AL baseline algorithms on acquired OOD by 14%, while simultaneously showing competitive performance on IND accuracy.
Roland Oruche, Marcos Zampieri, Prasad Calyam
DSAA3
2023 Knowledge Graph-based Embedding for Connecting Scholars in Academic Social Networks
abstract
In recent years, research tasks have increasingly involved using multi-disciplinary knowledge through collaborations of scholars from multiple fields. However, identifying a team of suitable collaborators from diverse fields for a given research task is a challenging and time-consuming process. In this paper, we propose a novel “ScholarTeamFinder” model that uses knowledge graph based link prediction to identify collaborators within an academic social network (ASN) to form a research team to address a multi-disciplinary research problem. Our approach involves building a heterogeneous knowledge graph within an ASN using entities such as scholars, publications, research grants, and the relationship among these entities. Following this, we use graph-based deep learning to learn the node embedding from the knowledge graph that can be used for scholar team recommendation. More specifically, we used the classical meth-path2vec as our base graph learning algorithm and improved its performance by considering semantic meaning of entities and encoding edge embeddings in the graph. Finally, we propose a beam-search algorithm for scholar team prediction based on our model embeddings. Our evaluation of ScholarTeamFinder is performed using large ASN datasets including a unique dataset (i.e., NSF award dataset) of federal grant awards collected over the last ten years and the scholars’ publication data, as well as three other widely used datasets (i.e., APS, SCHOLAT and Gowalla). Experiment results show that our model outperforms the state-of-the-art models across the different datasets.
Xiyao Cheng, Yuanxun Zhang, Harsh Joshi, Mayank Kejriwal, Prasad Calyam
DSAA5
2023 Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific Communities
abstract
Shortened time to knowledge discovery and adapting prior domain knowledge is a challenge for computational and data-intensive communities such as e.g., bioinformatics and neuroscience. The challenge for a domain scientist lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics when: investigating new methods, developing new tools, or integrating datasets. In this paper, we propose a novel "domain-specific topic model" (DSTM) to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplary scientific domains. Our DSTM is a generative model that extends the Latent Dirichlet Allocation (LDA) model and uses the Markov chain Monte Carlo (MCMC) algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include more than 25,000 of papers over the last ten years, featuring hundreds of tools and datasets that are commonly used in relevant studies. Evaluation experiments based on generalization and information retrieval metrics show that our model has better performance than the state-of-the-art baseline models for discovering highly-specific latent topics within a domain. Lastly, we demonstrate applications that benefit from our DSTM to discover intra-domain, cross-domain and trend knowledge patterns.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE Trans. Knowl. Data Eng.2
2022 Networked and Multimodal 3D Modeling of Cities for Collaborative Virtual Environments
abstract
3D city-scale models are useful in a number of applications, including education, city planning, navigation systems, artificial intelligence training, and simulations. However, final models need to be immersive and interactive, which requires a mixed reality (XR) environment design that combines e.g., a Cave Automatic Virtual Environment (CAVE) VR system with the Microsoft Hololens2 in a networked and multimodal setting. In this paper, we propose a pipeline to convert a city-scale point cloud into a finalized city-scale textured mesh in which, a number of XR devices can share the same environment and co-exist in a shared space for model interactions. Specifically, we use input point clouds obtained from wide area motion imagery systems or off-the-shelf drones pertaining to Albuquerque, New Mexico, but the pipeline is generalized so that other input can be used. Using four different traditional algorithms and an additional deep learning method, we create meshes for the model interactions. For each mesh produced, we map high-resolution textures onto them, producing a more accurate city, which is then passed into the shared/networked Unity environment. Ten participants provided their assessment of mesh quality and interactivity of the networked environment during exploration of different city reconstructions with the CAVE and laptop device modalities. Results on the perceptual immersive quality of the Point2Mesh deep learning meshes highlights the need for improvements to handle large city scale point clouds.
Benjamin Hall, Joseph Kessler, Osayamen Edo-Ohanba, Jaired Collins, Nick Allegreti, Ye Duan, Songjie Wang, Kannappan Palaniappan, Prasad Calyam
BDCAT10
2021 3D Modeling of Cities for Virtual Environments
abstract
Modeling and simulation of large urban regions is beneficial for a range of applications including intelligent transportation, smart cities, infrastructure planning, and training artificial intelligence for autonomous navigation systems including ground vehicles and aerial drones. Immersive environments including virtual reality (VR), augmented reality (AR), mixed reality (MR or XR) can be used to explore city scale regions for planning, design, training and operations. Virtual environments are in the midst of rapid change as innovations in display technologies, graphics processors and game engine software present new opportunities for incorporating modeling and simulation into engineering workflows. Game engine software like Unity with photorealistic rendering and realistic physics have plug-in support for a variety of virtual environments and typically model the scene as meshes. In this paper, we develop an end-to-end workflow for creating urban scale real world accurate synthetic environments that can be visualized in virtual environments including the Microsoft HoloLens head mounted display or the CAVE VR for multi-user interaction. Four meshing algorithms are evaluated for representation accuracy and city-scale meshes imported into Unity for assessing the quality of the immersive experience.
Calvin Davis, Jaired Collins, Joshua Fraser, Shizeng Yao, Emily Lattanzio, Bimal Balakrishnan, Ye Duan, Prasad Calyam, Kannappan Palaniappan
IEEE BigData9
2018 Fuzzy-Based Conversational Recommender for Data-intensive Science Gateway Applications
abstract
Neuro-scientists are increasingly relying on parallel and distributed computing resources for analysis and visualization of their neuron simulations. Although science gateways have democratized relevant high performance/throughput resources, users require expert knowledge about programming and infrastructure configuration that is beyond the repertoire of most neuroscience programs. These factors become deterrents for the successful adoption and the ultimate diffusion (i.e., systemic spread) of science gateways in the neuroscience community. In this paper, we present a novel intuitionistic fuzzy logic based conversational recommender that can provide guidance to users when using science gateways for research and education workflows. The users interact with a context-aware chatbot that is embedded within custom web-portals to obtain simulation tools/resources to accomplish their goals. In order to ensure user goals are met, the chatbot profiles a user's cyberinfrastructure and neuroscience domain proficiency level using a `usability quadrant' approach. Simulation of user queries for an exemplary neuroscience use case demonstrates that our chatbot can provide step-by-step navigational support and generate distinct responses based on user proficiency.
Arjun Ankathatti Chandrashekara, Radha Krishna Murthy Talluri, Sai Swathi Sivarathri, Reshmi Mitra, Prasad Calyam, Kerk F. Kee, Satish S. Nair
IEEE BigData5
2018 Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
abstract
Machine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today's innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE BigData2