Yuanxun Zhang

dblp:165/2229 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-authorTheory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Disentangling Complex Questions in LLMs via Multi-Hop Dependency Graphs
abstract
While Large language models (LLMs) have shown to exhibit remarkable performance in a wide range of NLP tasks, they often struggle to interpret and reason over multi-hop questions in open-domain question answering (ODQA) settings. While popular prompt approaches such as Chain-of-Thought and Plan-and-Solve facilitate more manageable questions for OQDA via task decomposition, these approaches are prone to generating erroneous and redundant intermediate steps in multi-hop queries due to limited capacity for modeling complex entity relationships. In this paper, we introduce a novel prompt approach for multi-hop QA viz., MoDeGraph (Multi-Hop Dependency Graphs), that is designed to steer LLMs to extract and model entity relationships in complex questions. MoDeGraph constructs a dependency graph from LLM-generated entity-relation triples to enable more coherent and human-like multi-step reasoning. Experimental results in knowledge-intensive tasks for multi-hop QA demonstrate our approach produces more coherent and faithful reasoning chains as well as consistent increase in QA performance across several benchmark datasets.
Roland Oruche, Alphaeus Dmonte, Vani Seth, Zian Zeng, Yuanxun Zhang, Marcos Zampieri, Prasad Calyam
CIKM5
2025 ScholarFinder: Knowledge Embedding Based Recommendations Using a Deep Embedded Clustering Model
abstract
Bold scientific research tasks need multi-disciplinary knowledge and collaborations that require finding scholars from particular domains with relevant knowledge. Given the variety of scholars and diversity, finding the appropriate scholar is an important and challenging problem for scientific communities. In this paper, we propose a “ScholarFinder” framework that uses contextual information (abstracts or publications) for embedding a scholar's knowledge in an unsupervised learning manner. Specifically, we implement an unsupervised embedding technique viz.,Variational AutoEncoder (VAE). For better feature representation learning, we also implement aVariational Deep Embedded Clustering (VDEC)method that further enhances downstream tasks (e.g., clustering, classification) accuracy, scalability, and performance. In addition, we incorporate a multi-task learning scheme into our VDEC model for improving the effectiveness of simultaneously learning both embedding and clustering. Subsequently, the downstream tasks can be built based on pre-trained scholars' knowledge embeddings to predict suitability of a scholar for a research task. Using a dataset involving a 20-year collection of federal grant awards, we have demonstrated how our pre-trained model improved the performance for downstream tasks. We have also investigated how our pre-trained model can be integrated into a knowledge graph to achieve better performance. Lastly, we show that our ScholarFinder model variants outperform state-of-the-art baseline models (i.e., XGBoost, GBDT, AdaBoost, DNN, GraphSAGE, DEC, VaDE) and recent LLM based models (i.e., Bert4Rec, OpenP5) by atleast 18%.
Yuanxun Zhang, Xiyao Cheng, Roland Oruche, Sai Swathi Sivarathri, Prasad Calyam
IEEE Trans. Big Data1
2024 Influence Role Recognition and LLM-Based Scholar Recommendation in Academic Social Networks
abstract
Identifying scholars and their relevant publications in interdisciplinary collaborations within an academic social network (ASN) can help drive new scientific knowledge discovery. This involves a challenging and time-consuming process, which requires scholar's influence role recognition in a scholar team for a given research task. In this paper, we propose a novel “ScholarInfluencer” recommendation system that: (a) uses a classification model combined with network analysis on a heterogeneous knowledge graph to recognize the scholar influencers within interdisciplinary teams of collaborators, and (b) features a large language model (LLM) to use influence role recognition results to support user queries to produce pertinent scholar and their publication recommendations. Our novel approach involves building a heterogeneous knowledge graph using diverse ASN datasets involving entities such as scholars, publications, research grants, and the relationship among these entities. We perform an evaluation of ScholarInfluencer using four widely-used ASN datasets (i.e., NSF, DBLP, Cora and CA-HepTh). Our experiment results show that our influence role recognition model outperforms the state-of-the-art models across the different datasets; especially in the case of the NSF dataset, our model outperforms by up to 13.6%. Further, we show how our recommendation model with role recognition outperforms the model without role recognition across the different datasets; especially in the case of the NSF dataset, our model outperforms by 7%.
Xiyao Cheng, Lakshmi Srinivas Edara, Yuanxun Zhang, Mayank Kejriwal, Prasad Calyam
DSAA3
2023 Knowledge Graph-based Embedding for Connecting Scholars in Academic Social Networks
abstract
In recent years, research tasks have increasingly involved using multi-disciplinary knowledge through collaborations of scholars from multiple fields. However, identifying a team of suitable collaborators from diverse fields for a given research task is a challenging and time-consuming process. In this paper, we propose a novel “ScholarTeamFinder” model that uses knowledge graph based link prediction to identify collaborators within an academic social network (ASN) to form a research team to address a multi-disciplinary research problem. Our approach involves building a heterogeneous knowledge graph within an ASN using entities such as scholars, publications, research grants, and the relationship among these entities. Following this, we use graph-based deep learning to learn the node embedding from the knowledge graph that can be used for scholar team recommendation. More specifically, we used the classical meth-path2vec as our base graph learning algorithm and improved its performance by considering semantic meaning of entities and encoding edge embeddings in the graph. Finally, we propose a beam-search algorithm for scholar team prediction based on our model embeddings. Our evaluation of ScholarTeamFinder is performed using large ASN datasets including a unique dataset (i.e., NSF award dataset) of federal grant awards collected over the last ten years and the scholars’ publication data, as well as three other widely used datasets (i.e., APS, SCHOLAT and Gowalla). Experiment results show that our model outperforms the state-of-the-art models across the different datasets.
Xiyao Cheng, Yuanxun Zhang, Harsh Joshi, Mayank Kejriwal, Prasad Calyam
DSAA2
2023 Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific Communities
abstract
Shortened time to knowledge discovery and adapting prior domain knowledge is a challenge for computational and data-intensive communities such as e.g., bioinformatics and neuroscience. The challenge for a domain scientist lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics when: investigating new methods, developing new tools, or integrating datasets. In this paper, we propose a novel "domain-specific topic model" (DSTM) to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplary scientific domains. Our DSTM is a generative model that extends the Latent Dirichlet Allocation (LDA) model and uses the Markov chain Monte Carlo (MCMC) algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include more than 25,000 of papers over the last ten years, featuring hundreds of tools and datasets that are commonly used in relevant studies. Evaluation experiments based on generalization and information retrieval metrics show that our model has better performance than the state-of-the-art baseline models for discovering highly-specific latent topics within a domain. Lastly, we demonstrate applications that benefit from our DSTM to discover intra-domain, cross-domain and trend knowledge patterns.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE Trans. Knowl. Data Eng.1
2021 Recommender-as-a-service with chatbot guided domain-science knowledge discovery in a science gateway
abstract
Scientists in disciplines such as neuroscience and bioinformatics are increasingly relying on science gateways for experimentation on voluminous data, as well as analysis and visualization in multiple perspectives. Though current science gateways provide easy access to computing resources, datasets and tools specific to the disciplines, scientists often use slow and tedious manual efforts to perform knowledge discovery to accomplish their research/education tasks. Recommender systems can provide expert guidance and can help them to navigate and discover relevant publications, tools, data sets, or even automate cloud resource configurations suitable for a given scientific task. To realize the potential of integration of recommenders in science gateways in order to spur research productivity, we present a novel "OnTimeRecommend" recommender system. The OnTimeRecommend comprises of several integrated recommender modules implemented as microservices that can be augmented to a science gateway in the form of a recommender-as-a-service. The guidance for use of the recommender modules in a science gateway is aided by a chatbot plug-in viz., Vidura Advisor. To validate our OnTimeRecommend, we integrate and show benefits for both novice and expert users in domain-specific knowledge discovery within two exemplar science gateways, one in neuroscience (CyNeuro) and the other in bioinformatics (KBCommons).
Komal Bhupendra Vekaria, Prasad Calyam, Sai Swathi Sivarathri, Songjie Wang, Yuanxun Zhang, Dong Xu 0002, Trupti Joshi, Satish S. Nair
Concurr. Comput. Pract. Exp.5
2021 Multi-Cloud Performance and Security Driven Federated Workflow Management
abstract
Federated multi-cloud resource allocation for data-intensive application workflows is generally performed based on performance or quality of service (i.e., QSpecs) considerations. At the same time, end-to-end security requirements of these workflows across multiple domains are considered as an afterthought due to lack of standardized formalization methods. Consequently, diverse/heterogenous domain resource and security policies cause inter-conflicts between application's security and performance requirements that lead to sub-optimal resource allocations. In this paper, we present a joint performance and security-driven federated resource allocation scheme for data-intensive scientific applications. In order to aid joint resource brokering among multi-cloud domains with diverse/heterogenous security postures, we first define and characterize a data-intensive application's security specifications (i.e., SSpecs). Then we describe an alignment technique inspired by Portunes Algebra to homogenize the various domain resource policies (i.e., RSpecs) along an application's workflow lifecycle stages. Using such formalization and alignment, we propose a near optimal cost-aware joint QSpecs-SSpecs-driven, RSpecs-compliant resource allocation algorithm for multi-cloud computing resource domain/location selection as well as network path selection. We implement our security formalization, alignment, and allocation scheme as a framework, viz., “OnTimeURB” and validate it in a multi-cloud environment with exemplar data-intensive application workflows involving distributed computing and remote instrumentation use cases with different performance and security requirements.
Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Ronny Bazan Antequera, Trupti Joshi, Tommi A. White, Dong Xu 0002
IEEE Trans. Cloud Comput.5
2018 Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
abstract
Machine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today's innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE BigData1
2018 Social Plane for Recommenders in Network Performance Expectation Management
abstract
Multi-domain end-to-end network performance monitoring federations such as perfSONAR are increasingly being used in Big Data application management. They rely on trustworthy collaborative measurement intelligence to identify and diagnose network anomaly events that impact application performance. Large volumes of end-to-end measurement traces are generated on a daily basis, and new Big Data analysis techniques are needed to isolate network-wide anomaly event(s) and to diagnose the root-cause(s). In addition, not all network operators and application users have enough knowledge and experience to understand the anomaly events. The lack of a platform for sharing knowledge and working collaboratively makes it difficult to isolate and diagnose network-wide anomaly events quickly and accurately. In this paper, we define a “social plane” that relies on recommended measurements based on “content-based filtering” and “collaborative filtering” approaches to enable network performance expectation management. Based on similarity analysis, the content-based filtering facilitates users to subscribe to useful measurements, and the collaborative filtering promotes users to share knowledge on anomaly symptoms. Using real perfSONAR measurements and synthetic events, we show the effectiveness of our social plane approach within a SoyKB Big Data application case study using social network creation and mingling of experts. Our experimental results show that our measurements recommendation scheme has high precision, recall, and accuracy, as well as efficiency in terms of the time taken for large volume measurement trace analysis.
Yuanxun Zhang, Prasad Calyam, Saptarshi Debroy, Sai Shreya Nuguri
IEEE Trans. Netw. Serv. Manag.1
2016 End-to-End Security Formalization and Alignment for Federated Workflow Management
abstract
Traditionally, the allocation and dynamic adaptation of federated cyberinfrastructure resources residing across multiple domains for data-intensive application workflows have been performance or quality of service-centric (i.e., QSpecs), often compromising the end-to-end security requirements of scientific workflows. Lack of standardized formalization methods of the workflows' end-to-end security requirements, and diverse/heterogenous domain resource and security policies make inter-conflict characterization between application's security and performance requirements non-trivial, and leads to sub-optimal resource allocation. In this paper, we present a joint security and performance-driven federated resource allocation and adaptation scheme to define and characterize a data-intensive scientific application's security specifications (i.e., SSpecs). In order to aid security-driven resource brokering among domains with diverse security postures, we describe an alignment technique inspired by Portunes Algebra to combine domain-specific resource policies (i.e., RSpecs) along the application workflow life cycle. We use standardized guidelines that help in compute/storage resource domain/location selection as well as network path selection based on both application QSpecs and SSpecs. We implement our security formalization and alignment methods as a framework, viz., "OnTimeURB" and apply it on an exemplar Distributed Computing workflow to show the benefits of joint QSpecs-SSpecs-driven, RSpecs-compliant federated workflow management.
Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Trupti Joshi, Dong Xu 0002
CLOUD5
2016 Network measurement recommendations for performance bottleneck correlation analysis
abstract
Multi-domain network performance monitoring (NPM) federations, such as perfSONAR rely on collaborative measurement intelligence to identify network anomaly events and diagnose performance bottlenecks affecting data-intensive science applications. In this paper, we present a novel measurement recommendation scheme to assist network operators and application users by recommending pertinent samples from a pool of measurement data involving multiple domains to detect and troubleshoot correlated network anomaly events. The recommendations are based on the principles of content-based filtering. Such recommendations are complimented with Bayesian Inference based domain reputation meta-information to strengthen the veracity information of the recommended traces. Using actual long-term and short-term perfSONAR traces, we analyze recommendation results and show: a) how the content-based filter recommends the most pertinent traces based on their attributes, and b) the time-variant characteristics of domain reputation. Finally, using synthetic traces, we show the effectiveness of our proposed measurements recommendation scheme in accurately identifying anomaly events for an exemplar use case, and also show how our content filter based recommendation scheme performs better in terms of false alarms in comparison to: a) recommendations that consider partial trace features for filtering, and b) greedy recommendation approaches based on random trace selection.
Yuanxun Zhang, Saptarshi Debroy, Prasad Calyam
LANMAN1
2016 PGen: large-scale genomic variations analysis workflow and browser in SoyKB
abstract
BACKGROUND: With the advances in next-generation sequencing (NGS) technology and significant reductions in sequencing costs, it is now possible to sequence large collections of germplasm in crops for detecting genome-scale genetic variations and to apply the knowledge towards improvements in traits. To efficiently facilitate large-scale NGS resequencing data analysis of genomic variations, we have developed "PGen", an integrated and optimized workflow using the Extreme Science and Engineering Discovery Environment (XSEDE) high-performance computing (HPC) virtual system, iPlant cloud data storage resources and Pegasus workflow management system (Pegasus-WMS). The workflow allows users to identify single nucleotide polymorphisms (SNPs) and insertion-deletions (indels), perform SNP annotations and conduct copy number variation analyses on multiple resequencing datasets in a user-friendly and seamless way. RESULTS: We have developed both a Linux version in GitHub ( https://github.com/pegasus-isi/PGen-GenomicVariations-Workflow ) and a web-based implementation of the PGen workflow integrated within the Soybean Knowledge Base (SoyKB), ( http://soykb.org/Pegasus/index.php ). Using PGen, we identified 10,218,140 single-nucleotide polymorphisms (SNPs) and 1,398,982 indels from analysis of 106 soybean lines sequenced at 15X coverage. 297,245 non-synonymous SNPs and 3330 copy number variation (CNV) regions were identified from this analysis. SNPs identified using PGen from additional soybean resequencing projects adding to 500+ soybean germplasm lines in total have been integrated. These SNPs are being utilized for trait improvement using genotype to phenotype prediction approaches developed in-house. In order to browse and access NGS data easily, we have also developed an NGS resequencing data browser ( http://soykb.org/NGS_Resequence/NGS_index.php ) within SoyKB to provide easy access to SNP and downstream analysis results for soybean researchers. CONCLUSION: PGen workflow has been optimized for the most efficient analysis of soybean data using thorough testing and validation. This research serves as an example of best practices for development of genomics data analysis workflows by integrating remote HPC resources and efficient data management with ease of use for biological users. PGen workflow can also be easily customized for analysis of data in other species.
Saad M. Khan, Juexin Wang, Mats Rynge, Yuanxun Zhang, Shiyuan Chen, João V. Maldonado dos Santos, Babu Valliyodan, Prasad Calyam, Nirav C. Merchant, Henry T. Nguyen, Dong Xu 0002, Trupti Joshi
BMC Bioinform.5
2016 Network-Wide Anomaly Event Detection and Diagnosis With perfSONAR
abstract
High-performance computing (HPC) environments supporting data-intensive applications need multidomain network performance measurements from open frameworks such as perfSONAR. Detected network-wide correlated anomaly events that impact data throughput performance need to be quickly and accurately notified along with a root-cause analysis for remediation. In this paper, we present a novel network anomaly events detection and diagnosis scheme for network-wide visibility that improves accuracy of root-cause analysis. We address analysis limitations in cases where there is absence of complete network topology information, and when measurement probes are mis-calibrated leading to erroneous diagnosis. Our proposed scheme fuses perfSONAR time-series path measurements data from multiple domains using principal component analysis (PCA) to transform data for accurate correlated and uncorrelated anomaly events detection. We quantify the certainty of such detection using a measurement data sanity checking that involves: 1) measurement data reputation analysis to qualify the measurement samples and 2) filter framework to prune potentially misleading samples. Lastly, using actual perfSONAR one-way delay measurement traces, we show our proposed scheme's effectiveness in diagnosing the root-cause of critical network performance anomaly events.
Yuanxun Zhang, Saptarshi Debroy, Prasad Calyam
IEEE Trans. Netw. Serv. Manag.1