VLDB 2026 Research / reviewers in the wild / expert
Xia Jiang
dblp:33/6067
· DBLP profile ↗
35ranked-venue papers
21as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
3 papers |
Mathematical optimization · 100% | |
| Artificial intelligence
3 papers |
Language models and text generation · 97% Probabilistic and Bayesian machine learning · 3% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Large Language Models as End-to-end Combinatorial Optimization Solvers · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model
LLM-based optimization |
0.9 | 1 | 2025 | Large Language Models as End-to-end Combinatorial Optimization Solvers · NeurIPS 2025 |
Rendering › parallel rendering
distributed rendering |
0.9 | 1 | 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture · IEEE Trans. Serv. Comput. 2025 |
Cloud and datacenter computing › cloud applications
cloud rendering |
0.9 | 1 | 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture · IEEE Trans. Serv. Comput. 2025 |
Mathematical optimization
combinatorial optimization |
0.9 | 1 | 2025 | Large Language Models as End-to-end Combinatorial Optimization Solvers · NeurIPS 2025 |
Mathematical optimization › continuous optimization
nonsmooth optimization |
0.9 | 1 | 2025 | Inexact proximal gradient algorithm with random reshuffling for nonsmooth optimization · Sci. China Inf. Sci. 2025 |
Mathematical optimization › continuous optimization › convex optimization › proximal methods
proximal gradient method |
0.9 | 1 | 2025 | Inexact proximal gradient algorithm with random reshuffling for nonsmooth optimization · Sci. China Inf. Sci. 2025 |
Mathematical optimization
solver code generation |
0.9 | 1 | 2025 | DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of Constraints · ICLR 2025 |
Mathematical optimization › combinatorial optimization
vehicle routing |
0.9 | 1 | 2025 | DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of Constraints · ICLR 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.3 | 1 | 2025 | DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of Constraints · ICLR 2025 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
random reshuffling |
0.3 | 1 | 2025 | Inexact proximal gradient algorithm with random reshuffling for nonsmooth optimization · Sci. China Inf. Sci. 2025 |
Mathematical optimization
stochastic optimization |
0.3 | 1 | 2025 | Inexact proximal gradient algorithm with random reshuffling for nonsmooth optimization · Sci. China Inf. Sci. 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network |
0.1 | 1 | 2006 | A Bayesian Network for Outbreak Detection and Prediction · AAAI 2006 |
Medical and health informatics › public health › public health informatics
outbreak detection |
0.0 | 1 | 2006 | A Bayesian Network for Outbreak Detection and Prediction · AAAI 2006 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.7self-debugging · 1.7retrieval-augmented generation · 1.7reinforcement learning · 1.7large language model · 1.7inexact proximal gradient · 0.9bayesian network · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive directional importance sampling with von Mises-Fisher mixture model for fuzzy reliability analysis
Xia Jiang, Kaichao Zhang |
Fuzzy Sets Syst. | 1 |
| 2025 | Single-Loop Variance-Reduced Stochastic Algorithm for Nonconvex-Concave Minimax OptimizationabstractNonconvex-concave (NC-C) finite-sum minimax problems have broad applications in decentralized optimization and various machine learning tasks. However, the nonsmooth nature of NC-C problems makes it challenging to design effective variance reduction techniques. Existing vanilla stochastic algorithms using uniform samples for gradient estimation often exhibit slow convergence rates and require bounded variance assumptions. In this paper, we develop a novel probabilistic variance reduction updating scheme and propose a single-loop algorithm called the probabilistic variance-reduced smoothed gradient descent-ascent (PVR-SGDA) algorithm. The proposed algorithm achieves an iteration complexity of ${\mathcal{O}}\left({{\varepsilon ^{ - 4}}}\right)$, surpassing the best-known rates of stochastic algorithms for NC-C minimax problems and matching the performance of the best deterministic algorithms in this context. Finally, we demonstrate the effectiveness of the proposed algorithm through numerical simulations. Xia Jiang, Linglingzhi Zhu, Taoli Zheng, Anthony Man-Cho So |
ICASSP | 1 |
| 2025 | DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of ConstraintsabstractThis paper proposes Decomposed Retrieval of Constraints (DRoC), a novel framework aimed at enhancing large language models (LLMs) in exploiting solvers to tackle vehicle routing problems (VRPs) with intricate constraints. While LLMs have shown promise in solving simple VRPs, their potential in addressing complex VRP variants is still suppressed, due to the limited embedded internal knowledge that is required to accurately reflect diverse VRP constraints. Our approach mitigates the issue by integrating external knowledge via a novel retrieval-augmented generation (RAG) approach. More specifically, the DRoC decomposes VRP constraints, externally retrieves information relevant to each constraint, and synergistically combines internal and external knowledge to benefit the program generation for solving VRPs. The DRoC also allows LLMs to dynamically select between RAG and self-debugging mechanisms, thereby optimizing program generation without the need for additional training. Experiments across 48 VRP variants exhibit the superiority of DRoC, with significant improvements in the accuracy rate and runtime error rate delivered by the generated programs. The DRoC framework has the potential to elevate LLM performance in complex optimization tasks, fostering the applicability of LLMs in industries such as transportation and logistics. Xia Jiang, Yaoxin Wu, Yingqian Zhang 0001 |
ICLR | 1 |
| 2025 | Large Language Models as End-to-end Combinatorial Optimization SolversabstractCombinatorial optimization (CO) problems, central to decision-making scenarios like logistics and manufacturing, are traditionally solved using problem-specific algorithms requiring significant domain expertise. While large language models (LLMs) have shown promise in automating CO problem solving, existing approaches rely on intermediate steps such as code generation or solver invocation, limiting their generality and accessibility. This paper introduces a novel framework that empowers LLMs to serve as end-to-end CO solvers by directly mapping natural language problem descriptions to solutions. We propose a two-stage training strategy: supervised fine-tuning (SFT) imparts LLMs with solution construction patterns from domain-specific solvers, while a feasibility-and-optimality-aware reinforcement learning (FOARL) process explicitly mitigates constraint violations and refines solution quality. Evaluation across seven NP-hard CO problems shows that our method achieves a high feasibility rate and reduces the average optimality gap to 1.03–8.20% by tuning a 7B-parameter LLM, surpassing both general-purpose LLMs (e.g., GPT-4o), reasoning models (e.g., DeepSeek-R1), and domain-specific heuristics. Our method establishes a unified language-based pipeline for CO without extensive code execution or manual architectural adjustments for different problems, offering a general and language-driven alternative to traditional solver design while maintaining relative feasibility guarantees. Xia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao, Yingqian Zhang 0001 |
NeurIPS | 1 |
| 2025 | Inexact proximal gradient algorithm with random reshuffling for nonsmooth optimization
Xia Jiang, Yanyan Fang, Xianlin Zeng, Jian Sun 0003, Jie Chen 0003 |
Sci. China Inf. Sci. | 1 |
| 2025 | A Multi-Scale Sparse Channel Transformer Network for image reconstruction of astronomical bright source contamination
Congcong Shen, Xia Jiang, A-Li Luo, Fuji Ren, Yuanlu Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | DSDNet: Target Detection Algorithm for SDSS Photometric Images Based on Convolutional Neural NetworksabstractThe Sloan Digital Sky Survey (SDSS), a large-scale astronomical survey project, has released a vast volume of photometric images. These images play a pivotal role in deriving fundamental parameters of celestial objects and investigating the structure of the universe. Nevertheless, in dense star fields, the characteristics of celestial sources are intricate, rendering traditional methods incapable of conducting precise analyses. To address this challenge, this paper introduces a new algorithm named DSDNet for detecting celestial sources in dense star fields and counting them. During the feature extraction phase, DSDNet generates larger feature maps, thereby preserving more information about small targets. By incorporating a convolutional attention module, the model's capacity to learn the features of celestial sources in dense star fields is augmented. Furthermore, to more effectively manage the blending phenomenon among sources, DSDNet integrates CNN and Transformer architectures, enhancing the model's ability to comprehend global features. Experimental findings show that DSDNet exhibits excellent performance, attaining an F1-score of 97.23%. This makes it a valuable resource for aiding SDSS in processing dense star field images. Xia Jiang, A-Li Luo, Fuji Ren |
IEEE Signal Process. Lett. | 3 |
| 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture
Yun Wang 0039, Bing Deng, Xia Jiang, Xuyan Hu, Dongjie Tang, Randy Xu, Yijin Sun, Zhengwei Qi |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Distributed Stochastic Proximal Algorithm With Random Reshuffling for Nonsmooth Finite-Sum OptimizationabstractThe nonsmooth finite-sum minimization is a fundamental problem in machine learning. This article develops a distributed stochastic proximal-gradient algorithm with random reshuffling to solve the finite-sum minimization over time-varying multiagent networks. The objective function is a sum of differentiable convex functions and nonsmooth regularization. Each agent in the network updates local variables by local information exchange and cooperates to seek an optimal solution. We prove that local variable estimates generated by the proposed algorithm achieve consensus and are attracted to a neighborhood of the optimal solution with an O((1/T)+(1/√T)) convergence rate, where T is the total number of iterations. Finally, some comparative simulations are provided to verify the convergence performance of the proposed algorithm. Xia Jiang, Xianlin Zeng, Jian Sun 0003, Jie Chen 0003, Lihua Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | An innovative ensemble model based on multiple neural networks and a novel heuristic optimization algorithm for COVID-19 forecasting
Zongxi Qu, Xia Jiang, Chunhua Niu |
Expert Syst. Appl. | 3 |
| 2023 | Learning the Policy for Mixed Electric Platoon Control of Automated and Human-Driven Vehicles at Signalized Intersection: A Random Search ApproachabstractThe upgrading and updating of vehicles have accelerated in the past decades. Out of the need for environmental friendliness and intelligence, electric vehicles (EVs) and connected and automated vehicles (CAVs) have become new components of transportation systems. This paper develops a reinforcement learning framework to implement adaptive control for an electric platoon composed of CAVs and human-driven vehicles (HDVs) at a signalized intersection. Firstly, a Markov Decision Process (MDP) model is proposed to describe the decision process of the mixed platoon. Novel state representation and reward function are designed for the model to consider the behavior of the whole platoon. Secondly, in order to deal with the delayed reward, an Augmented Random Search (ARS) algorithm is proposed. The control policy learned by the agent can guide the longitudinal motion of the CAV, which serves as the leader of the platoon. Finally, a series of simulations are carried out in simulation suite SUMO. Compared with several traditional reinforcement learning approaches, the proposed method can obtain a higher reward. Meanwhile, the simulation results demonstrate the effectiveness of the delay reward, which is designed to outperform distributed reward mechanism. Compared with some state-of-the-art optimization-based frameworks, the simulation analysis reveals that more energy can be saved for different sizes of platoon. Sensitivity analysis is also conducted by adjusting the relative importance of the optimization goal to show the flexibility of ARS and delay reward setting. On the premise that travel delay is not sacrificed, the proposed control method can save up to 53.64% electric energy. Xia Jiang, Jian Zhang 0011 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Distributed Stochastic Gradient Tracking Algorithm With Variance Reduction for Non-Convex OptimizationabstractThis article proposes a distributed stochastic algorithm with variance reduction for general smooth non-convex finite-sum optimization, which has wide applications in signal processing and machine learning communities. In distributed setting, a large number of samples are allocated to multiple agents in the network. Each agent computes local stochastic gradient and communicates with its neighbors to seek for the global optimum. In this article, we develop a modified variance reduction technique to deal with the variance introduced by stochastic gradients. Combining gradient tracking and variance reduction techniques, this article proposes a distributed stochastic algorithm, gradient tracking algorithm with variance reduction (GT-VR), to solve large-scale non-convex finite-sum optimization over multiagent networks. A complete and rigorous proof shows that the GT-VR algorithm converges to the first-order stationary points with$O({1}/{k})$convergence rate. In addition, we provide the complexity analysis of the proposed algorithm. Compared with some existing first-order methods, the proposed algorithm has a lower$\mathcal {O}(PM\epsilon ^{-1})$gradient complexity under some mild condition. By comparing state-of-the-art algorithms and GT-VR in numerical simulations, we verify the efficiency of the proposed algorithm. Xia Jiang, Xianlin Zeng, Jian Sun 0003, Jie Chen 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Developing and Evaluation of Computational Phenotypes of Metastatic Breast Cancer Using All of Us Data
Israel Dilan-Pantojas, Shyam Visweswaran, Michael J. Becich, Xia Jiang, Richard D. Boyce |
AMIA | 5 |
| 2022 | Distributed Solver for Discrete-Time Lyapunov Equations Over Dynamic Networks With Linear Convergence RateabstractThe problem of solving discrete-time Lyapunov equations (DTLEs) is investigated over multiagent network systems, where each agent has access to its local information and communicates with its neighbors. To obtain a solution to DTLE, a distributed algorithm with uncoordinated constant step sizes is proposed over time-varying topologies. The convergence properties and the range of constant step sizes of the proposed algorithm are analyzed. Moreover, a linear convergence rate is proved and the convergence performances over dynamic networks are verified by numerical simulations. Xia Jiang, Xianlin Zeng, Jian Sun 0003, Jie Chen 0003 |
IEEE Trans. Cybern. | 1 |
| 2020 | Leveraging Bayesian networks and information theory to learn risk factors for breast cancer metastasisabstractBACKGROUND: Even though we have established a few risk factors for metastatic breast cancer (MBC) through epidemiologic studies, these risk factors have not proven to be effective in predicting an individual's risk of developing metastasis. Therefore, identifying critical risk factors for MBC continues to be a major research imperative, and one which can lead to advances in breast cancer clinical care. The objective of this research is to leverage Bayesian Networks (BN) and information theory to identify key risk factors for breast cancer metastasis from data. METHODS: We develop the Markov Blanket and Interactive risk factor Learner (MBIL) algorithm, which learns single and interactive risk factors having a direct influence on a patient's outcome. We evaluate the effectiveness of MBIL using simulated datasets, and compare MBIL with the BN learning algorithms Fast Greedy Search (FGS), PC algorithm (PC), and CPC algorithm (CPC). We apply MBIL to learn risk factors for 5 year breast cancer metastasis using a clinical dataset we curated. We evaluate the learned risk factors by consulting with breast cancer experts and literature. We further evaluate the effectiveness of MBIL at learning risk factors for breast cancer metastasis by comparing it to the BN learning algorithms Necessary Path Condition (NPC) and Greedy Equivalent Search (GES). RESULTS: The averages of the Jaccard index for the simulated datasets containing 2000 records were 0.705, 0.272, 0.228, and 0.147 for MBIL, FGS, PC, and CPC respectively. MBIL, NPC, and GES all learned that grade and lymph_nodes_positive are direct risk factors for 5 year metastasis. Only MBIL and NPC found that surgical_margins is a direct risk factor. Only NPC found that invasive is a direct risk factor. MBIL learned that HER2 and ER interact to directly affect 5 year metastasis. Neither GES nor NPC learned that HER2 and ER are direct risk factors. DISCUSSION: The results involving simulated datasets indicated that MBIL can learn direct risk factors substantially better than standard Bayesian network learning algorithms. An application of MBIL to a real breast cancer dataset identified both single and interactive risk factors that directly influence breast cancer metastasis, which can be investigated further. Xia Jiang, Alan Wells, Adam Brufsky, Darshan Shetty, Kahmil Shajihan, Richard E. Neapolitan |
BMC Bioinform. | 1 |
| 2019 | Systematic discovery of the functional impact of somatic genome alterations in individual tumors through tumor-specific causal inferenceabstractCancer is mainly caused by somatic genome alterations (SGAs). Precision oncology involves identifying and targeting tumor-specific aberrations resulting from causative SGAs. We developed a novel tumor-specific computational framework that finds the likely causative SGAs in an individual tumor and estimates their impact on oncogenic processes, which suggests the disease mechanisms that are acting in that tumor. This information can be used to guide precision oncology. We report a tumor-specific causal inference (TCI) framework, which estimates causative SGAs by modeling causal relationships between SGAs and molecular phenotypes (e.g., transcriptomic, proteomic, or metabolomic changes) within an individual tumor. We applied the TCI algorithm to tumors from The Cancer Genome Atlas (TCGA) and estimated for each tumor the SGAs that causally regulate the differentially expressed genes (DEGs) in that tumor. Overall, TCI identified 634 SGAs that are predicted to cause cancer-related DEGs in a significant number of tumors, including most of the previously known drivers and many novel candidate cancer drivers. The inferred causal relationships are statistically robust and biologically sensible, and multiple lines of experimental evidence support the predicted functional impact of both the well-known and the novel candidate drivers that are predicted by TCI. TCI provides a unified framework that integrates multiple types of SGAs and molecular phenotypes to estimate which genome perturbations are causally influencing one or more molecular/cellular phenotypes in an individual tumor. By identifying major candidate drivers and revealing their functional impact in an individual tumor, TCI sheds light on the disease mechanisms of that tumor, which can serve to advance our basic knowledge of cancer biology and to support precision oncology that provides tailored treatment of individual tumors. Chunhui Cai, Gregory F. Cooper, Kevin N. Lu, Shuping Xu, Zhenlong Zhao, Xueer Chen, Adrian V. Lee, Nathan Clark, Vicky Chen, Songjian Lu, Lujia Chen 0001, Liyue Yu, Harry Hochheiser, Xia Jiang, Q. Jane Wang, Xinghua Lu 0001 |
PLoS Comput. Biol. | 16 |
| 2018 | Distributed Algorithm for Discrete-Time Lyapunov EquationsabstractThis paper investigates the problem of solving a unique solution to discrete-time Lyapunov equations (DTLE) using multi-agent networks. We propose a distributed algorithm where each agent only uses partial information of the matrices. The agents of the algorithm reach a consensus by exchanging information with their neighbors over an undirected connected graph. We provide convergence analysis and the convergence rate estimate for the proposed algorithm. Finally, convergence performance is verified by numerical simulations. Xia Jiang, Xianlin Zeng, Jian Sun 0003, Jie Chen 0003 |
ICARCV | 1 |
| 2018 | Using natural language processing and machine learning to identify breast cancer local recurrenceabstractBACKGROUND: Identifying local recurrences in breast cancer from patient data sets is important for clinical research and practice. Developing a model using natural language processing and machine learning to identify local recurrences in breast cancer patients can reduce the time-consuming work of a manual chart review. METHODS: We design a novel concept-based filter and a prediction model to detect local recurrences using EHRs. In the training dataset, we manually review a development corpus of 50 progress notes and extract partial sentences that indicate breast cancer local recurrence. We process these partial sentences to obtain a set of Unified Medical Language System (UMLS) concepts using MetaMap, and we call it positive concept set. We apply MetaMap on patients' progress notes and retain only the concepts that fall within the positive concept set. These features combined with the number of pathology reports recorded for each patient are used to train a support vector machine to identify local recurrences. RESULTS: We compared our model with three baseline classifiers using either full MetaMap concepts, filtered MetaMap concepts, or bag of words. Our model achieved the best AUC (0.93 in cross-validation, 0.87 in held-out testing). CONCLUSIONS: Compared to a labor-intensive chart review, our model provides an automated way to identify breast cancer local recurrences. We expect that by minimally adapting the positive concept set, this study has the potential to be replicated at other institutions with a moderately sized training dataset. Zexian Zeng, Sasa Espino, Ankita Roy, Xiaoyu Li 0006, Seema A. Khan, Susan E. Clare, Xia Jiang, Richard E. Neapolitan, Yuan Luo 0001 |
BMC Bioinform. | 7 |
| 2017 | An algorithm for direct causal learning of influences on patient outcomes
Chandramouli Rathnam, Xia Jiang |
Artif. Intell. Medicine | 3 |
| 2016 | Computational methods for ubiquitination site prediction using physicochemical properties of protein sequencesabstractBACKGROUND: Ubiquitination is a very important process in protein post-translational modification, which has been widely investigated by biology scientists and researchers. Different experimental and computational methods have been developed to identify the ubiquitination sites in protein sequences. This paper aims at exploring computational machine learning methods for the prediction of ubiquitination sites using the physicochemical properties (PCPs) of amino acids in the protein sequences. RESULTS: We first establish six different ubiquitination data sets, whose records contain both ubiquitination sites and non-ubiquitination sites in variant numbers of protein sequence segments. In particular, to establish such data sets, protein sequence segments are extracted from the original protein sequences used in four published papers on ubiquitination, while 531 PCP features of each extracted protein sequence segment are calculated based on PCP values from AAindex (Amino Acid index database) by averaging PCP values of all amino acids on each segment. Various computational machine-learning methods, including four Bayesian network methods (i.e., Naïve Bayes (NB), Feature Selection NB (FSNB), Model Averaged NB (MANB), and Efficient Bayesian Multivariate Classifier (EBMC)) and three regression methods (i.e., Support Vector Machine (SVM), Logistic Regression (LR), and Least Absolute Shrinkage and Selection Operator (LASSO)), are then applied to the six established segment-PCP data sets. Five-fold cross-validation and the Area Under Receiver Operating Characteristic Curve (AUROC) are employed to evaluate the ubiquitination prediction performance of each method. Results demonstrate that the PCP data of protein sequences contain information that could be mined by machine learning methods for ubiquitination site prediction. The comparative results show that EBMC, SVM and LR perform better than other methods, and EBMC is the only method that can get AUCs greater than or equal to 0.6 for the six established data sets. Results also show EBMC tends to perform better for larger data. CONCLUSIONS: Machine learning methods have been employed for the ubiquitination site prediction based on physicochemical properties of amino acids on protein sequences. Results demonstrate the effectiveness of using machine learning methodology to mine information from PCP data concerning protein sequences, as well as the superiority of EBMC, SVM and LR (especially EBMC) for the ubiquitination prediction compared to other methods. Binghuang Cai, Xia Jiang |
BMC Bioinform. | 2 |
| 2016 | Discovering causal interactions using Bayesian network scoring and information gainabstractBACKGROUND: The problem of learning causal influences from data has recently attracted much attention. Standard statistical methods can have difficulty learning discrete causes, which interacting to affect a target, because the assumptions in these methods often do not model discrete causal relationships well. An important task then is to learn such interactions from data. Motivated by the problem of learning epistatic interactions from datasets developed in genome-wide association studies (GWAS), researchers conceived new methods for learning discrete interactions. However, many of these methods do not differentiate a model representing a true interaction from a model representing non-interacting causes with strong individual affects. The recent algorithm MBS-IGain addresses this difficulty by using Bayesian network learning and information gain to discover interactions from high-dimensional datasets. However, MBS-IGain requires marginal effects to detect interactions containing more than two causes. If the dataset is not high-dimensional, we can avoid this shortcoming by doing an exhaustive search. RESULTS: We develop Exhaustive-IGain, which is like MBS-IGain but does an exhaustive search. We compare the performance of Exhaustive-IGain to MBS-IGain using low-dimensional simulated datasets based on interactions with marginal effects and ones based on interactions without marginal effects. Their performance is similar on the datasets based on marginal effects. However, Exhaustive-IGain compellingly outperforms MBS-IGain on the datasets based on 3 and 4-cause interactions without marginal effects. We apply Exhaustive-IGain to investigate how clinical variables interact to affect breast cancer survival, and obtain results that agree with judgements of a breast cancer oncologist. CONCLUSIONS: We conclude that the combined use of information gain and Bayesian network scoring enables us to discover higher order interactions with no marginal effects if we perform an exhaustive search. We further conclude that Exhaustive-IGain can be effective when applied to real data. Zexian Zeng, Xia Jiang, Richard E. Neapolitan |
BMC Bioinform. | 2 |
| 2016 | An informatics research agenda to support precision medicine: seven key areasabstractThe recent announcement of the Precision Medicine Initiative by President Obama has brought precision medicine (PM) to the forefront for healthcare providers, researchers, regulators, innovators, and funders alike. As technologies continue to evolve and datasets grow in magnitude, a strong computational infrastructure will be essential to realize PM's vision of improved healthcare derived from personal data. In addition, informatics research and innovation affords a tremendous opportunity to drive the science underlying PM. The informatics community must lead the development of technologies and methodologies that will increase the discovery and application of biomedical knowledge through close collaboration between researchers, clinicians, and patients. This perspective highlights seven key areas that are in need of further informatics research and innovation to support the realization of PM. Jessica D. Tenenbaum, Paul Avillach, Marge M. Benham-Hutchins, Matthew K. Breitenstein, Erin L. Crowgey, Mark A. Hoffman, Xia Jiang, Subha Madhavan, John E. Mattison, Radhakrishnan Nagarajan, Bisakha Ray, Dmitriy Shin, Shyam Visweswaran, Zhongming Zhao, Robert R. Freimuth |
J. Am. Medical Informatics Assoc. | 7 |
| 2015 | Evaluation of a two-stage framework for prediction using big genomic dataabstractWe are in the era of abundant 'big' or 'high-dimensional' data. These data afford us the opportunity to discover predictors of an event of interest, and to estimate occurrence of the event based on values of these predictors. For example, 'genome-wide association studies' examine millions of single-nucleotide polymorphisms (SNPs), along with disease status. We can learn SNPs that affect disease status from these data sets, and use the knowledge learned to predict disease likelihood. Owing to the large number of features, it is difficult for many prediction methods to use all the features directly. The ReliefF algorithm ranks a set of features in terms of how well they predict a target. It can be used to identify good predictors, which can then be provided to a prediction method. We compared the performance of eight prediction methods when predicting binary outcomes using high-dimensional discrete data sets. We performed two-stage prediction, where ReliefF is used in the first stage to identify good predictors. Bayesian network (BN)-based methods performed best overall. Furthermore, ReliefF did not improve their performance. The BN-based methods use the Bayesian Dirichlet Equivalent Uniform score to evaluate candidate models, and use BN inference algorithms to perform prediction. This score and these algorithms were developed for discrete variables. This perhaps explains why they perform better in this domain. Many prediction methods are available, and researchers have little reason for choosing one over the other in the domain of binary prediction using high-dimensional data sets. Our results indicate that the best choices overall are BN-based methods. Xia Jiang, Richard E. Neapolitan |
Briefings Bioinform. | 1 |
| 2014 | Understanding the SIMD Efficiency of Graph Traversal on GPU
Yichao Cheng, Hong An, Zhitao Chen, Xia Jiang |
ICA3PP (1) | 6 |
| 2014 | A novel artificial neural network method for biomedical prediction based on matrix pseudo-inversion
Binghuang Cai, Xia Jiang |
J. Biomed. Informatics | 2 |
| 2011 | Learning genetic epistasis using Bayesian network scoring criteriaabstractBACKGROUND: Gene-gene epistatic interactions likely play an important role in the genetic basis of many common diseases. Recently, machine-learning and data mining methods have been developed for learning epistatic relationships from data. A well-known combinatorial method that has been successfully applied for detecting epistasis is Multifactor Dimensionality Reduction (MDR). Jiang et al. created a combinatorial epistasis learning method called BNMBL to learn Bayesian network (BN) epistatic models. They compared BNMBL to MDR using simulated data sets. Each of these data sets was generated from a model that associates two SNPs with a disease and includes 18 unrelated SNPs. For each data set, BNMBL and MDR were used to score all 2-SNP models, and BNMBL learned significantly more correct models. In real data sets, we ordinarily do not know the number of SNPs that influence phenotype. BNMBL may not perform as well if we also scored models containing more than two SNPs. Furthermore, a number of other BN scoring criteria have been developed. They may detect epistatic interactions even better than BNMBL.Although BNs are a promising tool for learning epistatic relationships from data, we cannot confidently use them in this domain until we determine which scoring criteria work best or even well when we try learning the correct model without knowledge of the number of SNPs in that model. RESULTS: We evaluated the performance of 22 BN scoring criteria using 28,000 simulated data sets and a real Alzheimer's GWAS data set. Our results were surprising in that the Bayesian scoring criterion with large values of a hyperparameter called α performed best. This score performed better than other BN scoring criteria and MDR at recall using simulated data sets, at detecting the hardest-to-detect models using simulated data sets, and at substantiating previous results using the real Alzheimer's data set. CONCLUSIONS: We conclude that representing epistatic interactions using BN models and scoring them using a BN scoring criterion holds promise for identifying epistatic genetic variants in data. In particular, the Bayesian scoring criterion with large values of a hyperparameter α appears more promising than a number of alternatives. Xia Jiang, Richard E. Neapolitan, M. Michael Barmada, Shyam Visweswaran |
BMC Bioinform. | 1 |
| 2010 | A real-time temporal Bayesian architecture for event surveillance and its application to patient-specific multiple disease outbreak detection
Xia Jiang, Gregory F. Cooper |
Data Min. Knowl. Discov. | 1 |
| 2010 | A Bayesian network model for spatial event surveillance
Xia Jiang, Daniel B. Neill, Gregory F. Cooper |
Int. J. Approx. Reason. | 1 |
| 2010 | A Bayesian spatio-temporal method for disease outbreak detectionabstractA system that monitors a region for a disease outbreak is called a disease outbreak surveillance system. A spatial surveillance system searches for patterns of disease outbreak in spatial subregions of the monitored region. A temporal surveillance system looks for emerging patterns of outbreak disease by analyzing how patterns have changed during recent periods of time. If a non-spatial, non-temporal system could be converted to a spatio-temporal one, the performance of the system might be improved in terms of early detection, accuracy, and reliability. A Bayesian network framework is proposed for a class of space-time surveillance systems called BNST. The framework is applied to a non-spatial, non-temporal disease outbreak detection system called PC in order to create the spatio-temporal system called PCTS. Differences in the detection performance of PC and PCTS are examined. The results show that the spatio-temporal Bayesian approach performs well, relative to the non-spatial, non-temporal approach. Xia Jiang, Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Generalized AMOC Curves For Evaluation and Improvement of Event Surveillance
Xia Jiang, Gregory F. Cooper, Daniel B. Neill |
AMIA | 1 |
| 2009 | Bayesian prediction of an epidemic curve
Xia Jiang, Garrick L. Wallstrom, Gregory F. Cooper, Michael M. Wagner 0001 |
J. Biomed. Informatics | 1 |
| 2007 | A Recursive Algorithm for Spatial Cluster Detection
Xia Jiang, Gregory F. Cooper |
AMIA | 1 |
| 2006 | A Bayesian Network for Outbreak Detection and Prediction
Xia Jiang, Garrick L. Wallstrom |
AAAI | 1 |
| 2005 | Sea surface wind and cold tongue over the winter South China Sea
Qinyu Liu, Xia Jiang, Shang-Ping Xie, W. Timothy Liu |
IGARSS | 2 |
| 2004 | An information dissemination protocol for an ad hoc networkabstractPrevious research has illustrated that location-based routing protocols improve the effectiveness of mobile ad hoc network (MANET) routing. The goal of a location server, which may be used in conjunction with a location-based routing protocol, is to provide accurate location information on the mobile nodes in the network. In this paper, we propose an information dissemination protocol called LEAP (legend exchange and augmentation protocol), which we use to implement a location server for a MANET. We compare our legend-based location server to three other location service alternatives via extensive simulations, and illustrate that our LEAP implementation offers both effective accuracy and low overhead. Xia Jiang, Tracy Camp |
IPCCC | 1 |