Dongxiao Zhu

dblp:15/6233 · DBLP profile ↗
← Back
56ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-0225-7817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Trauma-Informed Data Donation: Integrating Expert and Donor Perspectives on Designing Against Re-Traumatization During Collection of Sexual Violence Data
abstract
Data donation has received attention as a more consensual means of collecting personal data for scientific inquiry and AI technology. Yet the nature of data often donated–such as harmful online messages and menstrual tracking logs–carries risk of retraumatization (the forced reliving of traumatic experience). While the well-being of data donors is considered in prior work, approaches to retraumatization remain ad hoc. We present Trauma-Informed Data Donation (TIDD): a context-specific, exploratory design framework for adapting the Trauma-Informed Approach (TIA) from the Public Health domain to data donation. TIDD was the product of a 2-year research through design process with experts on sexual violence and trauma, and observational interviews of data donors. We use a case study applying TIDD to our custom data donation platform, Ube, as an invitation for designers to consider how TIDD could be used as a malleable foundation for donation of data associated with other forms of trauma.
Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko
DIS8
2026 Not All Tokens Are Meant to Be Forgotten
abstract
Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they tend to memorize unwanted information, such as private or copyrighted content, raising significant privacy and legal concerns. Unlearning has emerged as a promising solution, but existing methods face a significant challenge of over-forgetting. This issue arises because they indiscriminately suppress the generation of all the tokens in forget samples, leading to a substantial loss of model utility. To overcome this challenge, we introduce the Targeted Information Forgetting (TIF) framework, which consists of (1) a flexible targeted information identifier designed to differentiate between unwanted words (UW) and general words (GW) in the forget samples, and (2) a novel Targeted Preference Optimization approach that leverages Logit Preference Loss to unlearn unwanted information associated with UW and Preservation Loss to retain general information in GW, effectively improving the unlearning process while mitigating utility degradation. Extensive experiments on the TOFU and MUSE benchmarks demonstrate that the proposed TIF framework enhances unlearning effectiveness while preserving model utility and achieving state-of-the-art results.
Xiangyu Zhou 0001, Yao Qiang, Saleh Zare Zade, Douglas Zytko, Prashant Khanduri, Dongxiao Zhu
AAAI6
2026 Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations
Ujunwa Mgboh, Rafi Ibn Sultan, Joshua Kim, Kundan Thind, Dongxiao Zhu
AIME (2)5
2025 GeoSAM: Fine-Tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation
abstract
In geographical image segmentation, performance is often constrained by the limited availability of training data and a lack of generalizability, particularly for segmenting mobility infrastructure such as roads, sidewalks, and crosswalks. Vision foundation models like the Segment Anything Model (SAM), pre-trained on millions of natural images, have demonstrated impressive zero-shot segmentation performance, providing a potential solution. However, SAM struggles with geographical images, such as aerial and satellite imagery, due to its training being confined to natural images and the narrow features and textures of these objects blending into their surroundings. To address these challenges, we propose Geographical SAM (GeoSAM), a SAM-based framework that fine-tunes SAM using automatically generated multi-modal prompts. Specifically, GeoSAM integrates point prompts from a pre-trained task-specific model as primary visual guidance, and text prompts generated by a large language model as secondary semantic guidance, enabling the model to better capture both spatial structure and contextual meaning. GeoSAM outperforms existing approaches for mobility infrastructure segmentation in both familiar and completely unseen regions by at least 5% in mIoU, representing a significant leap in leveraging foundation models to segment mobility infrastructure, including both road and pedestrian infrastructure in geographical images. The source code is publicly available.
Rafi Ibn Sultan, Chengyin Li, Hui Zhu 0016, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
ECAI6
2025 Automatic Calibration for Membership Inference Attack on Large Language Models
abstract
Membership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, existing methods often misinfer non-members as members, leading to a high false positive rate, or depend on additional reference models for probability calibration, which limits their practicality. To overcome these challenges, we introduce a novel framework called Automatic Calibration Membership Inference Attack (ACMIA), which utilizes a tunable temperature to calibrate output probabilities effectively. This approach is inspired by our theoretical insights into maximum likelihood estimation during the pre-training of LLMs. We introduce ACMIA in three configurations designed to accommodate different levels of model access and increase the probability gap between members and non-members, improving the reliability and robustness of membership inference. Extensive experiments on various open-source LLMs demonstrate that our proposed attack is highly effective, robust, and generalizable, surpassing state-of-the-art baselines across three widely used benchmarks. The source code is publicly available.
Saleh Zare Zade, Yao Qiang, Xiangyu Zhou 0001, Hui Zhu 0016, Mohammad Amin Roshani, Prashant Khanduri, Dongxiao Zhu
ECAI7
2025 Interpretability-Aware Vision Transformer
abstract
Vision Transformers (ViTs) have become prominent models for solving various vision tasks. However, the interpretability of ViTs has not kept pace with their promising performance. While there has been a surge of interest in developing post hoc solutions to explain ViTs’ outputs, these methods do not generalize to different downstream tasks and various transformer architectures. Furthermore, if ViTs are not properly trained with the given data and do not prioritize the region of interest, the post hoc methods become less effective. To overcome this limitation, we introduce a novel training procedure that inherently enhances ViT’s interpretability. Our interpretability-aware ViT (IA-ViT) draws inspiration from a fresh insight: both the class patch and image patches consistently generate predicted distributions and attention maps. IA-ViT is composed of a feature extractor, a predictor, and an interpreter, which are trained jointly with an interpretability-aware training objective. Consequently, the interpreter simulates the behavior of the predictor and provides a faithful explanation through its single-head self-attention mechanism. Our comprehensive experimental results demonstrate the effectiveness of IA-ViT in several image classification tasks, with both qualitative and quantitative evaluations of model performance and interpretability. Our code is available at: https://github.com/qiangyao1988/IA-ViT.
Yao Qiang, Chengyin Li, Hui Zhu 0016, Prashant Khanduri, Dongxiao Zhu
IJCNN5
2025 AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation
abstract
Segment Anything Model (SAM) is one of the pioneering prompt-based foundation models for image segmentation and has been rapidly adopted for various medical imaging applications. However, in clinical settings, cre-ating effective prompts is notably challenging and time-consuming, requiring the expertise of domain specialists such as physicians. This requirement significantly dimin-ishes SAM's primary advantage—its interactive capability with end users—in medical applications. Moreover, recent studies have indicated that SAM, originally designed for 2D natural images, performs suboptimally on 3D medical image segmentation tasks. This subpar performance is attributed to the domain gaps between natural and medical images and the disparities in spatial arrangements between 2D and 3D images, particularly in multi-organ segmentation applications. To overcome these challenges, we present a novel technique termed AutoProSAM. This method au-tomates 3D multi-organ CT-based segmentation by lever-aging SAM's foundational model capabilities without relying on domain experts for prompts. The approach utilizes parameter-efficient adaptation techniques to adapt SAMfor 3D medical imagery and incorporates an effective automatic prompt learning paradigm specific to this domain. By eliminating the need for manual prompts, it enhances SAM's capabilities for 3D medical image segmentation and achieves state-of-the-art (SOTA) performance in CT-based multi-organ segmentation tasks. The code is in this link.
Chengyin Li, Rafi Ibn Sultan, Prashant Khanduri, Yao Qiang, Chetty J. Indrin, Dongxiao Zhu
WACV6
2025 MulModSeg: Enhancing Unpaired Multi-Modal Medical Image Segmentation with Modality-Conditioned Text Embedding and Alternating Training
abstract
In the diverse field of medical imaging, automatic segmentation has numerous applications and must handle a wide variety of input domains, such as different types of Computed Tomography (CT) scans and Magnetic Resonance (MR) images. This heterogeneity challenges automatic segmentation algorithms to maintain consistent performance across different modalities due to the requirement for spatially aligned and paired images. Typically, segmentation models are trained using a single modality, which limits their ability to generalize to other types of input data without employing transfer learning techniques. Additionally, leveraging complementary information from different modalities to enhance segmentation precision often necessitates substantial modifications to popular encoder-decoder designs, such as introducing multiple branched encoding or decoding paths for each modality. In this work, we propose a simple Multi-Modal Segmentation (MulModSeg) strategy to enhance medical image segmentation across multiple modalities, specifically CT and MR. It incorporates two key designs: a modality-conditioned text embedding framework via a frozen text encoder that adds modality awareness to existing segmentation frameworks without significant structural modifications or computational overhead, and an alternating training procedure that facilitates the integration of essential features from unpaired CT and MR inputs. Through extensive experiments with both Fully Convolutional Network and Transformer-based backbones, MulModSeg consistently outperforms previous methods in segmenting abdominal multi-organ and cardiac substructures for both CT and MR modalities. The code is available in this link.
Chengyin Li, Hui Zhu 0016, Rafi Ibn Sultan, Hassan Bagher-Ebadian, Prashant Khanduri, Chetty J. Indrin, Kundan Thind, Dongxiao Zhu
WACV8
2025 Collective Consent: Who Needs to Consent to the Donation of Data Representing Multiple People?
abstract
Data donation is a growing form of personal data collection that foregrounds consent and conscious participation of the data donor. There remains little guidance on who must consent to data donation, particularly when the data represents multiple people. We provide empirical perspectives on this question through in-situ observation and interviews (N=18) with online daters who chose to donate messaging interactions with potential sexual partners for sexual violence research. Findings elucidate two diverging perspectives. Participants advocating for ''unilateral consent'' argued that consent of their messaging partner is not necessary, in part, because the anticipated benefit of data donation superseded consent. Participants advocating for ''collective consent'' wanted both messaging partners to consent to its donation, citing concerns for privacy of, and personal relationships with, the other person. Findings suggest that collective consent interfaces should be incorporated in data donation platforms, even if not strictly required by legal regulation, to improve donation of multi-person data.
Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko
Proc. ACM Hum. Comput. Interact.8
2024 MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural Networks
abstract
To better understand the output of deep neural networks (DNN), attribution based methods have been an important approach for model interpretability, which assign a score for each input dimension to indicate its importance towards the model outcome. Notably, the attribution methods use the ax- ioms of sensitivity and implementation invariance to ensure the validity and reliability of attribution results. Yet, the ex- isting attribution methods present challenges for effective in- terpretation and efficient computation. In this work, we in- troduce MFABA, an attribution algorithm that adheres to ax- ioms, as a novel method for interpreting DNN. Addition- ally, we provide the theoretical proof and in-depth analy- sis for MFABA algorithm, and conduct a large scale exper- iment. The results demonstrate its superiority by achieving over 101.5142 times faster speed than the state-of-the-art at- tribution algorithms. The effectiveness of MFABA is thor- oughly evaluated through the statistical analysis in compar- ison to other methods, and the full implementation package is open-source at: https://github.com/LMBTough/MFABA.
Huaming Chen, Jiayu Zhang 0001, Xinyi Wang 0005, Zhibo Jin, Minhui Xue 0001, Dongxiao Zhu, Kim-Kwang Raymond Choo
AAAI7
2024 "It's Not What We Were Trying to Get At, but I Think Maybe It Should Be": Learning How to Do Trauma-Informed Design with a Data Donation Platform for Online Dating Sexual Violence
abstract
A majority of people experience trauma, spurring calls to incorporate trauma-informed approaches (TIA) from public health and social work into technology design. While technologies touted as trauma-informed are starting to propagate the literature, there persists a gap in knowledge around how design teams apply TIA and qualify their technology as adhering to trauma-informed principles. We address this through a 12-month development project with trauma and sexual violence experts to produce Ube, a data donation platform for collecting online dating sexual consent data to improve sexual risk detection AI. Through analysis of design documentation we retrospectively articulate a trauma-informed design process that evolved through the course of Ube’s development, comprising three elements for integrating trauma-informed principles: design goals that adapt the definition of TIA to the application domain, design activities that map to trauma-informed principles, and consequent design choices. We conclude with methodological recommendations to improve trauma-informed design processes.
Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko
CHI8
2024 Fairness-Aware Vision Transformer via Debiased Self-Attention
Yao Qiang, Chengyin Li, Prashant Khanduri, Dongxiao Zhu
ECCV (37)4
2024 Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph
abstract
This paper introduces a novel information retrieval (IR) task of Conversational Entity Retrieval from a Knowledge Graph (CER-KG), which extends non-conversational entity retrieval from a knowledge graph (KG) to the conversational scenario. The user queries in CER-KG dialog turns may rely on the results of the preceding turns, which are KG entities. Similar to the conversational document IR, CER-KG can be viewed as a sequence of interrelated ranking tasks. To enable future research on CER-KG, we created QBLink-KG, a publicly available benchmark that was adapted from QBLink, a benchmark for text-based conversational reading comprehension of Wikipedia. As an initial approach to CER-KG, we experimented with Transformer- and LSTM-based query encoders in combination with the Neural Architecture for Conversational Entity Retrieval (NACER), our proposed feature-based neural architecture for entity ranking in CER-KG. NACER computes the ranking score of a candidate KG entity by taking into account diverse lexical and semantic matching signals between various KG components in its neighborhood, such as entities, categories, and literals, as well as entities in the results of the preceding turns in dialog history. The reported experimental results reveal the key challenges of CER-KG along with the possible directions for new approaches to this task.
Mona Zamiri, Yao Qiang, Fedor Nikolaev, Dongxiao Zhu, Alexander Kotov 0001
WWW4
2023 Learning Compact Features via In-Training Representation Alignment
abstract
Deep neural networks (DNNs) for supervised learning can be viewed as a pipeline of the feature extractor (i.e., last hidden layer) and a linear classifier (i.e., output layer) that are trained jointly with stochastic gradient descent (SGD) on the loss function (e.g., cross-entropy). In each epoch, the true gradient of the loss function is estimated using a mini-batch sampled from the training set and model parameters are then updated with the mini-batch gradients. Although the latter provides an unbiased estimation of the former, they are subject to substantial variances derived from the size and number of sampled mini-batches, leading to noisy and jumpy updates. To stabilize such undesirable variance in estimating the true gradients, we propose In-Training Representation Alignment (ITRA) that explicitly aligns feature distributions of two different mini-batches with a matching loss in the SGD training process. We also provide a rigorous analysis of the desirable effects of the matching loss on feature representation learning: (1) extracting compact feature representation; (2) reducing over-adaption on mini-batches via an adaptively weighting mechanism; and (3) accommodating to multi-modalities. Finally, we conduct large-scale experiments on both image and text classifications to demonstrate its superior performance to the strong baselines.
Xin Li 0081, Xiangrui Li, Deng Pan 0001, Yao Qiang, Dongxiao Zhu
AAAI5
2023 Negative Flux Aggregation to Estimate Feature Attributions
abstract
There are increasing demands for understanding deep neural networks' (DNNs) behavior spurred by growing security and/or transparency concerns. Due to multi-layer nonlinearity of the deep neural network architectures, explaining DNN predictions still remains as an open problem, preventing us from gaining a deeper understanding of the mechanisms. To enhance the explainability of DNNs, we estimate the input feature's attributions to the prediction task using divergence and flux. Inspired by the divergence theorem in vector analysis, we develop a novel Negative Flux Aggregation (NeFLAG) formulation and an efficient approximation algorithm to estimate attribution map. Unlike the previous techniques, ours doesn't rely on fitting a surrogate model nor need any path integration of gradients. Both qualitative and quantitative experiments demonstrate a superior performance of NeFLAG in generating more faithful attribution maps than the competing methods. Our code is available at https://github.com/xinli0928/NeFLAG.
Xin Li 0081, Deng Pan 0001, Chengyin Li, Yao Qiang, Dongxiao Zhu
IJCAI5
2023 FocalUNETR: A Focal Transformer for Boundary-Aware Prostate Segmentation Using CT Images
Chengyin Li, Yao Qiang, Rafi Ibn Sultan, Hassan Bagher-Ebadian, Prashant Khanduri, Indrin J. Chetty, Dongxiao Zhu
MICCAI (3)7
2022 Counterfactual Interpolation Augmentation (CIA): A Unified Approach to Enhance Fairness and Explainability of DNN
abstract
Bias in the training data can jeopardize fairness and explainability of deep neural network prediction on test data. We propose a novel bias-tailored data augmentation approach, Counterfactual Interpolation Augmentation (CIA), attempting to debias the training data by d-separating the spurious correlation between the target variable and the sensitive attribute. CIA generates counterfactual interpolations along a path simulating the distribution transitions between the input and its counterfactual example. CIA as a pre-processing approach enjoys two advantages: First, it couples with either plain training or debiasing training to markedly increase fairness over the sensitive attribute. Second, it enhances the explainability of deep neural networks by generating attribution maps via integrating counterfactual gradients. We demonstrate the superior performance of the CIA-trained deep neural network models using qualitative and quantitative experimental results. Our code is available at: https://github.com/qiangyao1988/CIA
Yao Qiang, Chengyin Li, Marco Brocanelli, Dongxiao Zhu
IJCAI4
2022 Tiny RNN Model with Certified Robustness for Text Classification
abstract
Mobile artificial intelligence has recently gained more attention due to the increasing computing power of mobile devices and applications in computer vision, natural language processing, and internet of things. Although large pre-trained language models (e.g., BERT, GPT) have recently achieved the state-of-the-art results on text classification tasks, they are not well suited for latency critical applications on mobile devices. Therefore, it is essential to design tiny models to reduce their memory and computing requirements. Model compression has shown promising results for this goal. However, some significant challenges are yet to be addressed, such as information loss and adversarial robustness. This paper attempts to tackle these challenges through a new training scheme that minimizes the information loss by maximizing the mutual information between the feature representations learned from the large and tiny models. In addition, we propose a certifiably robust defense method named GradMASK that masks a certain proportion of words in an input text. It can defend against both character-level perturbations and word substitution-based attacks. We perform extensive experiments demonstrating the effectiveness of our approach by comparing our tiny RNN models with compact RNNs (e.g., FastGRNN) and compressed RNNs (e.g., PRADO) in clean and adversarial test settings.
Yao Qiang, Supriya Tumkur Suresh Kumar, Marco Brocanelli, Dongxiao Zhu
IJCNN4
2022 AttCAT: Explaining Transformers via Attentive Class Activation Tokens
abstract
Transformers have improved the state-of-the-art in various natural language processing and computer vision tasks. However, the success of the Transformer model has not yet been duly explained. Current explanation techniques, which dissect either the self-attention mechanism or gradient-based attribution, do not necessarily provide a faithful explanation of the inner workings of Transformers due to the following reasons: first, attention weights alone without considering the magnitudes of feature values are not adequate to reveal the self-attention mechanism; second, whereas most Transformer explanation techniques utilize self-attention module, the skip-connection module, contributing a significant portion of information flows in Transformers, has not yet been sufficiently exploited in explanation; third, the gradient-based attribution of individual feature does not incorporate interaction among features in explaining the model's output. In order to tackle the above problems, we propose a novel Transformer explanation technique via attentive class activation tokens, aka, AttCAT, leveraging encoded features, their gradients, and their attention weights to generate a faithful and confident explanation for Transformer's output. Extensive experiments are conducted to demonstrate the superior performance of AttCAT, which generalizes well to different Transformer architectures, evaluation metrics, datasets, and tasks, to the baseline methods. Our code is available at: https://github.com/qiangyao1988/AttCAT.
Yao Qiang, Deng Pan 0001, Chengyin Li, Xin Li 0081, RhongHo Jang, Dongxiao Zhu
NeurIPS6
2022 Coupling User Preference with External Rewards to Enable Driver-centered and Resource-aware EV Charging Recommendation
Chengyin Li, Zheng Dong 0002, Nathan Fisher, Dongxiao Zhu
ECML/PKDD (4)4
2021 Improving Adversarial Robustness via Probabilistically Compact Loss with Logit Constraints
abstract
Convolutional neural networks (CNNs) have achieved state-of-the-art performance on various tasks in computer vision. However, recent studies demonstrate that these models are vulnerable to carefully crafted adversarial samples and suffer from a significant performance drop when predicting them. Many methods have been proposed to improve adversarial robustness (e.g., adversarial training and new loss functions to learn adversarially robust feature representations). Here we offer a unique insight into the predictive behavior of CNNs that they tend to misclassify adversarial samples into the most probable false classes. This inspires us to propose a new Probabilistically Compact (PC) loss with logit constraints which can be used as a drop-in replacement for cross-entropy (CE) loss to improve CNN's adversarial robustness. Specifically, PC loss enlarges the probability gaps between true class and false classes meanwhile the logit constraints prevent the gaps from being melted by a small perturbation. We extensively compare our method with the state-of-the-art using large scale datasets under both white-box and black-box attacks to demonstrate its effectiveness. The source codes are available at https://github.com/xinli0928/PC-LC.
Xin Li 0081, Xiangrui Li, Deng Pan 0001, Dongxiao Zhu
AAAI4
2021 Explaining Deep Neural Network Models with Adversarial Gradient Integration
abstract
Deep neural networks (DNNs) have became one of the most high performing tools in a broad range of machine learning areas. However, the multilayer non-linearity of the network architectures prevent us from gaining a better understanding of the models’ predictions. Gradient based attribution methods (e.g., Integrated Gradient (IG)) that decipher input features’ contribution to the prediction task have been shown to be highly effective yet requiring a reference input as the anchor for explaining model’s output. The performance of DNN model interpretation can be quite inconsistent with regard to the choice of references. Here we propose an Adversarial Gradient Integration (AGI) method that integrates the gradients from adversarial examples to the target example along the curve of steepest ascent to calculate the resulting contributions from all input features. Our method doesn’t rely on the choice of references, hence can avoid the ambiguity and inconsistency sourced from the reference selection. We demonstrate the performance of our AGI method and compare with competing methods in explaining image classification results. Code is available from https://github.com/pd90506/AGI.
Deng Pan 0001, Xin Li 0081, Dongxiao Zhu
IJCAI3
2021 Tackling ordinal regression problem for heterogeneous data: sparse and deep multi-task learning approaches
Lu Wang 0004, Dongxiao Zhu
Data Min. Knowl. Discov.2
2020 On the Learning Property of Logistic and Softmax Losses for Deep Neural Networks
abstract
Deep convolutional neural networks (CNNs) trained with logistic and softmax losses have made significant advancement in visual recognition tasks in computer vision. When training data exhibit class imbalances, the class-wise reweighted version of logistic and softmax losses are often used to boost performance of the unweighted version. In this paper, motivated to explain the reweighting mechanism, we explicate the learning property of those two loss functions by analyzing the necessary condition (e.g., gradient equals to zero) after training CNNs to converge to a local minimum. The analysis immediately provides us explanations for understanding (1) quantitative effects of the class-wise reweighting mechanism: deterministic effectiveness for binary classification using logistic loss yet indeterministic for multi-class classification using softmax loss; (2) disadvantage of logistic loss for single-label multi-class classification via one-vs.-all approach, which is due to the averaging effect on predicted probabilities for the negative class (e.g., non-target classes) in the learning process. With the disadvantage and advantage of logistic loss disentangled, we thereafter propose a novel reweighted logistic loss for multi-class classification. Our simple yet effective formulation improves ordinary logistic loss by focusing on learning hard non-target classes (target vs. non-target class in one-vs.-all) and turned out to be competitive with softmax loss. We evaluate our method on several benchmark datasets to demonstrate its effectiveness.
Xiangrui Li, Xin Li 0081, Deng Pan 0001, Dongxiao Zhu
AAAI4
2020 COVID-MobileXpert: On-Device COVID-19 Patient Triage and Follow-up using Chest X-rays
abstract
During the COVID-19 pandemic, there has been an emerging need for rapid, dedicated, and point-of-care COVID19 patient disposition techniques to optimize resource utilization and clinical workflow. In view of this need, we present COVIDMobileXpert: a lightweight deep neural network (DNN) based mobile app that can use chest X-ray (CXR) for COVID-19 case screening and radiological trajectory prediction. We design and implement a novel three-player knowledge transfer and distillation (KTD) framework including a pre-trained attending physician (AP) network that extracts CXR imaging features from a large scale of lung disease CXR images, a fine-tuned resident fellow (RF) network that learns the essential CXR imaging features to discriminate COVID-19 from pneumonia and/or normal cases with a small amount of COVID-19 cases, and a trained lightweight medical student (MS) network to perform on-device COVID-19 patient triage and follow-up. To tackle the challenge of vastly similar and dominant fore-and background in medical images, we employ novel loss functions and training schemes for the MS network to learn the robust features. We demonstrate the significant potential of COVID-MobileXpert for rapid deployment via extensive experiments with diverse MS architecture and tuning parameter settings. The source codes for cloud and mobile based implementations are available from the following url: https://github.com/xinli0928/COVID-Xray.
Xin Li 0081, Chengyin Li, Dongxiao Zhu
BIBM3
2020 Explainable Recommendation via Interpretable Feature Mapping and Evaluation of Explainability
abstract
Latent factor collaborative filtering (CF) has been a widely used technique for recommender system by learning the semantic representations of users and items. Recently, explainable recommendation has attracted much attention from research community. However, trade-off exists between explainability and performance of the recommendation where metadata is often needed to alleviate the dilemma. We present a novel feature mapping approach that maps the uninterpretable general features onto the interpretable aspect features, achieving both satisfactory accuracy and explainability in the recommendations by simultaneous minimization of rating prediction loss and interpretation loss. To evaluate the explainability, we propose two new evaluation metrics specifically designed for aspect-level explanation using surrogate ground truth. Experimental results demonstrate a strong performance in both recommendation and explaining explanation, eliminating the need for metadata. Code is available from https://github.com/pd90506/AMCF.
Deng Pan 0001, Xiangrui Li, Xin Li 0081, Dongxiao Zhu
IJCAI4
2020 Toward Tag-free Aspect Based Sentiment Analysis: A Multiple Attention Network Approach
abstract
Existing aspect based sentiment analysis (ABSA) approaches leverage various neural network models to extract the aspect sentiments via learning aspect-specific feature representations. However, these approaches heavily rely on manual tagging of user reviews according to the predefined aspects as the input, a laborious and time-consuming process. Moreover, the underlying methods do not explain how and why the opposing aspect level polarities in a user review lead to the overall polarity. In this paper, we tackle these two problems by designing and implementing a new Multiple-Attention Network (MAN) approach for more powerful ABSA without the need for aspect tags using two new tag-free data sets crawled directly from TripAdvisor (https://www.tripadvisor.com). With the Self- and Position-Aware attention mechanism, MAN is capable of extracting both aspect level and overall sentiments from the text reviews using the aspect level and overall customer ratings, and it can also detect the vital aspect(s) leading to the overall sentiment polarity among different aspects via a new aspect ranking scheme. We carry out extensive experiments to demonstrate the strong performance of MAN compared to other state-of-the-art ABSA approaches and the explainability of our approach by visualizing and interpreting attention weights in case studies.
Yao Qiang, Xin Li 0081, Dongxiao Zhu
IJCNN3
2019 Prioritization of Multi-Level Risk Factors for Obesity
abstract
Obesity has become a significant threat to health. Identifying and understanding the underlying obesity risk factors (ORFs) are crucial for optimizing prevention, intervention and treatment for obesity. Most existing methodological approaches to risk factor analysis are employed within the single task learning (STL) framework to learn a ranked list of ORFs for a whole population. However, obesity is a multi-faced health outcome. Some ORFs are highly specific to a certain subpopulation and others are universal to the entire population. Multi-task learning (MTL) framework offers a solution to connect multiple related tasks. Within the MTL framework, we implement two tailor-made models, i.e., multi-task feature learning (MTFL) and clustered multi-task learning (CMTL), to conduct ORFs analysis. The former is capable of finding the universal ORFs for all subpopulations without sacrificing the uniqueness of each subpopulation. The latter uncovers the grouping structure and conducts multi-level ORFs analysis simultaneously. Experiments on a public behavioral dataset demonstrate a superior performance of our methods in prioritizing multi-level ORFs.
Lu Wang 0004, Ming Dong 0001, Elizabeth Towner, Dongxiao Zhu
BIBM4
2019 A Deep Active Survival Analysis approach for precision treatment recommendations: Application of prostate cancer
Milad Zafar Nezhad, Najibesadat Sadati, Kai Yang 0005, Dongxiao Zhu
Expert Syst. Appl.4
2018 Proteus: network-aware web browsing on heterogeneous mobile systems
abstract
We present Proteus, a novel network-aware approach for optimizing web browsing on heterogeneous multi-core mobile systems. It employs machine learning techniques to predict which of the heterogeneous cores to use to render a given webpage and the operating frequencies of the processors. It achieves this by first learning offline a set of predictive models for a range of typical networking environments. A learnt model is then chosen at runtime to predict the optimal processor configuration, based on the web content, the network status and the optimization goal. We evaluate Proteus by implementing it into the open-source Chromium browser and testing it on two representative ARM big.LITTLE mobile multi-core platforms. We apply Proteus to the top 1,000 popular websites across seven typical network environments. Proteus achieves over 80% of best available performance. It obtains, on average, over 17% (up to 63%), 31% (up to 88%), and 30% (up to 91%) improvement respectively for load time, energy consumption and the energy delay product, when compared to two state-of-the-art approaches.
Jie Ren 0007, Jianbin Fang, Yansong Feng 0002, Dongxiao Zhu, Zhunchen Luo, Jie Zheng 0005, Zheng Wang 0001
CoNEXT5
2018 Robust feature selection via l2, 1-norm in finite mixture of regression
Xiangrui Li, Dongxiao Zhu
Pattern Recognit. Lett.2
2018 Multinomial classification with class-conditional overlapping sparse feature groups
Xiangrui Li, Dongxiao Zhu, Ming Dong 0001
Pattern Recognit. Lett.2
2017 Predictive deep network with leveraging clinical measure as auxiliary task
abstract
Deep neural networks (DNNs) have made impressive improvements for predictive modeling in various fields. Successful network building for predictive tasks usually requires abundant data, for effectively learning high-level, non-additive information from raw input features. The merit of high-level learning makes DNN promising in clinical research, but the need for abundant data hampers DNN applications. To address this challenge, we propose the use of Auxiliary-Task-Augmented Network (ATAN), a predictive model (for primary target) with introducing auxiliary tasks as regularization. ATAN leverages clinically relevant measures as auxiliary targets and learns the clinical relevance explicitly. We apply ATAN in a clinical dataset of hypertension collected from a vulnerable demographic population (African-American) to demonstrate its effectiveness.
Xiangrui Li, Dongxiao Zhu, Phillip Levy
BIBM2
2017 Multi-task Survival Analysis
abstract
Collecting labeling information of time-to-event analysis is naturally very time consuming, i.e., one has to wait for the occurrence of the event of interest, which may not always be observed for every instance. By taking advantage of censored instances, survival analysis methods internally consider more samples than standard regression methods, which partially alleviates this data insufficiency problem. Whereas most existing survival analysis models merely focus on a single survival prediction task, when there are multiple related survival prediction tasks, we may benefit from the tasks relatedness. Simultaneously learning multiple related tasks, multi-task learning (MTL) provides a paradigm to alleviate data insufficiency by bridging data from all tasks and improves generalization performance of all tasks involved. Even though MTL has been extensively studied, there is no existing work investigating MTL for survival analysis. In this paper, we propose a novel multi-task survival analysis framework that takes advantage of both censored instances and task relatedness. Specifically, based on two common used task relatedness assumptions, i.e., low-rank assumption and cluster structure assumption, we formulate two concrete models, COX-TRACE and COX-cCMTL, under the proposed framework, respectively. We develop efficient algorithms and demonstrate the performance of the proposed multi-task survival analysis models on the The Cancer Genome Atlas (TCGA) dataset. Our results show that the proposed approaches can significantly improve the prediction performance in survival analysis and can also discover some inherent relationships among different cancer types.
Lu Wang 0004, Yan Li 0052, Dongxiao Zhu, Jieping Ye
ICDM4
2017 SUBIC: A Supervised Bi-Clustering Approach for Precision Medicine
abstract
Traditional medicine typically applies one-sizefits-all treatment for the entire patient population whereas precision medicine develops tailored treatment schemes for different patient subgroups. The fact that some factors may be more significant for a specific patient subgroup motivates clinicians and medical researchers to develop new approaches to subgroup detection and analysis, which is an effective strategy to personalize treatment. In this study, we propose a novel patient subgroup detection method, called Supervised Biclustring (SUBIC) using convex optimization and apply our approach to detect patient subgroups and prioritize risk factors for hypertension (HTN) in a vulnerable demographic subgroup (African-American). Our approach not only finds patient subgroups with guidance of a clinically relevant target variable but also identifies and prioritizes risk factors by pursuing sparsity of the input variables and encouraging similarity among the input variables and between the input and target variables.
Milad Zafar Nezhad, Dongxiao Zhu, Najibesadat Sadati, Kai Yang 0005, Phillip Levy
ICMLA2
2017 Modeling Over-Dispersion for Network Data Clustering
abstract
Over-dispersed network data mining has emerged as a central theme in data science, evident by a sharp increase in the volume of real-world network data with imbalanced clusters.While most of existing clustering methods are designed for discovering the number of clusters and class specific connectivity patterns, few methods are available to uncover the imbalanced clusters,commonly existing in network communities and image segments.In this paper, we propose a generalized probabilistic modeling framework,SizeConnectivity, to estimate over-dispersed cluster size distribution together with class specific connectivity patterns from network data.We performed extensive synthetic and real-world experiments on clustering social network data and image data for detecting network communities and image segments.Our results demonstrate a superior performance of our SizeConnectivity clustering method in recovering the hidden structure of network data via modeling over-dispersion.
Lu Wang 0004, Dongxiao Zhu, Ming Dong 0001, Yan Li 0052
ICMLA2
2017 Object tracking via Dirichlet process-based appearance models
Raed Almomani, Ming Dong 0001, Dongxiao Zhu
Neural Comput. Appl.3
2016 SAFS: A deep feature selection approach for precision medicine
abstract
In this paper, we propose a new deep feature selection method based on deep architecture. Our method uses stacked auto-encoders for feature representation in higher-level abstraction. We developed and applied a novel feature learning approach to a specific precision medicine problem, which focuses on assessing and prioritizing risk factors for hypertension (HTN) in a vulnerable demographic subgroup (African-American). Our approach is to use deep learning to identify significant risk factors affecting left ventricular mass indexed to body surface area (LVMI) as an indicator of heart damage risk. The results show that our feature learning and representation approach leads to better results in comparison with others.
Milad Zafar Nezhad, Dongxiao Zhu, Xiangrui Li, Kai Yang 0005, Phillip Levy
BIBM2
2016 A Bayesian hierarchical appearance model for robust object tracking
abstract
In tracking, one of the major challenges comes from handling appearance variations caused by changes in scale, pose, illumination and occlusion. In this paper, we propose a novel Bayesian Hierarchical Appearance Model (BHAM) for robust object tracking. Our idea is to model the appearance of a target as a combination of multiple appearance models, each covering the target appearance changes under a given view angle. Specifically, target instances are modeled by Dirichlet Process and dynamically clustered based on their visual similarity. Thus, BHAM provides an infinite nonparametric mixture of distributions that can grow automatically with the complexity of the appearance data. We built an object tracking system by integrating BHAM with background subtraction and the KLT tracker. Our experimental results on real-world videos show that our system has superior performance when compared with several state-of-the-art trackers.
Raed Almomani, Ming Dong 0001, Dongxiao Zhu
ICME3
2016 Poisson-Markov Mixture Model and Parallel Algorithm for Binning Massive and Heterogenous DNA Sequencing Reads
Lu Wang 0004, Dongxiao Zhu, Yan Li 0052, Ming Dong 0001
ISBRA2
2015 Knowledge Discovery Using Big Data in Biomedical Systems
abstract
The papers in this special section were presented at the 13th International Workshop on Data Mining in Bioinformatics (BIOKDD’14) was organized in conjunction with the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining that was held on August 24, 2014 in New York, NY. It brought together international researchers in the interacting disciplines of data mining, systems biology, and bioinformatics at the Bloomberg Headquarters venue. The goal of this workshop is to encourage Knowledge Discovery and Data mining (KDD) researchers to take on the numerous challenges that Bioinformatics offers.
Sarath Chandra Janga, Dongxiao Zhu, Jake Yue Chen, Mohammed J. Zaki
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 GSCA: Reconstructing biological pathway topologies using a cultural algorithms approach
abstract
With the increasing availability of gene sets and pathway resources, novel approaches that combine both resources to reconstruct networks from gene sets are of interest. Currently, few computational approaches explore the search space of candidate networks using a parallel search. In particular, search agents employed by evolutionary computational approaches may better escape false peaks compared to previous approaches. It may also be hypothesized that gene sets may model signal transduction events, which refer to linear chains or cascades of reactions starting at the cell membrane and ending at the cell nucleus. These events may be indirectly observed as a set of unordered and overlapping gene sets. Thus, the goal is to reverse engineer the order information within each gene set to reconstruct the underlying source network using prior knowledge to limit the search space. We propose the Gene Set Cultural Algorithm (GSCA) to reconstruct networks from unordered gene sets. We introduce a robust heuristic based on the arborescence of a directed graph that performs well for random topological sort orderings across gene sets simulated for four E. coli networks and five Insilico networks from the DREAM3 and DREAM4 initiatives, respectively. Furthermore, GSCA performs favorably when reconstructing networks from randomly ordered gene sets for the aforementioned networks. Finally, we note that from a set of 23 gene sets discretized from a set of 300 S. cerevisiae expression profiles, GSCA reconstructs a network preserving most of the weak order information found in the KEGG Cell Cycle pathway, which was used as prior knowledge.
Thair Judeh, Thaer Jayyousi, Lipi R. Acharya, Robert G. Reynolds, Dongxiao Zhu
IEEE Congress on Evolutionary Computation5
2014 dSpliceType: A Multivariate Model for Detecting Various Types of Differential Splicing Events Using RNA-Seq
Dongxiao Zhu
ISBRA2
2012 Optimal structural inference of signaling pathways from unordered and overlapping gene sets
abstract
MOTIVATION: A plethora of bioinformatics analysis has led to the discovery of numerous gene sets, which can be interpreted as discrete measurements emitted from latent signaling pathways. Their potential to infer signaling pathway structures, however, has not been sufficiently exploited. Existing methods accommodating discrete data do not explicitly consider signal cascading mechanisms that characterize a signaling pathway. Novel computational methods are thus needed to fully utilize gene sets and broaden the scope from focusing only on pairwise interactions to the more general cascading events in the inference of signaling pathway structures. RESULTS: We propose a gene set based simulated annealing (SA) algorithm for the reconstruction of signaling pathway structures. A signaling pathway structure is a directed graph containing up to a few hundred nodes and many overlapping signal cascades, where each cascade represents a chain of molecular interactions from the cell surface to the nucleus. Gene sets in our context refer to discrete sets of genes participating in signal cascades, the basic building blocks of a signaling pathway, with no prior information about gene orderings in the cascades. From a compendium of gene sets related to a pathway, SA aims to search for signal cascades that characterize the optimal signaling pathway structure. In the search process, the extent of overlap among signal cascades is used to measure the optimality of a structure. Throughout, we treat gene sets as random samples from a first-order Markov chain model. We evaluated the performance of SA in three case studies. In the first study conducted on 83 KEGG pathways, SA demonstrated a significantly better performance than Bayesian network methods. Since both SA and Bayesian network methods accommodate discrete data, use a 'search and score' network learning strategy and output a directed network, they can be compared in terms of performance and computational time. In the second study, we compared SA and Bayesian network methods using four benchmark datasets from DREAM. In our final study, we showcased two context-specific signaling pathways activated in breast cancer. AVAILABILITY: Source codes are available from http://dl.dropbox.com/u/16000775/sa_sc.zip.
Lipi R. Acharya, Thair Judeh, Guangdi Wang, Dongxiao Zhu
Bioinform.4
2012 GSGS: A Computational Approach to Reconstruct Signaling Pathway Structures from Gene Sets
abstract
Reconstruction of signaling pathway structures is essential to decipher complex regulatory relationships in living cells. Existing approaches often rely on unrealistic biological assumptions and do not explicitly consider signal transduction mechanisms. Signal transduction events refer to linear cascades of reactions from cell surface to nucleus and characterize a signaling pathway. We propose a novel approach, Gene Set Gibbs Sampling, to reverse engineer signaling pathway structures from gene sets related to pathways. We hypothesize that signaling pathways are structurally an ensemble of overlapping linear signal transduction events which we encode as Information Flows (IFs). We infer signaling pathway structures from gene sets, referred to as Information Flow Gene Sets (IFGSs), corresponding to these events. Thus, an IFGS only reflects which genes appear in the underlying IF but not their ordering. GSGS offers a Gibbs sampling procedure to reconstruct the underlying signaling pathway structure by sequentially inferring IFs from the overlapping IFGSs related to the pathway. In the proof-of-concept studies, our approach is shown to outperform existing network inference approaches using data generated from benchmark networks in DREAM. We perform a sensitivity analysis to assess the robustness of our approach. Finally, we implement GSGS to reconstruct signaling mechanisms in breast cancer cells.
Lipi R. Acharya, Thair Judeh, Zhansheng Duan, Michael G. Rabbat, Dongxiao Zhu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2011 A Generalized Multivariate Approach to Pattern Discovery from Replicated and Incomplete Genome-Wide Measurements
abstract
Estimation of pairwise correlation from incomplete and replicated molecular profiling data is an ubiquitous problem in pattern discovery analysis, such as clustering and networking. However, existing methods solve this problem by ad hoc data imputation, followed by aveGation coefficient type approaches, which might annihilate important patterns present in the molecular profiling data. Moreover, these approaches do not consider and exploit the underlying experimental design information that specifies the replication mechanisms. We develop an Expectation-Maximization (EM) type algorithm to estimate the correlation structure using incomplete and replicated molecular profiling data with a priori known replication mechanism. The approach is sufficiently generalized to be applicable to any known replication mechanism. In case of unknown replication mechanism, it is reduced to the parsimonious model introduced previously. The efficacy of our approach was first evaluated by comprehensively comparing various bivariate and multivariate imputation approaches using simulation studies. Results from real-world data analysis further confirmed the superior performance of the proposed approach to the commonly used approaches, where we assessed the robustness of the method using data sets with up to 30 percent missing values.
Dongxiao Zhu, Lipi R. Acharya
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 svdPPCS: an effective singular value decomposition-based method for conserved and divergent co-expression gene module identification
abstract
BACKGROUND: Comparative analysis of gene expression profiling of multiple biological categories, such as different species of organisms or different kinds of tissue, promises to enhance the fundamental understanding of the universality as well as the specialization of mechanisms and related biological themes. Grouping genes with a similar expression pattern or exhibiting co-expression together is a starting point in understanding and analyzing gene expression data. In recent literature, gene module level analysis is advocated in order to understand biological network design and system behaviors in disease and life processes; however, practical difficulties often lie in the implementation of existing methods. RESULTS: Using the singular value decomposition (SVD) technique, we developed a new computational tool, named svdPPCS (SVD-based Pattern Pairing and Chart Splitting), to identify conserved and divergent co-expression modules of two sets of microarray experiments. In the proposed methods, gene modules are identified by splitting the two-way chart coordinated with a pair of left singular vectors factorized from the gene expression matrices of the two biological categories. Importantly, the cutoffs are determined by a data-driven algorithm using the well-defined statistic, SVD-p. The implementation was illustrated on two time series microarray data sets generated from the samples of accessory gland (ACG) and malpighian tubule (MT) tissues of the line W118 of M. drosophila. Two conserved modules and six divergent modules, each of which has a unique characteristic profile across tissue kinds and aging processes, were identified. The number of genes contained in these models ranged from five to a few hundred. Three to over a hundred GO terms were over-represented in individual modules with FDR < 0.1. One divergent module suggested the tissue-specific relationship between the expressions of mitochondrion-related genes and the aging process. This finding, together with others, may be of biological significance. The validity of the proposed SVD-based method was further verified by a simulation study, as well as the comparisons with regression analysis and cubic spline regression analysis plus PAM based clustering. CONCLUSIONS: svdPPCS is a novel computational tool for the comparative analysis of transcriptional profiling. It especially fits the comparison of time series data of related organisms or different tissues of the same organism under equivalent or similar experimental conditions. The general scheme can be directly extended to the comparisons of multiple data sets. It also can be applied to the integration of data sets from different platforms and of different sources.
Wensheng Zhang 0005, Andrea Edwards, Wei Fan 0001, Dongxiao Zhu, Kun Zhang 0012
BMC Bioinform.4
2009 A Generalized Multivariate Approach for Correlation-Based Pattern Discovery from Replicated Molecular Profiling Data
abstract
Correlation-based pattern discovery from replicated molecular profiling data enables essential data mining tasks, such as discovering biomolecule association networks and functional modules. Unfortunately, the existing approaches are not tailored to analyze replicated measurements, which is further confused by various replication mechanisms. With few exception, existing approaches average or summarize over replicates of diverse magnitude, which might wipe out important patterns of low magnitude and/or cancel out patterns of similar magnitude. The averaging or summarizing procedure, originally targeted for univariate differential expression analysis, has become a nuisance in multivariate correlation-based pattern discovery. Multivariate approaches that treat each replicate individually provide a promising alternative. Here we propose a multivariate parsimonious correlation model for replicated molecular profiling data with blind replication mechanisms, and a constrained (less parsimonious) correlation model explicitly considers the informed replication mechanisms. We derive a generalized formula for correlation-based pattern discovery for both blind and informed replication mechanisms. To promote it's use among the biomedical research community, we develop a correlation-based pattern discovery software with graphical user interface (GUI) for analyzing replicated molecular profiling data.
Dongxiao Zhu, Guorong Xu, Lipi R. Acharya
BIBM1
2009 Estimating an Optimal Correlation Structure from Replicated Molecular Profiling Data Using Finite Mixture Models
abstract
Estimating the correlation structure of a gene set is an ubiquitous problem in many pattern analyses of replicated molecular profiling data. However, the commonly used Maximum Likelihood Estimates (MLE) approaches, do not automatically accommodate replicated measurements. Often, an ad hoc step of preprocessing e.g. averaging, either weighted, un-weighted or something in between is needed, which might wipe out important patterns of low magnitude and/or cancel out patterns of similar magnitude. We treat each replicate individually as a random variable and design a finite mixture model to estimate an optimal correlation structure from replicated molecular profiling data. Assuming that the measurements are independently, identically distributed (i.i.d.) samples from a mixture of two multivariate normal distributions, one with a constrained set of parameters and the other with an unconstrained parameter structure, we employ an Expectation-Maximization (EM) algorithm to estimate component parameters. We carry out a comparative study, including both simulations and real-world data analysis, to assess the estimation of correlation structure using the proposed model and the constrained model given by the first component of the mixture. The two models were further tested for their performances in clustering real-world data. The mixture model approach is shown to have an overall better performance.
Lipi R. Acharya, Dongxiao Zhu
ICMLA2
2009 Semi-supervised gene shaving method for predicting low variation biological pathways from genome-wide data
abstract
BACKGROUND: The gene shaving algorithm and many other clustering algorithms identify gene clusters showing high variation across samples. However, gene expression in many signaling pathways show only modest and concordant changes that fail to be identified by these methods. The increasingly available signaling pathway prior knowledge provide new opportunity to solve this problem. RESULTS: We propose an innovative semi-supervised gene clustering algorithm, where the original gene shaving algorithm was extended and generalized so that prior knowledge of signaling pathways can be incorporated. Different from other methods, our method identifies gene clusters showing concerted and modest expression variation as well as strong expression correlation. Using available pathway gene sets as prior knowledge, whether complete or incomplete, our algorithm is capable of forming tightly regulated gene clusters showing modest variation across samples. We demonstrate the advantages of our algorithm over the original gene shaving algorithm using two microarray data sets. The stability of the gene clusters was accessed using a jackknife approach. CONCLUSION: Our algorithm represents one of the first clustering algorithms that is particularly designed to identify signaling pathways of low and concordant gene expression variation. The discriminating power is achieved by manufacturing a principal component enriched by signaling pathways.
Dongxiao Zhu
BMC Bioinform.1
2007 Improvement of Bayesian Network Inference Using a Relaxed Gene Ordering
abstract
Bayesian network structural learning from high throughput data has become a powerful tool in reconstructing signaling pathways. Recent bioinformatics research advocates the notion that signaling networks in the living cell are likely to be hierarchically organized. Genes resident in hierarchical layers constitute biological constraint, which can be readily used by many network structural learning algorithms to reduce the computational complexity. Based on the hierarchical constraint constructed by using breadth-first-search(BFS) on a manually assembled transcriptional regulation network inSaccharomycescerevisiae, we propose a new constrained Bayesian network structural learning algorithm that solves the NP-hard computational problem in a heuristic manner. We demonstrate the utility of our algorithm in constructing two important signaling pathways.
Dongxiao Zhu
ICMLA1
2007 Multivariate correlation estimator for inferring functional relationships from replicated genome-wide data
abstract
UNLABELLED: Estimating pairwise correlation from replicated genome-scale (a.k.a. OMICS) data is fundamental to cluster functionally relevant biomolecules to a cellular pathway. The popular Pearson correlation coefficient estimates bivariate correlation by averaging over replicates. It is not completely satisfactory since it introduces strong bias while reducing variance. We propose a new multivariate correlation estimator that models all replicates as independent and identically distributed (i.i.d.) samples from the multivariate normal distribution. We derive the estimator by maximizing the likelihood function. For small sample data, we provide a resampling-based statistical inference procedure, and for moderate to large sample data, we provide an asymptotic statistical inference procedure based on the Likelihood Ratio Test (LRT). We demonstrate advantages of the new multivariate correlation estimator over Pearson bivariate correlation estimator using simulations and real-world data analysis examples. AVAILABILITY: The estimator and statistical inference procedures have been implemented in an R package 'CORREP' that is available from CRAN [http://cran.r-project.org] and Bioconductor [http://www.bioconductor.org/]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dongxiao Zhu, Youjuan Li
Bioinform.1
2005 Gene coexpression network discovery with controlled statistical and biological significance
abstract
Many biological functions are executed as a module of coexpressed genes which can be conveniently viewed as a coexpression network. Genes are network vertices and significant pairwise coexpression are network edges. Traditional network discovery methods control either statistical significance or biological significance, but not both. We have designed and implemented a two-stage algorithm that controls both the statistical significance (false discovery rate, FDR) and the biological significance (minimum acceptable strength, MAS) of the discovered network. Based on the estimation of pairwise gene profile correlation, the algorithm provides an initial network discovery that controls only FDR, which is then followed by a second network discovery which controls both FDR and MAS. We illustrate the algorithm for discovery of coexpression networks for yeast galactose metabolism with controlled FDR and MAS.
Dongxiao Zhu, Alfred O. Hero III
ICASSP (5)1
2005 Network constrained clustering for gene microarray data
abstract
Many bioinformatics problems can be tackled from a fresh angle offered by the network perspective. Directly inspired by metabolic network structural studies, we propose an improved gene clustering approach for inferring gene signaling pathways. Based on the construction of co-expression networks that consists of both significantly linear and nonlinear gene associations together with controlled biological and statistical significance, we can make accurate discovery of many transitively coexpressed genes and similarly coexpressed genes. Our approach tends to group functionally related genes into a tight cluster. We illustrate our approach and compare it to the traditional clustering approaches on a retinal gene expression dataset. The clustering method has been implemented in an R package "GeneNT" that is freely available from: http://www-personal.umich.edu//sup /spl sim//zhud/gene nt.htm/.
Dongxiao Zhu, Alfred O. Hero III
ICASSP (5)1
2005 Network constrained clustering for gene microarray data
abstract
UNLABELLED: Many bioinformatics problems can be tackled from a fresh angle offered by the network perspective. Directly inspired by metabolic network structural studies, we propose an improved gene clustering approach for inferring gene signaling pathways from gene microarray data. Based on the construction of co-expression networks that consists of both significantly linear and non-linear gene associations together with controlled biological and statistical significance, our approach tends to group functionally related genes into tight clusters despite their expression dissimilarities. We illustrate our approach and compare it to the traditional clustering approaches on a yeast galactose metabolism dataset and a retinal gene expression dataset. Our approach greatly outperforms the traditional approach in rediscovering the relatively well known galactose metabolism pathway in yeast and in clustering genes of the photoreceptor differentiation pathway. AVAILABILITY: The clustering method has been implemented in an R package "GeneNT" that is freely available from: http://www.cran.org.
Dongxiao Zhu, Alfred O. Hero III, Ritu Khanna, Anand Swaroop
Bioinform.1
2005 Structural comparison of metabolic networks in selected single cell organisms
abstract
BACKGROUND: There has been tremendous interest in the study of biological network structure. An array of measurements has been conceived to assess the topological properties of these networks. In this study, we compared the metabolic network structures of eleven single cell organisms representing the three domains of life using these measurements, hoping to find out whether the intrinsic network design principle(s), reflected by these measurements, are different among species in the three domains of life. RESULTS: Three groups of topological properties were used in this study: network indices, degree distribution measures and motif profile measure. All of which are higher-level topological properties except for the marginal degree distribution. Metabolic networks in Archaeal species are found to be different from those in S. cerevisiae and the six Bacterial species in almost all measured higher-level topological properties. Our findings also indicate that the metabolic network in Archaeal species is similar to the exponential random network. CONCLUSION: If these metabolic network properties of the organisms studied can be extended to other species in their respective domains (which is likely), then the design principle(s) of Archaea are fundamentally different from those of Bacteria and Eukaryote. Furthermore, the functional mechanisms of Archaeal metabolic networks revealed in this study differentiate significantly from those of Bacterial and Eukaryotic organisms, which warrant further investigation.
Dongxiao Zhu, Zhaohui S. Qin
BMC Bioinform.1