Xin Dang

dblp:36/3974 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 51% Learning theory · 36% Time series and sequential data · 7%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
dependence measure
0.512021
Estimating Feature-Label Dependence Using Gini Distance Statistics · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Probabilistic and Bayesian machine learning
statistical dependence
0.512021
Estimating Feature-Label Dependence Using Gini Distance Statistics · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Data mining
anomaly detection
0.322015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture model learning
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining
clustering
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining › anomaly detection
outlier detection
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining › clustering
robust clustering
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Algorithms and data structures
kernel methods
0.112021
Estimating Feature-Label Dependence Using Gini Distance Statistics · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Time series and sequential data
anomaly detection
0.112009
Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Machine learning › Trustworthy machine learning
statistical depth
0.112009
Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Bioinformatics and computational biology
species identification
0.112007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
Bioinformatics and computational biology
taxonomy
0.112007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
Data mining › anomaly detection
novelty detection
0.112007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007

Methods — techniques the papers use, named apart from their topics

reproducing kernel hilbert space · 1.0gini distance covariance · 1.0distance covariance · 1.0rank-based scatter estimation · 0.4median-based location estimation · 0.4statistical depth functions · 0.1kernelized spatial depth · 0.1positive definite kernels · 0.1one-class learning · 0.1
YearPublicationVenuePosition
2026 ACA-Net: Adaptive cloud-aware network for remote sensing image thick cloud removal
Baopu Hou, Xin Dang, Jinguang Wang, Quankai Zhao, Hongjia Qu, Yuting Yang 0008, Xiaoxuan Chen, Bo Jiang 0014
Expert Syst. Appl.3
2023 Multivariate time series data imputation using attention-based mechanism
Jingqi Zhao, Chuitian Rong, Chunbin Lin, Xin Dang
Neurocomputing4
2021 Estimating Feature-Label Dependence Using Gini Distance Statistics
abstract
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RKHS). Two Gini distance based dependence measures are explored: Gini distance covariance and Gini distance correlation. Unlike Pearson covariance and correlation, which do not characterize independence, the above Gini distance based measures define dependence as well as independence of random variables. The test statistics are simple to calculate and do not require probability density estimation. Uniform convergence bounds and asymptotic bounds are derived for the test statistics. Comparisons with distance covariance statistics are provided. It is shown that Gini distance statistics converge faster than distance covariance statistics in the uniform convergence bounds, hence tighter upper bounds on both Type I and Type II errors. Moreover, the probability of Gini distance covariance statistic under-performing the distance covariance statistic in Type II error decreases to 0 exponentially with the increase of the sample size. Extensive experimental results are presented to demonstrate the performance of the proposed method.
Silu Zhang, Xin Dang, Dao Nguyen, Dawn Wilkins, Yixin Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Pareto cascade modeling of diffusion networks
abstract
Time plays an essential role in the diffusion of information, influence and disease over networks. Usually we are only able to collect cascade data in which an infection (receiving) time of each node is recorded but without any transmission information over the network. In this paper, we infer the transmission rates among nodes by Pareto distributions. Pareto modeling has several advantages. It is naturally motivated and has a nice interpretation. The scale parameter of a Pareto distribution naturally fits into the starting time of a transition, i.e., the infection time of a parent node in the cascade data is the starting point for a transition from the parent to its receiver. The shape parameter (alpha) serves as the transition rate. The larger the alpha is, the faster the transition is and there is a higher probability for disease or information to spread in a short time period. Pareto modeling is mathematically simple and computationally easy. It has explicit solutions for the optimization problem that maximizes time-dependent pairwise transmission likelihoods between all pairs of nodes. We present three modelings with a common transmission rate, with different transmission rates and with different infection rates. Experiments on real and synthetic data show that our models accurately estimate the transmission rates and perform better than the existing method.
Christopher Ma, Xin Dang, Yixin Chen 0002, Dawn Wilkins
IJCNN2
2018 Robust and Efficient Boosting Method Using the Conditional Risk
abstract
Well known for its simplicity and effectiveness in classification, AdaBoost, however, suffers from overfitting when class-conditional distributions have significant overlap. Moreover, it is very sensitive to noise that appears in the labels. This paper tackles the above limitations simultaneously via optimizing a modified loss function (i.e., the conditional risk). The proposed approach has the following two advantages. First, it is able to directly take into account label uncertainty with an associated label confidence. Second, it introduces a trustworthiness measure on training samples via the Bayesian risk rule, and hence the resulting classifier tends to have finite sample performance that is superior to that of the original AdaBoost when there is a large overlap between class conditional distributions. Theoretical properties of the proposed method are investigated. Extensive experimental results using synthetic data and real-world data sets from UCI machine learning repository are provided. The empirical study shows the high competitiveness of the proposed method in predication accuracy and robustness when compared with the original AdaBoost and several existing robust AdaBoost algorithms.
Zhi Xiao, Xin Dang
IEEE Trans. Neural Networks Learn. Syst.4
2017 An image reconstruction framework based on deep neural network for electrical impedance tomography
abstract
Electrical impedance tomography (EIT) reconstructs the internal impedance distribution by making voltage and current measurements on the object's boundary. The image reconstruction for EIT is a non-linear inverse problem. A generalized solutions based on an inverse operator is ill-conditioned and highly sensitive to the noise. In order to improve the quality of reconstructed images, this paper presents a new framework based on deep neural network (DNN) model. We apply the stacked autoencoder (SAE) and a logistic regression (LR) layer to constitute a 4-layer DNN model. This model is trained with simulation data to obtain the relationship between voltage measurements and the corresponding conductivity distribution, and then test the trained DNN model with untrained simulation data and experimental data, respectively. The output of the network is considered as the estimate of the conductivity distribution for image reconstruction. Both simulation and experimental results show the effectiveness of the proposed framework in improving the quality of reconstructed images.
Xiuyan Li, Xin Dang, Qi Wang 0040, Xiaojie Duan, Yukuan Sun
ICIP4
2017 Label confidence based AdaBoost algorithm
abstract
AdaBoost is a well-known simple and effective boosting algorithm for classification. It, however, suffers from the overfitting problem in the case of overlapping class distributions and is very sensitive to label noise. To tackle both problems simultaneously, we consider the conditional risk as the modified loss function. This modification leads to two advantages: it is able to directly take into account label uncertainty with an associated label confidence; it introduces a “trustworthiness” measure on training samples via the Bayesian risk rule, hence the resulting classifier tends to have superior finite sample performance than the original AdaBoost when there is a large overlap between class conditional distributions. We scrutinize the re-weighting procedure and classifier combination rule to show its adaptive ability on the mitigation of class noise and overfitting. Extensive experimental results using synthetic data and real-world data sets from UCI machine learning repository are provided. The empirical study shows high competitiveness of the proposed method in predication accuracy and robustness when compared with the original AdaBoost and several existing robust AdaBoost algorithms.
Xin Dang, Yixin Chen 0002
IJCNN2
2015 A generative Bayesian model to identify cancer driver genes
abstract
Cancer is a disease characterized largely by the accumulation of somatic mutations during the lifetime of a patient. Distinguishing driver mutations from passenger mutations had posed a challenge in modern cancer research. With the state of art of microarray technologies and clinical studies, a large numbers of candidate genes are extracted. Extracting informative genes out of them is essential. In our project we aim to find the cancer driver genes using somatic mutation data and protein protein interaction data. We developed a generative mixture model coupled with Bayesian parameter estimation to estimate background mutation rates and driver probabilities of each gene as well as the proportion of drivers among all sequenced genes. We choose suitable prior distributions for modelling both driver probabilities and background mutations of each gene. We apply our method to ovarian cancer data and numerically estimated the solution. Upon convergence, we are able to discover and identify some new candidate cancer driver genes.
Christopher Ma, Zhendong Zhao, Tina Gui, Yixin Chen 0002, Xin Dang, Dawn Wilkins
BIBM5
2015 Robust Model-Based Learning via Spatial-EM Algorithm
abstract
This paper presents a new robust EM algorithm for the finite mixture learning procedures. The proposed Spatial-EM algorithm utilizes median-based location and rank-based scatter estimators to replace sample mean and sample covariance matrix in each M step, hence enhancing stability and robustness of the algorithm. It is robust to outliers and initial values. Compared with many robust mixture learning methods, the Spatial-EM has the advantages of simplicity in implementation and statistical efficiency. We apply Spatial-EM to supervised and unsupervised learning scenarios. More specifically, robust clustering and outlier detection methods based on Spatial-EM have been proposed. We apply the outlier detection to taxonomic research on fish species novelty discovery. Two real datasets are used for clustering analysis. Compared with the regular EM and many other existing methods such as K-median, X-EM and SVM, our method demonstrates superior performance and high robustness.
Xin Dang, Henry L. Bart Jr., Yixin Chen 0002
IEEE Trans. Knowl. Data Eng.2
2014 Financial ratio selection for business failure prediction using soft set theory
Zhi Xiao, Xin Dang, Daoli Yang, Xianglei Yang
Knowl. Based Syst.3
2013 Rule based regression and feature selection for biological data
abstract
Regression is widely utilized in a variety of biological problems involving continuous outcomes. There are a number of methods for building regression models ranging from linear models to more complex nonlinear ones. While linear regression techniques can identify linear correlations between input and output, in many practical applications, the relations are nonlinear. These relations can be modeled by nonlinear regression techniques effectively. However, many models built with nonlinear techniques have limited interpretation, which is crucial in many biological problems. We propose a rule based regression algorithm that uses 1-norm regularized random forests. The proposed approach simultaneously extracts a small number of rules from generated random forests and eliminates unimportant features, and hence is able to provide a simple interpretation. We tested the approach on a seacoast chemical sensors dataset, a Stockori flowering time dataset, and three datasets from the UCI repository. The proposed approach is able to construct a significantly smaller set of regression rules using a subset of attributes while achieving prediction performance comparable to that of conventional random forests regression. It demonstrates high potential in terms of prediction performance and interpretation ease on studying nonlinear relationships of the subjects.
Shamitha Dissanayake, Sanjay Patel, Xin Dang, Todd Mlsna, Yixin Chen 0002, Dawn Wilkins
BIBM4
2012 The prediction for listed companies' financial distress by using multiple prediction methods with rough set and Dempster-Shafer evidence theory
Zhi Xiao, Xianglei Yang, Ying Pang, Xin Dang
Knowl. Based Syst.4
2012 Multiclass classification with potential function rules: Margin distribution and generalization
Fei Teng 0010, Yixin Chen 0002, Xin Dang
Pattern Recognit.3
2011 Leveraging domain information to restructure biological prediction
abstract
BACKGROUND: It is commonly believed that including domain knowledge in a prediction model is desirable. However, representing and incorporating domain information in the learning process is, in general, a challenging problem. In this research, we consider domain information encoded by discrete or categorical attributes. A discrete or categorical attribute provides a natural partition of the problem domain, and hence divides the original problem into several non-overlapping sub-problems. In this sense, the domain information is useful if the partition simplifies the learning task. The goal of this research is to develop an algorithm to identify discrete or categorical attributes that maximally simplify the learning task. RESULTS: We consider restructuring a supervised learning problem via a partition of the problem space using a discrete or categorical attribute. A naive approach exhaustively searches all the possible restructured problems. It is computationally prohibitive when the number of discrete or categorical attributes is large. We propose a metric to rank attributes according to their potential to reduce the uncertainty of a classification task. It is quantified as a conditional entropy achieved using a set of optimal classifiers, each of which is built for a sub-problem defined by the attribute under consideration. To avoid high computational cost, we approximate the solution by the expected minimum conditional entropy with respect to random projections. This approach is tested on three artificial data sets, three cheminformatics data sets, and two leukemia gene expression data sets. Empirical results demonstrate that our method is capable of selecting a proper discrete or categorical attribute to simplify the problem, i.e., the performance of the classifier built for the restructured problem always beats that of the original problem. CONCLUSIONS: The proposed conditional entropy based metric is effective in identifying good partitions of a classification problem, hence enhancing the prediction performance.
Xiaofei Nan, Zhengdong Zhao, Ronak Y. Patel, Haining Liu, Pankaj R. Daga, Robert J. Doerksen, Xin Dang, Yixin Chen 0002, Dawn Wilkins
BMC Bioinform.9
2009 Graph ranking for exploratory gene data analysis
abstract
BACKGROUND: Microarray technology has made it possible to simultaneously monitor the expression levels of thousands of genes in a single experiment. However, the large number of genes greatly increases the challenges of analyzing, comprehending and interpreting the resulting mass of data. Selecting a subset of important genes is inevitable to address the challenge. Gene selection has been investigated extensively over the last decade. Most selection procedures, however, are not sufficient for accurate inference of underlying biology, because biological significance does not necessarily have to be statistically significant. Additional biological knowledge needs to be integrated into the gene selection procedure. RESULTS: We propose a general framework for gene ranking. We construct a bipartite graph from the Gene Ontology (GO) and gene expression data. The graph describes the relationship between genes and their associated molecular functions. Under a species condition, edge weights of the graph are assigned to be gene expression level. Such a graph provides a mathematical means to represent both species-independent and species-dependent biological information. We also develop a new ranking algorithm to analyze the weighted graph via a kernelized spatial depth (KSD) approach. Consequently, the importance of gene and molecular function can be simultaneously ranked by a real-valued measure, KSD, which incorporates the global and local structure of the graph. Over-expressed and under-regulated genes also can be separately ranked. CONCLUSION: The gene-function bigraph integrates molecular function annotations into gene expression data. The relevance of genes is described in the graph (through a common function). The proposed method provides an exploratory framework for gene data analysis.
Cuilan Gao, Xin Dang, Yixin Chen 0002, Dawn Wilkins
BMC Bioinform.2
2009 Outlier Detection with the Kernelized Spatial Depth Function
abstract
Statistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, depth functions can measure the "extremeness" or "outlyingness" of a data point with respect to a given data set. Hence, they can detect outliers--observations that appear extreme relative to the rest of the observations. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. In this article, we propose a novel statistical depth, the kernelized spatial depth (KSD), which generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. We demonstrate this by the half-moon data and the ring-shaped data. Based on the KSD, we propose a novel outlier detection algorithm, by which an observation with a depth value less than a threshold is declared as an outlier. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. It applies to a one-class learning setting, in which "normal" observations are given as the training data, as well as to a missing label scenario, where the training set consists of a mixture of normal observations and outliers with unknown labels. We give upper bounds on the false alarm probability of a depth-based detector. These upper bounds can be used to determine the threshold. We perform extensive experiments on synthetic data and data sets from real applications. The proposed outlier detector is compared with existing methods. The KSD outlier detector demonstrates a competitive performance.
Yixin Chen 0002, Xin Dang, Hanxiang Peng, Henry L. Bart Jr.
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Depth-Based Novelty Detection and Its Application to Taxonomic Research
abstract
It is estimated that less than 10 percent of the world's species have been described, yet species are being lost daily due to human destruction of natural habitats. The job of describing the earth's remaining species is exacerbated by the shrinking number of practicing taxonomists and the very slow pace of traditional taxonomic research. In this article, we tackle, from a novelty detection perspective, one of the most important and challenging research objectives in taxonomy ­ new species identification. We propose a unique and efficient novelty detection framework based on statistical depth functions. Statistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, they can detect observations that appear extreme relative to the rest of the observations, i.e., novelty. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. We propose a novel statistical depth, the kernelized spatial depth (KSD) that generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. Observations with depth values less than a threshold are declared as novel. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. We give an upper bound on the false alarm probability of a depth-based detector, which can be used to determine the threshold. Experimental study demonstrates its excellent potential in new species discovery.
Yixin Chen 0002, Henry L. Bart Jr., Xin Dang, Hanxiang Peng
ICDM3
2007 Robust clustering in high dimensional data using statistical depths
abstract
BACKGROUND: Mean-based clustering algorithms such as bisecting k-means generally lack robustness. Although componentwise median is a more robust alternative, it can be a poor center representative for high dimensional data. We need a new algorithm that is robust and works well in high dimensional data sets e.g. gene expression data. RESULTS: Here we propose a new robust divisive clustering algorithm, the bisecting k-spatialMedian, based on the statistical spatial depth. A new subcluster selection rule, Relative Average Depth, is also introduced. We demonstrate that the proposed clustering algorithm outperforms the componentwise-median-based bisecting k-median algorithm for high dimension and low sample size (HDLSS) data via applications of the algorithms on two real HDLSS gene expression data sets. When further applied on noisy real data sets, the proposed algorithm compares favorably in terms of robustness with the componentwise-median-based bisecting k-median algorithm. CONCLUSION: Statistical data depths provide an alternative way to find the "center" of multivariate data sets and are useful and robust for clustering.
Yuanyuan Ding, Xin Dang, Hanxiang Peng, Dawn Wilkins
BMC Bioinform.2