VLDB 2026 Research / reviewers in the wild / expert
Qiang Wu 0003
dblp:87/2533-3
· DBLP profile ↗
21ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0002-4698-6966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 3 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Attention Deficit Hyperactivity Disorder Classification Based on Deep LearningabstractAttention Deficit Hyperactivity Disorder (ADHD) is a type of mental health disorder that can be seen from children to adults and affects patients' normal life. Accurate diagnosis of ADHD as early as possible is very important for the treatment of patients in clinical applications. Some traditional classification methods, although having been shown powerful in many other classification tasks, are not as successful in the application of ADHD classification. In this paper, we propose two novel deep learning approaches for ADHD classification based on functional magnetic resonance imaging. The first method incorporates independent component analysis with convolutional neural network. It first extracts independent components from each subject. The independent components are then fed into a convolutional neural network as input features to classify the ADHD patient from typical controls. The second method, called the correlation autoencoder method, uses correlations between regions of interest of the brain as the input of an autoencoder to learn latent features, which are then used in the classification task by a new neural network. These two methods use different ways to extract the inter-voxel information from fMRI, but both use convolutional neural networks to further extract predictive features for the classification task. Empirical experiments show that both methods are able to outperform the classical methods such as logistic regression, support vector machines, and other methods used in previous studies. Don Hong, Qiang Wu 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Optimality of regularized least squares ranking with imperfect kernels
Fangchao He, Lie Zheng, Qiang Wu 0003 |
Inf. Sci. | 4 |
| 2022 | Fast Rates of Gaussian Empirical Gain Maximization With Heavy-Tailed NoiseabstractIn a regression setup, we study in this brief the performance of Gaussian empirical gain maximization (EGM), which includes a broad variety of well-established robust estimation approaches. In particular, we conduct a refined learning theory analysis for Gaussian EGM, investigate its regression calibration properties, and develop improved convergence rates in the presence of heavy-tailed noise. To achieve these purposes, we first introduce a new weak moment condition that could accommodate the cases where the noise distribution may be heavy-tailed. Based on the moment condition, we then develop a novel comparison theorem that can be used to characterize the regression calibration properties of Gaussian EGM. It also plays an essential role in deriving improved convergence rates. Therefore, the present study broadens our theoretical understanding of Gaussian EGM. Shouyou Huang, Yunlong Feng, Qiang Wu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Robust pairwise learning with Huber loss
Shouyou Huang, Qiang Wu 0003 |
J. Complex. | 2 |
| 2021 | Optimal Rates of Distributed Regression with Imperfect KernelsabstractDistributed machine learning systems have been receiving increasing attentions for their efficiency to process large scale data. Many distributed frameworks have been proposed for different machine learning tasks. In this paper, we study the distributed kernel regression via the divide and conquer approach. The learning process consists of three stages. Firstly, the data is partitioned into multiple subsets. Then a base kernel regression algorithm is applied to each subset to learn a local regression model. Finally the local models are averaged to generate the final regression model for the purpose of predictive analytics or statistical inference. This approach has been proved asymptotically minimax optimal if the kernel is perfectly selected so that the true regression function lies in the associated reproducing kernel Hilbert space. However, this is usually, if not always, impractical because kernels that can only be selected via prior knowledge or a tuning process are hardly perfect. Instead it is more common that the kernel is good enough but imperfect in the sense that the true regression can be well approximated by but does not lie exactly in the kernel space. We show distributed kernel regression can still achieve capacity independent optimal rate in this case. To this end, we first establish a general framework that allows to analyze distributed regression with response weighted base algorithms by bounding the error of such algorithms on a single data set, provided that the error bounds have factored the impact of unexplained variance of the response variable. Then we perform a leave one out analysis of the kernel ridge regression and bias corrected kernel ridge regression, which in combination with the aforementioned framework allows us to derive sharp error bounds and capacity independent optimal rates for the associated distributed kernel regression algorithms. As a byproduct of the thorough analysis, we also prove the kernel ridge regression can achieve rates faster than $O(N^{-1})$ (where $N$ is the sample size) in the noise free setting which, to our best knowledge, are first observed and novel in regression learning. Qiang Wu 0003 |
J. Mach. Learn. Res. | 2 |
| 2021 | A Framework of Learning Through Empirical Gain MaximizationabstractWe develop in this letter a framework of empirical gain maximization (EGM) to address the robust regression problem where heavy-tailed noise or outliers may be present in the response variable. The idea of EGM is to approximate the density function of the noise distribution instead of approximating the truth function directly as usual. Unlike the classical maximum likelihood estimation that encourages equal importance of all observations and could be problematic in the presence of abnormal observations, EGM schemes can be interpreted from a minimum distance estimation viewpoint and allow the ignorance of those observations. Furthermore, we show that several well-known robust nonconvex regression paradigms, such as Tukey regression and truncated least square regression, can be reformulated into this new framework. We then develop a learning theory for EGM by means of which a unified analysis can be conducted for these well-established but not fully understood regression approaches. This new framework leads to a novel interpretation of existing bounded nonconvex loss functions. Within this new framework, the two seemingly irrelevant terminologies, the well-known Tukey's biweight loss for robust regression and the triweight kernel for nonparametric smoothing, are closely related. More precisely, we show that Tukey's biweight loss can be derived from the triweight kernel. Other frequently employed bounded nonconvex loss functions in machine learning, such as the truncated square loss, the Geman-McClure loss, and the exponential squared loss, can also be reformulated from certain smoothing kernels in statistics. In addition, the new framework enables us to devise new bounded nonconvex loss functions for robust learning. Yunlong Feng, Qiang Wu 0003 |
Neural Comput. | 2 |
| 2020 | Distributed Minimum Error Entropy AlgorithmsabstractMinimum Error Entropy (MEE) principle is an important approach in Information Theoretical Learning (ITL). It is widely applied and studied in various fields for its robustness to noise. In this paper, we study a reproducing kernel-based distributed MEE algorithm, DMEE, which is designed to work with both fully supervised data and semi-supervised data. The divide-and-conquer approach is employed, so there is no inter-node communication overhead. Similar as other distributed algorithms, DMEE significantly reduces the computational complexity and memory requirement on single computing nodes. With fully supervised data, our proved learning rates equal the minimax optimal learning rates of the classical pointwise kernel-based regressions. Under the semi-supervised learning scenarios, we show that DMEE exploits unlabeled data effectively, in the sense that first, under the settings with weak regularity assumptions, additional unlabeled data significantly improves the learning rates of DMEE. Second, with sufficient unlabeled data, labeled data can be distributed to many more computing nodes, that each node takes only O(1) labels, without spoiling the learning rates in terms of the number of labels. This conclusion overcomes the saturation phenomenon in unlabeled data size. It parallels a recent results for regularized least squares (Lin and Zhou, 2018), and suggests that an inflation of unlabeled data is a solution to the MEE learning problems with decentralized data source for the concerns of privacy protection. Our work refers to pairwise learning and non-convex loss. The theoretical analysis is achieved by distributed U-statistics and error decomposition techniques in integral operators. Xin Guo 0003, Ting Hu 0002, Qiang Wu 0003 |
J. Mach. Learn. Res. | 3 |
| 2019 | Machine learning-based microarray analyses indicate low-expression genes might collectively influence PAH diseaseabstractAccurately predicting and testing the types of Pulmonary arterial hypertension (PAH) of each patient using cost-effective microarray-based expression data and machine learning algorithms could greatly help either identifying the most targeting medicine or adopting other therapeutic measures that could correct/restore defective genetic signaling at the early stage. Furthermore, the prediction model construction processes can also help identifying highly informative genes controlling PAH, leading to enhanced understanding of the disease etiology and molecular pathways. In this study, we used several different gene filtering methods based on microarray expression data obtained from a high-quality patient PAH dataset. Following that, we proposed a novel feature selection and refinement algorithm in conjunction with well-known machine learning methods to identify a small set of highly informative genes. Results indicated that clusters of small-expression genes could be extremely informative at predicting and differentiating different forms of PAH. Additionally, our proposed novel feature refinement algorithm could lead to significant enhancement in model performance. To summarize, integrated with state-of-the-art machine learning and novel feature refining algorithms, the most accurate models could provide near-perfect classification accuracies using very few (close to ten) low-expression genes. Qiang Wu 0003, James West, Jiangping Bai |
PLoS Comput. Biol. | 2 |
| 2017 | Bias corrected regularization kernel network and its applicationsabstractRegularization kernel network (RKN) is an effective and widely used kernel method for nonlinear regression analysis. In this paper, I characterize its bias and propose an approach to correct the bias. This leads to a new method called bias corrected regularization kernel network (BCRKN). Theoretical characterizations and simulation studies are used to verify the effectiveness of this bias corrected method. I show that BCRKN has smaller asymptotic bias but slightly larger variance. In applications of learning with a single data set, RKN and BCRKN have comparable performance if the regularization parameter is appropriately tuned. But in applications of incremental learning with block wise streaming data, BCKRN shows to be more efficient due to bias reduction. Qiang Wu 0003 |
IJCNN | 1 |
| 2017 | Learning Theory of Distributed Regression with Bias Corrected Regularization Kernel NetworkabstractDistributed learning is an effective way to analyze big data. In distributed regression, a typical approach is to divide the big data into multiple blocks, apply a base regression algorithm on each of them, and then simply average the output functions learnt from these blocks. Since the average process will decrease the variance, not the bias, bias correction is expected to improve the learning performance if the base regression algorithm is a biased one. Regularization kernel network is an effective and widely used method for nonlinear regression analysis. In this paper we will investigate a bias corrected version of regularization kernel network. We derive the error bounds when it is applied to a single data set and when it is applied as a base algorithm in distributed regression. We show that, under certain appropriate conditions, the optimal learning rates can be reached in both situations. Zheng-Chu Guo, Lei Shi 0010, Qiang Wu 0003 |
J. Mach. Learn. Res. | 3 |
| 2015 | Sparse Representation in Kernel MachinesabstractWe study the properties of least square kernel regression with l1 coefficient regularization. The kernels can be flexibly chosen to be either positive definite or indefinite. Asymptotic learning rates are deduced under smoothness condition on the kernel. Sparse representation of the solution is characterized theoretically. Empirical simulations and real applications indicate that both good learning performance and sparse representation could be guaranteed. Qiang Wu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Learning theory approach to minimum error entropy criterion
Ting Hu 0002, Qiang Wu 0003, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 3 |
| 2012 | Empirical Mode Decomposition Analysis for Visual StylometryabstractIn this paper, we show how the tools of empirical mode decomposition (EMD) analysis can be applied to the problem of “visual stylometry,” generally defined as the development of quantitative tools for the measurement and comparisons of individual style in the visual arts. In particular, we introduce a new form of EMD analysis for images and show that it is possible to use its output as the basis for the construction of effective support vector machine (SVM)-based stylometric classifiers. We present the methodology and then test it on collections of two sets of digital captures of drawings: a set of authentic and well-known imitations of works attributed to the great Flemish artist Pieter Bruegel the Elder (1525-1569) and a set of works attributed to Dutch master Rembrandt van Rijn (1606-1669) and his pupils. Our positive results indicate that EMD-based methods may hold promise generally as a technique for visual stylometry. James Michael Hughes, Dong Mao, Daniel N. Rockmore, Yang Wang 0020, Qiang Wu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Estimating variable structure and dependence in multitask learning via gradientsabstractWe consider the problem of hierarchical or multitask modeling where we simultaneously learn the regression function and the underlying geometry and dependence between variables. We demonstrate how the gradients of the multiple related regression functions over the tasks allow for dimension reduction and inference of dependencies across tasks jointly and for each task individually. We provide Tikhonov regularization algorithms for both classification and regression that are efficient and robust for high-dimensional data, and a mechanism for incorporating a priori knowledge of task (dis)similarity into this framework. The utility of this method is illustrated on simulated and real data. Justin Guinney, Qiang Wu 0003, Sayan Mukherjee 0001 |
Mach. Learn. | 2 |
| 2010 | Learning Gradients: Predictive Models that Infer Geometry and Statistical Dependence
Qiang Wu 0003, Justin Guinney, Mauro Maggioni, Sayan Mukherjee 0001 |
J. Mach. Learn. Res. | 1 |
| 2008 | Localized Sliced Inverse RegressionabstractWe developed localized sliced inverse regression for supervised dimension reduction. It has the advantages of preventing degeneracy, increasing estimation accuracy, and automatic subclass discovery in classification problems. A semisupervised version is proposed for the use of unlabeled data. The utility is illustrated on simulated as well as real data sets. Qiang Wu 0003, Sayan Mukherjee 0001 |
NIPS | 1 |
| 2007 | Multi-kernel regularized classifiers
Qiang Wu 0003, Yiming Ying, Ding-Xuan Zhou |
J. Complex. | 1 |
| 2007 | Characterizing the Function Space for Bayesian Kernel Models
Natesh S. Pillai, Qiang Wu 0003, Sayan Mukherjee 0001, Robert L. Wolpert |
J. Mach. Learn. Res. | 2 |
| 2006 | Estimation of Gradients and Coordinate Covariation in ClassificationabstractWe introduce an algorithm that simultaneously estimates a classification function as well as its gradient in the supervised learning framework. The motivation for the algorithm is to find salient variables and estimate how they covary. An efficient implementation with respect to both memory and time is given. The utility of the algorithm is illustrated on simulated data as well as a gene expression data set. An error analysis is given for the convergence of the estimate of the classification function and its gradient to the true classification function and true gradient. Sayan Mukherjee 0001, Qiang Wu 0003 |
J. Mach. Learn. Res. | 2 |
| 2005 | SVM Soft Margin Classifiers: Linear Programming versus Quadratic ProgrammingabstractSupport vector machine (SVM) soft margin classifiers are important learning algorithms for classification problems. They can be stated as convex optimization problems and are suitable for a large data setting. Linear programming SVM classifiers are especially efficient for very large size samples. But little is known about their convergence, compared with the well-understood quadratic programming SVM classifier. In this article, we point out the difficulty and provide an error analysis. Our analysis shows that the convergence behavior of the linear programming SVM is almost the same as that of the quadratic programming SVM. This is implemented by setting a stepping-stone between the linear programming SVM and the classical 1-norm soft margin classifier. An upper bound for the misclassification error is presented for general probability distributions. Explicit learning rates are derived for deterministic and weakly separable distributions, and for distributions satisfying some Tsybakov noise condition. Qiang Wu 0003, Ding-Xuan Zhou |
Neural Comput. | 1 |
| 2004 | Support Vector Machine Soft Margin Classifiers: Error Analysis
Di-Rong Chen, Qiang Wu 0003, Yiming Ying, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 2 |