Aijun Zhang

dblp:75/1214 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorTheory of computation · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Stable Subsampling under Model Misspecification and Covariate Shift
abstract
The presence of covariate shift between training and test datasets, coupled with model misspecification, can lead to instability in regression predictions across diverse datasets. Meanwhile, training complex models with massive data imposes significant computational burden. In this article, we present a novel model-free subsampling algorithm for stable prediction, which employs uniform design and confounder balancing methods. Our subsampling algorithm aims to find the nearest neighbor subsampling points of uniform design with the goal of minimizing global stability loss, thereby reducing the data volume while achieving stable predictions. Theoretic analyses show that the uniform measure minimizes the maximum integrated mean square error (MIMSE) and the global stability loss evaluates the independence among variables in each candidate MIMSE-optimal subsampled sets. Simulation studies conducted on synthetic datasets, as well as applications on real datasets, demonstrate the superiority of our proposed method under model misspecification and covariate shift.
Jinjing Yang, Shaohua Xu, Zebin Yang 0001, Aijun Zhang
ACM Trans. Knowl. Discov. Data4
2024 DSISA: A New Neural Machine Translation Combining Dependency Weight and Neighbors
abstract
Most of the previous neural machine translations (NMT) rely on parallel corpus. Integrating explicitly prior syntactic structure information can improve the neural machine translation. In this article, we propose a Syntax Induced Self-Attention (SISA) which explores the influence of dependence relation between words through the attention mechanism and fine-tunes the attention allocation of the sentence through the obtained dependency weight. We present a new model, Double Syntax Induced Self-Attention (DSISA), which fuses the features extracted by SISA and a compact convolution neural network (CNN). SISA can alleviate long dependency in sentence, while CNN captures the limited context based on neighbors. DSISA utilizes two different neural networks to extract different features for richer semantic representation and replaces the first layer of Transformer encoder. DSISA not only makes use of the global feature of tokens in sentences but also the local feature formed with adjacent tokens. Finally, we perform simulation experiments that verify the performance of the new model on standard corpora.
Lingfang Li, Aijun Zhang, Mingxing Luo
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Model-Free Subsampling Method Based on Uniform Designs
abstract
Subsampling or subdata selection is a useful approach in large-scale statistical learning. Most existing studies focus on model-based subsampling methods which significantly depend on the model assumption. In this article, we consider the model-free subsampling strategy for generating subdata from the original full data. In order to measure the goodness of representation of a subdata with respect to the original data, we propose a criterion, generalized empirical$F$-discrepancy (GEFD), and study its theoretical properties in connection with the classical generalized$\ell _{2}$-discrepancy in the theory of uniform designs. These properties allow us to develop a kind of low-GEFD data-driven subsampling method based on the existing uniform designs. By simulation examples and a real case study, we show that the proposed subsampling method is superior to the random sampling method. Moreover, our method keeps robust under diverse model specifications while other popular model-based subsampling methods are under-performing. In practice, such a model-free property is more appealing than the model-based subsampling methods, where the latter may have poor performance when the model is misspecified, as demonstrated in our simulation studies. In addition, our method is orders of magnitude faster than other model-free subsampling methods, which makes it more applicable for subsampling of Big Data.
Aijun Zhang
IEEE Trans. Knowl. Data Eng.4
2024 A Sequential Stein's Method for Faster Training of Additive Index Models
Hengtao Zhang, Zebin Yang 0001, Agus Sudjianto, Aijun Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2023 ℓ0 Trend Filtering
abstract
The [Formula: see text] trend filtering ([Formula: see text]-TF) is a new effective tool for nonparametric regression with the power of automatic knot detection in function values or derivatives. It overcomes the drawback of [Formula: see text]-TF that is known to have bias issues. To solve the [Formula: see text]-TF problem, we propose an alternating minimization induced active set (AMIAS) search method based on the necessary optimality conditions derived from an augmented Lagrangian framework. The proposed method takes full advantage of the primal and dual variables with complementary supports, and decouples the high-dimensional problem into two subsystems on the active and inactive sets, respectively. A sequential AMIAS algorithm with warm start initialization is developed for efficient determination of the cardinality parameter, along with the output of solution paths. Theoretically, the oracle estimator of [Formula: see text]-TF is justified to behave like regression splines under the continuous time setting with mild conditions. Our numerical experiments include simulation studies for comparing [Formula: see text]-TF to [Formula: see text]-TF and free-knot splines on several synthetic examples, and a real data application of time series segmentation on Hong Kong PM2.5 indexes. History: Accepted by Antonio Frangioni, Area Editor for Design & Analysis of Algorithms – Continuous. Funding: This work was supported in part by Hong Kong General Research Fund [No. 17306519]. C. Wen’s research is partially supported by National Science Foundation of China [12171449] and Fundamental Research Funds for the Central Universities [WK3470000027, YD2040002019]. X. Wang’s research is partially supported by National Natural Science Foundation of China [Grants 72171216, 12231017, 71921001, and 71991474], and the National Key R&D Program of China [No. 2022YFA1003803]. Supplemental Material: The e-companion is available at https://doi.org/10.1287/ijoc.2021.0313 . The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2021.0313 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0313 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Canhong Wen, Aijun Zhang
INFORMS J. Comput.3
2023 Saliency Transfer Learning and Central-Cropping Network for Prostate Cancer Classification
Mengpei Jia, Jihao Luo, Aijun Zhang, Yongyong Chen, Peipei Shan, Binghui Zhao
Neural Process. Lett.5
2023 Single-Index Model Tree
abstract
In this paper, a novel single-index model tree (SIMTree) is proposed. It adopts the recursive partitioning strategy and each data segment is modeled by a single-index model (SIM), which is a flexible extension of linear regression with non-parametric link functions. The proposed SIMTree has 2 major advantages: a) with only a few leaf nodes, it can achieve competitive predictive performance compared to complicated black-box models; b) SIMs fitted on each local data segment are intrinsically interpretable. However, using conventional techniques to build such a SIMTree can be extremely time-consuming. SIM estimation typically involves iterative optimization via Newton-type algorithms; such a resource-intensive estimation procedure is repeatedly used for fitting leaf node SIMs and the search of optimal splits. To make the computation burden affordable, an effective training algorithm is proposed as enabled by the efficient utilization of Stein's lemma and several accelerating strategies in the tree construction algorithm. Moreover, a new Python package simtree is developed with elegant visualization modules that can further facilitate the model interpretation. Numerical results on extensive regression datasets show that SIMTree is an accurate and interpretable machine learning model.
Agus Sudjianto, Zebin Yang 0001, Aijun Zhang
IEEE Trans. Knowl. Data Eng.3
2022 Balance-Subsampled Stable Prediction Across Unknown Test Data
abstract
In data mining and machine learning, it is commonly assumed that training and test data share the same population distribution. However, this assumption is often violated in practice because of the sample selection bias, which might induce the distribution shift from training data to test data. Such a model-agnostic distribution shift usually leads to prediction instability across unknown test data. This article proposes a novel balance-subsampled stable prediction (BSSP) algorithm based on the theory of fractional factorial design. It isolates the clear effect of each predictor from the confounding variables. A design-theoretic analysis shows that the proposed method can reduce the confounding effects among predictors induced by the distribution shift, improving both the accuracy of parameter estimation and the stability of prediction across unknown test data. Numerical experiments on synthetic and real-world datasets demonstrate that our BSSP algorithm can significantly outperform the baseline methods for stable prediction across unknown test data.
Kun Kuang 0001, Hengtao Zhang, Runze Wu 0001, Fei Wu 0001, Yueting Zhuang, Aijun Zhang
ACM Trans. Knowl. Discov. Data6
2021 Hyperparameter Optimization via Sequential Uniform Designs
abstract
Hyperparameter optimization (HPO) plays a central role in the automated machine learning (AutoML). It is a challenging task as the response surfaces of hyperparameters are generally unknown, hence essentially a global optimization problem. This paper reformulates HPO as a computer experiment and proposes a novel sequential uniform design (SeqUD) strategy with three-fold advantages: a) the hyperparameter space is adaptively explored with evenly spread design points, without the need of expensive meta-modeling and acquisition optimization; b) the batch-by-batch design points are sequentially generated with parallel processing support; c) a new augmented uniform design algorithm is developed for the efficient real-time generation of follow-up design points. Extensive experiments are conducted on both global optimization tasks and HPO applications. The numerical results show that the proposed SeqUD strategy outperforms benchmark HPO methods, and it can be therefore a promising and competitive alternative to existing AutoML tools.
Zebin Yang 0001, Aijun Zhang
J. Mach. Learn. Res.2
2021 An effective SteinGLM initialization scheme for training multi-layer feedforward sigmoidal neural networks
Zebin Yang 0001, Hengtao Zhang, Agus Sudjianto, Aijun Zhang
Neural Networks4
2021 GAMI-Net: An explainable neural network based on generalized additive models with structured interactions
Zebin Yang 0001, Aijun Zhang, Agus Sudjianto
Pattern Recognit.2
2021 Enhancing Explainability of Neural Networks Through Architecture Constraints
abstract
Prediction accuracy and model explainability are the two most important objectives when developing machine learning algorithms to solve real-world problems. Neural networks are known to possess good prediction performance but suffer from a lack of model interpretability. In this article, we propose to enhance the explainability of neural networks through the following architecture constraints: 1) sparse additive subnetworks; 2) projection pursuit with orthogonality constraint; and 3) smooth function approximation. It leads to an enhanced explainable neural network (ExNN) with a superior balance between prediction performance and model interpretability. We derive sufficient identifiability conditions for the proposed ExNN model. The multiple parameters are simultaneously estimated by a modified minibatch gradient descent method based on the backpropagation algorithm for calculating the derivatives and the Cayley transform for preserving the projection orthogonality. Through simulation study under six different scenarios, we compare the proposed method to several benchmarks, including least absolute shrinkage and selection operator, support vector machine, random forest, extreme learning machine, and multilayer perceptron. It is shown that the proposed ExNN model keeps the flexibility of pursuing high prediction accuracy while attaining improved interpretability. Finally, a real data example is employed as a showcase application.
Zebin Yang 0001, Aijun Zhang, Agus Sudjianto
IEEE Trans. Neural Networks Learn. Syst.2
2019 Interval-valued data prediction via regularized artificial neural network
Zebin Yang 0001, Dennis K. J. Lin, Aijun Zhang
Neurocomputing3
2019 Mixed-level column augmented uniform designs
Aijun Zhang
J. Complex.3
2016 An optimized nonlinear grey Bernoulli model and its applications
Jianshan Lu, Weidong Xie, Hongbo Zhou 0009, Aijun Zhang
Neurocomputing4
2016 Exploring canonical correlation analysis with subspace and structured sparsity for web image annotation
Horace Ho-Shing Ip, Aijun Zhang
Image Vis. Comput.3
2007 The earth surface reflectance retrieval by exploiting the synergy of TERRA and AQUA MODIS data
abstract
The surface reflectance is of important geophysical parameter for many quantitative remote sensing applications. The inference of earth surface spectral reflectance using visible observations is complicated mainly because of scattering and absorption effects due to gases and aerosols. Surface reflectance product from MODIS data has been operationally offered, however, aerosol effect removal is based DDV algorithm which is restrictedly used for lower reflectance ground surface such as dense vegetation. In this paper we attempt a solution to the problem of retrieval of surface reflectance by exploiting the synergy of TERRA and AQUA MODIS data. Preliminary results of retrieval experiments and validation is encouraging and showed promising potential to address surface reflectance retrieval for land even for higher reflective surface.
Jiakui Tang, Aijun Zhang, Zhengmin He
IGARSS2
2007 An effective algorithm for generation of factorial designs with generalized minimum aberration
Kai-Tai Fang, Aijun Zhang, Runze Li 0001
J. Complex.2
2004 An experience on buffer analyzing in grid
abstract
Buffer analyzing is widely used for identifying areas surrounding geographic features. As the volume of spatial data increases rapidly, a high computational power is required. Over the past decade, grid has become a powerful tool to provide huge computational capabilities. It allocates the resource efficiently to applications. The objective of This work is to investigate the performance of buffer analysis in grid computing environments. In this paper, we discuss the parallel generation algorithm of line buffer zone. Our experiments were performed on a grid-computing environment, which is in the High-Throughput Spatial Information Processing Prototype System based on Grid platform in Institute of Remote Sensing Applications, Chinese Academy of Sciences. The computing nodes are heterogeneous. Performance results are also reported in This work.
Yincui Hu, Yong Xue, Guoyin Cai, Ying Luo 0006, Jianqin Wang, Jiakui Tang, Shaobo Zhong, Yanguang Wang, Xiaosong Sun, Aijun Zhang
IGARSS10
2004 A new approach to generate the look-up table for aerosol remote sensing on grid platform
abstract
Grid computing seeks to aggregate computing resources which are geographically distributed or heterogeneous and leverage on resources one don't own for oneself computational intensive applications. The procedure to generate the look-up table (LUT), which is very commonly used for aerosol remote sensing retrieval, is computational intensive even though the aims to take it are mainly to speedup the retrieval computation. This work focuses on realization of the compute-intensive look-up table generation on GCP-ARS (Grid Computation Platform for Aerosol Remote Sensing), which is one grid middleware we are developing based on Condor system. We discuss our approach to parameterization, task partitioning, generated methodology, and the collection of result. Experimental results obtained using Condor-pool consisted of commodity PCs are discussed.
Jiakui Tang, Yong Xue, Yanning Guan, Tong Yu 0007, Linxiang Liang, Yincui Hu, Ying Luo 0006, Guoyin Cai, Jianqin Wang, Shaobo Zhong, Yanguang Wang, Aijun Zhang
IGARSS12
2003 On Hadamard-Type Output Coding in Multiclass Learning
Aijun Zhang, Zhi-Li Wu, Chun-hung Li, Kai-Tai Fang
IDEAL1