VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Mao
dblp:232/4239
· DBLP profile ↗
23ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0002-9362-508XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Decentralized Median Value Estimation
Weidong Liu 0005, Xiaojun Mao |
IEEE Trans. Netw. | 2 |
| 2025 | Transfer Learning via Functional Balancing in Reproducing Kernel Hilbert SpacesabstractAs the availability of different data sources increases, transfer learning has become popular for improving estimation efficiency by incorporating those sources. In this paper, we propose a functional balancing transfer learning algorithm for observational studies integrating external summary information. We achieve efficiency improvement even when the model for the external summary information is misspecified, making the proposed algorithm more robust practically. The asymptotic properties of the proposed algorithm are investigated, and numerical studies demonstrate its advantages compared to alternative methods. Boyan Gu, Xiaojun Mao, Zhonglei Wang |
ICASSP | 3 |
| 2025 | A New Accelerated Off-Policy Stochastic Preconditioned TD(0) AlgorithmabstractIn this article, we consider policy evaluation in off-policy reinforcement learning and propose a novel procedure (Stochastic Preconditioned Temporal Difference (SPTD)) that achieves the optimal convergence rate under linear function approximation. The procedure has a linear computational complexity of the dimension of the feature space in each iteration. Under Markovian sampling, we establish finite-sample rates when the target policy can be different from the behavior policy for data generation. Our procedure is the first algorithm for the off-policy policy evaluation that has the optimal rate $\mathcal {O}(1/t)$O(1/t) under the mean square error. We also provide the first result on the asymptotic distribution and give the nearly optimal step size $\alpha _{t} = \mathcal {O}(t^{-2/3})$αt=O(t-2/3). The numerical performance of the procedure is studied in both on-policy and off-policy settings. Extensive numerical experiments demonstrate that our procedure uniformly outperforms existing methods. Weidong Liu 0005, Jiahua Ma, Xiaojun Mao, Kejie Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Group-Sparse Inductive Matrix Completion Through Transfer LearningabstractThe emergence of big data has enabled the creation of significant models by allowing the storage of large data volumes. Transfer learning is a machine learning technique that transfers knowledge between different domains by utilizing pretrained models from the source domain to optimize the target domain. In contrast, inductive matrix completion is a method that leverages side information from multiple sources to improve task performance. This paper explores inductive matrix completion within the transfer learning framework, with our proposed approach assuming group sparsity for the difference between the core matrices of the target and source domains. Theoretical guarantees of our method are investigated to demonstrate the gains achieved through transfer learning compared with standard inductive matrix completion. Several synthetic experiments are conducted to evaluate the performance of the proposed approach and existing methods, demonstrating that our method outperforms others. Xiaojun Mao, Hengfang Wang, Zhonglei Wang |
IEEE Trans. Inf. Theory | 1 |
| 2024 | SAM: A Self-Adaptive Attention Module for Context-Aware Recommendation SystemabstractRecently, textual information has been proven to positively affect recommendation systems. However, most of the existing methods only focus on representation learning of textual information in ratings, while potential selection bias induced by the textual information is ignored. In this work, we propose a novel and general self-adaptive module, the self-adaptive attention module (SAM), which adjusts the selection bias by capturing contextual information based on its representation. This module can be embedded into recommendation systems that contain learning components of contextual information. Experimental results on three real-world datasets demonstrate the effectiveness of our proposal, and the state-of-the-art models with SAM significantly outperform the original ones. Zhengpin Li, Xiaojun Mao, Jian Wang 0016, Zhongyu Wei |
ICASSP | 4 |
| 2024 | ST$_k$: A Scalable Module for Solving Top-k ProblemsabstractThe cost of ranking becomes significant in the new stage of deep learning. We propose ST$_k$, a fully differentiable module with a single trainable parameter, designed to solve the Top-k problem without requiring additional time or GPU memory. Due to its fully differentiable nature, ST$_k$ can be embedded end-to-end into neural networks and optimize the Top-k problems within a unified computational graph. We apply ST$_k$ to the Average Top-k Loss (AT$_k$), which inherently faces a Top-k problem. The proposed ST$_k$ Loss outperforms AT$_k$ Loss and achieves the best average performance on multiple benchmarks, with the lowest standard deviation. With the assistance of ST$_k$ Loss, we surpass the state-of-the-art (SOTA) on both CIFAR-100-LT and Places-LT leaderboards. Hanchen Xia, Weidong Liu 0005, Xiaojun Mao |
NeurIPS | 3 |
| 2024 | Distributed Estimation on Semi-Supervised Generalized Linear ModelabstractSemi-supervised learning is devoted to using unlabeled data to improve the performance of machine learning algorithms. In this paper, we study the semi-supervised generalized linear model (GLM) in the distributed setup. In the cases of single or multiple machines containing unlabeled data, we propose two distributed semi-supervised algorithms based on the distributed approximate Newton method. When the labeled local sample size is small, our algorithms still give a consistent estimation, while fully supervised methods fail to converge. Moreover, we theoretically prove that the convergence rate is greatly improved when sufficient unlabeled data exists. Therefore, the proposed method requires much fewer rounds of communications to achieve the optimal rate than its fully-supervised counterpart. In the case of the linear model, we prove the rate lower bound after one round of communication, which shows that rate improvement is essential. Finally, several simulation analyses and real data studies are provided to demonstrate the effectiveness of our method. J. Y. Tu, Weidong Liu 0005, Xiaojun Mao |
J. Mach. Learn. Res. | 3 |
| 2024 | Efficient and provable online reduced rank regression via online gradient descent
Weidong Liu 0005, Xiaojun Mao |
Mach. Learn. | 3 |
| 2024 | Multi-consensus decentralized primal-dual fixed point algorithm for distributed learning
Kejie Tang, Weidong Liu 0005, Xiaojun Mao |
Mach. Learn. | 3 |
| 2024 | Collective Matrix Completion via Graph ExtractionabstractCollective matrix completion (CMC) offers a straightforward approach to dealing with data with entries from various sources. Benefiting from the joint structure in the collective matrix, CMC often achieves fast convergence. However, since CMC conducts matrix-level operations, it neglects the entry-wise information that can potentially be very useful for matrix completion. In this paper, to capture the entry-wise information, we propose a method called graph collective matrix completion (GCoMC). Specifically, our method integrates a graph pattern extraction module into CMC via a relational graph convolutional network. Experiments on simulated and real-world datasets show that our method significantly outperforms some existing counterparts. Tong Zhan, Xiaojun Mao, Jian Wang 0016, Zhonglei Wang |
IEEE Signal Process. Lett. | 2 |
| 2024 | Efficient Sparse Least Absolute Deviation Regression With Differential PrivacyabstractIn recent years, privacy-preserving machine learning algorithms have attracted increasing attention because of their important applications in many scientific fields. However, in the literature, most privacy-preserving algorithms demand learning objectives to be strongly convex and Lipschitz smooth, which thus cannot cover a wide class of robust loss functions (e.g., quantile/least absolute loss). In this work, we aim to develop a fast privacy-preserving learning solution for a sparse robust regression problem. Our learning loss consists of a robust least absolute loss and an$\ell _{1}$sparse penalty term. To fast solve the non-smooth loss under a given privacy budget, we develop a Fast Robust And Privacy-Preserving Estimation (FRAPPE) algorithm for least absolute deviation regression. Our algorithm achieves a fast estimation by reformulating the sparse LAD problem as a penalized least square estimation problem and adopts a three-stage noise injection to guarantee the$(\epsilon,\delta)$-differential privacy. We show that our algorithm can achieve better privacy and statistical accuracy trade-off compared with the state-of-the-art privacy-preserving regression algorithms. In the end, we conduct experiments to verify the efficiency of our proposed FRAPPE algorithm. Weidong Liu 0005, Xiaojun Mao, Xin Zhang 0054 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Distributed Semi-Supervised Sparse Statistical InferenceabstractThe debiased estimator is a crucial tool in statistical inference for high-dimensional model parameters. However, constructing such an estimator involves estimating the high-dimensional inverse Hessian matrix, incurring significant computational costs. This challenge becomes particularly acute in distributed setups, where traditional methods necessitate computing a debiased estimator on every machine. This becomes unwieldy, especially with a large number of machines. In this paper, we delve into semi-supervised sparse statistical inference in a distributed setup. An efficient multi-round distributed debiased estimator, which integrates both labeled and unlabelled data, is developed. We will show that the additional unlabeled data helps to improve the statistical rate of each round of iteration. Our approach offers tailored debiasing methods forM-estimation and generalized linear models according to the specific form of the loss function. Our method also applies to a non-smooth loss like absolute deviation loss. Furthermore, our algorithm is computationally efficient since it requires only one estimation of a high-dimensional inverse covariance matrix. We demonstrate the effectiveness of our method by presenting simulation studies and real data applications that highlight the benefits of incorporating unlabeled data. J. Y. Tu, Weidong Liu 0005, Xiaojun Mao |
IEEE Trans. Inf. Theory | 3 |
| 2023 | Transductive Matrix Completion with Calibration for Multi-Task LearningabstractMulti-task learning has attracted much attention due to growing multi-purpose research with multiple related data sources. More- over, transduction with matrix completion is a useful method in multi-label learning. In this paper, we propose a transductive matrix completion algorithm that incorporates a calibration constraint for the features under the multi-task learning framework. The proposed algorithm recovers the incomplete feature matrix and target matrix simultaneously. Fortunately, the calibration information improves the completion results. In particular, we provide a statistical guarantee for the proposed algorithm, and the theoretical improvement induced by calibration information is also studied. Moreover, the proposed algorithm enjoys a sub-linear convergence rate. Several synthetic data experiments are conducted, which show the proposed algorithm out-performs other methods, especially when the target matrix is associated with the feature matrix in a nonlinear way. Hengfang Wang, Yasi Zhang, Xiaojun Mao, Zhonglei Wang |
ICASSP | 3 |
| 2023 | Byzantine-robust distributed sparse learning for M-estimation
J. Y. Tu, Weidong Liu 0005, Xiaojun Mao |
Mach. Learn. | 3 |
| 2022 | Applying Differential Privacy to Tensor CompletionabstractTensor completion aims at filling the missing or unobserved en-tries based on partially observed tensors. However, utilization of the observed tensors often raises serious privacy concerns in many practical scenarios. To address this issue, we propose a solid and unified framework that contains several approaches for applying differential privacy to the two most widely used tensor decomposition methods: i) CANDECOMP/PARAFAC and ii) Tucker decompositions. For each approach, we establish a rigorous privacy guarantee and meanwhile evaluate the privacy-accuracy trade-off. Experiments on synthetic datasets demonstrate that our proposal achieves high accuracy for tensor completion while ensuring strong privacy protections. Zhengpin Li, Xiaojun Mao, Jian Wang 0016 |
ICASSP | 3 |
| 2022 | Uncertainty Modeling in Generative Compressed SensingabstractCompressed sensing (CS) aims to recover a high-dimensional signal with structural priors from its low-dimensional linear measurements. Inspired by the huge success of deep neural networks in modeling the priors of natural signals, generative neural networks have been recently used to replace the hand-crafted structural priors in CS. However, the reconstruction capability of the generative model is fundamentally limited by the range of its generator, typically a small subset of the signal space of interest. To break this bottleneck and thus reconstruct those out-of-range signals, this paper presents a novel method called CS-BGM that can effectively expands the range of generator. Specifically, CS-BGM introduces uncertainties to the latent variable and parameters of the generator, while adopting the variational inference (VI) and maximum a posteriori (MAP) to infer them. Theoretical analysis demonstrates that expanding the range of generators is necessary for reducing the reconstruction error in generative CS. Extensive experiments show a consistent improvement of CS-BGM over the baselines. Yilang Zhang, Mengchu Xu, Xiaojun Mao, Jian Wang 0016 |
ICML | 3 |
| 2022 | Byzantine-tolerant distributed multiclass sparse linear discriminant analysisabstractCommunication cost and security issues are both important in large-scale distributed machine learning. In this paper, we investigate a multiclass sparse classification problem under two distributed systems. We propose two distributed multiclass sparse discriminant analysis algorithms based on mean-aggregation and median-aggregation under the normal distributed system or Byzantine failure system. Both of them are computation and communication efficient. Several theoretical results, including estimation error bounds, and support recovery, are established. With moderate initial estimators, our iterative estimators achieve a (near-)optimal rate and exact support recovery after a constant number of rounds. Experiments on both synthetic and real datasets are provided to demonstrate the effectiveness of our proposed methods. Yajie Bao, Weidong Liu 0005, Xiaojun Mao, Weijia Xiong |
UAI | 3 |
| 2022 | Compressive Sensing Approaches for Sparse Distribution Estimation Under Local PrivacyabstractRecent years, local differential privacy (LDP) has been adopted by many web service providers like Google [23], Apple [33] and Microsoft [15] to collect and analyse users’ data privately. In this paper, we consider the problem of discrete distribution estimation under local differential privacy constraints. Distribution estimation is one of the most fundamental estimation problems, which is widely studied in both non-private and private settings. In the local model, private mechanisms with provably optimal sample complexity are known. However, they are optimal only in the worst-case sense; their sample complexity is proportional to the size of the entire universe, which could be huge in practice. In this paper, we consider sparse or approximately sparse (e.g. highly skewed) distribution, and show that the number of samples needed could be significantly reduced. This problem has been studied recently [1], but they only consider strict sparse distributions and the high privacy regime. We propose new privatization mechanisms based on compressive sensing. Our methods work for approximately sparse distributions and medium privacy, and have optimal sample and communication complexity. Zhongzheng Xiong, Xiaojun Mao, Jian Wang 0016, Shan Ying, Zengfeng Huang |
WWW | 3 |
| 2021 | Matrix Completion with Model-free WeightingabstractIn this paper, we propose a novel method for matrix completion under general non-uniform missing structures. By controlling an upper bound of a novel balancing error, we construct weights that can actively adjust for the non-uniformity in the empirical risk without explicitly modeling the observation probabilities, and can be computed efficiently via convex optimization. The recovered matrix based on the proposed weighted empirical risk enjoys appealing theoretical guarantees. In particular, the proposed method achieves stronger guarantee than existing work in terms of the scaling with respect to the observation probabilities, under asymptotically heterogeneous missing settings (where entry-wise observation probabilities can be of different orders). These settings can be regarded as a better theoretical model of missing patterns with highly varying probabilities. We also provide a new minimax lower bound under a class of heterogeneous settings. Numerical experiments are also provided to demonstrate the effectiveness of the proposed method. Raymond K. W. Wong, Xiaojun Mao, Kwun Chuen Gary Chan |
ICML | 3 |
| 2021 | Variance Reduced Median-of-Means Estimator for Byzantine-Robust Distributed InferenceabstractThis paper develops an efficient distributed inference algorithm, which is robust against a moderate fraction of Byzantine nodes, namely arbitrary and possibly adversarial machines in a distributed learning system. In robust statistics, the median-of-means (MOM) has been a popular approach to hedge against Byzantine failures due to its ease of implementation and computational efficiency. However, the MOM estimator has the shortcoming in terms of statistical efficiency. The first main contribution of the paper is to propose a variance reduced median-of-means (VRMOM) estimator, which improves the statistical efficiency over the vanilla MOM estimator and is computationally as efficient as the MOM. Based on the proposed VRMOM estimator, we develop a general distributed inference algorithm that is robust against Byzantine failures. Theoretically, our distributed algorithm achieves a fast convergence rate with only a constant number of rounds of communications. We also provide the asymptotic normality result for the purpose of statistical inference. To the best of our knowledge, this is the first normality result in the setting of Byzantine-robust distributed learning. The simulation results are also presented to illustrate the effectiveness of our method. J. Y. Tu, Weidong Liu 0005, Xiaojun Mao |
J. Mach. Learn. Res. | 3 |
| 2020 | Automatic Generation of Electromyogram Diagnosis ReportabstractElectrophysiological tests, especially, electromyogram (EMG) and nerve conduction velocity (NCV) test are commonly used in clinical practice for diagnosis of muscle and nerve diseases. Report-writing of these tests can be problematic for under-experienced physicians and time-consuming for experienced physicians. In this paper, we apply several neural based natural language generation (NLG) methods to automatically generate diagnosis reports, a first attempt in this domain. Specifically, we use tabular diagnostic records of electrophysiological tests to generate Findings & Impression, which together constitute the diagnostic report. We further use gram-based metrics to evaluate our models and conduct a case study for the result. Qizheng Gu, Cong Nie, Ruixiang Zou, Wei Chen 0088, Chaojun Zheng, Dongqing Zhu, Xiaojun Mao, Zhongyu Wei, Dong Tian |
BIBM | 7 |
| 2020 | Median Matrix Completion: from Embarrassment to OptimalityabstractIn this paper, we consider matrix completion with absolute deviation loss and obtain an estimator of the median matrix. Despite several appealing properties of median, the non-smooth absolute deviation loss leads to computational challenge for large-scale data sets which are increasingly common among matrix completion problems. A simple solution to large-scale problems is parallel computing. However, embarrassingly parallel fashion often leads to inefficient estimators. Based on the idea of pseudo data, we propose a novel refinement step, which turns such inefficient estimators into a rate (near-)optimal matrix completion procedure. The refined estimator is an approximation of a regularized least median estimator, and therefore not an ordinary regularized empirical risk estimator. This leads to a non-standard analysis of asymptotic behaviors. Empirical results are also provided to confirm the effectiveness of the proposed method. Weidong Liu 0005, Xiaojun Mao, Raymond K. W. Wong |
ICML | 2 |
| 2020 | Distributed High-dimensional Regression Under a Quantile Loss FunctionabstractThis paper studies distributed estimation and support recovery for high-dimensional linear regression model with heavy-tailed noise. To deal with heavy-tailed noise whose variance can be infinite, we adopt the quantile regression loss function instead of the commonly used squared loss. However, the non-smooth quantile loss poses new challenges to high-dimensional distributed estimation in both computation and theoretical development. To address the challenge, we transform the response variable and establish a new connection between quantile regression and ordinary linear regression. Then, we provide a distributed estimator that is both computationally and communicationally efficient, where only the gradient information is communicated at each iteration. Theoretically, we show that, after a constant number of iterations, the proposed estimator achieves a near-oracle convergence rate without any restriction on the number of machines. Moreover, we establish the theoretical guarantee for the support recovery. The simulation analysis is provided to demonstrate the effectiveness of our method. Weidong Liu 0005, Xiaojun Mao, Zhuoyi Yang |
J. Mach. Learn. Res. | 3 |