VLDB 2026 Research / reviewers in the wild / expert
Hanfang Yang
dblp:60/11321
· DBLP profile ↗
14ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-0983-6758ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Locally Private Estimation with Public FeaturesabstractWe initiate the study of locally differentially private (LDP) learning with public features. We define semi-feature LDP, where some features are publicly available while the remaining ones, along with the label, require protection under local differential privacy. Under semi-feature LDP, we demonstrate that the mini-max convergence rate for non-parametric regression is significantly reduced compared to that of classical LDP. Then we propose HistOfTree, an estimator that fully leverages the information contained in both public and private features. Theoretically, HistOfTree reaches the mini-max optimal convergence rate. Empirically, HistOfTree achieves superior performance on both synthetic and real data. We also explore scenarios where users have the flexibility to select features for protection manually. In such cases, we propose an estimator and a data-driven parameter tuning strategy, leading to analogous theoretical and empirical results. Yuheng Ma 0001, Ke Jia, Hanfang Yang |
AISTATS | 3 |
| 2025 | ALTER: Augmentation for Large-Table-Based ReasoningabstractHan Zhang, Yuheng Ma, Hanfang Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yuheng Ma 0001, Hanfang Yang |
NAACL (Long Papers) | 3 |
| 2025 | Bagged Regularized k-Distances for Anomaly DetectionabstractWe consider the paradigm of unsupervised anomaly detection, which involves the identification of anomalies within a dataset in the absence of labeled examples. Though distance-based methods are top-performing for unsupervised anomaly detection, they suffer heavily from the sensitivity to the choice of the number of the nearest neighbors. In this paper, we propose a new distance-based algorithm called bagged regularized $k$-distances for anomaly detection (BRDAD), converting the unsupervised anomaly detection problem into a convex optimization problem. Our BRDAD algorithm selects the weights by minimizing the surrogate risk, i.e., the finite sample bound of the empirical risk of the bagged weighted $k$-distances for density estimation (BWDDE). This approach enables us to successfully address the sensitivity challenge of the hyperparameter choice in distance-based algorithms. Moreover, when dealing with large-scale datasets, the efficiency issues can be addressed by the incorporated bagging technique in our BRDAD algorithm. On the theoretical side, we establish fast convergence rates of the AUC regret of our algorithm and demonstrate that the bagging technique significantly reduces the computational complexity. On the practical side, we conduct numerical experiments to illustrate the insensitivity of the parameter selection of our algorithm compared with other state-of-the-art distance-based methods. Furthermore, our method achieves superior performance on real-world datasets with the introduced bagging technique compared to other approaches. Yuchao Cai, Hanfang Yang, Yuheng Ma 0001, Hanyuan Hang |
J. Mach. Learn. Res. | 2 |
| 2024 | Better Locally Private Sparse Estimation Given Multiple Samples Per UserabstractPrevious studies yielded discouraging results for item-level locally differentially private linear regression with $s$-sparsity assumption, where the minimax rate for $nm$ samples is $\mathcal{O}(sd / nm\varepsilon^2)$. This can be challenging for high-dimensional data, where the dimension $d$ is extremely large. In this work, we investigate user-level locally differentially private sparse linear regression. We show that with $n$ users each contributing $m$ samples, the linear dependency of dimension $d$ can be eliminated, yielding an error upper bound of $\mathcal{O}(s/ nm\varepsilon^2)$. We propose a framework that first selects candidate variables and then conducts estimation in the narrowed low-dimensional space, which is extendable to general sparse estimation problems with tight error bounds. Experiments on both synthetic and real datasets demonstrate the superiority of the proposed methods. Both the theoretical and empirical results suggest that, with the same number of samples, locally private sparse estimation is better conducted when multiple samples per user are available. Yuheng Ma 0001, Ke Jia, Hanfang Yang |
ICML | 3 |
| 2024 | Optimal Locally Private Nonparametric Classification with Public DataabstractIn this work, we investigate the problem of public data assisted non-interactive Local Differentially Private (LDP) learning with a focus on non-parametric classification. Under the posterior drift assumption, we for the first time derive the mini-max optimal convergence rate with LDP constraint. Then, we present a novel approach, the locally differentially private classification tree, which attains the mini-max optimal convergence rate. Furthermore, we design a data-driven pruning procedure that avoids parameter tuning and provides a fast converging estimator. Comprehensive experiments conducted on synthetic and real data sets show the superior performance of our proposed methods. Both our theoretical and experimental findings demonstrate the effectiveness of public data compared to private data, which leads to practical suggestions for prioritizing non-private data collection. Yuheng Ma 0001, Hanfang Yang |
J. Mach. Learn. Res. | 2 |
| 2023 | Extrapolated Random Tree for RegressionabstractIn this paper, we propose a novel tree-based algorithm named *Extrapolated Random Tree for Regression* (ERTR) that adapts to arbitrary smoothness of the regression function while maintaining the interpretability of the tree. We first put forward the *homothetic random tree for regression* (HRTR) that converges to the target function as the homothetic ratio approaches zero. Then ERTR uses a linear regression model to extrapolate HRTR estimations with different ratios to the ratio zero. From the theoretical perspective, we for the first time establish the optimal convergence rates for ERTR when the target function resides in the general Hölder space $C^{k,\alpha}$ for $k\in \mathbb{N}$, whereas the lower bound of the convergence rate of the random tree for regression (RTR) is strictly slower than ERTR in the space $C^{k,\alpha}$ for $k\geq 1$. This shows that ERTR outperforms RTR for the target function with high-order smoothness due to the extrapolation. In the experiments, we compare ERTR with state-of-the-art tree algorithms on real datasets to show the superior performance of our model. Moreover, promising improvements are brought by using the extrapolated trees as base learners in the extension of ERTR to ensemble methods. Yuchao Cai, Yuheng Ma 0001, Hanfang Yang |
ICML | 4 |
| 2023 | Decision Tree for Locally Private Estimation with Public DataabstractWe propose conducting locally differentially private (LDP) estimation with the aid of a small amount of public data to enhance the performance of private estimation. Specifically, we introduce an efficient algorithm called Locally differentially Private Decision Tree (LPDT) for LDP regression. We first use the public data to grow a decision tree partition and then fit an estimator according to the partition privately. From a theoretical perspective, we show that LPDT is $\varepsilon$-LDP and has a mini-max optimal convergence rate under a mild assumption of similarity between public and private data, whereas the lower bound of the convergence rate of LPDT without public data is strictly slower, which implies that the public data helps to improve the convergence rates of LDP estimation. We conduct experiments on both synthetic and real-world data to demonstrate the superior performance of LPDT compared with other state-of-the-art LDP regression methods. Moreover, we show that LPDT remains effective despite considerable disparities between public and private data. Yuheng Ma 0001, Yuchao Cai, Hanfang Yang |
NeurIPS | 4 |
| 2022 | Exploring Coarse-grained Pre-guided Attention to Assist Fine-grained Attention Reinforcement Learning AgentsabstractRecently, people have applied the attention mechanism to deep reinforcement learning (DRL), which commits to helping agents focus on crucial factors to learn the task more effectively. However, there is still some margin between the current attention methods and natural human attention since evidence suggests that human attention can be pre-guided before they perform a task, allowing humans to quickly catch areas of important factors at the beginning of the task and then gradually refine fine-grained attention to learn the details during training. This allows humans to use their attention more efficiently. In this paper, we propose an attention method that mimics human attention for DRL in the Atari Games. The proposed method contains a fusion attention module, for which we build a simulated human coarse-grained pre-guided (SHCP) attention module to assist the original fine-grained attention of RL agents. The proposed SHCP attention module contains information about key objects for game tasks and is implemented as a coarse-grained attention region. The experimental results demonstrate that our method can quickly boost performance in the early stages and then outperform the current state-of-the-art fine-grained attention methods significantly in sample efficiency, just like human attention. Further analysis shows that, with fusion attention, agents can not only capture rich features of pre-guided attention but also extend to more improved features after training, which suggests the pre-guided attention signal acts as a good initializer. Therefore, we consider our work reveals a potential and promising direction that combines human attention signals to affect agents' behavior via attention mechanisms. Yang Liu 0482, Xingrui Wang, Hanfang Yang |
IJCNN | 4 |
| 2022 | Under-bagging Nearest Neighbors for Imbalanced ClassificationabstractIn this paper, we propose an ensemble learning algorithm called under-bagging $k$-nearest neighbors (under-bagging $k$-NN) for imbalanced classification problems. On the theoretical side, by developing a new learning theory analysis, we show that with properly chosen parameters, i.e., the number of nearest neighbors $k$, the expected sub-sample size $s$, and the bagging rounds $B$, optimal convergence rates for under-bagging $k$-NN can be achieved under mild assumptions w.r.t. the arithmetic mean (AM) of recalls. Moreover, we show that with a relatively small $B$, the expected sub-sample size $s$ can be much smaller than the number of training data $n$ at each bagging round, and the number of nearest neighbors $k$ can be reduced simultaneously, especially when the data are highly imbalanced, which leads to substantially lower time complexity and roughly the same space complexity. On the practical side, we conduct numerical experiments to verify the theoretical results on the benefits of the under-bagging technique by the promising AM performance and efficiency of our proposed algorithm. Hanyuan Hang, Yuchao Cai, Hanfang Yang, Zhouchen Lin |
J. Mach. Learn. Res. | 3 |
| 2021 | Boosting Few-shot Abstractive Summarization with Auxiliary TasksabstractFor summarization in niche domains, data is not enough to fine-tune the large pre-trained model. In order to alleviate the few-shot problem, we design several auxiliary tasks to assist the main task---abstractive summarization. In this paper, we employ BART as the base sequence-to-sequence model and incorporate the main and auxiliary tasks under the multi-task framework. We transform all the tasks in the format of machine reading comprehension [19]. Moreover, we utilize the task-specific adapter to effectively share knowledge across tasks and the adaptive weight mechanism to adjust the contribution of auxiliary tasks to the main task. Experiments show the effectiveness of our method for few-shot datasets. We also propose to firstly pre-train the model on unlabeled datasets, and the methods proposed in this paper can further improve the model performance. Qiwei Bi, Hanfang Yang |
CIKM | 3 |
| 2020 | Boosted Histogram Transform for RegressionabstractIn this paper, we propose a boosting algorithm for regression problems called \emph{boosted histogram transform for regression} (BHTR) based on histogram transforms composed of random rotations, stretchings, and translations. From the theoretical perspective, we first prove fast convergence rates for BHTR under the assumption that the target function lies in the spaces $C^{0,\alpha}$. Moreover, if the target function resides in the subspace $C^{1,\alpha}$, by establishing the upper bound of the convergence rate for the boosted regressor, i.e. BHTR, and the lower bound for base regressors, i.e. histogram transform regressors (HTR), we manage to explain the benefits of the boosting procedure. In the experiments, compared with other state-of-the-art algorithms such as gradient boosted regression tree (GBRT), Breiman’s forest, and kernel-based methods, our BHTR algorithm shows promising performance on both synthetic and real datasets. Yuchao Cai, Hanyuan Hang, Hanfang Yang, Zhouchen Lin |
ICML | 3 |
| 2018 | Deep Sparse Informative Transfer SoftMax for Cross-Domain Image Classification
Hanfang Yang, Bo Yao 0007, Zijing Tan, Haocheng Tang, Yingjie Tian 0002 |
DASFAA (2) | 1 |
| 2015 | An Informative Logistic Regression for Cross-Domain Image Classification
Guangtang Zhu, Hanfang Yang, Guichun Zhou |
ICVS | 2 |
| 2013 | Understanding and Predicting Interestingness of VideosabstractThe amount of videos available on the Web is growing explosively. While some videos are very interesting and receive high rating from viewers, many of them are less interesting or even boring. This paper conducts a pilot study on the understanding of human perception of video interestingness, and demonstrates a simple computational method to identify more interesting videos. To this end we first construct two datasets of Flickr and YouTube videos respectively. Human judgements of interestingness are collected and used as the ground-truth for training computational models. We evaluate several off-the-shelf visual and audio features that are potentially useful for predicting interestingness on both datasets. Results indicate that audio and visual features are equally important and the combination of both modalities shows very promising results. Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001, Yingbin Zheng, Hanfang Yang |
AAAI | 6 |