VLDB 2026 Research / reviewers in the wild / expert
Zijun Gao
dblp:239/5605
· DBLP profile ↗
22ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SegRap2025: A benchmark of gross tumor volume and lymph node clinical target volume Segmentation for Radiotherapy Planning of nasopharyngeal carcinoma
Litingyu Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu 0001, Yicheng Wu 0001, Jin Ye 0002, Linhao Li, Yiwen Ye, Yong Xia 0001, Elias Tappeiner, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Shichuan Zhang, Shaoting Zhang 0001, Wenjun Liao, Guotai Wang |
Medical Image Anal. | 7 |
| 2026 | Fast Online Channel Estimation in Massive MIMO: A Zero-Shot Self-Supervised Approach
Zijun Gao, Wenqiang Yi, Fatma Benkhelifa, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Trustworthy assessment of heterogeneous treatment effect estimator via analysis of relative errorabstractAccurate heterogeneous treatment effect (HTE) estimation is essential for personalized recommendations, making it important to evaluate and compare HTE estimators. Traditional assessment methods are inapplicable due to missing counterfactuals. Current HTE evaluation methods rely on additional estimation or matching on test data, often ignoring the uncertainty introduced and potentially leading to incorrect conclusions. We propose incorporating uncertainty quantification into HTE estimator comparisons. In addition, we suggest shifting the focus to the estimation and inference of the relative error between methods rather than their absolute errors. Methodology-wise, we develop a relative error estimator based on the efficient influence function and establish its asymptotic distribution for inference. Compared to absolute error-based methods, the relative error estimator (1) is less sensitive to the error of nuisance function estimators, satisfying a "global double robustness" property, and (2) its confidence intervals are often narrower, making it more powerful for determining the more accurate HTE estimator. Through extensive empirical study of the ACIC challenge benchmark datasets, we show that the relative error-based method more effectively identifies the better HTE estimator with statistical confidence, even with a moderately large test dataset or inaccurate nuisance estimators. Zijun Gao |
AISTATS | 1 |
| 2025 | Bridging Multiple Worlds: Multi-marginal Optimal Transport for Causal Partial-identification ProblemabstractUnder the prevalent potential outcome model in causal inference, each unit is associated with multiple potential outcomes but at most one of which is observed, leading to many causal quantities being only partially identified. The inherent missing data issue echoes the multi-marginal optimal transport (MOT) problem, where marginal distributions are known, but how the marginals couple to form the joint distribution is unavailable. In this paper, we cast the causal partial identification problem in the framework of MOT with $K$ margins and $d$-dimensional outcomes and obtain the exact partial identified set. In order to estimate the partial identified set via MOT, statistically, we establish a convergence rate of the plug-in MOT estimator for the $\ell_2$ cost function stemming from the variance minimization problem and prove it is minimax optimal for arbitrary $K$ and $d \le 4$. We also extend the convergence result to general quadratic objective functions. Numerically, we demonstrate the efficacy of our method over synthetic datasets and several real-world datasets where our proposal consistently outperforms the baseline by a significant margin (over 70%). In addition, we provide efficient off-the-shelf implementations of MOT with general objective functions. Zijun Gao, Shu Ge, Jian Qian |
AISTATS | 1 |
| 2025 | Meta-ZSN2N: Zero-Shot Learning for Channel Estimation in Massive MIMO SystemsabstractIn massive MIMO systems, traditional channel estimation techniques often suffer from noise sensitivity and high computational complexity. Recently proposed deep supervised learning–based estimators have improved accuracy yet require large labeled datasets and exhibit poor generalization in dynamic channel conditions. Consequently, self-supervised methods have emerged, avoiding extensive label collection and enabling immediate online deployment. However, existing self-supervised frameworks typically rely on large networks with long run times, demanding substantial computational resources. In this work, we present a lightweight self-supervised channel estimation framework, Meta-ZSN2N. It first leverages a traditional estimator, then applies a specialized downsampling step, and finally refines the results via a lightweight two-layer neural network, resulting in a significantly simplified model and a substantially reduced runtime. To further accelerate online inference and boost generalization, we integrate a Meta-SGD module into our design. Simulation results indicate that our proposed lightweight method not only surpasses traditional estimators but also outperforms large learning networks with millions of parameters in terms of efficiency and adaptability. Zijun Gao, Wenqiang Yi, Fatma Benkhelifa, Arumugam Nallanathan |
GLOBECOM | 1 |
| 2025 | SAGEPhos: Sage Bio-Coupled and Augmented Fusion for Phosphorylation Site DetectionabstractPhosphorylation site prediction based on kinase-substrate interaction plays a vital role in understanding cellular signaling pathways and disease mechanisms. Computational methods for this task can be categorized into kinase-family-focused and individual kinase-targeted approaches. Individual kinase-targeted methods have gained prominence for their ability to explore a broader protein space and provide more precise target information for kinase inhibitors. However, most existing individual kinase-based approaches focus solely on sequence inputs, neglecting crucial structural information. To address this limitation, we introduce SAGEPhos (Structure-aware kinAse-substrate bio-coupled and bio-auGmented nEtwork for Phosphorylation site prediction), a novel framework that modifies the semantic space of main protein inputs using auxiliary inputs at two distinct modality levels. At the inter-modality level, SAGEPhos introduces a Bio-Coupled Modal Fusion method, distilling essential kinase sequence information to refine task-oriented local substrate feature space, creating a shared semantic space that captures crucial kinase-substrate interaction patterns. Within the substrate's intra-modality domain, it focuses on Bio-Augmented Fusion, emphasizing 2D local sequence information while selectively incorporating 3D spatial information from predicted structures to complement the sequence space. Moreover, to address the lack of structural information in current datasets, we contribute a new, refined phosphorylation site prediction dataset, which incorporates crucial structural elements and will serve as a new benchmark for the field. Experimental results demonstrate that SAGEPhos significantly outperforms baseline methods, notably achieving almost 10\% and 12\% improvements in prediction accuracy and AUC-ROC, respectively. We further demonstrate our algorithm's robustness and generalization through stable results across varied data partitions and significant improvements in zero-shot scenarios. These results underscore the effectiveness of constructing a larger and more precise protein space in advancing the state-of-the-art in phosphorylation site prediction. We release the SAGEPhos models and code at https://github.com/ZhangJJ26/SAGEPhos. Jingjie Zhang, Hanqun Cao, Zijun Gao, Chunbin Gu |
ICLR | 3 |
| 2025 | Tightening Causal Bounds via Covariate-Aware Optimal TransportabstractCausal estimands can vary significantly depending on the relationship between outcomes in treatment and control groups, leading to wide partial identification (PI) intervals that impede decision making. Incorporating covariates can substantially tighten these bounds, but requires determining the range of PI over probability models consistent with the joint distributions of observed covariates and outcomes in treatment and control groups. This problem is known to be equivalent to a conditional optimal transport (COT) optimization task, which is more challenging than standard optimal transport (OT) due to the additional conditioning constraints. In this work, we study a tight relaxation of COT that effectively reduces it to standard OT, leveraging its well-established computational and theoretical foundations. Our relaxation incorporates covariate information and ensures narrower PI intervals for any value of the penalty parameter, while becoming asymptotically exact as a penalty increases to infinity. This approach preserves the benefits of covariate adjustment in PI and results in a data-driven estimator for the PI set that is easy to implement using existing OT packages. We analyze the convergence rate of our estimator and demonstrate the effectiveness of our approach through extensive simulations, highlighting its practical use and superior performance compared to existing methods. Sirui Lin, Zijun Gao, Jose H. Blanchet, Peter W. Glynn |
ICML | 2 |
| 2025 | Dynamic Gradient Sparsification Training for Few-Shot Fine-Tuning of CT Lymph Node Segmentation Foundation Model
Zijun Gao, Wenjun Liao, Shichuan Zhang, Guotai Wang, Xiangde Luo |
MICCAI (5) | 2 |
| 2025 | Protein Inverse Folding From Structure FeedbackabstractThe inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications.
Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model.
Given a target protein structure, we begin by sampling candidate sequences from the inverse‐folding model, then predict the three‐dimensional structure of each sequence with the folding model to generate pairwise structural‐preference labels.
These labels are used to fine‐tune the inverse‐folding model under the DPO objective.
Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity.
Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5\% with regard to the baseline model.
This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization. Junde Xu, Zijun Gao, Xinyi Zhou 0010, Xingyi Cheng, Guangyong Chen, Pheng-Ann Heng, Jiezhong Qiu |
NeurIPS | 2 |
| 2025 | Efficient method for detecting targets from remote sensing images based on global attention mechanismabstractAbstract Remote sensing image target detection provides an effective and accurate data analysis tool for many application areas. Due to complex backgrounds, large differences in target scales, and missed detection of small targets, remote sensing image target detection is challenging. In order to enhance the model's understanding of the global information of remote sensing images, this paper proposes the GFA module. This module can establish the global contextual connection of remote sensing images to provide rich context to help understand the complex scene and background in which the target is located, without being limited to local information. Additionally, it focuses on channel information for enhanced target feature extraction. For the purpose of alleviating the serious imbalance in foreground–background samples that is present in single‐level target detection models. The loss function is reconstructed based on focal loss by redefining the balance factor α and focus factor γ , so that it can be dynamically adjusted during network training. Meanwhile, EIoU is used to further enhance the bounding box regression capability. Affine transformations were also used to augment the dataset in order to assist the model in adjusting to real‐world situations. The proposed method is experimentally validated on the publicly available HRRSD dataset. In comparison with YOLO v5, the mAP of the detection results improved by 2.7%. Compared with YOLO v8 and YOLO v10, the mAP improved by 3.2% and 3.3%. The model achieves an FPS of 40.1, an optimal balance between speed and accuracy. Further, experiments are conducted using the NWPU VHR‐10 dataset and the RSOD dataset, both of which demonstrated that the proposed method outperforms other target detection methods and improves remote sensing target detection performance. Zijun Gao, Jingwen Su, Zhankui Song |
IET Image Process. | 1 |
| 2025 | Data Augmentation for Time-Series Classification: An Extensive Empirical Study and Comprehensive SurveyabstractBackground: Data Augmentation (DA) has become a critical approach in Time Series Classification (TSC), primarily for its capacity to expand training datasets, enhance model robustness, introduce diversity, and reduce overfitting. However, the current landscape of DA in TSC is plagued with fragmented literature reviews, nebulous methodological taxonomies, inadequate evaluative measures, and a dearth of accessible and user-oriented tools. Objectives: This study addresses these challenges through a comprehensive examination of DA methodologies within the TSC domain. Methods: Our research began with an extensive literature review spanning a decade, revealing significant gaps in existing surveys and necessitating a detailed analysis of over 100 scholarly articles to identify more than 60 distinct DA techniques. This rigorous review led to the development of a novel taxonomy tailored to the specific needs of DA in TSC, categorizing techniques into five primary categories to guide researchers in selecting appropriate methods with greater clarity. In response to the lack of comprehensive evaluations of foundational DA techniques, we conducted a thorough empirical study, testing nearly 20 DA strategies across 15 diverse datasets representing all types within the UCR time-series repository. To improve practical use, we have consolidated most of these methods into a unified Python Library, whose user-friendly interface facilitates experimenting with various augmentation techniques, offering practitioners and researchers a more convenient tool for innovation than currently available options. Results: Using ResNet and LSTM architectures, we employed a multifaceted evaluation approach, including metrics such as Accuracy, Method Ranking, and Residual Analysis, resulting in a benchmark accuracy of 84.98 ± 16.41% in ResNet and 82.41 ± 18.71% in LSTM. Our investigation underscored the inconsistent efficacies of DA techniques, for instance, methods like RGWs (with an average rank of 7.13 and average accuracy of 83.42 ± 17.53% in LSTM) and Random Permutation significantly improved model performance, whereas others, like EMD, were less effective. Furthermore, we found that the intrinsic characteristics of datasets significantly influence the success of DA methods, leading to targeted recommendations based on empirical evidence to help practitioners select the most suitable DA techniques for specific datasets. Conclusions: In essence, this research presents an integrative perspective on the contemporary landscape of data augmentation for time series classification, combining theoretical frameworks with empirical evidence. The revelations and resources introduced herein are positioned to catalyze continued progress in this domain, fortifying machine learning models against the challenges posed by data limitations, and enhancing their generalizability and robustness. Zijun Gao, Haibao Liu |
J. Artif. Intell. Res. | 1 |
| 2024 | Gradient-based Parameter Selection for Efficient Fine-TuningabstractWith the growing size of pre-trained models, full fine-tuning and storing all the parameters for various down-stream tasks is costly and infeasible. In this paper, we propose a new parameter-efficient fine-tuning method, Gradient-based Parameter Selection (GPS), demonstrating that only tuning a few selected parameters from the pre-trained model while keeping the remainder of the model frozen can generate similar or better performance compared with the full model fine-tuning method. Different from the existing popular and state-of-the-art parameter-efficient fine-tuning approaches, our method does not in-troduce any additional parameters and computational costs during both the training and inference stages. Another ad-vantage is the model-agnostic and non-destructive property, which eliminates the need for any other design specific to a particular model. Compared with the full fine-tuning, GPS achieves 3.33% (91.78% vs. 88.45%, FGVC) and 9.61% (73.1% vs. 65.57%, VTAB) improvement of the accu-racy with tuning only 0.36% parameters of the pre-trained model on average over 24 image classification tasks; it also demonstrates a significant improvement of 17% and 16.8% in mDice and mIoU, respectively, on medical image segmentation task. Moreover, GPS achieves state-of-the-art performance compared with existing PEFT meth-ods. The code will be available in https://github.com/FightingFighting/GPS.git. Zhi Zhang 0009, Qizhe Zhang, Zijun Gao, Renrui Zhang, Ekaterina Shutova, Shiji Zhou, Shanghang Zhang |
CVPR | 3 |
| 2024 | Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation PerspectiveabstractVision-language models (VLMs) pre-trained on extensive datasets can inadvertently learn biases by correlating gender information with specific objects or scenarios.Current methods, which focus on modifying inputs and monitoring changes in the model's output probability scores, often struggle to comprehensively understand bias from the perspective of model components.We propose a framework that incorporates causal mediation analysis to measure and map the pathways of bias generation and propagation within VLMs.Our framework is applicable to a wide range of vision-language and multimodal tasks.In this work, we apply it to the object detection task and implement it on the GLIP model.This approach allows us to identify the direct effects of interventions on model bias and the indirect effects of interventions on bias mediated through different model components.Our results show that image features are the primary contributors to bias, with significantly higher impacts than text features, specifically accounting for 32.57% and 12.63% of the bias in the MSCOCO and PASCAL-SENTENCE datasets, respectively.Notably, the image encoder's contribution surpasses that of the text encoder and the deep fusion encoder.Further experimentation confirms that contributions from both language and vision modalities are aligned and non-conflicting.Consequently, focusing on blurring gender representations within the image encoder which contributes most to the model bias, reduces bias efficiently by 22.03% and 9.04% in the MSCOCO and PASCAL-SENTENCE datasets, respectively, with minimal performance loss or increased computational demands. 1 Zhaotian Weng, Zijun Gao, Jerone Theodore Alexander Andrews, Jieyu Zhao 0001 |
EMNLP | 2 |
| 2024 | An Uncertainty-Guided Tiered Self-training Framework for Active Source-Free Domain Adaptation in Prostate Segmentation
Xiangde Luo, Zijun Gao, Guotai Wang |
MICCAI (9) | 3 |
| 2024 | Learning using privileged information with logistic regression on acute respiratory distress syndrome detectionabstractThe advanced learning paradigm, learning using privileged information (LUPI), leverages information in training that is not present at the time of prediction. In this study, we developed privileged logistic regression (PLR) models under the LUPI paradigm to detect acute respiratory distress syndrome (ARDS), with mechanical ventilation variables or chest x-ray image features employed in the privileged domain and electronic health records in the base domain. In model training, the objective of privileged logistic regression was designed to incorporate data from the privileged domain and encourage knowledge transfer across the privileged and base domains. An asymptotic analysis was also performed, yielding sufficient conditions under which the addition of privileged information increases the rate of convergence in the proposed model. Results for ARDS detection show that PLR models achieve better classification performances than logistic regression models trained solely on the base domain, even when privileged information is partially available. Furthermore, PLR models demonstrate performance on par with or superior to state-of-the-art models under the LUPI paradigm. As the proposed models are effective, easy to interpret, and highly explainable, they are ideal for other clinical applications where privileged information is at least partially available. Zijun Gao, Shuyang Cheng, Emily Wittrup, Jonathan Gryak, Kayvan Najarian |
Artif. Intell. Medicine | 1 |
| 2024 | JOA-GAN: An improved single-image super-resolution network for remote sensing based on GANabstractAbstract Image super‐resolution (SR) has been widely applied in remote sensing to generate high‐resolution (HR) images without increasing hardware costs. However, SR is a severe ill‐posed problem. As deep learning advances, existing methods have solved this problem to a certain extent. However, the complex spatial distribution of remote sensing images still poses a challenge in effectively extracting abundant high‐frequency details from the images. Here, a single‐image super‐resolution (SISR) network based on the generative adversarial network (GAN) for remote sensing is presented, called JOA‐GAN. Firstly, a joint‐attention module (JOA) is proposed to focus the network on high‐frequency regions in remote sensing images to enhance the quality of image reconstruction. In the generator network, a multi‐scale densely connected feature extraction block (ERRDB) is proposed, which acquires features at different scales using MSconv blocks containing multi‐scale convolutions and automatically adjusts the features by JOA. In the discriminator network, the relative discriminator is used to compute the relative probability instead of the absolute probability, which helps the network learn clearer and more realistic texture details. JOA‐GAN is compared with other advanced methods, and the results demonstrate that JOA‐GAN has improved objective evaluation metrics and achieved superior visual effects. Zijun Gao, Zhankui Song |
IET Image Process. | 1 |
| 2024 | DAF-Retinex: Preserve the image detailed features and restore the reflected imageabstractAbstract Currently, deep learning methods for low‐light image enhancement tasks mainly focus on the illumination of images, while neglecting the problems of image noise and feature loss. To address this issue, this paper proposes a novel low‐light image enhancement network called DAF‐Retinex, based on the Retinex‐Net. To address the issue of image noise, different from traditional image denoise methods, this paper utilizes a fully convolutional neural network to denoise the reflection component, additionally, a denoising loss function is introduced to suppress noise. For preserving image details and extracting features, this paper creatively introduces self‐calibrated convolutions into low‐light image enhancement tasks, furthermore, a feature augmented attention block consisting of feature‐guided attention (FGA) is designed for feature learning to effectively enhance image illumination and extract image detail features. Experimental results demonstrate that the proposed algorithm in this paper effectively removes image noise and extracts detailed features, resulting in visually improved outcomes. On public datasets, the average improvement in objective evaluation metrics of image quality such as PSNR, SSIM, and NIQE are 1.13%, 4.12%, and 1.28%, respectively. Shiyu Huang 0005, Zijun Gao |
IET Image Process. | 2 |
| 2023 | Long-Tail Cross Modal HashingabstractExisting Cross Modal Hashing (CMH) methods are mainly designed for balanced data, while imbalanced data with long-tail distribution is more general in real-world. Several long-tail hashing methods have been proposed but they can not adapt for multi-modal data, due to the complex interplay between labels and individuality and commonality information of multi-modal data. Furthermore, CMH methods mostly mine the commonality of multi-modal data to learn hash codes, which may override tail labels encoded by the individuality of respective modalities. In this paper, we propose LtCMH (Long-tail CMH) to handle imbalanced multi-modal data. LtCMH firstly adopts auto-encoders to mine the individuality and commonality of different modalities by minimizing the dependency between the individuality of respective modalities and by enhancing the commonality of these modalities. Then it dynamically combines the individuality and commonality with direct features extracted from respective modalities to create meta features that enrich the representation of tail labels, and binaries meta features to generate hash codes. LtCMH significantly outperforms state-of-the-art baselines on long-tail datasets and holds a better (or comparable) performance on datasets with balanced labels. Zijun Gao, Jun Wang 0035, Guoxian Yu, Zhongmin Yan, Carlotta Domeniconi, Jinglin Zhang 0001 |
AAAI | 1 |
| 2022 | LinCDE: Conditional Density Estimation via Lindsey's MethodabstractConditional density estimation is a fundamental problem in statistics, with scientific and practical applications in biology, economics, finance and environmental studies, to name a few. In this paper, we propose a conditional density estimator based on gradient boosting and Lindsey's method (LinCDE). LinCDE admits flexible modeling of the density family and can capture distributional characteristics like modality and shape. In particular, when suitably parametrized, LinCDE will produce smooth and non-negative density estimates. Furthermore, like boosted regression trees, LinCDE does automatic feature selection. We demonstrate LinCDE's efficacy through extensive simulations and three real data examples. Zijun Gao, Trevor J. Hastie |
J. Mach. Learn. Res. | 1 |
| 2021 | Motion-based camera localization system in colonoscopy videosabstractOptical colonoscopy is an essential diagnostic and prognostic tool for many gastrointestinal diseases, including cancer screening and staging, intestinal bleeding, diarrhea, abdominal symptom evaluation, and inflammatory bowel disease assessment. However, the evaluation, classification, and quantification of findings from colonoscopy are subject to inter-observer variation. Automated assessment of colonoscopy is of interest considering the subjectivity present in qualitative human interpretations of colonoscopy findings. Localization of the camera is essential to interpreting the meaning and context of findings for diseases evaluated by colonoscopy. In this study, we propose a camera localization system to estimate the relative location of the camera and classify the colon into anatomical segments. The camera localization system begins with non-informative frame detection and removal. Then a self-training end-to-end convolutional neural network is built to estimate the camera motion, where several strategies are proposed to improve its robustness and generalization on endoscopic videos. Using the estimated camera motion a camera trajectory can be derived and a relative location index calculated. Based on the estimated location index, anatomical colon segment classification is performed by constructing a colon template. The proposed motion estimation algorithm was evaluated on an external dataset containing the ground truth for camera pose. The experimental results show that the performance of the proposed method is superior to other published methods. The relative location index estimation and anatomical region classification were further validated using colonoscopy videos collected from routine clinical practice. This validation yielded an average accuracy in classification of 0.754, which is substantially higher than the performances obtained using location indices built from other methods. Heming Yao, Ryan W. Stidham, Zijun Gao, Jonathan Gryak, Kayvan Najarian |
Medical Image Anal. | 3 |
| 2020 | Minimax Optimal Nonparametric Estimation of Heterogeneous Treatment EffectsabstractA central goal of causal inference is to detect and estimate the treatment effects of a given treatment or intervention on an outcome variable of interest, where a member known as the heterogeneous treatment effect (HTE) is of growing popularity in recent practical applications such as the personalized medicine. In this paper, we model the HTE as a smooth nonparametric difference between two less smooth baseline functions, and determine the tight statistical limits of the nonparametric HTE estimation as a function of the covariate geometry. In particular, a two-stage nearest-neighbor-based estimator throwing away observations with poor matching quality is near minimax optimal. We also establish the tight dependence on the density ratio without the usual assumption that the covariate densities are bounded away from zero, where a key step is to employ a novel maximal inequality which could be of independent interest. Zijun Gao, Yanjun Han |
NeurIPS | 1 |
| 2019 | Batched Multi-armed Bandits ProblemabstractIn this paper, we study the multi-armed bandit problem in the batched setting where the employed policy must split data into a small number of batches. While the minimax regret for the two-armed stochastic bandits has been completely characterized in \cite{perchet2016batched}, the effect of the number of arms on the regret for the multi-armed case is still open. Moreover, the question whether adaptively chosen batch sizes will help to reduce the regret also remains underexplored. In this paper, we propose the BaSE (batched successive elimination) policy to achieve the rate-optimal regrets (within logarithmic factors) for batched multi-armed bandits, with matching lower bounds even if the batch sizes are determined in an adaptive manner. Zijun Gao, Yanjun Han, Zhimei Ren, Zhengqing Zhou |
NeurIPS | 1 |