Haoyue Bai 0001

dblp:150/3371-1 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-8139-0431ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 OoDBench+: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization
abstract
Deep learning has demonstrated remarkable generalization capability with independent and identically distributed (i.i.d.) training and test data, however, it often struggles with data drawn from different, albeit causally related, distributions. This problem is generally known as Out-of-Distribution (OoD) generalization. While there is a plethora of algorithms proposed for OoD generalization, the current understanding of the data commonly employed to evaluate these algorithms remains relatively naive. In this study, we identify two distinct types of distribution shifts, namely diversity shift and correlation shift, that are ubiquitous in various OoD datasets. We propose a quantifiable formal definition for the two shifts and show that the performance of OoD algorithms is upper bounded by them. To validate our theoretical insight, we evaluate a number of OoD generalization algorithms across two groups of datasets from both classification and object detection areas, each dominated by one of the shifts, exposing the strengths of the algorithms against one shift as well as their limitations against the other. We further proved that all performance degradations according to data distribution shifts can be attributed to these two types of shifts defined in our paper. The benchmark integrates existing datasets and algorithms from different research areas that seem unrelated into a coherent picture, which may serve as a foundation for future OoD generalization research.
Nanyang Ye 0001, Kaican Li, Fan Wu 0006, Jundong Zhou, Haoyue Bai 0001, Runpeng Yu 0001, Lanqing Hong, Fengwei Zhou, Zhenguo Li, Jun Zhu 0001, Xinbing Wang, Chenghu Zhou
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content
abstract
The recent proliferation of photorealistic images created by generative models has sparked both excitement and concern, as these images are increasingly indistinguishable from real ones to the human eye. While offering new creative and commercial possibilities, the potential for misuse, such as in misinformation and fraud, highlights the need for effective detection methods. Current detection approaches often rely on access to model weights or require extensive collections of real image datasets, limiting their scalability and practical application in real-world scenarios. In this work, we introduce a novel black-box detection framework that requires only API access, sidestepping the need for model weights or large auxiliary datasets. Our approach leverages a corrupt-and-recover strategy: by masking part of an image and assessing the model’s ability to reconstruct it, we measure the likelihood that the image was generated by the model itself. For black-box models that do not support masked-image inputs, we incorporate a cost-efficient surrogate model trained to align with the target model’s distribution, enhancing detection capability. Our framework demonstrates strong performance, outperforming baseline methods by 4.31% in mean average precision across eight diffusion model variant datasets.
Haoyue Bai 0001, Yiyou Sun, Wei Cheng 0002
CVPR1
2024 HYPO: Hyperspherical Out-Of-Distribution Generalization
abstract
Out-of-distribution (OOD) generalization is critical for machine learning models deployed in the real world. However, achieving this can be fundamentally challenging, as it requires the ability to learn invariant features across different domains or environments. In this paper, we propose a novel framework HYPO (HYPerspherical OOD generalization) that provably learns domain-invariant representations in a hyperspherical space. In particular, our hyperspherical learning algorithm is guided by intra-class variation and inter-class separation principles—ensuring that features from the same class (across different training domains) are closely aligned with their class prototypes, while different class prototypes are maximally separated. We further provide theoretical justifications on how our prototypical learning objective improves the OOD generalization bound. Through extensive experiments on challenging OOD benchmarks, we demonstrate that our approach outperforms competitive baselines and achieves superior performance. Code is available at https://github.com/deeplearning-wisc/hypo.
Haoyue Bai 0001, Yifei Ming, Julian Katz-Samuels, Yixuan Li 0001
ICLR1
2024 AHA: Human-Assisted Out-of-Distribution Generalization and Detection
abstract
Modern machine learning models deployed often encounter distribution shifts in real-world applications, manifesting as covariate or semantic out-of-distribution (OOD) shifts. These shifts give rise to challenges in OOD generalization and OOD detection. This paper introduces a novel, integrated approach AHA (Adaptive Human-Assisted OOD learning) to simultaneously address both OOD generalization and detection through a human-assisted framework by labeling data in the wild. Our approach strategically labels examples within a novel maximum disambiguation region, where the number of semantic and covariate OOD data roughly equalizes. By labeling within this region, we can maximally disambiguate the two types of OOD data, thereby maximizing the utility of the fixed labeling budget. Our algorithm first utilizes a noisy binary search algorithm that identifies the maximal disambiguation region with high probability. The algorithm then continues with annotating inside the identified labeling region, reaping the full benefit of human feedback. Extensive experiments validate the efficacy of our framework. We observed that with only a few hundred human annotations, our method significantly outperforms existing state-of-the-art methods that do not involve human assistance, in both OOD generalization and OOD detection.
Haoyue Bai 0001, Jifan Zhang, Robert D. Nowak
NeurIPS1
2023 Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and Detection
abstract
Modern machine learning models deployed in the wild can encounter both covariate and semantic shifts, giving rise to the problems of out-of-distribution (OOD) generalization and OOD detection respectively. While both problems have received significant research attention lately, they have been pursued independently. This may not be surprising, since the two tasks have seemingly conflicting goals. This paper provides a new unified approach that is capable of simultaneously generalizing to covariate shifts while robustly detecting semantic shifts. We propose a margin-based learning framework that exploits freely available unlabeled data in the wild that captures the environmental test-time OOD distributions under both covariate and semantic shifts. We show both empirically and theoretically that the proposed margin constraint is the key to achieving both OOD generalization and detection. Extensive experiments show the superiority of our framework, outperforming competitive baselines that specialize in either OOD generalization or OOD detection. Code is publicly available at https://github.com/deeplearning-wisc/scone.
Haoyue Bai 0001, Gregory Canal, Xuefeng Du, Jeongyeol Kwon, Robert D. Nowak, Yixuan Li 0001
ICML1
2022 OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization
abstract
Deep learning has achieved tremendous success with independent and identically distributed (i. i.d.) data. However, the performance of neural networks often degenerates drastically when encountering out-of-distribution (OoD) data, i.e., when training and test data are sampled from different distributions. While a plethora of algorithms have been proposed for OoD generalization, our understanding of the data used to train and evaluate these algorithms remains stagnant. In this work, we first identify and measure two distinct kinds of distribution shifts that are ubiquitous in various datasets. Next, through extensive experiments, we compare OoD generalization algorithms across two groups of benchmarks, each dominated by one of the distribution shifts, revealing their strengths on one shift as well as limitations on the other shift. Overall, we position existing datasets and algorithms from different research areas seemingly unconnected into the same coherent picture. It may serve as a foothold that can be resorted to by future OoD generalization research. Our code is available at https://github.com/ynysjtulood_bench.
Nanyang Ye 0001, Kaican Li, Haoyue Bai 0001, Runpeng Yu 0001, Lanqing Hong, Fengwei Zhou, Zhenguo Li, Jun Zhu 0001
CVPR3
2022 A survey on deep learning-based single image crowd counting: Network design, loss function and supervisory signal
Haoyue Bai 0001, Jiageng Mao, Shueng-Han Gary Chan
Neurocomputing1
2021 DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic Augmentation
abstract
While deep learning demonstrates its strong ability to handle independent and identically distributed (IID) data, it often suffers from out-of-distribution (OoD) generalization, where the test data come from another distribution (w.r.t. the training one). Designing a general OoD generalization framework for a wide range of applications is challenging, mainly due to different kinds of distribution shifts in the real world, such as the shift across domains or the extrapolation of correlation. Most of the previous approaches can only solve one specific distribution shift, leading to unsatisfactory performance when applied to various OoD benchmarks. In this work, we propose DecAug, a novel decomposed feature representation and semantic augmentation approach for OoD generalization. Specifically, DecAug disentangles the category-related and context-related features by orthogonalizing the two gradients (w.r.t. intermediate features) of losses for predicting category and context labels, where category-related features contain causal information of the target object, while context-related features cause distribution shifts between training and test data. Furthermore, we perform gradient-based augmentation on context-related features to improve the robustness of learned representations. Experimental results show that DecAug outperforms other state-of-the-art methods on various OoD datasets, which is among the very few methods that can deal with different types of OoD generalization challenges.
Haoyue Bai 0001, Lanqing Hong, Fengwei Zhou, Nanyang Ye 0001, Han-Jia Ye, Shueng-Han Gary Chan, Zhenguo Li
AAAI1
2021 NAS-OoD: Neural Architecture Search for Out-of-Distribution Generalization
abstract
Recent advances on Out-of-Distribution (OoD) generalization reveal the robustness of deep learning models against distribution shifts. However, existing works focus on OoD algorithms, such as invariant risk minimization, domain generalization, or stable learning, without considering the influence of deep model architectures on OoD generalization, which may lead to sub-optimal performance. Neural Architecture Search (NAS) methods search for architecture based on its performance on the training data, which may result in poor generalization for OoD tasks. In this work, we propose robust Neural Architecture Search for OoD generalization (NAS-OoD), which optimizes the architecture with respect to its performance on generated OoD data by gradient descent. Specifically, a data generator is learned to synthesize OoD data by maximizing losses computed by different neural architectures, while the goal for architecture search is to find the optimal architecture parameters that minimize the synthetic OoD data losses. The data generator and the neural architecture are jointly optimized in an end-to-end manner, and the minimax training process effectively discovers robust architectures that generalize well for different distribution shifts. Extensive experimental results show that NAS-OoD achieves superior performance on various OoD generalization benchmarks with deep models having a much fewer number of parameters. In addition, on a real industry dataset, the proposed NAS-OoD method reduces the error rate by more than 70% compared with the state-of-the-art method, demonstrating the proposed method’s practicality for real applications.
Haoyue Bai 0001, Fengwei Zhou, Lanqing Hong, Nanyang Ye 0001, Shueng-Han Gary Chan, Zhenguo Li
ICCV1
2021 Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection
abstract
We present a flexible and high-performance framework, named Pyramid R-CNN, for two-stage 3D object detection from point clouds. Current approaches generally rely on the points or voxels of interest for RoI feature extraction on the second stage, but cannot effectively handle the sparsity and non-uniform distribution of those points, and this may result in failures in detecting objects that are far away. To resolve the problems, we propose a novel second-stage module, named pyramid RoI head, to adaptively learn the features from the sparse points of interest. The pyramid RoI head consists of three key components. Firstly, we propose the RoI-grid Pyramid, which mitigates the sparsity problem by extensively collecting points of interest for each RoI in a pyramid manner. Secondly, we propose RoI-grid Attention, a new operation that can encode richer information from sparse points by incorporating conventional attention-based and graph-based point operators into a unified formulation. Thirdly, we propose the Density-Aware Radius Prediction (DARP) module, which can adapt to different point density levels by dynamically adjusting the focusing range of RoIs. Combining the three components, our pyramid RoI head is robust to the sparse and imbalanced circumstances, and can be applied upon various 3D backbones to consistently boost the detection performance. Extensive experiments show that Pyramid R-CNN outperforms the state-of-the-art 3D detection models by a large margin on both the KITTI dataset and the Waymo Open dataset.
Jiageng Mao, Minzhe Niu, Haoyue Bai 0001, Xiaodan Liang, Hang Xu 0004, Chunjing Xu
ICCV3
2021 Voxel Transformer for 3D Object Detection
abstract
We present Voxel Transformer (VoTr), a novel and effective voxel-based Transformer backbone for 3D object detection from point clouds. Conventional 3D convolutional backbones in voxel-based 3D detectors cannot efficiently capture large context information, which is crucial for object recognition and localization, owing to the limited receptive fields. In this paper, we resolve the problem by introducing a Transformer-based architecture that enables long-range relationships between voxels by self-attention. Given the fact that non-empty voxels are naturally sparse but numerous, directly applying standard Transformer on voxels is non-trivial. To this end, we propose the sparse voxel module and the submanifold voxel module, which can operate on the empty and non-empty voxel positions effectively. To further enlarge the attention range while maintaining comparable computational overhead to the convolutional counterparts, we propose two attention mechanisms for multi-head attention in those two modules: Local Attention and Dilated Attention, and we further propose Fast Voxel Query to accelerate the querying process in multi-head attention. VoTr contains a series of sparse and submanifold voxel modules, and can be applied in most voxel-based detectors. Our proposed VoTr shows consistent improvement over the convolutional baselines while maintaining computational efficiency on the KITTI dataset and the Waymo Open dataset.
Jiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 0001, Jiashi Feng, Xiaodan Liang, Hang Xu 0004, Chunjing Xu
ICCV4
2021 Designing a Partially Understandable Neural Network through Semantic Embedding
abstract
In recent years, the convolution neural network (CNN) has been successfully applied in numerous fields as a machine learning model. However, such neural models are still considered to be “black box” for most tasks. The fundamental issue underlying this problem is that the information and knowledge learned by a neural network during the training process are unknown and unpredictable. In this study, we attempted to design a partially understandable neural network through semantic embedding. Firstly, we selected several understandable feature extraction operators as expected information. Secondly, we embedded these operators into the hierarchical layers of a neural network at the beginning of its training process. Finally, these embedded operators were only involved in forward calculation and remained unchanged in the error back-propagation of the training process. We applied our method to the ResNet and DenseNet models in image classification tasks. In the experiments, our new models achieved almost the same performance as the original ones, but decreased performance was exhibited when shutting the embedded parts, and the features in the embedded parts were understandable. The experiments verified that the embedded parts not only contributed to the classification tasks, but also caused the neural network model to be partially understandable.
Chengfu Tang, Dawei Dai, Haoyue Bai 0001, Guoyin Wang 0001, Shuyin Xia, Feng Hu 0001
IJCNN3