EDBT 2026 Demo / reviewers in the wild / expert
Nanyang Ye 0001
dblp:175/2581-1
· DBLP profile ↗
44ranked-venue papers
16as first author
35since 2021 · last 2026
0000-0003-3129-3953ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 11 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 first-author · 15 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Proactive Risk-Awareness of Autonomous Driving via Trajectory Monitoring
Xianfei Li, Nanyang Ye 0001 |
ICPR (4) | 4 |
| 2026 | OoDBench+: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationabstractDeep learning has demonstrated remarkable generalization capability with independent and identically distributed (i.i.d.) training and test data, however, it often struggles with data drawn from different, albeit causally related, distributions. This problem is generally known as Out-of-Distribution (OoD) generalization. While there is a plethora of algorithms proposed for OoD generalization, the current understanding of the data commonly employed to evaluate these algorithms remains relatively naive. In this study, we identify two distinct types of distribution shifts, namely diversity shift and correlation shift, that are ubiquitous in various OoD datasets. We propose a quantifiable formal definition for the two shifts and show that the performance of OoD algorithms is upper bounded by them. To validate our theoretical insight, we evaluate a number of OoD generalization algorithms across two groups of datasets from both classification and object detection areas, each dominated by one of the shifts, exposing the strengths of the algorithms against one shift as well as their limitations against the other. We further proved that all performance degradations according to data distribution shifts can be attributed to these two types of shifts defined in our paper. The benchmark integrates existing datasets and algorithms from different research areas that seem unrelated into a coherent picture, which may serve as a foundation for future OoD generalization research. Nanyang Ye 0001, Kaican Li, Fan Wu 0006, Jundong Zhou, Haoyue Bai 0001, Runpeng Yu 0001, Lanqing Hong, Fengwei Zhou, Zhenguo Li, Jun Zhu 0001, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | CFSM: A Novel Causal Feature Selection Module for Two-Dimensional Out-of-Distribution GeneralizationabstractIn real-world scenarios, training and test data are often collected in diverse settings, leading to domain shifts arising from evolving environments and selection bias. While causality-inspired methods have shown promising results in tackling the out-of-distribution (OOD) generalization issue, prior methods treat the discovered differences across domains as confounding variables. While effective in handling domain differences (i.e., unseen environmental features in test data), they may fail when confronted with intricate spurious correlations in real-world datasets. In this study, we first analyze this limitation to inadequate modeling of causal intervention and derive the OOD generalization bound to explain the challenges it introduces. To address this problem, we propose a modified causal intervention approach to mitigate various types of confounders. Motivated by the mathematical formulation of our modified causal intervention, we introduce the Causal Feature Selection Module (CFSM) to suppress model weights on both domain-differences features and spurious correlation features. Integrated within the Base Feature Extraction Module, In-Sample Module, and Cross-Sample Module (B-I-C architecture), CFSM collectively neutralizes the confounding effects arising from both domain discrepancies and correlation distinctions, thereby achieving causal feature selection. Under mild assumptions, we prove that the proposed CFSM method can achieve strictly lower OOD errors. Further experiments conducted on various benchmark datasets demonstrate the effectiveness of the proposed method. Compared to previous deconfounding methods, our method not only mitigates the effect of domain-differences features but also the hard-to-identify spurious correlation features, achieving significant improvements in two-dimensional OOD generalization. Weihan Yin, Yiyao Yang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven FrameworkabstractThis paper explores the application of large language models (LLMs) in nursing and elderly care, focusing on AI-driven patient monitoring and interaction. We introduce a novel Chinese nursing dataset and implement incremental pre-training (IPT) and supervised fine-tuning (SFT) techniques to enhance LLM performance in specialized tasks. Using LangChain, we develop an interactable nursing assistant capable of real-time care and personalized interventions. Experimental results demonstrate significant improvements, paving the way for AI-driven solutions to meet the growing demands of healthcare in aging populations. Qiao Sun 0003, Jiexin Xie, Nanyang Ye 0001, Qinying Gu, Shijie Guo |
COLING | 3 |
| 2025 | Decision SpikeFormer: Spike-Driven Transformer for Decision MakingabstractOffline reinforcement learning (RL) enables policy training solely on pre-collected data, avoiding direct environment interaction—a crucial benefit for energy-constrained embodied AI applications. Although Artificial Neural Networks (ANN)-based methods perform well in offline RL, their high computational and energy demands motivate exploration of more efficient alternatives. Spiking Neural Networks (SNNs) show promise for such tasks, given their low power consumption. In this work, we introduce DSFormer, the first spike-driven transformer model designed to tackle offline RL via sequence modeling. Unlike existing SNN transformers focused on spatial dimensions for vision tasks, we develop Temporal Spiking Self-Attention (TSSA) and Positional Spiking Self-Attention (PSSA) in DSFormer to capture the temporal and positional dependencies essential for sequence modeling in RL. Additionally, we propose Progressive Threshold-dependent Batch Normalization (PTBN), which combines the benefits of LayerNorm and BatchNorm to preserve temporal dependencies while maintaining the spiking nature of SNNs. Comprehensive results in the D4RL benchmark show DSFormer’s superiority over both SNN and ANN counterparts, achieving 78.4% energy savings, highlighting DSFormer’s advantages not only in energy efficiency but also in competitive performance. Code and models are public at project page. Qinying Gu, Nanyang Ye 0001 |
CVPR | 3 |
| 2025 | OODD: Test-time Out-of-Distribution Detection with Dynamic DictionaryabstractOut-of-distribution (OOD) detection remains challenging for deep learning models, particularly when test-time OOD samples differ significantly from training outliers. We propose OODD, a novel test-time OOD detection method that dynamically maintains and updates an OOD dictionary without fine-tuning. Our approach leverages a priority queue-based dictionary that accumulates representative OOD features during testing, combined with an informative inlier sampling strategy for in-distribution (ID) samples. To ensure stable performance during early testing, we propose a dual OOD stabilization mechanism that leverages strategically generated outliers derived from ID data. To our best knowledge, extensive experiments on the OpenOOD benchmark demonstrate that OODD significantly outperforms existing methods, achieving a 26.0% improvement in FPR95 on CIFAR-100 Far OOD detection compared to the state-of-the-art approach. Furthermore, we present an optimized variant of the KNN-based OOD detection framework that achieves a 3x speedup while maintaining detection performance. Our code is available at https://github.com/zxk1212/OODD. Zewen Sun, Hengyu Liu 0007, Qinying Gu, Nanyang Ye 0001 |
CVPR | 6 |
| 2025 | Conformalized Causal Learning for Uncertainty-Aware Mineral Prospectivity Mapping
Evelyn Jessica Jaya, Qinying Gu, Xinbing Wang, Nanyang Ye 0001 |
ICANN (4) | 4 |
| 2025 | ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian SplattingabstractAs 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. While steganography has advanced significantly in common 3D media like meshes and Neural Radiance Fields (NeRF), research into steganography for 3D- GS representations remains largely unexplored. To address this gap, we propose ConcealGS, a novel 3D steganography method that embeds implicit information into the explicit 3D representation of Gaussian Splatting. By introducing a consistency strategy for the decoder and a gradient optimization approach, ConcealGS overcomes limitations of NeRF-based models, enhancing both the robustness of implicit information and the quality of 3D reconstruction. Extensive evaluations across various potential application scenarios demonstrate that ConcealGS successfully recovers implicit information with negligible impact on rendering quality, offering a groundbreaking approach for embedding invisible yet recoverable information into 3D models. This work paves the way for advanced copyright protection and secure data transmission in the evolving landscape of 3D content creation and distribution. Code is available at https://github.com/zxk1212/ConcealGS. Hengyu Liu 0007, Chenxin Li, Yining Sun, Wuyang Li, Yifan Liu 0010, Yiyang Lin, Yixuan Yuan, Nanyang Ye 0001 |
ICASSP | 9 |
| 2025 | Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion ModelsabstractGiven a style-reference image as the additional image condition, text-to-image diffusion models have demonstrated impressive capabilities in generating images that possess the content of text prompts while adopting the visual style of the reference image. However, current state-of-the-art methods often struggle to disentangle content and style from style-reference images, leading to issues such as content leakages. To address this issue, we propose a masking-based method that efficiently decouples content from style without the need of tuning any model parameters. By simply masking specific elements in the style reference's image features, we uncover a critical yet under-explored principle: guiding with appropriately-selected fewer conditions (e.g., dropping several image feature elements) can efficiently avoid unwanted content flowing into the diffusion models, enhancing the style transfer performances of text-to-image diffusion models. In this paper, we validate this finding both theoretically and experimentally. Extensive experiments across various styles demonstrate the effectiveness of our masking-based method and support our theoretical results. Xinbing Wang, Chenghu Zhou, Qinying Gu, Nanyang Ye 0001 |
ICLR | 5 |
| 2025 | Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View LearningabstractImproving the generalization of multi-camera 3D object detection is essential for safe autonomous driving in the real world. In this paper, we consider a realistic yet more challenging scenario, which aims to improve the generalization when only single source data available for training, as gathering diverse domains of data and collecting annotations is time-consuming and labor-intensive. To this end, we propose the Fourier Cross-View Learning (FCVL) framework including Fourier Hierarchical Augmentation (FHiAug), an augmentation strategy in the frequency domain to boost domain diversity, and Fourier Cross-View Semantic Consistency Loss to facilitate the model to learn more domain-invariant features from adjacent perspectives. Furthermore, we provide theoretical guarantees via augmentation graph theory. To the best of our knowledge, this is the first study to explore generalizable multi-camera 3D object detection with a single source. Extensive experiments on various testing domains have demonstrated that our approach achieves the best performance across various domain generalization methods. Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
ICML | 5 |
| 2025 | Tri-AutoAug: Single Domain Generalization for Bird's-Eye-View 3D Object Detection Through Pixel-2D-3D FeaturesabstractWith the increasing popularity of autonomous driving based on the Bird's-Eye-View (BEV) representation, improving the generalization of such detection models is key for safe real-world applications. However, a realistic yet challenging scenario: Single Domain Generalization (SDG) for BEV, is still under-explored. A key ingredient for SDG is to increase data diversity via common image augmentation or adversarial data generation first. However, common image-level augmentation is not sufficient enough to ensure domain diversity in most part of latent space. The adversarial generation has the problem of unstable training or mode collapsing as well. To address these limitations, we present Tri-level Automatic Augmentation (Tri-AutoAug), a simple yet effective method to enlarge the diversity and quantity of data from image and 2D features and facilitate the model to learn more domain-invariant features in BEV space. Besides, Tri-AutoAug can automatically learn augmentation strategies to avoid spending too much time manually adjusting hyperparameters and maximize the benefit of Tri-level Augmentation. To the best of our knowledge, this is the first study to explore automatic augmentation for SDG BEV. Extensive experiments on NuScenes-C including eight testing domains have demonstrated that our approach can achieve the best performance across various domain generalization methods. More importantly, we evaluate the proposed method in real-world autonomous driving scenarios. Tri-AutoAug improves the out-of-distribution (ood) performance by 8.54% (mAP), which demonstrates that Tri-AutoAug provides a practical and feasible solution for the applications of 3D detectors in the real world. The code is available at https://github.com/ClaireTunlTri-AutoAug. Xianfei Li, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
ICRA | 6 |
| 2025 | Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningabstractOffline reinforcement learning (offline RL) is increasingly approached as a sequence modeling task, with methods leveraging advanced architectures like Transformers to capture trajectory dependencies. Despite significant progress, the mechanisms underlying their effectiveness and limitations remain insufficiently understood. We conduct a thorough analysis on the representative Decision Transformer (DT) model using an entropy analysis and identify the inconsistencies in state-action-reward ($\langle s, a, R \rangle$) distributions causing attention ``dispersal". To address this, we propose a hierarchical framework that decomposes sequence modeling into intra-step relational modeling—handled by a Token Merger that fuses each $\langle s, a, R \rangle$ triplet—and inter-step modeling—handled by a Token Mixer across timesteps. We investigate several Token Merger designs and validate their effectiveness across various offline RL methods.
Furthermore, our theoretical analysis and experimental results suggest that while Token Mixers are important, lightweight architecture can also achieve even better performance to more complex ones. We therefore propose a parameter-free Average Pooling Token Mixer, which, combined with a convolutional Token Merger, forms our final model, Decision HiFormer (DHi). DHi achieves a \textbf{73.6\%} improvement in inference speed and an \textbf{9.3\%} gain in policy performance on the D4RL benchmark compared to DT. DHi also generalizes well to real-world robotic manipulation tasks, offering both practical benefits and insights into sequence-based policy design for offline RL. Code and models are public at \href{https://wei-nijuan.github.io/DecisionHiFormer/}{project page}. Leiyu Wang, Heyue Li, Luoyi Fan, Nanyang Ye 0001, Qinying Gu |
NeurIPS | 7 |
| 2025 | Δ Energy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
Xinbing Wang, Qinying Gu, Nanyang Ye 0001 |
NeurIPS | 5 |
| 2025 | Bayes-CAL: Robust Cross-Modal Alignment by Bayesian Approach for Few-Shot OoD Generalization
Weihan Yin, Fan Wu 0006, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | InfoBound: A Provable Information-Bounds Inspired Framework for Both OoD Generalization and OoD DetectionabstractIn real-world scenarios, distribution shifts give rise to the importance of two problems: out-of-distribution (OoD) generalization, which focuses on models' generalization ability against covariate shifts (i.e., the changes of environments), and OoD detection, which aims to be aware of semantic shifts (i.e., test-time unseen classes). Real-world testing environments often involve a combination of both covariate and semantic shifts. While numerous methods have been proposed to address these critical issues, only a few works tackled them simultaneously. Moreover, prior works often improve one problem but sacrifice the other. To overcome these limitations, we delve into boosting OoD detection and OoD generalization from the perspective of information theory, which can be easily applied to existing models and different tasks. Building upon the theoretical bounds for mutual information and conditional entropy, we provide a unified approach, composed of Mutual Information Minimization (MI-Min) and Conditional Entropy Maximizing (CE-Max). Extensive experiments and comprehensive evaluations on multi-label image classification and object detection have demonstrated the superiority of our method. It successfully mitigates trade-offs between the two challenges compared to competitive baselines. Zichao Nie, Yuan Gao 0050, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | G-NAS: Generalizable Neural Architecture Search for Single Domain Generalization Object DetectionabstractIn this paper, we focus on a realistic yet challenging task, Single Domain Generalization Object Detection (S-DGOD), where only one source domain's data can be used for training object detectors, but have to generalize multiple distinct target domains. In S-DGOD, both high-capacity fitting and generalization abilities are needed due to the task's complexity. Differentiable Neural Architecture Search (NAS) is known for its high capacity for complex data fitting and we propose to leverage Differentiable NAS to solve S-DGOD. However, it may confront severe over-fitting issues due to the feature imbalance phenomenon, where parameters optimized by gradient descent are biased to learn from the easy-to-learn features, which are usually non-causal and spuriously correlated to ground truth labels, such as the features of background in object detection data. Consequently, this leads to serious performance degradation, especially in generalizing to unseen target domains with huge domain gaps between the source domain and target domains. To address this issue, we propose the Generalizable loss (G-loss), which is an OoD-aware objective, preventing NAS from over-fitting by using gradient descent to optimize parameters not only on a subset of easy-to-learn features but also the remaining predictive features for generalization, and the overall framework is named G-NAS. Experimental results on the S-DGOD urban-scene datasets demonstrate that the proposed G-NAS achieves SOTA performance compared to baseline methods. Codes are available at https://github.com/wufan-cse/G-NAS. Fan Wu 0006, Jinling Gao, Lanqing Hong, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
AAAI | 6 |
| 2024 | Domain Invariant Learning for Gaussian Processes and Bayesian ExplorationabstractOut-of-distribution (OOD) generalization has long been a challenging problem that remains largely unsolved. Gaussian processes (GP), as popular probabilistic model classes, especially in the small data regime, presume strong OOD generalization abilities. Surprisingly, their OOD generalization abilities have been under-explored before compared with other lines of GP research. In this paper, we identify that GP is not free from the problem and propose a domain invariant learning algorithm for Gaussian processes (DIL-GP) with a min-max optimization on the likelihood. DIL-GP discovers the heterogeneity in the data and forces invariance across partitioned subsets of data. We further extend the DIL-GP to improve Bayesian optimization's adaptability on changing environments. Numerical experiments demonstrate the superiority of DIL-GP for predictions on several synthetic and real-world datasets. We further demonstrate the effectiveness of the DIL-GP Bayesian optimization method on a PID parameters tuning experiment for a quadrotor. The full version and source code are available at: https://github.com/Billzxl/DIL-GP. Xilong Zhao, Siyuan Bian, Yaoyun Zhang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
AAAI | 8 |
| 2024 | MiniConGTS: A Near Ultimate Minimalist Contrastive Grid Tagging Scheme for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims to co-extract the sentiment triplets in a given corpus.Existing approaches within the pretraining-finetuning paradigm tend to either meticulously craft complex tagging schemes and classification heads, or incorporate external semantic augmentation to enhance performance.In this study, we, for the first time, re-evaluate the redundancy in tagging schemes and the internal enhancement in pretrained representations.We propose a method to improve and utilize pretrained representations by integrating a minimalist tagging scheme and a novel token-level contrastive learning strategy.The proposed approach demonstrates comparable or superior performance compared to stateof-the-art techniques while featuring a more compact design and reduced computational overhead.Additionally, we are the first to formally evaluate GPT-4's performance in fewshot learning and Chain-of-Thought scenarios for this task.The results demonstrate that the pretraining-finetuning paradigm remains highly effective even in the era of large language models. Qiao Sun 0003, Liujia Yang, Minghao Ma, Nanyang Ye 0001, Qinying Gu |
EMNLP | 4 |
| 2024 | CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD DetectionabstractRecent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair their ability to handle distribution shifts. In real-world scenarios, machine learning systems inevitably encounter both covariate shifts (e.g., changes in image styles) and semantic shifts (e.g., test-time unseen classes). This highlights the importance of enhancing out-of-distribution (OOD) generalization on covariate shifts and simultaneously detecting semantic-shifted unseen classes. Thus a critical but underexplored question arises: How to improve VL-PTMs’ generalization ability to closed-set OOD data, while effectively detecting open-set unseen classes during fine-tuning? In this paper, we propose a novel objective function of OOD detection that also serves to improve OOD generalization. We show that minimizing the gradient magnitude of energy scores on training data leads to domain-consistent Hessians of classification loss, a strong indicator for OOD generalization revealed by theoretical analysis. Based on this finding, we have developed a unified fine-tuning framework that allows for concurrent optimization of both tasks. Extensive experiments have demonstrated the superiority of our method. The code is available at https://github.com/LinLLLL/CRoFT. Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
ICML | 6 |
| 2024 | Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization
Weihan Yin, Yiyao Yang, Fan Wu 0006, Zhaoyu Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
Int. J. Comput. Vis. | 9 |
| 2024 | OoD-Control: Generalizing Control in Unseen EnvironmentsabstractGeneralizing out-of-distribution (OoD) is critical but challenging in real applications such as unmanned aerial vehicle (UAV) flight control. Previous machine learning-based control has shown promise in dealing with complex real-world environments but suffers huge performance degradation facing OoD scenarios, posing risks to the stability and safety of UAVs. In this paper, we found that the introduced random noises during training surprisingly yield theoretically guaranteed performances via a proposed functional optimization framework. More encouragingly, this framework does not involve common Lyapunov assumptions used in this field, making it more widely applicable. With this framework, the upperbound for control error is induced. We also proved that the induced random noises can lead to lower OoD control errors. Based on our theoretical analysis, we further propose OoD-Control to generalize control in unseen environments. Numerical experiments demonstrate the superiority of the proposed algorithm, surpassing previous state-of-the-art by 65% under challenging unseen environments. We further extend to outdoor real-world experiments and found that the control error is reduced by 50% approximately. Nanyang Ye 0001, Zhaoyu Zeng, Jundong Zhou, Yuxiao Duan, Haoqi Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Certifiable Out-of-Distribution GeneralizationabstractMachine learning methods suffer from test-time performance degeneration when faced with out-of-distribution (OoD) data whose distribution is not necessarily the same as training data distribution. Although a plethora of algorithms have been proposed to mitigate this issue, it has been demonstrated that achieving better performance than ERM simultaneously on different types of distributional shift datasets is challenging for existing approaches. Besides, it is unknown how and to what extent these methods work on any OoD datum without theoretical guarantees. In this paper, we propose a certifiable out-of-distribution generalization method that provides provable OoD generalization performance guarantees via a functional optimization framework leveraging random distributions and max-margin learning for each input datum. With this approach, the proposed algorithmic scheme can provide certified accuracy for each input datum's prediction on the semantic space and achieves better performance simultaneously on OoD datasets dominated by correlation shifts or diversity shifts. Our code is available at https://github.com/ZlatanWilliams/StochasticDisturbanceLearning. Nanyang Ye 0001, Jia Wang 0031, Zhaoyu Zeng, Jiayao Shao, Chensheng Peng, Bikang Pan, Kaican Li, Jun Zhu 0001 |
AAAI | 1 |
| 2023 | Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution GeneralizationabstractRecent advances in large pre-trained models showed promising results in few-shot learning. However, their generalization ability on two-dimensional Out-of-Distribution (OoD) data, i.e., correlation shift and diversity shift, has not been thoroughly investigated. Researches have shown that even with a significant amount of training data, few methods can achieve better performance than the standard empirical risk minimization method (ERM) in OoD generalization. This few-shot OoD generalization dilemma emerges as a challenging direction in deep neural network generalization research, where the performance suffers from overfitting on few-shot examples and OoD generalization errors. In this paper, leveraging a broader supervision source, we explore a novel Bayesian cross-modal image-text alignment learning method (Bayes-CAL) to address this issue. Specifically, the model is designed as only text representations are fine-tuned via a Bayesian modelling approach with gradient orthogonalization loss and invariant risk minimization (IRM) loss. The Bayesian approach is essentially introduced to avoid overfitting the base classes observed during training and improve generalization to broader unseen classes. The dedicated loss is introduced to achieve better image-text alignment by disentangling the causal and non-casual parts of image features. Numerical experiments demonstrate that Bayes-CAL achieved state-of-the-art OoD generalization performances on two-dimensional distribution shifts. Moreover, compared with CLIP-like models, Bayes-CAL yields more stable generalization performances on unseen classes. Our code is available at https://github.com/LinLLLL/BayesCAL. Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
AAAI | 4 |
| 2023 | An Annealing Mechanism for Adversarial Training AccelerationabstractDespite the empirical success in various domains, it has been revealed that deep neural networks are vulnerable to maliciously perturbed input data that can dramatically degrade their performance. These are known as adversarial attacks. To counter adversarial attacks, adversarial training formulated as a form of robust optimization has been demonstrated to be effective. However, conducting adversarial training brings much computational overhead compared with standard training. In order to reduce the computational cost, we propose an annealing mechanism, annealing mechanism for adversarial training acceleration (Amata), to reduce the overhead associated with adversarial training. The proposed Amata is provably convergent, well-motivated from the lens of optimal control theory, and can be combined with existing acceleration methods to further enhance performance. It is demonstrated that, on standard datasets, Amata can achieve similar or better robustness with around 1/3-1/2 the computational time compared with traditional methods. In addition, Amata can be incorporated into other adversarial training acceleration algorithms (e.g., YOPO, Free, Fast, and ATTA), which leads to a further reduction in computational time on large-scale problems. Nanyang Ye 0001, Qianxiao Li, Xiaoyun Zhou 0001, Zhanxing Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | OoDHDR-Codec: Out-of-Distribution Generalization for HDR Image CompressionabstractRecently, deep learning has been proven to be a promising approach in standard dynamic range (SDR) image compression. However, due to the wide luminance distribution of high dynamic range (HDR) images and the lack of large standard datasets, developing a deep model for HDR image compression is much more challenging. To tackle this issue, we view HDR data as distributional shifts of SDR data and the HDR image compression can be modeled as an out-of-distribution generalization (OoD) problem. Herein, we propose a novel out-of-distribution (OoD) HDR image compression framework (OoDHDR-codec). It learns the general representation across HDR and SDR environments, and allows the model to be trained effectively using a large set of SDR datases supplemented with much fewer HDR samples. Specifically, OoDHDR-codec consists of two branches to process the data from two environments. The SDR branch is a standard blackbox network. For the HDR branch, we develop a hybrid system that models luminance masking and tone mapping with white-box modules and performs content compression with black-box neural networks. To improve the generalization from SDR training data on HDR data, we introduce an invariance regularization term to learn the common representation for both SDR and HDR compression. Extensive experimental results show that the OoDHDR codec achieves strong competitive in-distribution performance and state-of-the-art OoD performance. To the best of our knowledge, our proposed approach is the first work to model HDR compression as OoD generalization problems and our OoD generalization algorithmic framework can be applied to any deep compression model in addition to the network architectural choice demonstrated in the paper. Code available at https://github.com/caolinfeng/OoDHDR-codec. Linfeng Cao, Aofan Jiang, Huaying Wu, Nanyang Ye 0001 |
AAAI | 5 |
| 2022 | Regularization Penalty Optimization for Addressing Data Quality Variance in OoD AlgorithmsabstractDue to the poor generalization performance of traditional empirical risk minimization (ERM) in the case of distributional shift, Out-of-Distribution (OoD) generalization algorithms receive increasing attention. However, OoD generalization algorithms overlook the great variance in the quality of training data, which significantly compromises the accuracy of these methods. In this paper, we theoretically reveal the relationship between training data quality and algorithm performance, and analyze the optimal regularization scheme for Lipschitz regularized invariant risk minimization. A novel algorithm is proposed based on the theoretical results to alleviate the influence of low quality data at both the sample level and the domain level. The experiments on both the regression and classification benchmarks validate the effectiveness of our method with statistical significance. Runpeng Yu 0001, Hong Zhu 0003, Kaican Li, Lanqing Hong, Rui Zhang 0003, Nanyang Ye 0001, Shao-Lun Huang, Xiuqiang He 0001 |
AAAI | 6 |
| 2022 | OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationabstractDeep learning has achieved tremendous success with independent and identically distributed (i. i.d.) data. However, the performance of neural networks often degenerates drastically when encountering out-of-distribution (OoD) data, i.e., when training and test data are sampled from different distributions. While a plethora of algorithms have been proposed for OoD generalization, our understanding of the data used to train and evaluate these algorithms remains stagnant. In this work, we first identify and measure two distinct kinds of distribution shifts that are ubiquitous in various datasets. Next, through extensive experiments, we compare OoD generalization algorithms across two groups of benchmarks, each dominated by one of the distribution shifts, revealing their strengths on one shift as well as limitations on the other shift. Overall, we position existing datasets and algorithms from different research areas seemingly unconnected into the same coherent picture. It may serve as a foothold that can be resorted to by future OoD generalization research. Our code is available at https://github.com/ynysjtulood_bench. Nanyang Ye 0001, Kaican Li, Haoyue Bai 0001, Runpeng Yu 0001, Lanqing Hong, Fengwei Zhou, Zhenguo Li, Jun Zhu 0001 |
CVPR | 1 |
| 2022 | Achieving adversarial robustness via sparsity
Ningyi Liao, Shufan Wang, Liyao Xiang, Nanyang Ye 0001, Pengzhi Chu |
Mach. Learn. | 4 |
| 2021 | DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic AugmentationabstractWhile deep learning demonstrates its strong ability to handle independent and identically distributed (IID) data, it often suffers from out-of-distribution (OoD) generalization, where the test data come from another distribution (w.r.t. the training one). Designing a general OoD generalization framework for a wide range of applications is challenging, mainly due to different kinds of distribution shifts in the real world, such as the shift across domains or the extrapolation of correlation. Most of the previous approaches can only solve one specific distribution shift, leading to unsatisfactory performance when applied to various OoD benchmarks. In this work, we propose DecAug, a novel decomposed feature representation and semantic augmentation approach for OoD generalization. Specifically, DecAug disentangles the category-related and context-related features by orthogonalizing the two gradients (w.r.t. intermediate features) of losses for predicting category and context labels, where category-related features contain causal information of the target object, while context-related features cause distribution shifts between training and test data. Furthermore, we perform gradient-based augmentation on context-related features to improve the robustness of learned representations. Experimental results show that DecAug outperforms other state-of-the-art methods on various OoD datasets, which is among the very few methods that can deal with different types of OoD generalization challenges. Haoyue Bai 0001, Lanqing Hong, Fengwei Zhou, Nanyang Ye 0001, Han-Jia Ye, Shueng-Han Gary Chan, Zhenguo Li |
AAAI | 5 |
| 2021 | Amata: An Annealing Mechanism for Adversarial Training AccelerationabstractDespite the empirical success in various domains, it has been revealed that deep neural networks are vulnerable to maliciously perturbed input data that much degrade their performance. This is known as adversarial attacks. To counter adversarial attacks, adversarial training formulated as a form of robust optimization has been demonstrated to be effective. However, conducting adversarial training brings much computational overhead compared with standard training. In order to reduce the computational cost, we propose an annealing mechanism, Amata, to reduce the overhead associated with adversarial training. The proposed Amata is provably convergent, well-motivated from the lens of optimal control theory and can be combined with existing acceleration methods to further enhance performance. It is demonstrated that on standard datasets, Amata can achieve similar or better robustness with around 1/3 to 1/2 the computational time compared with traditional methods. In addition, Amata can be incorporated into other adversarial training acceleration algorithms (e.g. YOPO, Free, Fast, and ATTA), which leads to further reduction in computational time on large-scale problems. Nanyang Ye 0001, Qianxiao Li, Xiaoyun Zhou 0001, Zhanxing Zhu |
AAAI | 1 |
| 2021 | Adversarial Invariant LearningabstractThough machine learning algorithms are able to achieve pattern recognition from the correlation between data and labels, the presence of spurious features in the data decreases the robustness of these learned relationships with respect to varied testing environments. This is known as out-of-distribution (OoD) generalization problem. Recently, invariant risk minimization (IRM) attempts to tackle this issue by penalizing predictions based on the unstable spurious features in the data collected from different environments. However, similar to domain adaptation or domain generalization, a prevalent non-trivial limitation in these works is that the environment information is assigned by human specialists, i.e. a priori, or determined heuristically. However, an inappropriate group partitioning can dramatically deteriorate the OoD generalization and this process is expensive and time-consuming. To deal with this issue, we propose a novel theoretically principled min-max framework to iteratively construct a worst-case splitting, i.e. creating the most challenging environment splittings for the backbone learning paradigm (e.g. IRM) to learn the robust feature representation. We also design a differentiable training strategy to facilitate the feasible gradient- based computation. Numerical experiments show that our algorithmic framework has achieved superior and stable performance in various datasets, such as Colored MNIST and Punctuated Stanford sentiment treebank (SST). Furthermore, we also find our algorithm to be robust even to a strong data poisoning attack. To the best of our knowledge, this is one of the first to adopt differentiable environment splitting method to enable stable predictions across environments without environment index information, which achieves the state-of-the-art performance on datasets with strong spurious correlation, such as Colored MNIST. Nanyang Ye 0001, Jingxuan Tang, Huayu Deng, Xiaoyun Zhou 0001, Qianxiao Li, Zhenguo Li, Guang-Zhong Yang, Zhanxing Zhu |
CVPR | 1 |
| 2021 | BayesFT: Bayesian Optimization for Fault Tolerant Neural Network ArchitectureabstractTo deploy deep learning algorithms on resource-limited scenarios, an emerging device-resistive random access memory (ReRAM) has been regarded as promising via analog computing. However, the practicability of ReRAM is primarily limited due to the weight drifting of ReRAM neural networks due to multi-factor reasons, including manufacturing, thermal noises, and etc. In this paper, we propose a novel Bayesian optimization method for fault tolerant neural network architecture (BayesFT). For neural architecture search space design, instead of conducting neural architecture search on the whole feasible neural architecture search space, we first systematically explore the weight drifting tolerance of different neural network components, such as dropout, normalization, number of layers, and activation functions in which dropout is found to be able to improve the neural network robustness to weight drifting. Based on our analysis, we propose an efficient search space by only searching for dropout rates for each layer. Then, we use Bayesian optimization to search for the optimal neural architecture robust to weight drifting. Empirical experiments demonstrate that our algorithmic framework has outperformed the state-of-the-art methods by up to 10 times on various tasks, such as image classification and object detection. Nanyang Ye 0001, Jingbiao Mei, Zhicheng Fang, Huaying Wu, Xiaoyao Liang |
DAC | 1 |
| 2021 | DeepIC: Coding for Interference Channels via Deep LearningabstractThe two-user interference channel is a model for multi one-to-one communications, where two transmitters wish to communicate with their corresponding receivers via a shared wireless medium. Two most common and simple coding schemes are Time Division (TD) and Treating Interference as Noise (TIN). Interestingly, it is shown that there exists an asymptotic scheme, called Han-Kobayashi scheme, that performs better than TD and TIN. However, Han-Kobayashi scheme has impractically high complexity and is designed for asymptotic settings, which leads to a gap between information theory and practice. In this paper, we focus on designing practical codes for interference channels. As it is challenging to analytically design practical codes with feasible complexity, we apply deep learning to learn codes for interference channels. We demonstrate that DeepIC, a convolutional neural network-based code with an iterative decoder, outperforms TD and TIN by a noticeable margin for two-user Additive White Gaussian Noise channels with moderate amount of interference. Karl Chahine, Nanyang Ye 0001, Hyeji Kim |
GLOBECOM | 2 |
| 2021 | NAS-OoD: Neural Architecture Search for Out-of-Distribution GeneralizationabstractRecent advances on Out-of-Distribution (OoD) generalization reveal the robustness of deep learning models against distribution shifts. However, existing works focus on OoD algorithms, such as invariant risk minimization, domain generalization, or stable learning, without considering the influence of deep model architectures on OoD generalization, which may lead to sub-optimal performance. Neural Architecture Search (NAS) methods search for architecture based on its performance on the training data, which may result in poor generalization for OoD tasks. In this work, we propose robust Neural Architecture Search for OoD generalization (NAS-OoD), which optimizes the architecture with respect to its performance on generated OoD data by gradient descent. Specifically, a data generator is learned to synthesize OoD data by maximizing losses computed by different neural architectures, while the goal for architecture search is to find the optimal architecture parameters that minimize the synthetic OoD data losses. The data generator and the neural architecture are jointly optimized in an end-to-end manner, and the minimax training process effectively discovers robust architectures that generalize well for different distribution shifts. Extensive experimental results show that NAS-OoD achieves superior performance on various OoD generalization benchmarks with deep models having a much fewer number of parameters. In addition, on a real industry dataset, the proposed NAS-OoD method reduces the error rate by more than 70% compared with the state-of-the-art method, demonstrating the proposed method’s practicality for real applications. Haoyue Bai 0001, Fengwei Zhou, Lanqing Hong, Nanyang Ye 0001, Shueng-Han Gary Chan, Zhenguo Li |
ICCV | 4 |
| 2021 | Semi-supervised Vein Segmentation of Ultrasound Images for Autonomous VenipunctureabstractVenipuncture is an indispensable procedure for both diagnosis and treatment. In this paper, unlike existing solutions that fully or partially rely on professional assistance, a compact robotic system integrating both novel hardware and software developments is introduced. The hardware consists of a set of units to facilitate the supporting, positioning, puncturing, and imaging functionalities. To achieve full automation, a novel deep learning framework — semi-ResNeXt-Unet for semi-supervised vein segmentation from ultrasound images is proposed. The depth information of vein is calculated and enables the automated navigation for the puncturing unit. The algorithm is validated on 40 volunteers, and the proposed semi-ResNeXt-Unet improves the dice similarity coefficient (DSC) by 5.36%, decreases the centroid error by 1.38 pixels and decreases the failure rate by 5.60%, compared to fully-supervised ResNeXt-Unet. Bolin Lai, Nanyang Ye 0001, Zhongyuan Ren, Xiaoyun Zhou 0001, Peng Qi 0001 |
IROS | 6 |
| 2019 | Predicting Visible Image Differences Under Varying Display Brightness and Viewing DistanceabstractNumerous applications require a robust metric that can predict whether image differences are visible or not. However, the accuracy of existing white-box visibility metrics, such as HDR-VDP, is often not good enough. CNN-based black-box visibility metrics have proven to be more accurate, but they cannot account for differences in viewing conditions, such as display brightness and viewing distance. In this paper, we propose a CNN-based visibility metric, which maintains the accuracy of deep network solutions and accounts for viewing conditions. To achieve this, we extend the existing dataset of locally visible differences (LocVis) with a new set of measurements, collected considering aforementioned viewing conditions. Then, we develop a hybrid model that combines white-box processing stages for modeling the effects of luminance masking and contrast sensitivity, with a black-box deep neural network. We demonstrate that the novel hybrid model can handle the change of viewing conditions correctly and outperforms state-of-the-art metrics. Nanyang Ye 0001, Krzysztof Wolski, Rafal Mantiuk |
CVPR | 1 |
| 2019 | Visibility Metric for Visually Lossless Image CompressionabstractEncoding images in a visually lossless manner helps to achieve the best trade-off between image compression performance and quality and so that compression artifacts are invisible to the majority of users. Visually lossless encoding can often be achieved by manually adjusting compression quality parameters of existing lossy compression methods, such as JPEG or WebP. But the required compression quality parameter can also be determined automatically using visibility metrics. However, creating an accurate visibility metric is challenging because of the complexity of the human visual system and the effort needed to collect the required data. In this paper, we investigate how to train an accurate visibility metric for visually lossless compression from a relatively small dataset. Our experiments show that prediction error can be reduced by 40% compared with the state-of-theart, and that our proposed method can save between 25%-75% of storage space compared with the default quality parameter used in commercial software. We demonstrate how the visibility metric can be used for visually lossless image compression and for benchmarking image compression encoders. Nanyang Ye 0001, María Pérez-Ortiz 0001, Rafal Mantiuk |
PCS | 1 |
| 2018 | Trained Perceptual Transform for Quality Assessment of High Dynamic Range Images and VideoabstractIn this paper, we propose a trained perceptually transform for quality assessment of high dynamic range (HDR) images and video. The transform is used to convert absolute luminance values found in HDR images into perceptually uniform units, which can be used with any standard-dynamic-range metric. The new transform is derived by fitting the parameters of a previously proposed perceptual encoding function to 4 different HDR subjective quality assessment datasets using Bayesian optimization. The new transform combined with a simple peak signal-to-noise ratio measure achieves better prediction performance in cross-dataset validation than existing transforms. We provide Matlab code for our metric11https://github.com/ynyCL/T-PT-metric. Nanyang Ye 0001, María Pérez-Ortiz 0001, Rafal Mantiuk |
ICIP | 1 |
| 2018 | Stochastic Fractional Hamiltonian Monte CarloabstractIn this paper, we propose a novel stochastic fractional Hamiltonian Monte Carlo approach which generalizes the Hamiltonian Monte Carlo method within the framework of fractional calculus and L\'evy diffusion. Due to the large ``jumps'' introduced by L\'evy noise and momentum term, the proposed dynamics is capable of exploring the parameter space more efficiently and effectively. We have shown that the fractional Hamiltonian Monte Carlo could sample the multi-modal and high-dimensional target distribution more efficiently than the existing methods driven by Brownian diffusion. We further extend our method for optimizing deep neural networks. The experimental results show that the proposed stochastic fractional Hamiltonian Monte Carlo for training deep neural networks could converge faster than other popular optimization schemes and generalize better. Nanyang Ye 0001, Zhanxing Zhu |
IJCAI | 1 |
| 2018 | Bayesian Adversarial LearningabstractDeep neural networks have been known to be vulnerable to adversarial attacks, raising lots of security concerns in the practical deployment. Popular defensive approaches can be formulated as a (distributionally) robust optimization problem, which minimizes a ``point estimate'' of worst-case loss derived from either per-datum perturbation or adversary data-generating distribution within certain pre-defined constraints. This point estimate ignores potential test adversaries that are beyond the pre-defined constraints. The model robustness might deteriorate sharply in the scenario of stronger test adversarial data. In this work, a novel robust training framework is proposed to alleviate this issue, Bayesian Robust Learning, in which a distribution is put on the adversarial data-generating distribution to account for the uncertainty of the adversarial data-generating process. The uncertainty directly helps to consider the potential adversaries that are stronger than the point estimate in the cases of distributionally robust optimization. The uncertainty of model parameters is also incorporated to accommodate the full Bayesian framework. We design a scalable Markov Chain Monte Carlo sampling strategy to obtain the posterior distribution over model parameters. Various experiments are conducted to verify the superiority of BAL over existing adversarial training methods. The code for BAL is available at \url{https://tinyurl.com/ycxsaewr }. Nanyang Ye 0001, Zhanxing Zhu |
NeurIPS | 1 |
| 2018 | Dataset and Metrics for Predicting Local Visible DifferencesabstractA large number of imaging and computer graphics applications require localized information on the visibility of image distortions. Existing image quality metrics are not suitable for this task as they provide a single quality value per image. Existing visibility metrics produce visual difference maps, and are specifically designed for detecting just noticeable distortions but their predictions are often inaccurate. In this work, we argue that the key reason for this problem is the lack of large image collections with a good coverage of possible distortions that occur in different applications. To address the problem, we collect an extensive dataset of reference and distorted image pairs together with user markings indicating whether distortions are visible or not. We propose a statistical model that is designed for the meaningful interpretation of such data, which is affected by visual search and imprecision of manual marking. We use our dataset for training existing metrics and we demonstrate that their performance significantly improves. We show that our dataset with the proposed statistical model can be used to train a new CNN-based metric, which outperforms the existing solutions. We demonstrate the utility of such a metric in visually lossless JPEG compression, super-resolution and watermarking. Krzysztof Wolski, Daniele Giunchi, Nanyang Ye 0001, Piotr Didyk, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Anthony Steed, Rafal Mantiuk |
ACM Trans. Graph. | 3 |
| 2017 | Langevin Dynamics with Continuous Tempering for Training Deep Neural NetworksabstractMinimizing non-convex and high-dimensional objective functions is challenging, especially when training modern deep neural networks. In this paper, a novel approach is proposed which divides the training process into two consecutive phases to obtain better generalization performance: Bayesian sampling and stochastic optimization. The first phase is to explore the energy landscape and to capture the `fat'' modes; and the second one is to fine-tune the parameter learned from the first phase. In the Bayesian learning phase, we apply continuous tempering and stochastic approximation into the Langevin dynamics to create an efficient and effective sampler, in which the temperature is adjusted automatically according to the designed ``temperature dynamics''. These strategies can overcome the challenge of early trapping into bad local minima and have achieved remarkable improvements in various types of neural networks as shown in our theoretical analysis and empirical experiments. Nanyang Ye 0001, Zhanxing Zhu, Rafal Mantiuk |
NIPS | 1 |
| 2016 | Mouse calibration aided real-time gaze estimation based on boost Gaussian Bayesian learningabstractIn this paper, we propose a novel gaze estimation method to evaluate the attention span of users upon on-screen content via a single webcam. Our method is based on supervised descent method for eye region of interest (ROI) extraction. Then, boost Gaussian Bayesian regressors are applied to learn a robust mapping from the input eye ROI to gaze coordinates. To get enough training samples, we implant our scheme as a plug-in into web browsers for data collection from users without bothering. To improve accuracy, we also introduce mouse click to help train the regressors. Experiment results show that our method outperforms the existing method and can provide gaze estimation data for user behaviour analysis in real-time implementation. Nanyang Ye 0001, Xiaoming Tao 0001, Linhao Dong, Ning Ge 0001 |
ICIP | 1 |
| 2015 | Efficient Multi-Cell Clustering for Coordinated Multi-Point Transmission with Blossom Tree AlgorithmabstractCoordinated multi-point(CoMP) transmission clustering schemes could provide significant gains of system performance, such as throughput and cell- edge user data rates. Due to limitations of the backhaul communication and signal processing capability of base stations(BSs), the intrinsic problem of CoMP is that the selection of which BSs shall cooperate as only a few of BSs can be grouped in a cluster. However, approximating the theoretical performance bound of this clustering problem in CoMP at present is seldom discussed due to its inherent combinatorial complexity. In this paper, a novel efficient multi-cell clustering scheme based on blossom tree algorithm is proposed for cellular networks, incorporating CoMP with two cells in each cluster. With blossom tree algorithm, the proposed scheme can find out the optimal clustering strategy and help the CoMP transmission reach its theoretical performance bound on data rate in real-time computing(milliseconds in MATLAB simulation for one clustering). The simulation results show that our proposed method outperforms the existing dynamic greedy method in terms of cell edge users' average achievable data rate. Besides, it can also maintain high performance when extended to larger clusters in that with 4-cell clustering, the proposed method can reach 23.8% higher data rates than dynamic greedy method. Nanyang Ye 0001, Linhao Dong, Xiaoming Tao 0001, Ning Ge 0001 |
VTC Fall | 1 |