Boyuan Feng

dblp:227/2946 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
22since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Accelerating neural network training: An analysis of the AlgoPerf competition
abstract
The goal of the AlgoPerf: Training Algorithms competition is to evaluate practical speed-ups in neural network training achieved solely by improving the underlying training algorithms. In the external tuning ruleset, submissions must provide workload-agnostic hyperparameter search spaces, while in the self-tuning ruleset they must be completely hyperparameter-free. In both rulesets, submissions are compared on time-to-result across multiple deep learning workloads, training on fixed hardware. This paper presents the inaugural AlgoPerf competition's results, which drew 18 diverse submissions from 10 teams. Our investigation reveals several key findings: (1) The winning submission in the external tuning ruleset, using Distributed Shampoo, demonstrates the effectiveness of non-diagonal preconditioning over popular methods like Adam, even when compared on wall-clock runtime. (2) The winning submission in the self-tuning ruleset, based on the Schedule Free AdamW algorithm, demonstrates a new level of effectiveness for completely hyperparameter-free training algorithms. (3) The top-scoring submissions were surprisingly robust to workload changes. We also discuss the engineering challenges encountered in ensuring a fair comparison between different training algorithms. These results highlight both the significant progress so far, and the considerable room for further improvements.
Priya Kasimbeg, Frank Schneider 0001, Runa Eschenhagen, Juhan Bae, Chandramouli Shama Sastry, Mark Saroufim, Boyuan Feng, Less Wright, Edward Z. Yang, Zachary Nado, Sourabh Medapati, Philipp Hennig, Michael G. Rabbat, George E. Dahl
ICLR7
2025 GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Tong Geng, Ang Li 0006, Yufei Ding 0001
USENIX ATC2
2025 D3BT: Dynamic 3D Body Transformer for Body Fat Percentage Assessment
abstract
3D body scan has been adopted for body composition assessment due to its ability to accurately capture body shape measurements. However, the complexity of mesh representation and the lack of fine-shape descriptors limit its applications in body fat percentage analysis. Most studies rely on algorithms applied to anthropometric values derived from 3D scans, such as multiple girth measurements, which fail to account for the body's detailed shape. To address these issues, we explore the feasibility of using point cloud representation. However, few existing point-based methods are aimed at the human body or regression tasks. In this study, we introduce a new model, D3BT, which utilizes a transformer-based network on the body point cloud to efficiently learn shape information for regional and global fat percentage regression tasks. The model dynamically divides the points into voxels for enhanced transformer training, providing higher density and better alignment across different subjects, which is more suitable for body shape learning. We evaluate various models for predicting body fat percentage from 3D body scans, using ground truth data from dual-energy X-ray absorptiometry (DXA) reports. Compared to traditional methods that depend on anthropometric measurements and other point-based approaches, the proposed model shows superior results. In extensive experiments, the model reduces the Root Mean Square Error (RMSE) by an average of 10.30% and achieves an average R-squared score of 0.86.
Yijiang Zheng, Zhuoxin Long, Boyuan Feng, Ruting Cheng, Khashayar Vaziri, James K. Hahn
IEEE J. Biomed. Health Informatics3
2024 ZENO: A Type-based Optimization Framework for Zero Knowledge Neural Network Inference
abstract
Zero knowledge Neural Networks draw increasing attention for guaranteeing computation integrity and privacy of neural networks (NNs) based on zero-knowledge Succinct Non-interactive ARgument of Knowledge (zkSNARK) security scheme. However, the performance of zkSNARK NNs is far from optimal due to the million-scale circuit computation with heavy scalar-level dependency. In this paper, we propose a type-based optimizing framework for efficient zero-knowledge NN inference, namely ZENO (ZEro knowledge Neural network Optimizer). We first introduce ZENO language construct to maintain high-level semantics and the type information (e.g., privacy and tensor) for allowing more aggressive optimizations. We then propose privacy-type driven and tensor-type driven optimizations to further optimize the generated zkSNARK circuit. Finally, we design a set of NN-centric system optimizations to further accelerate zkSNARK NNs. Experimental results show that ZENO achieves up to 8.5× end-to-end speedup than state-of-the-art zkSNARK NNs. We reduce proof time for VGG16 from 6 minutes to 48 seconds, which makes zkSNARK NNs practical.
Boyuan Feng, Zheng Wang 0075, Yufei Ding 0001
ASPLOS (1)1
2024 OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
Zheng Wang 0075, Boyuan Feng, Guyue Huang, Dheevatsa Mudigere, Bharath Muthiah, Ang Li 0006, Yufei Ding 0001
USENIX ATC3
2023 On Adversarial Robustness of Point Cloud Semantic Segmentation
abstract
Recent research efforts on 3D point cloud semantic segmentation (PCSS) have achieved outstanding performance by adopting neural networks. However, the robustness of these complex models have not been systematically analyzed. Given that PCSS has been applied in many safety-critical applications like autonomous driving, it is important to fill this knowledge gap, especially, how these models are affected under adversarial samples. As such, we present a comparative study of PCSS robustness. First, we formally define the attacker's objective under performance degradation and object hiding. Then, we develop new attack by whether to bound the norm. We evaluate different attack options on two datasets and three PCSS models. We found all the models are vulnerable and attacking point color is more effective. With this study, we call the attention of the research community to develop new approaches to harden PCSS models.
Jiacen Xu 0001, Zhe Zhou 0001, Boyuan Feng, Yufei Ding 0001, Zhou Li 0001
DSN3
2023 MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
Boyuan Feng, Zheng Wang 0075, Tong Geng, Kevin J. Barker, Ang Li 0006, Yufei Ding 0001
OSDI2
2023 TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUs
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Yufei Ding 0001
USENIX ATC2
2023 ScLSTM: single-cell type detection by siamese recurrent network and hierarchical clustering
abstract
MOTIVATION: Categorizing cells into distinct types can shed light on biological tissue functions and interactions, and uncover specific mechanisms under pathological conditions. Since gene expression throughout a population of cells is averaged out by conventional sequencing techniques, it is challenging to distinguish between different cell types. The accumulation of single-cell RNA sequencing (scRNA-seq) data provides the foundation for a more precise classification of cell types. It is crucial building a high-accuracy clustering approach to categorize cell types since the imbalance of cell types and differences in the distribution of scRNA-seq data affect single-cell clustering and visualization outcomes. RESULT: To achieve single-cell type detection, we propose a meta-learning-based single-cell clustering model called ScLSTM. Specifically, ScLSTM transforms the single-cell type detection problem into a hierarchical classification problem based on feature extraction by the siamese long-short term memory (LSTM) network. The similarity matrix derived from the improved sigmoid kernel is mapped to the siamese LSTM feature space to analyze the differences between cells. ScLSTM demonstrated superior classification performance on 8 scRNA-seq data sets of different platforms, species, and tissues. Further quantitative analysis and visualization of the human breast cancer data set validated the superiority and capability of ScLSTM in recognizing cell types.
Hanjing Jiang, Yabing Huang, Qianpeng Li, Boyuan Feng
BMC Bioinform.4
2022 QGTC: accelerating quantized graph neural networks via GPU tensor core
abstract
Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never been realized on modern GPU platforms. To this end, we propose the first Tensor Core (TC) based computing framework, QGTC, to support any-bitwidth computation for QGNNs on GPUs. We introduce a novel quantized low-bit arithmetic design based on the low-bit data representation and bit-decomposed computation. We craft a novel TC-tailored CUDA kernel design by incorporating 3D-stacked bit compression, zero-tile jumping, and non-zero tile reuse technique to improve the performance systematically. We incorporate an effective bandwidth-optimized subgraph packing strategy to maximize the transferring efficiency between CPU host and GPU device. We integrate QGTC with Pytorch for better programmability and extensibility. Extensive experiments demonstrate that QGTC can achieve evident inference speedup (on average 2.7X) compared with the state-of-the-art DGL framework across diverse settings.
Boyuan Feng, Yufei Ding 0001
PPoPP2
2022 EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table
abstract
Deep learning Recommendation Models (DLRMs) plays an important role in various application domains. However, existing DLRM training systems require a large number of GPUs due to the memory-intensive embedding tables. To this end, we propose EL-Rec, an efficient computing framework harnessing the Tensor-train (TT) technique to democratize the training of large-scale DLRMs with limited GPU resources. Specifically, EL-Rec optimizes TT decomposition based on key computation primitives of embedding tables and implements a high-performance compressed embedding table which is a drop-in replacement of Pytorch API. EL-Rec introduces an index reordering technique to harvest the performance gains from both local and global information of training inputs. EL-Rec also highlights a pipeline training paradigm to eliminate the communication overhead between the host memory and the training worker. Comprehensive experiments demonstrate that EL-Rec can handle the largest publicly available DLRM dataset with a single GPU and achieves 3× speedup over the state-of-the-art DLRM frameworks.
Zheng Wang 0075, Boyuan Feng, Dheevatsa Mudigere, Bharath Muthiah, Yufei Ding 0001
SC3
2022 Faith: An Efficient Framework for Transformer Verification on GPUs
Boyuan Feng, Tianqi Tang 0001, Zhaodong Chen 0001, Zheng Wang 0075, Yuan Xie 0001, Yufei Ding 0001
USENIX ATC1
2022 STPAcc: Structural TI-Based Pruning for Accelerating Distance-Related Algorithms on CPU-FPGA Platforms
abstract
As a promising solution to boost the performance of distance-related algorithms (e.g.,$K$-means and KNN), FPGA-based acceleration attracts lots of attention, but also comes with numerous challenges. In this work, we propose,STPAcc, an optimization framework based on structural triangle-inequality (TI)-based pruning (STP) for accelerating distance-related algorithms on CPU-FPGA platforms. STPAcc provides a domain-specific language to unify distance-related algorithms effectively, a structural TI-based pruning strategy to remove unnecessary distance computations, a coarse-grained workload partitioning and mapping strategy to fully exploit the potentials of the CPU-FPGA platform, and fine-grained hardware optimizations to further improve performance on the FPGA. Intensive experiments show that STPAcc designs achieve$31.42\times $speedup and$99.63\times $better energy efficiency on average over standard CPU-based implementations.
Boyuan Feng, Gushu Li, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 UAG: Uncertainty-aware Attention Graph Neural Network for Defending Adversarial Attacks
abstract
With the increasing popularity of graph-based learning, graph neural networks (GNNs) emerge as the essential tool for gaining insights from graphs. However, unlike the conventional CNNs that have been extensively explored and exhaustively tested, people are still worrying about the GNNs' robustness under the critical settings, such as financial services. The main reason is that existing GNNs usually serve as a black-box in predicting and do not provide the uncertainty on the predictions. On the other side, the recent advancement of Bayesian deep learning on CNNs has demonstrated its success of quantifying and explaining such uncertainties to fortify CNN models. Motivated by these observations, we propose UAG, the first systematic solution to defend adversarial attacks on GNNs through identifying and exploiting hierarchical uncertainties in GNNs. UAG develops a Bayesian uncertainty technique to explicitly capture uncertainties in GNNs and further employs an uncertainty-aware attention technique to defend adversarial attacks on GNNs. Intensive experiments show that our proposed defense approach outperforms the state-of-the-art solutions by a significant margin.
Boyuan Feng, Yufei Ding 0001
AAAI1
2021 TiAcc: Triangle-inequality based Hardware Accelerator for K-means on FPGAs
abstract
K-means is one of the most important unsuper-vised learning algorithms. In this paper, we present TiAcc, a triangle-inequality based K-means hardware accelerator on FPGAs. TiAcc highlights itself with an algorithm-hardware co-design strategy tailored for K-means clustering. Specifically, TiAcc leverages a novel triangle-inequality based filtering to eliminate unnecessary distance computations without changing the final clustering results. Meanwhile, it employs a pipeline decoupling approach to mitigate the irregularity of the remaining computations, and an efficient hardware architecture design to fully exploit the pipeline and parallel processing capability of FPGAs. Moreover, TiAcc provides parameterized configuration knobs that can minimize the manual efforts in the arduous hardware design process and provides flexibility to optimize hardware designs for a variety of datasets with different sizes and dimensionalities. Intensive experiments show that TiAcc achieves an average 4.94× speedup and significant energy efficiency (average 74.22 ×) compared with an optimized K-means running on a server-grade Xeon CPU.
Boyuan Feng, Gushu Li, Georgios Tzimpragos, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001
CCGRID2
2021 An Efficient Quantitative Approach for Optimizing Convolutional Neural Networks
abstract
With the increasing popularity of deep learning, Convolutional Neural Networks (CNNs) have been widely applied in various domains, such as image classification and object detection, and achieve stunning success in terms of their high accuracy over the traditional statistical methods. To exploit the potentials of CNN models, a huge amount of research and industry efforts have been devoted to optimizing CNNs. Among these endeavors, CNN architecture design has attracted tremendous attention because of its great potential of improving model accuracy or reducing model complexity. However, existing work either introduces repeated training overhead in the search process or lacks an interpretable metric to guide the design.
Boyuan Feng, Xueqiao Peng, Yufei Ding 0001
CIKM2
2021 Saga: Sparse Adversarial Attack on EEG-Based Brain Computer Interface
abstract
With the recent advancement of the Brain-Computer Interface (BCI), Electroencephalogram (EEG) analytics gain a lot of research attention from various domains. Understanding the vulnerabilities of EEG analytics is important for safely applying this emerging technology in our daily life. Recent studies show that EEG analytics are vulnerable to adversarial attacks when adding small perturbations on the EEG data. However, fewer research efforts have been devoted to the robustness of EEG analytics under sparse perturbations that attack only small portions of the data. In this paper, we conduct the first in-depth study on the robustness of EEG analytics under sparse perturbations and propose the first Sparse Adversarial eeG Attack, SAGA, to identify weakness of EEG analytics. Specifically, by viewing EEG data as time series collected from several channels, we design an adaptive mask to uniformly represent diverse sparsity in adversarial attacks. We further introduce a PGD-based iterative solver to automatically select the time steps and channels under the given sparsity constraints and effectively identify the adversarial examples on EEG data. Extensive experiments show that SAGA can effectively generate sparse perturbations and introduces a 77.02% accuracy drop on average by only perturbing 5% channels and time steps.
Boyuan Feng, Yufei Ding 0001
ICASSP1
2021 DSXplore: Optimizing Convolutional Neural Networks via Sliding-Channel Convolutions
abstract
As the key advancement of the convolutional neural networks (CNNs), depthwise separable convolutions (DSCs) are becoming one of the most popular techniques to reduce the computations and parameters size of CNNs meanwhile maintaining the model accuracy. It also brings profound impact to improve the applicability of the compute- and memory-intensive CNNs to a broad range of applications, such as mobile devices, which are generally short of computation power and memory. However, previous research in DSCs are largely focusing on compositing the limited existing DSC designs, thus, missing the opportunities to explore more potential designs that can achieve better accuracy and higher computation/parameter reduction. Besides, the off-the-shelf convolution implementations offer limited computing schemes, therefore, lacking support for DSCs with different convolution patterns.To this end, we introduce, DSXplore, the first optimized design for exploring DSCs on CNNs. Specifically, at the algorithm level, DSXplore incorporates a novel factorized kernel-sliding-channel convolution (SCC), featured with input-channel overlapping to balance the accuracy performance and the reduction of computation and memory cost. SCC also offers enormous space for design exploration by introducing adjustable kernel parameters. Further, at the implementation level, we carry out an optimized GPU-implementation tailored for SCC by leveraging several key techniques, such as the input-centric backward design and the channel-cyclic optimization. Intensive experiments on different datasets across mainstream CNNs show the advantages of DSXplore in balancing accuracy and computation/parameter reduction over the standard convolution and the existing DSCs.
Boyuan Feng, Yufei Ding 0001
IPDPS2
2021 GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUs
Boyuan Feng, Gushu Li, Shuangchen Li, Lei Deng 0003, Yuan Xie 0001, Yufei Ding 0001
OSDI2
2021 EGEMM-TC: accelerating scientific computing on tensor cores with extended precision
abstract
Nvidia Tensor Cores achieve high performance with half-precision matrix inputs tailored towards deep learning workloads. However, this limits the application of Tensor Cores especially in the area of scientific computing with high precision requirements. In this paper, we build Emulated GEMM on Tensor Cores (EGEMM-TC) to extend the usage of Tensor Cores to accelerate scientific computing applications without compromising the precision requirements. First, EGEMM-TC employs an extendable workflow of hardware profiling and operation design to generate a lightweight emulation algorithm on Tensor Cores with extended-precision. Second, EGEMM-TC exploits a set of Tensor Core kernel optimizations to achieve high performance, including the highly-efficient tensorization to exploit the Tensor Core memory architecture and the instruction-level optimizations to coordinate the emulation computation and memory access. Third, EGEMM-TC incorporates a hardware-aware analytic model to offer large flexibility for automatic performance tuning across various scientific computing workloads and input datasets. Extensive evaluations show that EGEMM-TC can achieve on average 3.13× and 11.18× speedup over the cuBLAS kernels and the CUDA-SDK kernels on CUDA Cores, respectively. Our case study on several scientific computing applications further confirms that EGEMM-TC can generalize the usage of Tensor Cores and achieve about 1.8× speedup compared to the hand-tuned, highly-optimized implementations running on CUDA Cores.
Boyuan Feng, Guoyang Chen, Weifeng Zhang 0003, Yuan Xie 0001, Yufei Ding 0001
PPoPP1
2021 APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor cores
abstract
Over the years, accelerating neural networks with quantization has been widely studied. Unfortunately, prior efforts with diverse precisions (e.g., 1-bit weights and 2-bit activations) are usually restricted by limited precision support on GPUs (e.g., int1 and int4). To break such restrictions, we introduce the first Arbitrary Precision Neural Network framework (APNN-TC)1 to fully exploit quantization benefits on Ampere GPU Tensor Cores. Specifically, APNN-TC first incorporates a novel emulation algorithm to support arbitrary short bit-width computation with int1 compute primitives and XOR/AND Boolean operations. Second, APNN-TC integrates arbitrary precision layer designs to efficiently map our emulation algorithm to Tensor Cores with novel batching strategies and specialized memory organization. Third, APNN-TC embodies a novel arbitrary precision NN design to minimize memory access across layers and further improve performance. Extensive evaluations show that APNN-TC can achieve significant speedup over CUTLASS kernels and various NN models, such as ResNet and VGG.
Boyuan Feng, Tong Geng, Ang Li 0006, Yufei Ding 0001
SC1
2021 Palleon: A Runtime System for Efficient Video Processing toward Dynamic Class Skew
Boyuan Feng, Gushu Li, Yuan Xie 0001, Yufei Ding 0001
USENIX ATC1
2020 SGQuant: Squeezing the Last Bit on Graph Neural Networks with Specialized Quantization
abstract
With the increasing popularity of graph-based learning, Graph Neural Networks (GNNs) win lots of attention from research and industry field because of their high accuracy. However, existing GNNs suffer from high memory footprints (e.g., node embedding features). This high memory footprint hurdles the potential applications towards memory-constrained devices, such as the widely-deployed IoT devices. To this end, we propose a specialized GNN quantization scheme, SGQuant, to systematically reduce the GNN memory consumption. Specifically, we first propose a GNN-tailored quantization algorithm design and a GNN quantization fine-tuning scheme to reduce memory consumption while maintaining accuracy. Then, we investigate the multi-granularity quantization strategy that operates at different levels (components, graph topology, and layers) of GNN computation. Moreover, we offer an automatic bit-selecting (ABS) to pinpoint the most appropriate quantization bits for the above multi-granularity quantizations. Intensive experiments show that SGQuant can effectively reduce the memory footprint from 4.25× to 31.9× compared with the original full-precision GNNs while limiting the accuracy drop to 0.4% on average.
Boyuan Feng, Xueqiao Peng, Yufei Ding 0001
ICTAI1
2020 A Close Look at Multi-tenant Parallel CNN Inference for Autonomous Driving
Yitong Huang, Yu Zhang 0086, Boyuan Feng, Yanyong Zhang, Yufei Ding 0001
NPC3
2020 Domain-adversarial multi-task framework for novel therapeutic property prediction of compounds
abstract
MOTIVATION: With the rapid development of high-throughput technologies, parallel acquisition of large-scale drug-informatics data provides significant opportunities to improve pharmaceutical research and development. One important application is the purpose prediction of small-molecule compounds with the objective of specifying the therapeutic properties of extensive purpose-unknown compounds and repurposing the novel therapeutic properties of FDA-approved drugs. Such a problem is extremely challenging because compound attributes include heterogeneous data with various feature patterns, such as drug fingerprints, drug physicochemical properties and drug perturbation gene expressions. Moreover, there is a complex non-linear dependency among heterogeneous data. In this study, we propose a novel domain-adversarial multi-task framework for integrating shared knowledge from multiple domains. The framework first uses an adversarial strategy to learn target representations and then models non-linear dependency among several domains. RESULTS: Experiments on two real-world datasets illustrate that our approach achieves an obvious improvement over competitive baselines. The novel therapeutic properties of purpose-unknown compounds that we predicted have been widely reported or brought to clinics. Furthermore, our framework can integrate various attributes beyond the three domains examined herein and can be applied in industry for screening significant numbers of small-molecule drug candidates. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/JohnnyY8/DAMT-Model. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lingwei Xie, Zhongnan Zhang, Kunhui Lin, Xiaochen Bo, Boyuan Feng, Kun Wan 0001, Yufei Ding 0001
Bioinform.7
2019 KPynq: A Work-Efficient Triangle-Inequality Based K-Means on FPGA
abstract
K-means is a popular but computation-intensive algorithm for unsupervised learning. To address this issue, we present KPynq, a work-efficient triangle-inequality based K-means on FPGA for handling large-size, high-dimension datasets. KPynq leverages an algorithm-level optimization to balance the performance and computation irregularity, and a hardware architecture design to fully exploit the pipeline and parallel processing capability of various FPGAs. In the experiment, KPynq consistently outperforms the CPU-based standard K-means in terms of its speedup (up to 4.2×) and significant energy efficiency (up to 218×).
Zhaorui Zeng, Boyuan Feng, Lei Deng 0003, Yufei Ding 0001
FCCM3
2019 Reconciling Feature-Reuse and Overfitting in DenseNet with Specialized Dropout
abstract
Recently convolutional neural networks (CNNs) achieve great accuracy in visual recognition tasks. DenseNets become one of the most popular CNN models due to its effectiveness in the feature-reuse. However, like other CNN models, DenseNets also face the overfitting problem if not more severe. Existing dropout methods can be applied but not effective. In particular, the property of the feature-reuse in DenseNets will be impeded, and the dropout effect will be weakened by the spatial correlation inside feature maps. To address these problems, we craft the design of a specialized dropout method from three aspects, the dropout location, the dropout granularity, and the dropout probability. The insights attained here could potentially be applied as a general approach for boosting the accuracy of other CNN models with similar shortcut connections. Experimental results show that DenseNets with our specialized dropout method yield better accuracies compared to vanilla DenseNets and state-of-the-art CNN models, and such accuracy boost increases with the model depth.
Kun Wan 0001, Boyuan Feng, Yufei Ding 0001, Lingwei Xie
ICTAI3