Borui Zhao

dblp:260/5192 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
18since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Asymmetric Decision-Making in Online Knowledge Distillation: Unifying Consensus and Divergence
abstract
Online Knowledge Distillation (OKD) methods represent a streamlined, one-stage distillation training process that obviates the necessity of transferring knowledge from a pretrained teacher network to a more compact student network. In contrast to existing logits-based OKD methods, this paper presents an innovative approach to leverage intermediate spatial representations. Our analysis of the intermediate features from both teacher and student models reveals two pivotal insights: (1) the similar features between students and teachers are predominantly focused on the foreground objects. (2) teacher models emphasize foreground objects more than students. Building on these findings, we propose Asymmetric Decision-Making (ADM) to enhance feature consensus learning for student models while continuously promoting feature diversity in teacher models. Specifically, Consensus Learning for student models prioritizes spatial features with high consensus relative to teacher models. Conversely, Divergence Learning for teacher models highlights spatial features with lower similarity compared to student models, indicating superior performance by teacher models in these regions. Consequently, ADM facilitates the student models to catch up with the feature learning process of the teacher models. Extensive experiments demonstrate that ADM consistently surpasses existing OKD methods across various online knowledge distillation settings and also achieves superior results when transferred to offline knowledge distillation, semantic segmentation and diffusion distillation tasks.
Zhaowei Chen, Borui Zhao, Yuchen Ge, Renjie Song, Jiajun Liang
ICML2
2025 Efficient Hierarchical Federated Services for Heterogeneous Mobile Edge
abstract
As 6G networks actively advance edge intelligence, Federated Learning (FL) emerges as a key technology that enables data sharing while preserving data privacy and fostering collaboration among edge devices for intelligent service learning. However, the multi-dimensional heterogeneous and hierarchical network architecture brings many challenges to FL deployment, including selecting appropriate nodes for model training and designing effective methods for model aggregation. Compared with most studies that focus on solving individual problems within 6G, this paper proposes an efficient deployment scheme named hierarchical heterogeneous FL (HHFL), which comprehensively considers various influencing factors. First, the deployment of HHFL over 6G is modeled amid the heterogeneity of communications, computation, and data. An optimization problem is then formulated, aiming to minimize deployment costs in terms of latency and energy consumption. Subsequently, to tackle this optimization challenge, we design an intelligent FL deployment framework, consisting of a hierarchical aggregation deployment (HAD) component for hierarchical FL aggregation structure construction and an adaptive node selection (ANS) component for selecting diverse clients based on multi-dimensional discrepancy criteria. Experimental results demonstrate that our proposed framework not only adapts to various application requirements but also outperforms existing technologies by achieving superior learning performance, reduced latency, and lower energy consumption.
Shengyuan Liang, Qimei Cui, Xueqing Huang, Borui Zhao, Yan-Zhao Hou, Xiaofeng Tao 0001
IEEE Trans. Serv. Comput.4
2025 Multi-Layer Collaborative Federated Learning architecture for 6G Open RAN
Borui Zhao, Qimei Cui, Wei Ni 0001, Xueqi Li 0004, Shengyuan Liang
Wirel. Networks1
2023 Curriculum Temperature for Knowledge Distillation
abstract
Most existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search. In general, the temperature controls the discrepancy between two distributions and can faithfully determine the difficulty level of the distillation task. Keeping a constant temperature, i.e., a fixed level of task difficulty, is usually sub-optimal for a growing student during its progressive learning stages. In this paper, we propose a simple curriculum-based technique, termed Curriculum Temperature for Knowledge Distillation (CTKD), which controls the task difficulty level during the student's learning career through a dynamic and learnable temperature. Specifically, following an easy-to-hard curriculum, we gradually increase the distillation loss w.r.t. the temperature, leading to increased distillation difficulty in an adversarial manner. As an easy-to-use plug-in technique, CTKD can be seamlessly integrated into existing knowledge distillation frameworks and brings general improvements at a negligible additional computation cost. Extensive experiments on CIFAR-100, ImageNet-2012, and MS-COCO demonstrate the effectiveness of our method.
Zheng Li 0028, Xiang Li 0041, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo 0001, Jun Li 0027, Jian Yang 0003
AAAI4
2023 Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data
abstract
Semi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these methods all suffer from the waste of complicated examples since all pseudo-labels have to be selected by a high threshold to filter out noisy ones. Hence, the examples with ambiguous predictions will not contribute to the training phase. For better leveraging all unlabeled examples, we propose two novel techniques: Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL). EML incorporates the prediction distribution of non-target classes into the optimization objective to avoid competition with target class, and thus generating more high-confidence predictions for selecting pseudo-label. ANL introduces the additional negative pseudo-label for all unlabeled data to leverage low-confidence examples. It adaptively allocates this label by dynamically evaluating the top-k performance of the model. EML and ANL do not introduce any additional parameter and hyperparameter. We integrate these techniques with FixMatch, and develop a simple yet powerful framework called FullMatch. Extensive experiments on several common SSL benchmarks (CIFAR-10/100, SVHN, STL-10 and ImageNet) demonstrate that FullMatch exceeds FixMatch by a large margin. Integrated with FlexMatch (an advanced FixMatch-based framework), we achieve state-of-the-art performance. Source code is available at https://github.com/megvii-research/FullMatch.
Xin Tan 0002, Borui Zhao, Zhaowei Chen, Renjie Song, Jiajun Liang, Xuequan Lu
CVPR3
2023 DOT: A Distillation-Oriented Trainer
abstract
Knowledge distillation transfers knowledge from a large model to a small one via task and distillation losses. In this paper, we observe a trade-off between task and distillation losses, i.e., introducing distillation loss limits the convergence of task loss. We believe that the trade-off results from the insufficient optimization of distillation loss. The reason is: The teacher has a lower task loss than the student, and a lower distillation loss drives the student more similar to the teacher, then a better-converged task loss could be obtained. To break the trade-off, we propose the Distillation-Oriented Trainer (DOT). DOT separately considers gradients of task and distillation losses, then applies a larger momentum to distillation loss to accelerate its optimization. We empirically prove that DOT breaks the trade-off, i.e., both losses are sufficiently optimized. Extensive experiments validate the superiority of DOT. Notably, DOT achieves a +2.59% accuracy improvement on ImageNet-1k for the ResNet50-MobileNetV1 pair. Conclusively, DOT greatly benefits the student’s optimization properties in terms of loss convergence and model generalization. https://github.com/megvii-research/mdistiller.
Borui Zhao, Quan Cui, Renjie Song, Jiajun Liang
ICCV1
2023 Cumulative Spatial Knowledge Distillation for Vision Transformers
abstract
Distilling knowledge from convolutional neural networks (CNNs) is a double-edged sword for vision transformers (ViTs). It boosts the performance since the image-friendly local-inductive bias of CNN helps ViT learn faster and better, but leading to two problems: (1) Network designs of CNN and ViT are completely different, which leads to different semantic levels of intermediate features, making spatial-wise knowledge transfer methods (e.g., feature mimicking) inefficient. (2) Distilling knowledge from CNN limits the network convergence in the later training period since ViT’s capability of integrating global information is suppressed by CNN’s local-inductive-bias supervision.To this end, we present Cumulative Spatial Knowledge Distillation (CSKD). CSKD distills spatial-wise knowledge to all patch tokens of ViT from the corresponding spatial responses of CNN, without introducing intermediate features. Furthermore, CSKD exploits a Cumulative Knowledge Fusion (CKF) module, which introduces the global response of CNN and increasingly emphasizes its importance during the training. Applying CKF leverages CNN’s local inductive bias in the early training period and gives full play to ViT’s global capability in the later one. Extensive experiments and analysis on ImageNet-1k and downstream datasets demonstrate the superiority of our CSKD. Code: https://github.com/Zzzzz1/CSKD
Borui Zhao, Renjie Song, Jiajun Liang
ICCV1
2023 QBox: Partial Transfer Learning With Active Querying for Object Detection
abstract
Object detection requires plentiful data annotated with bounding boxes for model training. However, in many applications, it is difficult or even impossible to acquire a large set of labeled examples for the target task due to the privacy concern or lack of reliable annotators. On the other hand, due to the high-quality image search engines, such as Flickr and Google, it is relatively easy to obtain resource-rich unlabeled datasets, whose categories are a superset of those of target data. In this article, to improve the target model with cost-effective supervision from source data, we propose a partial transfer learning approach QBox to actively query labels for bounding boxes of source images. Specifically, we design two criteria, i.e., informativeness and transferability, to measure the potential utility of a bounding box for improving the target model. Based on these criteria, QBox actively queries the labels of the most useful boxes from the source domain and, thus, requires fewer training examples to save the labeling cost. Furthermore, the proposed query strategy allows annotators to simply labeling a specific region, instead of the whole image, and, thus, significantly reduces the labeling difficulty. Extensive experiments are performed on various partial transfer benchmarks and a real COVID-19 detection task. The results validate that QBox improves the detection accuracy with lower labeling cost compared to state-of-the-art query strategies for object detection.
Ying-Peng Tang, Xiu-Shen Wei, Borui Zhao, Sheng-Jun Huang
IEEE Trans. Neural Networks Learn. Syst.3
2022 Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal Information
abstract
Fine-grained image classification is a challenging computer vision task where various species share similar visual appearances, resulting in misclassification if merely based on visual clues. Therefore, it is helpful to leverage additional information, e.g., the locations and dates for data shooting, which can be easily accessible but rarely exploited. In this paper, we first demonstrate that existing multimodal methods fuse multiple features only on a single dimension, which essentially has insufficient help in feature discrimination. To fully explore the potential of multimodal information, we propose a dynamic MLP on top of the image representation, which interacts with multimodal features at a higher and broader dimension. The dynamic MLP is an efficient structure parameterized by the learned embeddings of variable locations and dates. It can be regarded as an adaptive nonlinear projection for generating more discriminative image representations in visual tasks. To our best knowledge, it is the first attempt to explore the idea of dynamic networks to exploit multimodal information in fine-grained image classification tasks. Extensive experiments demonstrate the effectiveness of our method. The t-SNE algorithm visually indicates that our technique improves the recognizability of image representations that are visually similar but with different categories. Furthermore, among published works across multiple fine-grained datasets, dynamic MLP consistently achieves SOTA results11https://paperswithcode.com/dataset/inaturalist and takes third place in the iNaturalist challenge at FGVC822https://www.kaggle.com/c/inaturalist-2021/leaderboard. Code is available at httpsr//glthub.com/megvii-research/DynamicMLPForFinegrained.
Lingfeng Yang, Xiang Li 0041, Renjie Song, Borui Zhao, Juntian Tao, Jiajun Liang, Jian Yang 0003
CVPR4
2022 Decoupled Knowledge Distillation
abstract
State-of-the-art distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. To provide a novel viewpoint to study logit distillation, we re-formulate the classical KD loss into two parts, i.e., target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). We empirically investigate and prove the effects of the two parts: TCKD transfers knowledge concerning the “difficulty” of training samples, while NCKD is the prominent reason why logit distillation works. More importantly, we reveal that the classical KD loss is a coupled formulation, which (1) suppresses the effectiveness of NCKD and (2) limits the flexibility to balance these two parts. To address these issues, we present Decoupled Knowledge Distillation (DKD), enabling TCKD and NCKD to play their roles more efficiently and flexibly. Compared with complex feature-based methods, our DKD achieves comparable or even better results and has better training efficiency on CIFAR-100, ImageNet, and MS-COCO datasets for image classification and object detection tasks. This paper proves the great potential of logit distillation, and we hope it will be helpful for future research. The code is available at https://github.com/megviiresearch/mdistiller.
Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, Jiajun Liang
CVPR1
2022 Discriminability-Transferability Trade-Off: An Information-Theoretic Perspective
Quan Cui, Bingchen Zhao, Borui Zhao, Renjie Song, Boyan Zhou, Jiajun Liang, Osamu Yoshie
ECCV (26)4
2022 Efficient One Pass Self-distillation with Zipf's Label Smoothing
Jiajun Liang, Linze Li 0001, Zhaodong Bing, Borui Zhao, Haoqiang Fan
ECCV (11)4
2022 Federated Learning-based Heterogeneous Load Prediction and Slicing for 5G Systems and Beyond
abstract
Network Slicing (NS) is an enabling technology to support vertical industries in 5G-and-beyond (B5G) systems. The management of Radio Access Network (RAN) slices re-lies on extensive awareness of network load status and data analysis. With increasingly diversified services and ubiquitous user data, centralized slice management is unsustainable. This paper presents a new distributed load prediction and slicing framework based on Federated Learning (FL). Specifically, we predict the slice-level traffic by designing a Federated Long Short Time Memory (Fed-LSTM) algorithm with a tailored loss function to reduce Service Level Agreement (SLA) violation rate. Given the predicted traffic load, we model slice resource orchestration as an improved two-dimensional Polygon Knapsack (2D- PK) problem, split slice traffic at different granularities, and solve the problem using the Maximal Rectangles Bottom-Left (MAXRECTS-BL) algorithm. Experimental results on real-world captured dataset show that our approach can achieve significant improvement in prediction accuracy, slicing SLA violation and resource utilization, compared to existing techniques.
Liyuan Pu, Qimei Cui, Borui Zhao, Wei Ni 0001, Ming Ai, Xiaofeng Tao 0001
GLOBECOM4
2022 RecursiveMix: Mixed Learning with History
abstract
Mix-based augmentation has been proven fundamental to the generalization of deep vision models. However, current augmentations only mix samples from the current data batch during training, which ignores the possible knowledge accumulated in the learning history. In this paper, we propose a recursive mixed-sample learning paradigm, termed ``RecursiveMix'' (RM), by exploring a novel training strategy that leverages the historical input-prediction-label triplets. More specifically, we iteratively resize the input image batch from the previous iteration and paste it into the current batch while their labels are fused proportionally to the area of the operated patches. Furthermore, a consistency loss is introduced to align the identical image semantics across the iterations, which helps the learning of scale-invariant feature representations. Based on ResNet-50, RM largely improves classification accuracy by $\sim$3.2% on CIFAR-100 and $\sim$2.8% on ImageNet with negligible extra computation/storage costs. In the downstream object detection task, the RM-pretrained model outperforms the baseline by 2.1 AP points and surpasses CutMix by 1.4 AP points under the ATSS detector on COCO. In semantic segmentation, RM also surpasses the baseline and CutMix by 1.9 and 1.1 mIoU points under UperNet on ADE20K, respectively. Codes and pretrained models are available at https://github.com/implus/RecursiveMix.
Lingfeng Yang, Xiang Li 0041, Borui Zhao, Renjie Song, Jian Yang 0003
NeurIPS3
2022 Delving deep into spatial pooling for squeeze-and-excitation networks
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan
Pattern Recognit.4
2022 SST: Spatial and Semantic Transformers for Multi-Label Image Recognition
abstract
Multi-label image recognition has attracted considerable research attention and achieved great success in recent years. Capturing label correlations is an effective manner to advance the performance of multi-label image recognition. Two types of label correlations were principally studied, i.e., the spatial and semantic correlations. However, in the literature, previous methods considered only either of them. In this work, inspired by the great success of Transformer, we propose a plug-and-play module, named the Spatial and Semantic Transformers (SST), to simultaneously capture spatial and semantic correlations in multi-label images. Our proposal is mainly comprised of two independent transformers, aiming to capture the spatial and semantic correlations respectively. Specifically, our Spatial Transformer is designed to model the correlations between features from different spatial positions, while the Semantic Transformer is leveraged to capture the co-existence of labels without manually defined rules. Other than methodological contributions, we also prove that spatial and semantic correlations complement each other and deserve to be leveraged simultaneously in multi-label image recognition. Benefitting from the Transformer's ability to capture long-range correlations, our method remarkably outperforms state-of-the-art methods on four popular multi-label benchmark datasets. In addition, extensive ablation studies and visualizations are provided to validate the essential components of our method.
Quan Cui, Borui Zhao, Renjie Song, Xiaoqin Zhang 0002, Osamu Yoshie
IEEE Trans. Image Process.3
2022 A Lightweight Encoder-Decoder Path for Deep Residual Networks
abstract
In this article, we present a novel lightweight path for deep residual neural networks. The proposed method integrates a simple plug-and-play module, i.e., a convolutional encoder-decoder (ED), as an augmented path to the original residual building block. Due to the abstract design and ability of the encoding stage, the decoder part tends to generate feature maps where highly semantically relevant responses are activated, while irrelevant responses are restrained. By a simple elementwise addition operation, the learned representations derived from the identity shortcut and original transformation branch are enhanced by our ED path. Furthermore, we exploit lightweight counterparts by removing a portion of channels in the original transformation branch. Fortunately, our lightweight processing does not cause an obvious performance drop but brings a computational economy. By conducting comprehensive experiments on ImageNet, MS-COCO, CUB200-2011, and CIFAR, we demonstrate the consistent accuracy gain obtained by our ED path for various residual architectures, with comparable or even lower model complexity. Concretely, it decreases the top-1 error of ResNet-50 and ResNet-101 by 1.22% and 0.91% on the task of ImageNet classification and increases the mmAP of Faster R-CNN with ResNet-101 by 2.5% on the MS-COCO object detection task. The code is available at https://github.com/Megvii-Nanjing/ED-Net.
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan, Yang Yu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 HCE: Hierarchical Context Embedding for Region-Based Object Detection
abstract
State-of-the-art two-stage object detectors apply a classifier to a sparse set of object proposals, relying on region-wise features extracted by RoIPool or RoIAlign as inputs. The region-wise features, in spite of aligning well with the proposal locations, may still lack the crucial context information which is necessary for filtering out noisy background detections, as well as recognizing objects possessing no distinctive appearances. To address this issue, we present a simple but effective Hierarchical Context Embedding (HCE) framework, which can be applied as a plug-and-play component, to facilitate the classification ability of a series of region-based detectors by mining contextual cues. Specifically, to advance the recognition of context-dependent object categories, we propose an image-level categorical embedding module which leverages the holistic image-level context to learn object-level concepts. Then, novel RoI features are generated by exploiting hierarchically embedded context information beneath both whole images and interested regions, which are also complementary to conventional RoI features. Moreover, to make full use of our hierarchical contextual RoI features, we propose the early-and-late fusion strategies (i.e., feature fusion and confidence fusion), which can be combined to boost the classification accuracy of region-based detectors. Comprehensive experiments demonstrate that our HCE framework is flexible and generalizable, leading to significant and consistent improvements upon various region-based detectors, including FPN, Cascade R-CNN, Mask R-CNN and PA-FPN. With simple modification, our HCE framework can be conveniently adapted to fit the structure of one-stage detectors, and achieve improved performance for SSD, RetinaNet and EfficientDet.
Xin Jin 0023, Borui Zhao, Xiaoqin Zhang 0002, Yanwen Guo 0001
IEEE Trans. Image Process.3
2020 Hierarchical Context Embedding for Region-Based Object Detection
Xin Jin 0023, Borui Zhao, Xiu-Shen Wei, Yanwen Guo 0001
ECCV (21)3