Haofei Zhang

dblp:270/0826 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 17 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
abstract
Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated TPG, existing image-to-image models appear to be natural candidates for this task. However, when applied directly, these methods often produce unfaithful results, failing to preserve fine-grained details due to feature confusion between complex textile patterns and the inherent non-rigid texture distortions in clothing images. In this paper, we propose a novel method, SLDDM-TPG, for faithful and high-fidelity TPG. Our method consists of two stages: (1) a latent disentangled network (LDN) that resolves feature confusion in clothing representations and constructs a multi-dimensional, independent clothing feature space; and (2) a semi-supervised latent diffusion model (S-LDM), which receives guidance signals from LDN and generates faithful results through semi-supervised diffusion training, combined with our designed fine-grained alignment strategy. Extensive evaluations show that SLDDM-TPG reduces FID by 4.1 and improves SSIM by up to 0.116 on our CTP-HD dataset, and also demonstrate good generalization on the VITON-HD dataset.
Chenggong Hu, Yi Wang 0068, Mengqi Xue, Haofei Zhang, Jie Song 0011
AAAI4
2026 D3-RSMDE: 40× Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
abstract
Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy and efficiency. Although using Vision Transformer (ViT) backbones for dense prediction is fast, they often exhibit poor perceptual quality. Conversely, diffusion models offer high fidelity but at a prohibitive computational cost. To overcome these limitations, we propose Depth Detail Diffusion for Remote Sensing Monocular Depth Estimation (D³-RSMDE), an efficient framework designed to achieve an optimal balance between speed and quality. Our framework first leverages a ViT-based module to rapidly generate a high-quality preliminary depth map construction, which serves as a structural prior, effectively replacing the time-consuming initial structure generation stage of diffusion models. Based on this prior, we propose a Progressive Linear Blending Refinement (PLBR) strategy, which uses a lightweight U-Net to refine the details in only a few iterations. The entire refinement step operates efficiently in a compact latent space supported by a Variational Autoencoder (VAE). Extensive experiments demonstrate that D³-RSMDE achieves a notable 11.85% reduction in the Learned Perceptual Image Patch Similarity (LPIPS) perceptual metric over leading models like Marigold, while also achieving over a 40× speedup in inference and maintaining VRAM usage comparable to lightweight ViT models.
Zunlei Feng, Haofei Zhang, Mingli Song, Jie Song 0011
AAAI4
2026 Enhancing Attention Patterns in Vision Transformers for Robustness
Haofei Zhang, Hanyang Yuan, Haoze Jiang, Jiacong Hu, Shengxuming Zhang, Mingli Song
ICIC (7)2
2026 RAIN: An embarrassingly simple approach to debiasing attribution evaluation
Jiarui Duan, Haofei Zhang, Mengqi Xue, Huiqiong Wang, Mingli Song
Comput. Vis. Image Underst.3
2026 A three-stage model for infrared small target detection with spatial and semantic feature fusion
Sixiang Ji, Haofei Zhang, Jingmin Zhang, Chun Fei, Xiaoyang Wang 0005, Juanxiu Liu, Ping Zhang 0023
Expert Syst. Appl.2
2025 LFBC: A Lifecycle-Managed False Bubble Flow Control Scheme for Torus Networks
Haofei Zhang, Youmeng Li
ICA3PP (1)1
2025 From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition
abstract
Multilingual automatic speech recognition (ASR) models greatly facilitate recognizing low-resource languages by sharing representations across similar languages. However, the commonly adopted modeling units, e.g., character-level modeling, lack language-specific information, resulting in a susceptible word prediction to phonemes and characters. Recently, subword-level modeling has demonstrated significant effectiveness for monolingual automatic recognition systems, while it is adverse to cross-lingual feature sharing. In this paper, we propose a novel low-resource ASR method that leverages the advantages of two different modeling units. Specifically, a character-level ASR model is trained on the multilingual dataset for modeling the short-term speech and learning general speech knowledge from relevant languages. Afterwards, we convert the character-level prediction into subwords for learning contextual information of the target language. Extensive experiments on Uyghur with Kazakh and Kyrgyz as auxiliary languages have shown that our proposed method significantly reduces word error rate (WER).
Haofei Zhang, Huiqiong Wang, Mingli Song
ICASSP2
2025 Dataset Ownership Verification in Contrastive Pre-trained Models
abstract
High-quality open-source datasets, which necessitate substantial efforts for curation, has become the primary catalyst for the swift progress of deep learning. Concurrently, protecting these datasets is paramount for the well-being of the data owner. Dataset ownership verification emerges as a crucial method in this domain, but existing approaches are often limited to supervised models and cannot be directly extended to increasingly popular unsupervised pre-trained models. In this work, we propose the first dataset ownership verification method tailored specifically for self-supervised pre-trained models by contrastive learning. Its primary objective is to ascertain whether a suspicious black-box backbone has been pre-trained on a specific unlabeled dataset, aiding dataset owners in upholding their rights. The proposed approach is motivated by our empirical insights that when models are trained with the target dataset, the unary and binary instance relationships within the embedding space exhibit significant variations compared to models trained without the target dataset. We validate the efficacy of this approach across multiple contrastive pre-trained models including SimCLR, BYOL, SimSiam, MOCO v3, and DINO. The results demonstrate that our method rejects the null hypothesis with a $p$-value markedly below $0.05$, surpassing all previous methodologies. Our code is available at https://github.com/xieyc99/DOV4CL.
Yuechen Xie, Mengqi Xue, Haofei Zhang, Xingen Wang, Bingde Hu, Genlang Chen, Mingli Song
ICLR4
2025 Reinforced Model Merging
abstract
The success of large language models has garnered widespread attention for model merging techniques, especially training-free methods which combine model capabilities within the parameter space. However, two challenges remain: (1) uniform treatment of all parameters leads to performance degradation; (2) search-based algorithms are often inefficient. In this paper, we present an innovative framework termed Reinforced Model Merging (RMM), which encompasses an environment and agent tailored for merging tasks. These components interact to execute layer-wise merging actions, aiming to search the optimal merging architecture. Notably, RMM operates without any gradient computations on the original models, rendering it feasible for edge devices. Furthermore, by utilizing data subsets during the evaluation process, we addressed the bottleneck in the reward feedback phase, thereby accelerating RMM by up to 100 times. Extensive experiments demonstrate that RMM achieves state-of-the-art performance across various vision and NLP datasets and effectively overcomes the limitations of the existing baseline methods. Our code is available at https://github.com/WuDiHJQ/Reinforced-Model-Merging.
Jingwen Ye, Shunyu Liu 0001, Haofei Zhang, Jie Song 0011, Zunlei Feng, Mingli Song
ICME4
2025 Coordinate-aware thermal infrared tracking via natural language modeling
Miao Yan, Ping Zhang 0023, Haofei Zhang, Ruqian Hao, Juanxiu Liu, Xiaoyang Wang 0005
Expert Syst. Appl.3
2025 From One to Many: Portable Model Construction With Independent Network Units
abstract
Artificial Intelligence of Things (AIoT) devices are highly versatile, operating in diverse environments, which necessitates local fine-tuning of deployed models to ensure compatibility with specific conditions. However, traditional fine-tuning methods often rely on cloud-based collaborative training, which is impractical for complete on-device deployment and incurs high re-tuning costs when the environment changes. To address these challenges, we propose a portable and modular approach named Portable model construction with Independent Network Units (POINT), which enables efficient customization for diverse tasks on the AIoT devices. POINT combines cloud-based model management with feature reuse to efficiently adapt pre-trained models using minimal local data. This enables offline finetuning directly on end-side devices without relying on extensive computational resources. Compared to Parameter-Efficient Fine-Tuning (PEFT) approaches like Low-Rank Adaptation (LoRA), POINT reduces the number of trainable parameters by up to 98.94%. During deployment, task-relevant models are selected from a centralized cloud repository and integrated as the backbone of the target model. Only a lightweight task-specific head is trained for downstream tasks, making POINT exceptionally lightweight and suitable for end-side customization. To mitigate performance degradation as the number of selected models increases, POINT incorporates a dynamic optimization strategy, which balances resource constraints and model performance by adaptively managing the learning process. Extensive experiments demonstrate that POINT achieves competitive results with only 100 images per category and 20 epochs of training, significantly reducing computational and resource requirements. Compared to traditional methods, POINT offers an efficient, scalable, and resource-conscious solution for deploying AI models in diverse and constrained AIoT environments.
Zhaocheng Lu, Haofei Zhang, Jiabin Xia, Jingwen Ye, Mingli Song
IEEE Internet Things J.2
2025 A Survey of Neural Trees: Co-Evolving Neural Networks and Decision Trees
abstract
Neural networks (NNs) and decision trees (DTs) are both popular models of machine learning, yet coming with mutually exclusive advantages and limitations. To bring the best of the two worlds, a variety of approaches are proposed to integrate NNs and DTs explicitly or implicitly. In this survey, these approaches are organized in a school which we term neural trees (NTs). This survey aims to present a comprehensive review of NTs and explore in detail how they enhance the model interpretability. Our first contribution is a detailed taxonomy of NTs, which characterizes the seamless integration and co-evolution of NNs and DTs. Subsequently, we analyze NTs in terms of their interpretability and performance and suggest potential solutions to the remaining challenges. Finally, this survey concludes with a discussion about other considerations like conditional computation and promising directions toward this field. A list of papers reviewed in this survey, along with their corresponding codes, is available at: https://github.com/ zju-vipa/awesome-neural-trees.
Haoling Li, Jie Song 0011, Mengqi Xue, Haofei Zhang, Mingli Song
IEEE Trans. Neural Networks Learn. Syst.4
2024 On the Concept Trustworthiness in Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs), which break down the reasoning process into the input-to-concept mapping and the concept-to-label prediction, have garnered significant attention due to their remarkable interpretability achieved by the interpretable concept bottleneck. However, despite the transparency of the concept-to-label prediction, the mapping from the input to the intermediate concept remains a black box, giving rise to concerns about the trustworthiness of the learned concepts (i.e., these concepts may be predicted based on spurious cues). The issue of concept untrustworthiness greatly hampers the interpretability of CBMs, thereby hindering their further advancement. To conduct a comprehensive analysis on this issue, in this study we establish a benchmark to assess the trustworthiness of concepts in CBMs. A pioneering metric, referred to as concept trustworthiness score, is proposed to gauge whether the concepts are derived from relevant regions. Additionally, an enhanced CBM is introduced, enabling concept predictions to be made specifically from distinct parts of the feature map, thereby facilitating the exploration of their related regions. Besides, we introduce three modules, namely the cross-layer alignment (CLA) module, the cross-image alignment (CIA) module, and the prediction alignment (PA) module, to further enhance the concept trustworthiness within the elaborated CBM. The experiments on five datasets across ten architectures demonstrate that without using any concept localization annotations during training, our model improves the concept trustworthiness by a large margin, meanwhile achieving superior accuracy to the state-of-the-arts. Our code is available at https://github.com/hqhQAQ/ProtoCBM.
Qihan Huang, Jie Song 0011, Haofei Zhang, Mingli Song
AAAI4
2024 Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
abstract
Due to the inability to receive signals from the Global Navigation Satellite System (GNSS) in extreme conditions, achieving accurate and robust navigation for Unmanned Aerial Vehicles (UAVs) is a challenging task. Recently emerged, vision-based navigation has been a promising and feasible alternative to GNSS-based navigation. However, existing vision-based techniques are inadequate in addressing flight deviation caused by environmental disturbances and inaccurate position predictions in practical settings. In this paper, we present a novel angle robustness navigation paradigm to deal with flight deviation in point-to-point navigation tasks. Additionally, we propose a model that includes the Adaptive Feature Enhance Module, Cross-knowledge Attention-guided Module and Robust Task-oriented Head Module to accurately predict direction angles for high-precision navigation. To evaluate the vision-based navigation methods, we collect a new dataset termed as UAV_AR368. Furthermore, we design the Simulation Flight Testing Instrument (SFTI) using Google Earth to simulate different flight environments, thereby reducing the expenses associated with real flight testing. Experiment results demonstrate that the proposed model outperforms the state-of-the-art by achieving improvements of 26.0% and 45.6% in the success rate of arrival under ideal and disturbed circumstances, respectively.
Zunlei Feng, Haofei Zhang, Yang Gao 0001, Jie Lei 0002, Mingli Song
AAAI3
2024 RS-SAM: Integrating Multi-scale Information for Enhanced Remote Sensing Image Segmentation
Enkai Zhang, Anda Cao, Haofei Zhang, Huiqiong Wang, Mingli Song
ACCV (8)5
2024 On the Evaluation Consistency of Attribution-Based Explanations
Jiarui Duan, Haoling Li, Haofei Zhang, Hao Jiang 0014, Mengqi Xue, Mingli Song, Jie Song 0011
ECCV (70)3
2024 BDFC:A New Flow Control Mechanism for Torus Networks
Haofei Zhang, Youmeng Li
ICA3PP (6)1
2024 ProtoPFormer: Concentrating on Prototypical Parts in Vision Transformers for Interpretable Image Recognition
Mengqi Xue, Qihan Huang, Haofei Zhang, Jie Song 0011, Mingli Song, Canghong Jin
IJCAI3
2024 LG-CAV: Train Any Concept Activation Vector with Language Guidance
abstract
Concept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV.
Qihan Huang, Jie Song 0011, Mengqi Xue, Haofei Zhang, Bingde Hu, Huiqiong Wang, Hao Jiang 0014, Xingen Wang, Mingli Song
NeurIPS4
2023 Generalization Matters: Loss Minima Flattening via Parameter Hybridization for Efficient Online Knowledge Distillation
abstract
Most existing online knowledge distillation (OKD) techniques typically require sophisticated modules to produce diverse knowledge for improving students' generalization ability. In this paper, we strive to fully utilize multi-model settings instead of well-designed modules to achieve a distillation effect with excellent generalization performance. Generally, model generalization can be reflected in the flatness of the loss landscape. Since averaging parameters of multiple models can find flatter minima, we are inspired to extend the process to the sampled convex combinations of multi-student models in OKD. Specifically, by linearly weighting students' parameters in each training batch, we construct a Hybrid-Weight Model (HWM) to represent the parameters surrounding involved students. The supervision loss of HWM can estimate the landscape's curvature of the whole region around students to measure the generalization explicitly. Hence we integrate HWM's loss into students' training and propose a novel OKD framework via parameter hybridization (OKDPH) to promote flatter minima and obtain robust solutions. Considering the redundancy of parameters could lead to the collapse of HWM, we further introduce a fusion operation to keep the high similarity of students. Compared to the state-of-the-art (SOTA) OKD methods and SOTA methods of seeking flat minima, our OKDPH achieves higher performance with fewer parameters, benefiting OKD with lightweight and robust characteristics. Our code is publicly available at https://github.com/tianlizhang/OKDPH.
Tianli Zhang, Mengqi Xue, Haofei Zhang, Yu Wang 0176, Lechao Cheng, Jie Song 0011, Mingli Song
CVPR4
2023 Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype Networks
abstract
Part-prototype networks (e.g., ProtoPNet, ProtoTree, and ProtoPool) have attracted broad research interest for their intrinsic interpretability and comparable accuracy to non-interpretable counterparts. However, recent works find that the interpretability from prototypes is fragile, due to the semantic gap between the similarities in the feature space and that in the input space. In this work, we strive to address this challenge by making the first attempt to quantitatively and objectively evaluate the interpretability of the part-prototype networks. Specifically, we propose two evaluation metrics, termed as "consistency score" and "stability score", to evaluate the explanation consistency across images and the explanation robustness against perturbations, respectively, both of which are essential for explanations taken into practice. Furthermore, we propose an elaborated part-prototype network with a shallow-deep feature alignment (SDFA) module and a score aggregation (SA) module to improve the interpretability of prototypes. We conduct systematical evaluation experiments and provide substantial discussions to uncover the interpretability of existing part-prototype networks. Experiments on three benchmarks across nine architectures demonstrate that our model achieves significantly superior performance to the state of the art, in both the accuracy and interpretability. Our code is available at https://github.com/hqhQAQ/EvalProtoPNet.
Qihan Huang, Mengqi Xue, Wenqi Huang 0002, Haofei Zhang, Jie Song 0011, Yongcheng Jing, Mingli Song
ICCV4
2023 Schema Inference for Interpretable Image Classification
Haofei Zhang, Mengqi Xue, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song
ICLR1
2023 Improving Expressivity of GNNs with Subgraph-specific Factor Embedded Normalization
abstract
Graph Neural Networks~(GNNs) have emerged as a powerful category of learning architecture for handling graph-structured data. However, existing GNNs typically ignore crucial structural characteristics in node-induced subgraphs, which thus limits their expressiveness for various downstream tasks. In this paper, we strive to strengthen the representative capabilities of GNNs by devising a dedicated plug-and-play normalization scheme, termed as SUbgraph-sPEcific FactoR Embedded Normalization (SuperNorm), that explicitly considers the intra-connection information within each node-induced subgraph. To this end, we embed the subgraph-specific factor at the beginning and the end of the standard BatchNorm, as well as incorporate graph instance-specific statistics for improved distinguishable capabilities. In the meantime, we provide theoretical analysis to support that, with the elaborated SuperNorm, an arbitrary GNN is at least as powerful as the 1-WL test in distinguishing non-isomorphism graphs. Furthermore, the proposed SuperNorm scheme is also demonstrated to alleviate the over-smoothing phenomenon. Experimental results related to predictions of graph, node, and link properties on the eight popular datasets demonstrate the effectiveness of the proposed method. The code is available at https://github.com/chenchkx/SuperNorm.
Kai-Xuan Chen 0001, Shunyu Liu 0001, Tongtian Zhu, Ji Qiao, Yingjie Tian 0002, Tongya Zheng, Haofei Zhang, Zunlei Feng, Jingwen Ye, Mingli Song
KDD8
2023 Constituent Attention for Vision Transformers
Haoling Li, Mengqi Xue, Jie Song 0011, Haofei Zhang, Wenqi Huang 0002, Lingyu Liang, Mingli Song
Comput. Vis. Image Underst.4
2023 Knowledge Amalgamation for Object Detection With Transformers
abstract
Knowledge amalgamation (KA) is a novel deep model reusing task aiming to transfer knowledge from several well-trained teachers to a multi-talented and compact student. Currently, most of these approaches are tailored for convolutional neural networks (CNNs). However, there is a tendency that Transformers, with a completely different architecture, are starting to challenge the domination of CNNs in many computer vision tasks. Nevertheless, directly applying the previous KA methods to Transformers leads to severe performance degradation. In this work, we explore a more effective KA scheme for Transformer-based object detection models. Specifically, considering the architecture characteristics of Transformers, we propose to dissolve the KA into two aspects: sequence-level amalgamation (SA) and task-level amalgamation (TA). In particular, a hint is generated within the sequence-level amalgamation by concatenating teacher sequences instead of redundantly aggregating them to a fixed-size one as previous KA approaches. Besides, the student learns heterogeneous detection tasks through soft targets with efficiency in the task-level amalgamation. Extensive experiments on PASCAL VOC and COCO have unfolded that the sequence-level amalgamation significantly boosts the performance of students, while the previous methods impair the students. Moreover, the Transformer-based students excel in learning amalgamated knowledge, as they have mastered heterogeneous detection tasks rapidly and achieved superior or at least comparable performance to those of the teachers in their specializations.
Haofei Zhang, Feng Mao, Mengqi Xue, Gongfan Fang, Zunlei Feng, Jie Song 0011, Mingli Song
IEEE Trans. Image Process.1
2022 Up to 100x Faster Data-Free Knowledge Distillation
abstract
Data-free knowledge distillation (DFKD) has recently been attracting increasing attention from research communities, attributed to its capability to compress a model only using synthetic data. Despite the encouraging results achieved, state-of-the-art DFKD methods still suffer from the inefficiency of data synthesis, making the data-free training process extremely time-consuming and thus inapplicable for large-scale tasks. In this work, we introduce an efficacious scheme, termed as FastDFKD, that allows us to accelerate DFKD by a factor of orders of magnitude. At the heart of our approach is a novel strategy to reuse the shared common features in training data so as to synthesize different data instances. Unlike prior methods that optimize a set of data independently, we propose to learn a meta-synthesizer that seeks common features as the initialization for the fast data synthesis. As a result, FastDFKD achieves data synthesis within only a few steps, significantly enhancing the efficiency of data-free training. Experiments over CIFAR, NYUv2, and ImageNet demonstrate that the proposed FastDFKD achieves 10x and even 100x acceleration while preserving performances on par with state of the art. Code is available at https://github.com/zju-vipa/Fast-Datafree.
Gongfan Fang, Kanya Mo, Xinchao Wang, Jie Song 0011, Shitao Bei, Haofei Zhang, Mingli Song
AAAI6
2022 Meta-attention for ViT-backed Continual Learning
abstract
Continual learning is a longstanding research topic due to its crucial role in tackling continually arriving tasks. Up to now, the study of continual learning in computer vision is mainly restricted to convolutional neural networks (CNNs). However, recently there is a tendency that the newly emerging vision transformers (ViTs) are gradually dominating the field of computer vision, which leaves CNN-based continual learning lagging behind as they can suffer from severe performance degradation if straightforwardly applied to ViTs. In this paper, we study ViT-backed continual learning to strive for higher performance riding on recent advances of ViTs. Inspired by mask-based continual learning methods in CNNs, where a mask is learned per task to adapt the pre-trained ViT to the new task, we propose MEta-ATtention (MEAT), i.e., attention to self-attention, to adapt a pre-trained ViT to new tasks without sacrificing performance on already learned tasks. Unlike prior mask-based methods like Piggyback, where all parameters are associated with corresponding masks, MEAT leverages the characteristics of ViTs and only masks a portion of its parameters. It renders MEAT more efficient and effective with less overhead and higher accuracy. Extensive experiments demonstrate that MEAT exhibits significant superiority to its state-of-the-art CNN counterparts, with 4.0 ∼ 6.0% absolute boosts in accuracy. Our code has been released at https://github.com/zju-vipa/MEAT-TIL.
Mengqi Xue, Haofei Zhang, Jie Song 0011, Mingli Song
CVPR2
2022 Bootstrapping ViTs: Towards Liberating Vision Transformers from Pre-training
abstract
Recently, vision Transformers (ViTs) are developing rapidly and starting to challenge the domination of con-volutional neural networks (CNNs) in the realm of computer vision (CV). With the general-purpose Transformer architecture replacing the hard-coded inductive biases of convolution, ViTs have surpassed CNNs, especially in data-sufficient circumstances. However, ViTs are prone to over-fit on small datasets and thus rely on large-scale pre-training, which expends enormous time. In this paper, we strive to liberate ViTs from pre-training by introducing CNNs' in- ductive biases back to ViTs while preserving their network architectures for higher upper bound and setting up more suitable optimization objectives. To begin with, an agent CNN is designed based on the given ViT with inductive bi-ases. Then a bootstrapping training algorithm is proposed to jointly optimize the agent and ViT with weight sharing, during which the ViT learns inductive biases from the intermediate features of the agent. Extensive experiments on CIFAR-10/100 and ImageNet-1k with limited training data have shown encouraging results that the inductive biases help ViTs converge significantly faster and outperform conventional CNNs with even fewer parameters. Our code is publicly available at https://github.com/zhfeing/Bootstrapping-ViTs-pytorch.
Haofei Zhang, Jiarui Duan, Mengqi Xue, Jie Song 0011, Mingli Song
CVPR1
2022 Learn decision trees with deep visual primitives
Mengqi Xue, Haofei Zhang, Qihan Huang, Jie Song 0011, Mingli Song
J. Vis. Commun. Image Represent.2
2021 Tree-Like Decision Distillation
abstract
Knowledge distillation pursues a diminutive yet well-behaved student network by harnessing the knowledge learned by a cumbersome teacher model. Prior methods achieve this by making the student imitate shallow behaviors, such as soft targets, features, or attention, of the teacher. In this paper, we argue that what really matters for distillation is the intrinsic problem-solving process captured by the teacher. By dissecting the decision process in a layer-wise manner, we found that the decision-making procedure in the teacher model is conducted in a coarse-to-fine manner, where coarse-grained discrimination (e.g., animal vs vehicle) is attained in early layers, and fine-grained dis-crimination (e.g., dog vs cat, car vs truck) in latter layers. Motivated by this observation, we propose a new distillation method, dubbed as Tree-like Decision Distillation (TDD), to endow the student with the same problem-solving mechanism as that of the teacher. Extensive experiments demonstrated that TDD yields competitive performance compared to state of the arts. More importantly, it enjoys better interpretability due to its interpretable decision distillation instead of dark knowledge distillation.
Jie Song 0011, Haofei Zhang, Xinchao Wang, Mengqi Xue, Dacheng Tao, Mingli Song
CVPR2
2021 Secrecy-Oriented Optimization of Sparse Code Multiple Access for Simultaneous Wireless Information and Power Transfer in 6G Aerial Access Networks
abstract
This article focuses on the simultaneous wireless information and power transfer (SWIPT) systems, which provide both the power supply and the communications for Internet‐of‐Things (IoT) devices in the sixth‐generation (6G) network. Due to the extremely stringent requirements on reliability, speed, and security in the 6G network, aerial access networks (AANs) are deployed to extend the coverage of wireless communications and guarantee robustness. Moreover, sparse code multiple access (SCMA) is implemented on the SWIPT system to further promote the spectrum efficiency. To improve the speed and security of SWIPT systems in 6G AANs, we have developed an optimization algorithm of SCMA to maximize the secrecy sum rate (SSR). Specifically, a power‐splitting (PS) strategy is applied by each user to coordinate its energy harvesting and information decoding. Hence, the SSR maximization problems in the SCMA system are formulated in terms of the PS and resource allocation, under the constraints on the minimum rates and minimum harvested energy of individual users. Then, a successive convex approximation method is introduced to transform the nonconvex problems to the convex ones, which are then solved by an iterative algorithm. In addition, we investigate the SSR performance of the SCMA system supported by our optimization methods, when the impacts from different perspectives are considered. Our studies and simulation results show that the SCMA system supported by our proposed optimization algorithms significantly outperforms the legacy system with uniform power allocation and fixed PS.
Jingmin Zhang, Xiaokui Yue, Haofei Zhang, Tao Ni 0005, Wensheng Lin
Wirel. Commun. Mob. Comput.4