EDBT 2026 Demo / reviewers in the wild / expert
Qiming Zhang 0001
dblp:47/11423-1
· DBLP profile ↗
22ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-0060-0543ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code GenerationabstractJiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao, Jing Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi Zhang 0080, Qiming Zhang 0001, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao 0001, Jing Zhang 0037 |
ACL (1) | 4 |
| 2025 | MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target DetectionabstractIn the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available. Mingjin Zhang, Yuanjun Ouyang, Fei Gao 0006, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037 |
AAAI | 5 |
| 2025 | Semi-supervised Infrared Small Target Detection with Thermodynamic-Inspired Uneven Perturbation and Confidence AdaptationabstractSingle-frame Infrared Small Target (SIRST) detection has made significant advancements, but it still faces challenges due to limited labeled data and the foreground-background class imbalance. To address these issues, we introduce a novel Semi-Supervised SIRST Detection (S^3D) pipeline in this paper. First, drawing inspiration from thermodynamics, we propose augmenting infrared images using both chromatically and spatially uneven perturbations. This dual-stream perturbation enhances the diversity and balance of infrared samples, contributing to the robustness of detection models. Additionally, we develop a confidence-adaptive matching method to maintain weighted consistency among perturbed unlabeled samples. Second, to tackle class imbalance in labeled data, we compel the model to generate discriminative predictions for challenging, misclassified examples while down-weighting well-classified examples. We achieve this by modifying the standard cross-entropy loss to squeeze the detector and truncating the loss on well-classified examples. Our innovative Truncated Squeeze (TS) loss focuses on learning discriminative representations for difficult cases and prevents over-optimization for simpler ones. To assess the effectiveness of the perturbation techniques and loss functions, we apply them to various SIRST detectors and conduct comprehensive experiments on two benchmark datasets. Notably, our proposed methods consistently and significantly improve accuracy. Remarkably, our approach achieves over 98% performance of the state-of-the-art fully-supervised method using only 1/8 of the labeled samples. Mingjin Zhang, Wenteng Shang, Fei Gao 0006, Qiming Zhang 0001, Fengqin Lu, Jing Zhang 0037 |
AAAI | 4 |
| 2025 | Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive ImagingabstractThe objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method. Mingjin Zhang, Longyi Li, Fei Gao 0006, Qiming Zhang 0001, Jie Guo 0009 |
IJCAI | 4 |
| 2025 | Computational Fluid Dynamic Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) aims to identify and locate small targets amidst background noise. It is highly valuable in various practical application domains, such as maritime rescue and early warning systems deployed in challenging conditions such as harsh weather, low illumination, and long imaging distances. Different from existing works that either adopt well-designed backbone networks or devise specific modules to improve them from different aspects, in this article, we formulate the learning process of IRSTD from a novel perspective, i.e., the mechanism of pixel movement. Considering that the movement of pixels passing through the layers of the network for IRSTD can be analogized to the flow of particles in a fluid dynamic system, we propose a computational fluid dynamic network (CFD-Net) derived from computational fluid dynamics. Technically, we leverage the superiority of the unilateral difference equation with third-order accuracy and devise a unilateral differential residual structure as the backbone of CFD-Net. This design ensures that the pixel stream only flows in the forward direction. In addition, a switch-controlled multidirectional treatment tank (SMTT) is introduced to CFD-Net to dynamically guide the pixel stream to the appropriate path for different targets with varying shapes and orientations, facilitating learning robust target representation and improving detection performance. The proposed CFD-Net is evaluated on the IRSTD-1k and SIRST datasets and is found to outperform existing state-of-the-art (SOTA) methods. Mingjin Zhang, Ke Yue, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionabstractMulti-view camera-based 3D object detection has become popular due to its low cost, but accurately inferring 3D geometry solely from camera data remains challenging and may lead to inferior performance. Although distilling precise 3D geometry knowledge from LiDAR data could help tackle this challenge, the benefits of LiDAR information could be greatly hindered by the significant modality gap between different sensory modalities. To address this issue, we propose a Simulated multi-modal Distillation (SimDistill) method by carefully crafting the model architecture and distillation strategy. Specifically, we devise multi-modal architectures for both teacher and student models, including a LiDAR-camera fusion-based teacher and a simulated fusion-based student. Owing to the ``identical'' architecture design, the student can mimic the teacher to generate multi-modal features with merely multi-view images as input, where a geometry compensation module is introduced to bridge the modality gap. Furthermore, we propose a comprehensive multi-modal distillation scheme that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space. Incorporating them together, our SimDistill can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment. Extensive experiments validate the effectiveness and superiority of SimDistill over state-of-the-art methods, achieving an improvement of 4.8% mAP and 4.1% NDS over the baseline detector. The source code will be released at https://github.com/ViTAE-Transformer/SimDistill. Haimei Zhao, Qiming Zhang 0001, Shanshan Zhao 0001, Zhe Chen 0013, Jing Zhang 0037, Dacheng Tao |
AAAI | 2 |
| 2024 | LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
Jing Zhang 0037, Di Wang 0023, Qiming Zhang 0001, Zengmao Wang, Bo Du 0001 |
IJCAI | 4 |
| 2024 | Unleashing the Power of Generic Segmentation Model: A Simple Baseline for Infrared Small Target DetectionabstractRecent advancements in deep learning have greatly advanced the field of infrared small object detection (IRSTD). Despite their remarkable success, a notable gap persists between these IRSTD methods and generic segmentation approaches in natural image domains. This gap primarily arises from the significant modality differences and the limited availability of infrared data. In this study, we aim to bridge this divergence by investigating the adaptation of generic segmentation models, such as the Segment Anything Model (SAM), to IRSTD tasks. Our investigation reveals that many generic segmentation models can achieve comparable performance to state-of-the-art IRSTD methods. However, their full potential in IRSTD remains untapped. To address this, we propose a simple, lightweight, yet effective baseline model for segmenting small infrared objects. Through appropriate distillation strategies, we empower smaller student models to outperform state-of-the-art methods, even surpassing fine-tuned teacher results. Furthermore, we enhance the model's performance by introducing a novel query design comprising dense and sparse queries to effectively encode multi-scale features. Through extensive experimentation across four popular IRSTD datasets, our model demonstrates significantly improved performance in both accuracy and throughput compared to existing approaches, surpassing SAM and Semantic-SAM by over 14 IoU on NUDT and 4 IoU on IRSTD1k. The source code and models will be released at https://github.com/O937-blip/SimIR. Mingjin Zhang, Chi Zhang 0080, Qiming Zhang 0001, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037 |
ACM Multimedia | 3 |
| 2024 | ViTPose++: Vision Transformer for Generic Body Pose EstimationabstractIn this paper, we show the surprisingly good properties of plain vision transformers for body pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training paradigm, and transferability of knowledge between models, through a simple baseline model dubbed ViTPose. ViTPose employs the plain and non-hierarchical vision transformer as an encoder to encode features and a lightweight decoder to decode body keypoints in either a top-down or a bottom-up manner. It can be scaled to 1B parameters by taking the advantage of the scalable model capacity and high parallelism, setting a new Pareto front for throughput and performance. Besides, ViTPose is very flexible regarding the attention type, input resolution, and pre-training and fine-tuning strategy. Based on the flexibility, a novel ViTPose++ model is proposed to deal with heterogeneous body keypoint categories via knowledge factorization, i.e., adopting task-agnostic and task-specific feed-forward networks in the transformer. We also demonstrate that the knowledge of large ViTPose models can be easily transferred to small ones via a simple knowledge token. Our largest single model ViTPose-G sets a new record on the MS COCO test set without model ensemble. Furthermore, our ViTPose++ model achieves state-of-the-art performance simultaneously on a series of body pose estimation tasks, including MS COCO, AI Challenger, OCHuman, MPII for human keypoint detection, COCO-Wholebody for whole-body keypoint detection, as well as AP-10K and APT-36K for animal keypoint detection, without sacrificing inference speed. Yufei Xu, Jing Zhang 0037, Qiming Zhang 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Vision Transformer With Quadrangle AttentionabstractWindow-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is data-agnostic, constrains the flexibility of transformers to adapt to objects of varying sizes, shapes, and orientations. To address this issue, we propose a novel quadrangle attention (QA) method that extends the window-based attention to a general quadrangle formulation. Our method employs an end-to-end learnable quadrangle regression module that predicts a transformation matrix to transform default windows into target quadrangles for token sampling and attention calculation, enabling the network to model various targets with different shapes and orientations and capture rich context information. We integrate QA into plain and hierarchical vision transformers to create a new architecture named QFormer, which offers minor code modifications and negligible extra computational cost. Extensive experiments on public benchmarks demonstrate that QFormer outperforms existing representative vision transformers on various vision tasks, including classification, object detection, semantic segmentation, and pose estimation. The code will be made publicly available at QFormer. Qiming Zhang 0001, Jing Zhang 0037, Yufei Xu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolutionabstractSingle hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between spectral features. This results in inadequate utilization of spectral information and artifacts after upsampling. To address this issue, we propose ES-SAformer, an ESSA attention-embedded Transformer network for single-HSI-SR with an iterative refining structure. Specifically, we first introduce a robust and spectral-friendly similarity metric, i.e., the spectral correlation coefficient of the spectrum (SCC), to replace the original attention matrix and incorporates inductive biases into the model to facilitate training. Built upon it, we further utilize the kernelizable attention technique with theoretical support to form a novel efficient SCC-kernel-based self-attention (ESSA) and reduce attention computation to linear complexity. ESSA enlarges the receptive field for features after upsampling without bringing much computation and allows the model to effectively utilize spatial-spectral information from different scales, resulting in the generation of more natural high-resolution images. Without the need for pretraining on large-scale datasets, our experiments demonstrate ESSA’s effectiveness in both visual quality and quantitative results. The code will be released at ESSAformer. Mingjin Zhang, Chi Zhang 0080, Qiming Zhang 0001, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037 |
ICCV | 3 |
| 2023 | ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond
Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2023 | Advancing Plain Vision Transformer Toward Remote Sensing Foundation ModelabstractLarge-scale vision foundation models have made significant progress in visual tasks on natural images, with vision transformers (ViTs) being the primary choice due to their good scalability and representation ability. However, large-scale models in remote sensing (RS) have not yet been sufficiently explored. In this article, we resort to plain ViTs with about 100 million parameters and make the first attempt to propose large vision models tailored to RS tasks and investigate how such large models perform. To handle the large sizes and objects of arbitrary orientations in RS images, we propose a new rotated varied-size window attention to replace the original full attention in transformers, which can significantly reduce the computational cost and memory footprint while learning better object representation by extracting rich context from the generated diverse windows. Experiments on detection tasks show the superiority of our model over all state-of-the-art models, achieving 81.24% mean average precision (mAP) on the DOTA-V1.0 dataset. The results of our models on downstream classification and segmentation tasks also show competitive performance compared to existing advanced methods. Further experiments show the advantages of our models in terms of computational complexity and data efficiency in transferring. The code and models will be released athttps://github.com/ViTAE-Transformer/Remote-Sensing-RVSA. Di Wang 0023, Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Boost Spectrum Prediction With Temporal-Frequency Fusion Network via Transfer LearningabstractModeling and predicting the radio spectrum is vital for spectrum management, such as spectrum sharing and anomaly detection. Nevertheless, the precise spectrum prediction is challenging due to the interference from both intra-spectrum and external factors. To tackle these complex internal and external correlations, we develop a model named TF$^2$AN, consisting of three components: 1) a robust signal detection algorithm based on image processing, 2) an attention-based Long Short-term Memory network to capture the temporal-frequency correlations, 3) a generalized fusion module to take the heterogeneous external factors into account. This structure shows prominent effectiveness for spectrum prediction on a single monitoring station with sufficient data. However, when the data derived from a single station is insufficient, the performance of the deep learning model will decline a lot. Considering that more than one monitoring station is deployed in practice, the new challenge becomes how to enhance our model by leveraging the data from multiple stations or frequency bands. Therefore, we further propose T-TF$^2$AN, a transfer learning-based framework for data augmentation and knowledge sharing in spectrum prediction. Compared to TF$^2$AN, better performance is achieved. Besides, the model interpretability and training efficiency are also discussed with two case studies, respectively. Kehan Li 0001, Chao Li 0062, Jiming Chen 0001, Qiming Zhang 0001, Zebo Liu, Shibo He |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | RegionCL: Exploring Contrastive Region Pairs for Self-supervised Representation Learning
Yufei Xu, Qiming Zhang 0001, Jing Zhang 0037, Dacheng Tao |
ECCV (33) | 2 |
| 2022 | VSA: Learning Varied-Size Window Attention in Vision Transformers
Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Dacheng Tao |
ECCV (25) | 1 |
| 2022 | ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationabstractAlthough no specific domain knowledge is considered in the design, plain vision transformers have shown excellent performance in visual recognition tasks. However, little effort has been made to reveal the potential of such simple structures for pose estimation tasks. In this paper, we show the surprisingly good capabilities of plain vision transformers for pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training paradigm, and transferability of knowledge between models, through a simple baseline model called ViTPose. Specifically, ViTPose employs plain and non-hierarchical vision transformers as backbones to extract features for a given person instance and a lightweight decoder for pose estimation. It can be scaled up from 100M to 1B parameters by taking the advantages of the scalable model capacity and high parallelism of transformers, setting a new Pareto front between throughput and performance. Besides, ViTPose is very flexible regarding the attention type, input resolution, pre-training and finetuning strategy, as well as dealing with multiple pose tasks. We also empirically demonstrate that the knowledge of large ViTPose models can be easily transferred to small ones via a simple knowledge token. Experimental results show that our basic ViTPose model outperforms representative methods on the challenging MS COCO Keypoint Detection benchmark, while the largest model sets a new state-of-the-art. The code and models are available at https://github.com/ViTAE-Transformer/ViTPose. Yufei Xu, Jing Zhang 0037, Qiming Zhang 0001, Dacheng Tao |
NeurIPS | 3 |
| 2021 | ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasabstractTransformers have shown great potential in various computer vision tasks owing to their strong capability in modeling long-range dependency using the self-attention mechanism. Nevertheless, vision transformers treat an image as 1D sequence of visual tokens, lacking an intrinsic inductive bias (IB) in modeling local visual structures and dealing with scale variance. Alternatively, they require large-scale training data and longer training schedules to learn the IB implicitly. In this paper, we propose a new Vision Transformer Advanced by Exploring intrinsic IB from convolutions, i.e., ViTAE. Technically, ViTAE has several spatial pyramid reduction modules to downsample and embed the input image into tokens with rich multi-scale context by using multiple convolutions with different dilation rates. In this way, it acquires an intrinsic scale invariance IB and is able to learn robust feature representation for objects at various scales. Moreover, in each transformer layer, ViTAE has a convolution block in parallel to the multi-head self-attention module, whose features are fused and fed into the feed-forward network. Consequently, it has the intrinsic locality IB and is able to learn local features and global dependencies collaboratively. Experiments on ImageNet as well as downstream tasks prove the superiority of ViTAE over the baseline transformer and concurrent works. Source code and pretrained models will be available at https://github.com/Annbless/ViTAE. Yufei Xu, Qiming Zhang 0001, Jing Zhang 0037, Dacheng Tao |
NeurIPS | 2 |
| 2020 | Grapy-ML: Graph Pyramid Mutual Learning for Cross-Dataset Human ParsingabstractHuman parsing, or human body part semantic segmentation, has been an active research topic due to its wide potential applications. In this paper, we propose a novel GRAph PYramid Mutual Learning (Grapy-ML) method to address the cross-dataset human parsing problem, where the annotations are at different granularities. Starting from the prior knowledge of the human body hierarchical structure, we devise a graph pyramid module (GPM) by stacking three levels of graph structures from coarse granularity to fine granularity subsequently. At each level, GPM utilizes the self-attention mechanism to model the correlations between context nodes. Then, it adopts a top-down mechanism to progressively refine the hierarchical features through all the levels. GPM also enables efficient mutual learning. Specifically, the network weights of the first two levels are shared to exchange the learned coarse-granularity information across different datasets. By making use of the multi-granularity labels, Grapy-ML learns a more discriminative feature representation and achieves state-of-the-art performance, which is demonstrated by extensive experiments on the three popular benchmarks, e.g. CIHP dataset. The source code is publicly available at https://github.com/Charleshhy/Grapy-ML. Haoyu He 0001, Jing Zhang 0037, Qiming Zhang 0001, Dacheng Tao |
AAAI | 3 |
| 2019 | Structured Pruning for Efficient ConvNets via Incremental RegularizationabstractParameter pruning is a promising approach for CNN compression and acceleration by eliminating redundant model parameters with tolerable performance degrade. Despite its effectiveness, existing regularization-based parameter pruning methods usually drive weights towards zero with large and constant regularization factors, which neglects the fragility of the expressiveness of CNNs, and thus calls for a more gentle regularization scheme so that the networks can adapt during pruning. To achieve this, we propose a new and novel regularization-based pruning method, named IncReg, to incrementally assign different regularization factors to different weights based on their relative importance. Empirical analysis on CIFAR-10 dataset verifies the merits of IncReg. Further extensive experiments with popular CNNs on CIFAR-10 and ImageNet datasets show that IncReg achieves comparable to even better results compared with state-of-the-arts. Our source codes and trained models are available here: https://github.com/mingsun-tse/caffe_increg. Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Lu Yu 0003, Haoji Hu |
IJCNN | 2 |
| 2019 | Category Anchor-Guided Unsupervised Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaptation (UDA) aims to enhance the generalization capability of a certain model from a source domain to a target domain. UDA is of particular significance since no extra effort is devoted to annotating target domain samples. However, the different data distributions in the two domains, or \emph{domain shift/discrepancy}, inevitably compromise the UDA performance. Although there has been a progress in matching the marginal distributions between two domains, the classifier favors the source domain features and makes incorrect predictions on the target domain due to category-agnostic feature alignment. In this paper, we propose a novel category anchor-guided (CAG) UDA model for semantic segmentation, which explicitly enforces category-aware feature alignment to learn shared discriminative features and classifiers simultaneously. First, the category-wise centroids of the source domain features are used as guided anchors to identify the active features in the target domain and also assign them pseudo-labels. Then, we leverage an anchor-based pixel-level distance loss and a discriminative loss to drive the intra-category features closer and the inter-category features further apart, respectively. Finally, we devise a stagewise training mechanism to reduce the error accumulation and adapt the proposed model progressively. Experiments on both the GTA5$\rightarrow $Cityscapes and SYNTHIA$\rightarrow $Cityscapes scenarios demonstrate the superiority of our CAG-UDA model over the state-of-the-art methods. The code is available at \url{https://github.com/RogerZhangzz/CAG\_UDA}. Qiming Zhang 0001, Jing Zhang 0037, Wei Liu 0005, Dacheng Tao |
NeurIPS | 1 |
| 2018 | Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Haoji Hu |
BMVC | 2 |