Maying Shen

dblp:195/2178 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
12since 2021 · last 2025
0009-0000-9416-680XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
abstract
Vision language models (VLMs) respond to user-crafted text prompts and visual inputs, and are applied to numerous real-world problems. VLMs integrate visual modalities with large language models (LLMs), which are well known to be prompt-sensitive. Hence, it is crucial to determine whether VLMs inherit this instability to varying prompts. We therefore investigate which prompt variations VLMs are most sensitive to and which VLMs are most agnostic to prompt variations. To this end, we introduce PARC (Prompt Analysis via Reliability and Calibration), a VLM prompt sensitivity analysis framework built on three pillars: (1) plausible prompt variations in both the language and vision domain, (2) a novel model reliability score with built-in guarantees, and (3) a calibration step that enables dataset-and prompt-spanning prompt variation analysis. Regarding prompt variations, PARC’s evaluation shows that VLMs mirror LLM language prompt sensitivity in the vision domain, and most destructive variations change the expected answer. Regarding models, outstandingly robust VLMs among 22 evaluated models come from the InternVL2 family. We further find indications that prompt sensitivity is linked to training data. https://github.com/NVlabs/PARC
Jenny Schmalfuss, Nadine Chang, Vibashan VS, Maying Shen, Andrés Bruhn, José M. Álvarez 0004
CVPR4
2025 MDP: Multidimensional Vision Model Pruning with Latency Constraint
abstract
Current structural pruning methods face two significant limitations: (i) they often limit pruning to finer-grained levels like channels, making aggressive parameter reduction challenging, and (ii) they focus heavily on parameter and FLOP reduction, with existing latency-aware methods frequently relying on simplistic, suboptimal linear models that fail to generalize well to transformers, where multiple interacting dimensions impact latency. In this paper, we address both limitations by introducing Multi-Dimensional Pruning(MDP), a novel paradigm that jointly optimizes across a variety of pruning granularities—including channels, query/key, heads, embeddings, and blocks. MDP employs an advanced latency modeling technique to accurately capture latency variations across all prunable dimensions, achieving an optimal balance between latency and accuracy. By reformulating pruning as a Mixed-Integer Nonlinear Program (MINLP), MDP efficiently identifies the optimal pruned structure across all prunable dimensions while respecting latency constraints. This versatile framework supports both CNNs and transformers. Extensive experiments demonstrate that MDP significantly outperforms previous methods, especially at high pruning ratios. On ImageNet, MDP achieves a 28% speed increase with a +1.4 Top-1 accuracy improvement over prior work like HALP for ResNet50 pruning. Against the latest transformer pruning method, Isomorphic, MDP delivers an additional 37% acceleration with a +0.7 Top-1 accuracy improvement.
Xinglong Sun, Barath Lakshmanan, Maying Shen, Shiyi Lan, Jingde Chen, José M. Álvarez 0004
CVPR3
2025 SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation
Maying Shen, Nadine Chang, Sifei Liu, José M. Álvarez 0004
KDD (1)1
2025 Advancing Weight and Channel Sparsification with Enhanced Saliency
abstract
Pruning aims to accelerate and compress models by removing redundant parameters, identified by specifically designed importance scores which are usually imperfect. This removal is irreversible, often leading to subpar performance in pruned models. Dynamic sparse training, while attempting to adjust sparse structures during training for continual reassessment and refinement, has several limitations including criterion inconsistency between pruning and growth, unsuitability for structured sparsity, and short-sighted growth strategies. Our paper introduces an efficient, innovative paradigm to enhance a given importance criterion for either unstructured or structured sparsity. Our method separates the model into an active structure for exploitation and an exploration space for potential updates. During exploitation, we optimize the active structure, whereas in exploration, we reevaluate and reintegrate parameters from the exploration space through a pruning and growing step consistently guided by the same given importance criterion. To prepare for exploration, we briefly “reactivate” all parameters in the exploration space and train them for a few iterations while keeping the active part frozen, offering a preview of the potential performance gains from reintegrating these parameters. We show on various datasets and configurations that existing importance criterion even simple as magnitude can be enhanced with ours to achieve state-of-the-art performance and training cost reductions. Notably, on ImageNet with ResNet50, ours achieves an$+1.3$increase in Top-1 accuracy over prior art at 90% ERK [49] sparsity. Compared with the SOTA latency pruning method HALP [58], we reduced its training cost by over 70% while attaining a faster and more accurate pruned model.
Xinglong Sun, Maying Shen, Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004
WACV2
2024 Adaptive Sharpness-Aware Pruning for Robust Sparse Networks
abstract
Robustness and compactness are two essential attributes of deep learning models that are deployed in the real world. The goals of robustness and compactness may seem to be at odds, since robustness requires generalization across domains, while the process of compression exploits specificity in one domain. We introduce \textit{Adaptive Sharpness-Aware Pruning (AdaSAP)}, which unifies these goals through the lens of network sharpness. The AdaSAP method produces sparse networks that are robust to input variations which are \textit{unseen at training time}. We achieve this by strategically incorporating weight perturbations in order to optimize the loss landscape. This allows the model to be both primed for pruning and regularized for improved robustness. AdaSAP improves the robust accuracy of pruned models on image classification by up to +6\% on ImageNet C and +4\% on ImageNet V2, and on object detection by +4\% on a corrupted Pascal VOC dataset, over a wide range of compression ratios, pruning criteria, and network architectures, outperforming recent pruning art by large margins.
Anna Bair, Hongxu Yin, Maying Shen, Pavlo Molchanov 0001, José M. Álvarez 0004
ICLR3
2023 Global Vision Transformer Pruning with Hessian-Aware Saliency
abstract
Transformers yield state-of-the-art results across many tasks. However, their heuristically designed architecture impose huge computational costs during inference. This work aims on challenging the common design philosophy of the Vision Transformer (ViT) model with uniform dimension across all the stacked blocks in a model stage, where we redistribute the parameters both across transformer blocks and between different structures within the block via the first systematic attempt on global structural pruning. Dealing with diverse ViT structural components, we derive a novel Hessian-based structural pruning criteria comparable across all layers and structures, with latency-aware regularization for direct latency reduction. Performing iterative pruning on the DeiT-Base model leads to a new architecture family called NViT (Novel ViT), with a novel parameter redistribution that utilizes parameters more efficiently. On ImageNet-1K, NViT-Base achieves a$2.6\times FLOPs$reduction,$5.1\times$parameter reduction, and$1.9\times run$-time speedup over the DeiT-Base model in a near lossless manner. Smaller NViT variants achieve more than 1% accuracy gain at the same throughput of the DeiT Small/Tiny variants, as well as a lossless$3.3\times parameter$reduction over the SWIN-Small model. These results outperform prior art by a large margin. Further analysis is provided on the parameter redistribution insight of NViT, where we show the high prunability of ViT models, distinct sensitivity within ViT block, and unique parameter distribution trend across stacked ViT blocks. Our insights provide viability for a simple yet effective parameter redistribution rule towards more efficient ViTs for off-the-shelf performance boost.
Huanrui Yang, Hongxu Yin, Maying Shen, Pavlo Molchanov 0001, Hai Li 0001, Jan Kautz
CVPR3
2023 Augmenting Legacy Networks for Flexible Inference
abstract
On intelligent vehicles, Deep Neural Networks (DNNs) may run on devices whose computational load varies over time. Within the context of variable network architectures, that can be used to constrain the inference latency for real-time deployment with varying system resources, we introduce LeAF (Legacy Augmentation for Flexible inference), a novel paradigm to augment a pre-trained DNN with trainable, shallow execution paths that can run in place of the legacy ones. While preserving the legacy DNN weights, LeAF allows changing the DNN architecture with minimal overhead to effectively adapt to different system performance targets. LeAF-ResNet-50 has less than 14% storage overhead over the legacy DNN; its accuracy varies from the legacy 76.1% to 70.15% (up to 5% better than Slimmable [1] with a latency that is 37% better than OFA [2] on an A100 GPU with batch size 256). Our analysis shows the importance of considering not only the target device, but also the batch size and the temporal dynamic of the DNN configuration to optimize the performances of variable architecture DNNs, LeAF in particular.
Jason Clemons, Iuri Frosio, Maying Shen, José M. Álvarez 0004, Stephen W. Keckler
IV3
2023 Hardware-Aware Latency Pruning for Real-Time 3D Object Detection
abstract
3D Object detection is a fundamental task in vision-based autonomous driving. Deep learning perception models achieve an outstanding performance at the expense of continuously increasing resource needs and, as such, increasing training costs. As inference time is still a priority, developers usually adopt a training pipeline where they first start using a compact architecture that yields a good trade-off between accuracy and latency. This architecture is usually found either by searching manually or by using neural architecture search approaches. Then, train the model and use light optimization techniques such as quantization to boost the model’s performance. In contrast, in this paper, we advocate for starting on a much larger model and then applying aggressive optimization to adapt the model to the resource-constraints. Our results on large-scale settings for 3D object detection demonstrate the benefits of initially focusing on maximizing the model’s accuracy and then achieving the latency requirements using network pruning.
Maying Shen, Joshua Chen, Justin Hsu, Xinglong Sun, Oliver Knieps, Carmen Maxim, José M. Álvarez 0004
IV1
2022 When to Prune? A Policy towards Early Structural Pruning
abstract
Pruning enables appealing reductions in network memory footprint and time complexity. Conventional post-training pruning techniques lean towards efficient inference while overlooking the heavy computation for training. Recent exploration of pre-training pruning at initialization hints on training cost reduction via pruning, but suffers noticeable performance degradation. We attempt to combine the benefits of both directions and propose a policy that prunes as early as possible during training without hurting performance. Instead of pruning at initialization, our method exploits initial dense training for few epochs to quickly guide the architecture, while constantly evaluating dominant sub-networks via neuron importance ranking. This unveils dominant sub-networks whose structures turn stable, allowing conventional pruning to be pushed earlier into the training. To do this early, we further introduce an Early Pruning Indicator (EPI) that relies on sub-network architectural similarity and quickly triggers pruning when the sub-network's architecture stabilizes. Through extensive experiments on ImageNet, we show that EPI empowers a quick tracking of early training epochs suitable for pruning, offering same efficacy as an otherwise “oracle” grid-search that scans through epochs and requires orders of magnitude more compute. Our method yields 1.4% top-l accuracy boost over state-of-the-art pruning counterparts, cuts down training cost on GPU by 2.4x, hence offers a new efficiency-accuracy boundary for network pruning during training.
Maying Shen, Pavlo Molchanov 0001, Hongxu Yin, José M. Álvarez 0004
CVPR1
2022 Soft Masking for Cost-Constrained Channel Pruning
Ryan Humble, Maying Shen, Jorge Albericio Latorre, Eric Darve, José M. Álvarez 0004
ECCV (11)2
2022 Structural Pruning via Latency-Saliency Knapsack
abstract
Structural pruning can simplify network architecture and improve inference speed. We propose Hardware-Aware Latency Pruning (HALP) that formulates structural pruning as a global resource allocation optimization problem, aiming at maximizing the accuracy while constraining latency under a predefined budget on targeting device. For filter importance ranking, HALP leverages latency lookup table to track latency reduction potential and global saliency score to gauge accuracy drop. Both metrics can be evaluated very efficiently during pruning, allowing us to reformulate global structural pruning under a reward maximization problem given target constraint. This makes the problem solvable via our augmented knapsack solver, enabling HALP to surpass prior work in pruning efficacy and accuracy-efficiency trade-off. We examine HALP on both classification and detection tasks, over varying networks, on ImageNet and VOC datasets, on different platforms. In particular, for ResNet-50/-101 pruning on ImageNet, HALP improves network throughput by $1.60\times$/$1.90\times$ with $+0.3\%$/$-0.2\%$ top-1 accuracy changes, respectively. For SSD pruning on VOC, HALP improves throughput by $1.94\times$ with only a $0.56$ mAP drop. HALP consistently outperforms prior art, sometimes by large margins. Project page at \url{https://halp-neurips.github.io/}.
Maying Shen, Hongxu Yin, Pavlo Molchanov 0001, Jianna Liu, José M. Álvarez 0004
NeurIPS1
2021 Optimal Quantization Using Scaled Codebook
abstract
We study the problem of quantizing N sorted, scalar datapoints with a fixed codebook containing K entries that are allowed to be rescaled. The problem is defined as finding the optimal scaling factor α and the datapoint assignments into the α-scaled codebook to minimize the squared error between original and quantized points. Previously, the globally optimal algorithms for this problem were derived only for certain codebooks (binary and ternary) or under the assumption of certain distributions (Gaussian, Laplacian). By studying the properties of the optimal quantizer, we derive an $\mathcal{O}\left( {NK\log K} \right)$ algorithm that is guaranteed to find the optimal quantization parameters for any fixed codebook regardless of data distribution. We apply our algorithm to synthetic and real-world neural network quantization problems and demonstrate the effectiveness of our approach.
Yerlan Idelbayev, Pavlo Molchanov 0001, Maying Shen, Hongxu Yin, Miguel Á. Carreira-Perpiñán, José M. Álvarez 0004
CVPR3
2018 Anomaly detection based on Nearest Neighbor search with Locality-Sensitive B-tree
Maying Shen, Xinghao Jiang, Tanfeng Sun
Neurocomputing1
2017 Anomaly Detection by Analyzing the Pedestrian Behavior and the Dynamic Changes of Behavior
Maying Shen, Xinghao Jiang, Tanfeng Sun
ICIC (1)1