Ajay Jaiswal

dblp:30/9707 · also Ajay Kumar Jaiswal · DBLP profile ↗
← Back
36ranked-venue papers
14as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 12 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Improving Software Project Cost Estimation and Planning Accuracy Using Genetic Algorithms and Fuzzy Logic
abstract
Accurate software project cost estimation is essential for efficient resource allocation, risk management, and scheduling. Accuracy and generality are frequently lacking in traditional estimate models. To improve prediction accuracy, this study suggests a hybrid framework (Fuzzy + GA) that combines fuzzy logic and genetic algorithms. Evaluations were performed using reference datasets from Desharnais, Kitchenham, and Maxwell with RMSE values of 0.4531, 0.0312, and 0.0416 and R squared scores of 0.7513, 0.9512, and 0.9142, respectively. The model outperformed current techniques and demonstrated notable gains.
Ajay Jaiswal, Jagdish Raikwal, Ratnesh Litoriya
Int. J. Softw. Eng. Knowl. Eng.1
2025 Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense Framework
abstract
Yuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu, Xinyu Zhao, Qi Lin, Yu Cheng, Andrew Kwong, Zhichao Cao, Tianlong Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhen Tan 0001, Ajay Jaiswal, Huaizhi Qu, Yu Cheng 0001, Andrew Kwong, Zhichao Cao 0002, Tianlong Chen 0001
EMNLP3
2025 SEBRA : Debiasing through Self-Guided Bias Ranking
abstract
Ranking samples by fine-grained estimates of spuriosity (the degree to which spurious cues are present) has recently been shown to significantly benefit bias mitigation, over the traditional binary biased-vs-unbiased partitioning of train sets. However, this spuriousity ranking comes with the requirement of human supervision. In this paper, we propose a debiasing framework based on our novel Self-Guided Bias Ranking (Sebra), that mitigates biases via an automatic ranking of data points by spuriosity within their respective classes. Sebra leverages a key local symmetry in Empirical Risk Minimization (ERM) training -- the ease of learning a sample via ERM inversely correlates with its spuriousity; the fewer spurious correlations a sample exhibits, the harder it is to learn, and vice versa. However, globally across iterations, ERM tends to deviate from this symmetry. Sebra dynamically steers ERM to correct this deviation, facilitating the sequential learning of attributes in increasing order of difficulty, ie, decreasing order of spuriosity. As a result, the sequence in which Sebra learns samples naturally provides spuriousity rankings. We use the resulting fine-grained bias characterization in a contrastive learning framework to mitigate biases from multiple sources. Extensive experiments show that Sebra consistently outperforms previous state-of-the-art unsupervised debiasing techniques across multiple standard benchmarks, including UrbanCars, BAR, and CelebA.
Adarsh K, Abhra Chaudhuri, Ajay Jaiswal, Ziquan Liu, Xiatian Zhu, Lu Yin 0006
ICLR3
2025 From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
abstract
Large Language Models (LLMs) matrices can often be expressed in low-rank format with potential to relax memory and compute resource requirements. Unlike previous works which pivot around developing novel matrix decomposition algorithms, in this work we focus to study the emerging non-uniform low-rank properties across weight matrices in LLMs through the lens of stabilizing gradient subspace. \textit{Firstly,} we provide a theoretical framework to understand the stabilization of gradient subspaces through Hessian analysis. \textit{Secondly,} we empirically establish a consequential relationship between the gradient dynamics and low-rank expressiveness of weight matrices. Our findings reveal that different LLM components exhibit varying levels of converged low-rank structure, necessitating a non-uniform rank reduction across them to minimize performance drop due to compression. In view of that, we present \textit{Weight Low-Rank Projection} \textbf{(WeLore)} that unifies weight compression and memory-efficient fine-tuning as ONE, in a data-agnostic and one-shot way. Going beyond only as a compression technique, WeLore categorizes weight matrices into Low-rank Components (LRCs) and Non-Low-rank Components (N-LRCs) based on their ability to express themselves as low-rank. Our gradient dynamics perspective illustrate that \textit{LRCs tend to have better finetuning capabilities} and their standalone finetuning can closely mimic (sometimes outperform) the training loss trajectory and performance of full-finetuning with notable memory and compute footprint reduction. All codes and checkpoints will be released.
Ajay Jaiswal, Yifan Wang 0035, Lu Yin 0006, Shiwei Liu 0003, Runjin Chen, Ananth Grama, Yuandong Tian, Zhangyang Wang
ICML1
2025 GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
abstract
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002
NeurIPS6
2025 AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
abstract
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overlooks the structural diversity of LLMs and the varying spectral properties across modules. In this paper, we introduce AlphaDecay, a simple yet effective method that adaptively assigns different weight decay strengths to each module of an LLM. Our approach is guided by Heavy-Tailed Self-Regularization (HT-SR) theory, which analyzes the empirical spectral density (ESD) of weight correlation matrices to quantify “heavy-tailedness.” Modules exhibiting more pronounced heavy-tailed ESDs, reflecting stronger feature learning, are assigned weaker decay, while modules with lighter-tailed spectra receive stronger decay. Our method leverages tailored weight decay assignments to balance the module-wise differences in spectral properties, leading to improved performance. Extensive pre-training tasks with various model sizes from 60M to 1B demonstrate that AlphaDecay achieves better perplexity and generalization than conventional uniform decay and other adaptive decay baselines. The code is available at https://github.com/hed-ucas/AlphaDecay.
Songjun Tu, Ajay Jaiswal, Li Shen 0008, Ganzhao Yuan, Shiwei Liu 0003, Lu Yin 0006
NeurIPS3
2025 Real-numbered singular matrix transformation for non-invertible and cancelable biometric templates
Onkar Singh, Ajay Jaiswal
Appl. Intell.2
2025 Random permutation-based linear regression for cancelable biometrics
abstract
Abstract The security of biometric data in biometric‐based authentication systems is a significant concern. Cancellable biometrics aim to generate templates that can be replaced by new templates if compromised. We propose a new approach for generating cancellable biometric templates based on linear regression with random permutation. Our approach generates a virtual image for every biometric image by applying linear regression. In the next step, the cancellable biometric template is produced by randomly permuting each virtual image depending on a key assigned to each individual. If the template is compromised, it can be cancelled, and a new template can be generated by altering the key. Our method has shown superior performance compared to existing random permutation‐based methods in terms of authentication accuracy across six databases, encompassing face, iris, and ear, even when dealing with low‐resolution images. It performed well on challenging databases like UBIRIS and Georgia Tech, demonstrating the robustness of the proposed approach.
Onkar Singh, Ajay Jaiswal, Nitin Kumar 0001
Expert Syst. J. Knowl. Eng.2
2024 Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
abstract
Abhinav Bandari, Lu Yin, Cheng-Yu Hsieh, Ajay Kumar Jaiswal, Tianlong Chen, Li Shen, Ranjay Krishna, Shiwei Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Abhinav Bandari, Lu Yin 0006, Cheng-Yu Hsieh, Ajay Jaiswal, Tianlong Chen 0001, Li Shen 0008, Ranjay Krishna, Shiwei Liu 0003
EMNLP4
2024 FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
abstract
Autoregressive Large Language Models (e.g., LLaMa, GPTs) are omnipresent achieving remarkable success in language understanding and generation.However, such impressive capability typically comes with a substantial model size, which presents significant challenges for autoregressive token-by-token generation.To mitigate computation overload incurred during generation, several early-exit and layer-dropping strategies have been proposed.Despite some promising success due to the redundancy across LLMs layers on metrics like Rough-L/BLUE, our careful knowledgeintensive evaluation unveils issues such as generation collapse, hallucination, and noticeable performance drop even at the trivial exit ratio of ∼ 10-15% of layers.We attribute these errors primarily to ineffective handling of the KV cache through state copying during early exit.In this work, we observe the saturation of computationally expensive feed-forward blocks of LLM layers and propose FFN-SkipLLM, which is a novel fine-grained skip strategy for autoregressive LLMs.FFN-SkipLLM leverages an input-adaptive feed-forward skipping approach that can skip ∼ 25-30% of FFN blocks of LLMs with marginal change in performance on knowledge-intensive generation tasks without any requirement to handle the KV cache.Our extensive experiments and ablation studies across benchmarks like MT-Bench, Factoid-QA, and variable-length text summarization illustrate how our simple and easy-touse method can facilitate faster autoregressive decoding.
Ajay Jaiswal, Bodun Hu, Lu Yin 0006, Yeonju Ro, Tianlong Chen 0001, Shiwei Liu 0003, Aditya Akella
EMNLP1
2024 Compressing LLMs: The Truth is Rarely Pure and Never Simple
abstract
Despite their remarkable achievements, modern Large Language Models (LLMs) encounter exorbitant computational and memory footprints. Recently, several works have shown significant success in *training-free* and *data-free* compression (pruning and quantization) of LLMs achieving 50-60\% sparsity and reducing the bit-width down to 3 or 4 bits per weight, with negligible perplexity degradation over the uncompressed baseline. As recent research efforts are focused on developing increasingly sophisticated compression methods, our work takes a step back, and re-evaluates the effectiveness of existing SoTA compression methods, which rely on a fairly simple and widely questioned metric, perplexity (even for dense LLMs). We introduce **K**nowledge-**I**ntensive **C**ompressed LLM Benchmar**K** **(LLM-KICK)**, a collection of carefully-curated tasks to re-define the evaluation protocol for compressed LLMs, which have significant alignment with their dense counterparts, and perplexity fail to capture subtle change in their true capabilities. LLM-KICK unveils many favorable merits and unfortunate plights of current SoTA compression methods: all pruning methods suffer significant performance degradation, sometimes at trivial sparsity ratios (*e.g.*, 25-30\%), and fail for N:M sparsity on knowledge-intensive tasks; current quantization methods are more successful than pruning; yet, pruned LLMs even at $\geq 50$\% sparsity are robust in-context retrieval and summarization systems; among others. LLM-KICK is designed to holistically access compressed LLMs' ability for language understanding, reasoning, generation, in-context retrieval, in-context summarization, *etc.* We hope our study can foster the development of better LLM compression methods. The reproduced codes are available at https://github.com/VITA-Group/llm-kick.
Ajay Jaiswal, Zhe Gan, Xianzhi Du, Zhangyang Wang, Yinfei Yang
ICLR1
2024 Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
Lu Yin 0006, Ajay Jaiswal, Shiwei Liu 0003, Souvik Kundu 0009, Zhangyang Wang
ICML2
2024 Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
abstract
Large Language Models (LLMs), renowned for their remarkable performance across diverse domains, present a challenge due to their colossal model size when it comes to practical deployment. In response to this challenge, efforts have been directed toward the application of traditional network pruning techniques to LLMs, uncovering a massive number of parameters can be pruned in one-shot without hurting performance. Building upon insights gained from pre-LLM models, particularly BERT-level language models, prevailing LLM pruning strategies have consistently adhered to the practice of uniformly pruning all layers at equivalent sparsity levels, resulting in robust performance. However, this observation stands in contrast to the prevailing trends observed in the field of vision models, where non-uniform layerwise sparsity typically yields substantially improved results. To elucidate the underlying reasons for this disparity, we conduct a comprehensive analysis of the distribution of token features within LLMs. In doing so, we discover a strong correlation with the emergence of outliers, defined as features exhibiting significantly greater magnitudes compared to their counterparts in feature dimensions. Inspired by this finding, we introduce a novel LLM pruning methodology that incorporates a tailored set of **non-uniform layerwise sparsity ratios** specifically designed for LLM pruning, termed as **O**utlier **W**eighed **L**ayerwise sparsity (**OWL**). The sparsity ratio of OWL is directly proportional to the outlier ratio observed within each layer, facilitating a more effective alignment between layerwise weight sparsity and outlier ratios. Our empirical evaluation, conducted across the LLaMA-V1/V2, Vicuna, OPT, and Mistral, spanning various benchmarks, demonstrates the distinct advantages offered by OWL over previous methods. For instance, OWL exhibits a remarkable performance gain, surpassing the state-of-the-art Wanda and SparseGPT by **61.22** and **6.80** perplexity at a high sparsity level of 70%, respectively, while delivering **2.6$\times$** end-to-end inference speed-up in the DeepSparse inference engine. Code is available at https://github.com/luuyin/OWL.git.
Lu Yin 0006, You Wu 0001, Zhenyu Zhang 0015, Cheng-Yu Hsieh, Yaqing Wang 0007, Yiling Jia, Gen Li 0012, Ajay Jaiswal, Mykola Pechenizkiy, Michael Bendersky, Zhangyang Wang, Shiwei Liu 0003
ICML8
2024 LLaGA: Large Language and Graph Assistant
abstract
Graph Neural Networks (GNNs) have empowered the advance in graph-structured data analysis. Recently, the rise of Large Language Models (LLMs) like GPT-4 has heralded a new era in deep learning. However, their application to graph data poses distinct challenges due to the inherent difficulty of translating graph structures to language. To this end, we introduce the the **L**arge **L**anguage **a**nd **G**raph **A**ssistant (**LLaGA**), an innovative model that effectively integrates LLM capabilities to handle the complexities of graph-structured data. LLaGA retains the general-purpose nature of LLMs while adapting graph data into a format compatible with LLM input. LLaGA achieves this by reorganizing graph nodes to structure-aware sequences and then mapping these into the token embedding space through a versatile projector. LLaGA excels in versatility, generalizability and interpretability, allowing it to perform consistently well across different datasets and tasks, extend its ability to unseen datasets or tasks, and provide explanations for graphs. Our extensive experiments across popular graph benchmarks show that LLaGA delivers outstanding performance across four datasets and three tasks using one single model, surpassing state-of-the-art graph models in both supervised and zero-shot scenarios.
Runjin Chen, Tong Zhao 0003, Ajay Jaiswal, Neil Shah, Zhangyang Wang
ICML3
2024 Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
abstract
Compressing high-capability Large Language Models (LLMs) has emerged as a favored strategy for resource-efficient inferences. While state-of-the-art (SoTA) compression methods boast impressive advancements in preserving benign task performance, the potential risks of compression in terms of safety and trustworthiness have been largely neglected. This study conducts the first, thorough evaluation of **three (3) leading LLMs** using **five (5) SoTA compression techniques** across **eight (8) trustworthiness dimensions**. Our experiments highlight the intricate interplay between compression and trustworthiness, revealing some interesting patterns. We find that quantization is currently a more effective approach than pruning in achieving efficiency and trustworthiness simultaneously. For instance, a 4-bit quantized model retains the trustworthiness of its original counterpart, but model pruning significantly degrades trustworthiness, even at 50% sparsity. Moreover, employing quantization within a moderate bit range could unexpectedly improve certain trustworthiness dimensions such as ethics and fairness. Conversely, extreme quantization to very low bit levels (3 bits) tends to reduce trustworthiness significantly. This increased risk cannot be uncovered by looking at benign performance alone, in turn, mandating comprehensive trustworthiness evaluation in practice. These findings culminate in practical recommendations for simultaneously achieving high utility, efficiency, and trustworthiness in LLMs. Code and models are available at https://decoding-comp-trust.github.io.
Junyuan Hong, Jinhao Duan, Zhangheng Li, Chulin Xie, Kelsey Lieberman, James Diffenderfer, Brian R. Bartoldson, Ajay Jaiswal, Kaidi Xu, Bhavya Kailkhura, Dan Hendrycks, Dawn Song, Zhangyang Wang, Bo Li 0026
ICML9
2024 Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At Once
abstract
Sparse Neural Networks (SNNs) have received voluminous attention for mitigating the explosion in computational costs and memory footprints of modern deep neural networks. Despite their popularity, most state-of-the-art training approaches seek to find a single high-quality sparse subnetwork with a preset sparsity pattern and ratio, making them inadequate to satiate platform and resource variability. Recently proposed approaches attempt to jointly train multiple subnetworks (we term as “sparse co-training") with a fixed sparsity pattern, to allow switching sparsity ratios subject to resource requirements. In this work, we take one more step forward and expand the scope of sparse co-training to cover diverse sparsity patterns and multiple sparsity ratios at once. We introduce Sparse Cocktail, the first sparse co-training framework that co-trains a suite of sparsity patterns simultaneously, loaded with multiple sparsity ratios which facilitate harmonious switch across various sparsity patterns and ratios at inference depending on the hardware availability. More specifically, Sparse Cocktail alternatively trains subnetworks generated from different sparsity patterns with a gradual increase in sparsity ratios across patterns and relies on an unified mask generation process and the Dense Pivot Co-training to ensure the subnetworks of different patterns orchestrate their shared parameters without canceling each other’s performance. Experiment results on image classification, object detection, and instance segmentation illustrate the favorable effectiveness and flexibility of Sparse Cocktail, pointing to a promising direction for sparse co-training. Codes will be released.
Zhangheng Li, Shiwei Liu 0003, Tianlong Chen 0001, Ajay Jaiswal, Zhenyu Zhang 0015, Dilin Wang, Raghuraman Krishnamoorthi, Shiyu Chang, Zhangyang Wang
ICML4
2024 FarSight: A Physics-Driven Whole-Body Biometric System at Large Distance and Altitude
abstract
Whole-body biometric recognition is an important area of research due to its vast applications in law enforcement, border security, and surveillance. This paper presents the end-to-end design, development and evaluation of FarSight, an innovative software system designed for whole-body (fusion of face, gait and body shape) biometric recognition. FarSight accepts videos from elevated platforms and drones as input and outputs a candidate list of identities from a gallery. The system is designed to address several challenges, including (i) low-quality imagery, (ii) large yaw and pitch angles, (iii) robust feature extraction to accommodate large intra-person variabilities and large inter-person similarities, and (iv) the large domain gap between training and test sets. FarSight combines the physics of imaging and deep learning models to enhance image restoration and biometric feature encoding. We test FarSight’s effectiveness using the newly acquired IARPA Biometric Recognition and Identification at Altitude and Range (BRIAR) dataset. Notably, FarSight demonstrated a substantial performance increase on the BRIAR dataset, with gains of +11.82% Rank-20 identification and +11.30% TAR@1% FAR.
Feng Liu 0037, Ryan Ashbaugh, Nicholas Chimitt, Najmul Hassan, Ali Hassani 0001, Ajay Jaiswal, Zhiyuan Mao, Christopher Perry, Yiyang Su, Pegah Varghaei, Kai Wang 0058, Stanley H. Chan, Arun Ross, Humphrey Shi, Zhangyang Wang, Xiaoming Liu 0002
WACV6
2024 Bug severity prediction using LDA and sentiment scores: A CNN approach
abstract
Abstract The crucial part of the software development cycle is software maintenance. The demands included in the software management are fault fixes and request to change or bring a new feature. If priority is not given to these demands, then it may lead to customer dissatisfaction, inefficient planning, and software failure as well. Therefore, it is important to study the severity of the bug reports to maintain the efficiency of the software. Various research has been conducted in the past to predict the severity of the paper using text mining focusing only on the content of the bug reports. The sentiment of the user while reporting a bug also plays a vital role. In this study, we will be focusing on two aspects, that is, sentiment and content to improve the prediction. We propose a prediction model based on LDA to study the content aspect and emotion analysis to study the sentiment aspect. The model is validated on the datasets collected from the Eclipse project using Convolutional Neural Network (CNN). The results show that the CNN model effectively utilizes the content and sentiment aspect of the data to handle the severity prediction. CNN has weight sharing feature that decreases the number of parameters used for training. It also improves generalization and overfitting is avoided The Accuracy, Precision, Recall, and F‐measure are improved when both aspects are taken into account rather than considering only content.
Ritu Bibyan, Sameer Anand, Ajay Jaiswal, Anu G. Aggarwal 0001
Expert Syst. J. Knowl. Eng.3
2024 Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
Gregory Holste, Yiliang Zhou, Song Wang 0026, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang 0003, Dongkyun Kim, Trong-Hieu Nguyen Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jong Bin Ryu, Feng Hong 0004, Arsh Verma, Yosuke Yamagishi, Hyeryeong Seo, Myungjoo Kang, Leo A. Celi, Zhiyong Lu, Ronald M. Summers, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.4
2023 Physics-Driven Turbulence Image Restoration with Stochastic Refinement
abstract
Image distortion by atmospheric turbulence is a stochastic degradation, which is a critical problem in long-range optical imaging systems. A number of research has been conducted during the past decades, including model-based and emerging deep-learning solutions with the help of synthetic data. Although fast and physics-grounded simulation tools have been introduced to help the deep-learning models adapt to real-world turbulence conditions recently, the training of such models only relies on the synthetic data and ground truth pairs. This paper proposes the Physics-integrated Restoration Network (PiRN) to bring the physics-based simulator directly into the training process to help the network to disentangle the stochasticity from the degradation and the underlying image. Furthermore, to overcome the "average effect" introduced by deterministic models and the domain gap between the synthetic and real-world degradation, we further introduce PiRN with Stochastic Refinement (PiRN-SR) to boost its perceptual quality. Overall, our PiRN and PiRN-SR improve the generalization to real-world unknown turbulence conditions and provide a state-of-the-art restoration in both pixel-wise accuracy and perceptual quality. Our codes are available at https://github.com/VITA-Group/PiRN.
Ajay Jaiswal, Xingguang Zhang, Stanley H. Chan, Zhangyang Wang
ICCV1
2023 Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers
Tianlong Chen 0001, Zhenyu Zhang 0015, Ajay Jaiswal, Shiwei Liu 0003, Zhangyang Wang
ICLR3
2023 Sparsity May Cry: Let Us Fail (Current) Sparse Neural Networks Together!
Shiwei Liu 0003, Tianlong Chen 0001, Zhenyu Zhang 0015, Xuxi Chen, Tianjin Huang, Ajay Jaiswal, Zhangyang Wang
ICLR6
2023 Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate Communication
abstract
Graphs are omnipresent and GNNs are a powerful family of neural networks for learning over graphs. Despite their popularity, scaling GNNs either by deepening or widening suffers from prevalent issues of $\textit{unhealthy gradients, over-smoothening, information squashing}$, which often lead to sub-standard performance. In this work, we are interested in exploring a principled way to scale GNNs capacity without deepening or widening, which can improve its performance across multiple small and large graphs. Motivated by the recent intriguing phenomenon of model soups, which suggest that fine-tuned weights of multiple large-language pre-trained models can be merged to a better minima, we argue to exploit the fundamentals of model soups to mitigate the aforementioned issues of memory bottleneck and trainability during GNNs scaling. More specifically, we propose not to deepen or widen current GNNs, but instead present $\textbf{first data-centric perspective}$ of model soups to build powerful GNNs by dividing giant graph data to build independently and parallelly trained multiple comparatively weaker GNNs without any intermediate communication, and $\textit{combining their strength}$ using a greedy interpolation soup procedure to achieve state-of-the-art performance. Moreover, we provide a wide variety of model soup preparation techniques by leveraging state-of-the-art graph sampling and graph partitioning approaches that can handle large graph data structures. Our extensive experiments across many real-world small and large graphs, illustrate the effectiveness of our approach and point towards a promising orthogonal direction for GNN scaling. Codes are available at: https://github.com/VITA-Group/graph_ladling
Ajay Jaiswal, Shiwei Liu 0003, Tianlong Chen 0001, Ying Ding 0001, Zhangyang Wang
ICML1
2023 Instant Soup: Cheap Pruning Ensembles in A Single Pass Can Draw Lottery Tickets from Large Models
abstract
Large pre-trained transformers have been receiving explosive attention in the past few years, due to their acculturation for numerous downstream applications via fine-tuning, but their exponentially increasing parameter counts are becoming a primary hurdle to even just fine-tune them without industry-standard hardware. Recently, Lottery Ticket Hypothesis (LTH) and its variants, have been exploited to prune these large pre-trained models generating subnetworks which can achieve similar performance as their dense counterparts, but LTH pragmatism is enormously inhibited by repetitive full training and pruning routine of iterative magnitude pruning (IMP) which worsens with increasing model size. Motivated by the recent observations of model soups, which suggest that fine-tuned weights of multiple models can be merged to a better minima, we propose **Instant Soup Pruning (ISP)** to generate lottery ticket quality subnetworks, using a fraction of the original IMP cost by replacing the expensive intermediate pruning stages of IMP with computationally efficient weak mask generation and aggregation routine. More specifically, during the mask generation stage, ISP takes a small handful of iterations using varying training protocols and data subsets to generate many weak and noisy subnetworks, and superpose them to average out the noise creating a high-quality denoised subnetwork. Our extensive experiments and ablation on two popular large-scale pre-trained models: $\texttt{CLIP} (unexplored in pruning till date)$ and $\texttt{BERT}$ across multiple benchmark vision $\texttt{\{MNIST, SVHN, Cars, GTSRB, CIFAR-10, CIFAR-100\}}$ and language datasets $\texttt{\{MNLI, QNLI, QQP, SST, ...\}}$ validate the effectiveness of ISP compared to several state-of-the-art pruning methods. Additionally, we show that ISP can be easily modified with minimal overhead to produce benefits comparable to model soups, without the prerequisite to generate multiple candidates fine-tuned models. Codes are available at: https://github.com/VITA-Group/instant_soup.
Ajay Jaiswal, Shiwei Liu 0003, Tianlong Chen 0001, Ying Ding 0001, Zhangyang Wang
ICML1
2023 Outline, Then Details: Syntactically Guided Coarse-To-Fine Code Generation
abstract
For a complicated algorithm, its implementation by a human programmer usually starts with outlining a rough control flow followed by iterative enrichments, eventually yielding carefully generated syntactic structures and variables in a hierarchy. However, state-of-the-art large language models generate codes in a single pass, without intermediate warm-ups to reflect the structured thought process of "outline-then-detail". Inspired by the recent success of chain-of-thought prompting, we propose ChainCoder, a program synthesis language model that generates Python code progressively, i.e. from coarse to fine in multiple passes. We first decompose source code into layout frame components and accessory components via abstract syntax tree parsing to construct a hierarchical representation. We then reform our prediction target into a multi-pass objective, each pass generates a subsequence, which is concatenated in the hierarchy. Finally, a tailored transformer architecture is leveraged to jointly encode the natural language descriptions and syntactically aligned I/O data samples. Extensive evaluations show that ChainCoder outperforms state-of-the-arts, demonstrating that our progressive generation eases the reasoning procedure and guides the language model to generate higher-quality solutions. Our codes are available at: https://github.com/VITA-Group/ChainCoder.
Wenqing Zheng, S. P. Sharan, Ajay Jaiswal, Yihan Xi, Dejia Xu, Zhangyang Wang
ICML3
2023 How Does Pruning Impact Long-Tailed Multi-label Medical Image Classifiers?
Gregory Holste, Ziyu Jiang, Ajay Jaiswal, Maria Hanna, Shlomo Minkowitz, Alan C. Legasto, Joanna G. Escalon, Sharon Steinberger, Mark Bittman, Thomas C. Shen, Ying Ding 0001, Ronald M. Summers, George Shih, Yifan Peng 0002, Zhangyang Wang
MICCAI (5)3
2023 The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that Matter
abstract
Large pre-trained transformers are $\textit{show-stealer}$ in modern-day deep learning, and it becomes crucial to comprehend the parsimonious patterns that exist within them as they grow in scale. With exploding parameter counts, Lottery Ticket Hypothesis (LTH) and its variants, have lost their pragmatism in sparsifying them due to high computation and memory bottleneck of repetitive $\textit{train-prune-retrain}$ routine of iterative magnitude pruning (IMP) which worsens with increasing model size. In this paper, we comprehensively study $\textit{induced sparse patterns}$ across multiple large pre-trained vision and language transformers. We propose the existence of -- $\textbf{essential sparsity}$ defined with a $\textbf{sharp dropping point}$ beyond which the performance declines much faster w.r.t the rise of sparsity level, when we directly remove weights with the smallest magnitudes in $\textbf{one-shot}$. We also present an intriguing emerging phenomenon of $\textbf{abrupt sparsification}$ during the pre-training of BERT, i.e., BERT suddenly becomes heavily sparse in pre-training after certain iterations. Moreover, our observations also indicate a $\textbf{counter-intuitive}$ finding that BERT trained with a larger amount of pre-training data tends to have a better ability to condense knowledge in comparatively relatively fewer parameters. Lastly, we investigate the effect of the pre-training loss on essential sparsity and discover that self-supervised learning (SSL) objectives trigger stronger emergent sparsification properties than supervised learning (SL). All our codes will be publicly available.
Ajay Jaiswal, Shiwei Liu 0003, Tianlong Chen 0001, Zhangyang Wang
NeurIPS1
2023 Attend Who is Weak: Pruning-assisted Medical Image Localization under Sophisticated and Implicit Imbalances
abstract
Deep neural networks (DNNs) have rapidly become a de facto choice for medical image understanding tasks. However, DNNs are notoriously fragile to the class imbalance in image classification. We further point out that such imbalance fragility can be amplified when it comes to more sophisticated tasks such as pathology localization, as imbalances in such problems can have highly complex and often implicit forms of presence. For example, different pathology can have different sizes or colors (w.r.t.the background), different underlying demographic distributions, and in general different difficulty levels to recognize, even in a meticulously curated balanced distribution of training data. In this paper, we propose to use pruning to automatically and adaptively identify hard-to-learn (HTL) training samples, and improve pathology localization by attending them explicitly, during training in supervised, semi-supervised, and weakly-supervised settings. Our main inspiration is drawn from the recent finding that deep classification models have difficult-to-memorize samples and those may be effectively exposed through network pruning [15] - and we extend such observation beyond classification for the first time. We also present an interesting demographic analysis which illustrates HTLs ability to capture complex demographic imbalances. Our extensive experiments on the Skin Lesion Localization task in multiple training settings by paying additional attention to HTLs show significant improvement of localization performance by ~2-3%.
Ajay Jaiswal, Tianlong Chen 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001, Zhangyang Wang
WACV1
2022 Single Frame Atmospheric Turbulence Mitigation: A Benchmark Study and a New Physics-Inspired Transformer Model
Zhiyuan Mao, Ajay Jaiswal, Zhangyang Wang, Stanley H. Chan
ECCV (19)2
2022 RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy Medical Imaging
abstract
AI-powered Medical Imaging has recently achieved enormous attention due to its ability to provide fast-paced healthcare diagnoses. However, it usually suffers from a lack of high-quality datasets due to high annotation cost, interobserver variability, human annotator error, and errors in computer-generated labels. Deep learning models trained on noisy labelled datasets are sensitive to the noise type and lead to less generalization on the unseen samples. To address this challenge, we propose a Robust Stochastic Knowledge Distillation (RoS-KD) framework which mimics the notion of learning a topic from multiple sources to ensure deterrence in learning noisy information. More specifically, RoS-KD learns a smooth, well-informed, and robust student manifold by distilling knowledge from multiple teachers trained on overlapping subsets of training data. Our extensive experiments on popular medical imaging classification tasks (cardiopulmonary disease and lesion classification) using real-world datasets, show the performance benefit of RoS-KD, its ability to distill knowledge from many popular large networks (ResNet-50, DenseNet-121, MobileNetV2) in a comparatively small network, and its robustness to adversarial attacks (PGD, FSGM). More specifically, RoS-KD achieves >2% and > 4% improvement on F1-score for lesion classification and cardiopulmonary disease classification tasks, respectively, when the underlying student is ResNet-18 against recent competitive knowledge distillation baseline. Additionally, on cardiopulmonary disease classification task, RoS-KD outperforms most of the SOTA baselines by ~1% gain in AUC score.
Ajay Jaiswal, Kumar Ashutosh, Justin F. Rousseau, Yifan Peng 0002, Zhangyang Wang, Ying Ding 0001
ICDM1
2022 Training Your Sparse Neural Network Better with Any Mask
abstract
Pruning large neural networks to create high-quality, independently trainable sparse masks, which can maintain similar performance to their dense counterparts, is very desirable due to the reduced space and time complexity. As research effort is focused on increasingly sophisticated pruning methods that leads to sparse subnetworks trainable from the scratch, we argue for an orthogonal, under-explored theme: improving training techniques for pruned sub-networks, i.e. sparse training. Apart from the popular belief that only the quality of sparse masks matters for sparse training, in this paper we demonstrate an alternative opportunity: one can carefully customize the sparse training techniques to deviate from the default dense network training protocols, consisting of introducing “ghost" neurons and skip connections at the early stage of training, and strategically modifying the initialization as well as labels. Our new sparse training recipe is generally applicable to improving training from scratch with various sparse masks. By adopting our newly curated techniques, we demonstrate significant performance gains across various popular datasets (CIFAR-10, CIFAR-100, TinyImageNet), architectures (ResNet-18/32/104, Vgg16, MobileNet), and sparse mask options (lottery ticket, SNIP/GRASP, SynFlow, or even randomly pruning), compared to the default training protocols, especially at high sparsity levels. Codes will be publicly available.
Ajay Jaiswal, Tianlong Chen 0001, Ying Ding 0001, Zhangyang Wang
ICML1
2022 Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great Again
abstract
Despite the enormous success of Graph Convolutional Networks (GCNs) in modeling graph-structured data, most of the current GCNs are shallow due to the notoriously challenging problems of over-smoothening and information squashing along with conventional difficulty caused by vanishing gradients and over-fitting. Previous works have been primarily focused on the study of over-smoothening and over-squashing phenomena in training deep GCNs. Surprisingly, in comparison with CNNs/RNNs, very limited attention has been given to understanding how healthy gradient flow can benefit the trainability of deep GCNs. In this paper, firstly, we provide a new perspective of gradient flow to understand the substandard performance of deep GCNs and hypothesize that by facilitating healthy gradient flow, we can significantly improve their trainability, as well as achieve state-of-the-art (SOTA) level performance from vanilla-GCNs. Next, we argue that blindly adopting the Glorot initialization for GCNs is not optimal, and derive a topology-aware isometric initialization scheme for vanilla-GCNs based on the principles of isometry. Additionally, contrary to ad-hoc addition of skip-connections, we propose to use gradient-guided dynamic rewiring of vanilla-GCNs with skip connections. Our dynamic rewiring method uses the gradient flow within each layer during training to introduce on-demand skip-connections adaptively. We provide extensive empirical evidence across multiple datasets that our methods improve gradient flow in deep vanilla-GCNs and significantly boost their performance to comfortably compete and outperform many fancy state-of-the-art methods. Codes are available at: https://github.com/VITA-Group/GradientGCN.
Ajay Jaiswal, Peihao Wang, Tianlong Chen 0001, Justin F. Rousseau, Ying Ding 0001, Zhangyang Wang
NeurIPS1
2022 Fiber Bragg grating sensors driven structural health monitoring by using multimedia-enabled iot and big data technology
Ambarish G. Mohapatra, Jaideep Talukdar, Tarini Ch. Mishra, Sameer Anand, Ajay Jaiswal, Ashish Khanna, Deepak Gupta 0002
Multim. Tools Appl.5
2021 Using Radiomics as Prior Knowledge for Thorax Disease Classification and Localization in Chest X-rays
Yan Han 0001, Chongyan Chen, Liyan Tang, Mingquan Lin, Ajay Jaiswal, Song Wang 0026, Ahmed H. Tewfik, George Shih, Ying Ding 0001, Yifan Peng 0002
AMIA5
2021 SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
abstract
Computer-aided diagnosis plays a salient role in more accessible and accurate cardiopulmonary diseases classification and localization on chest radiography. Millions of people get affected and die due to these diseases without an accurate and timely diagnosis. Recently proposed contrastive learning heavily relies on data augmentation, especially positive data augmentation. However, generating clinically-accurate data augmentations for medical images is extremely difficult because the common data augmentation methods in computer vision, such as sharp, blur, and crop operations, can severely alter the clinical settings of medical images. In this paper, we proposed a novel and simple data augmentation method based on patient metadata and supervised knowledge to create clinically accurate positive and negative augmentations for chest X-rays. We introduce an end-to-end framework, SCALP, which extends the self-supervised contrastive approach to a supervised setting. Specifically, SCALP pulls together chest X-rays from the same patient (positive keys) and pushes apart chest X-rays from different patients (negative keys). In addition, it uses ResNet-50 along with the triplet-attention mechanism to identify cardiopulmonary diseases, and Grad-CAM++ to highlight the abnormal regions. Our extensive experiments demonstrate that SCALP outperforms existing baselines with significant margins in both classification and localization tasks. Specifically, the average classification AUCs improve from 82.8% (SOTA using DenseNet-121) to 83.9% (SCALP using ResNet-50), while the localization results improve on average by 3.7% over different IoU thresholds.
Ajay Jaiswal, Cyprian Zander, Yan Han 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001
ICDM1
2012 A Hybrid of Principal Component Analysis and Partial Least Squares for Face Recognition across Pose
Ajay Jaiswal, Nitin Kumar 0001, R. K. Agrawal 0001
CIARP1