VLDB 2026 Research / reviewers in the wild / expert
Jayaraman J. Thiagarajan
dblp:16/7803
· DBLP profile ↗
95ranked-venue papers
20as first author
38since 2021 · last 2025
0000-0002-8517-5816ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 11 first-author · 19 since 2021Artificial intelligence and machine learning · 35 · 8 first-author · 20 since 2021Systems, architecture and hardware · 12 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ModelX : A Novel Transfer Learning Approach Across Heterogeneous DatasetsabstractLeveraging an existing performance model to predict the runtime of a new application on a new system can save days and weeks of data collection time. However, knowledge transfer between High Performance Computing (HPC) systems can be challenging due to data heterogeneity caused by differences in data collection methods, architectural or application-specific individuality. This results in (1) sets of performance features that have significantly different names, orders, or the number of performance features that do not match between two datasets (heterogeneous domains), or (2) distribution shifts between datasets although their feature names match (homogeneous domains). While existing transfer learning techniques can handle mild distribution shifts, they fail to transfer knowledge when the source and target features do not match. This work introduces a novel transfer learning methodology-Cross Prediction Model (ModelX), which overcomes the large distribution discrepancy between homogeneous domains and enables transfer learning between heterogeneous domains. Extensive evaluations show that ModelX outperforms traditional transfer learning methods for all experiments using 11 HPC and 4 Machine Learning (ML) datasets. To the best of our knowledge, this is the first methodology to enable knowledge transfer between two heterogeneous domains with no matching features. Finally, we demonstrate an application of ModelX to an HPC job scheduling scenario using real-world job traces where it helps to reduce the job turnaround time of a set of jobs by 71%. Arunavo Dey, Neil Antony, Aakash Dhakal, Kowshik Thopalli, Jayaraman J. Thiagarajan, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Jae-Seung Yeom, Tanzima Z. Islam |
HPDC | 5 |
| 2025 | On The Role of Prompt Construction In Enhancing Efficacy and Efficiency of LLM-Based Tabular Data GenerationabstractLLM-based data generation for real-world tabular data can be challenged by the lack of sufficient semantic context in feature names used to describe columns. We hypothesize that enriching prompts with even minimal contextual information, such as a brief explanation of what each feature represents can improve both the quality and efficiency of data generation. To test this, we investigate three prompt construction methods: Expert-guided, LLM-guided, and Novel-Mapping, with the latter two being automated approaches. Using the GReaT framework, our experiments show that context-enriched prompts significantly enhance the quality of the generated data while improving training efficiency. Notably, the LLM-guided method performed on par with expert-guided approaches, demonstrating its effectiveness as a scalable alternative. Banooqa H. Banday, Kowshik Thopalli, Tanzima Z. Islam, Jayaraman J. Thiagarajan |
ICASSP | 4 |
| 2025 | Leveraging Registers in Vision Transformers for Robust AdaptationabstractVision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence of high-norm tokens in ViTs, which can interfere with unsupervised object discovery. To address this, the use of "registers" which are additional tokens that isolate high norm patch tokens while capturing global image-level information has been proposed. While registers have been studied extensively for object discovery, their generalization properties particularly in out-of-distribution (OOD) scenarios, remains underexplored. In this paper, we examine the utility of register token embeddings in providing additional features for improving generalization and anomaly rejection. To that end, we propose a simple method that combines the special CLS token embedding commonly employed in ViTs with the average-pooled register embeddings to create feature representations which are subsequently used for training a downstream classifier. We find that this enhances OOD generalization and anomaly rejection, while maintaining in-distribution (ID) performance. Extensive experiments across multiple ViT backbones trained with and without registers reveal consistent improvements of 2-4% in top-1 OOD accuracy and a 2-3% reduction in false positive rates for anomaly detection. Importantly, these gains are achieved without additional computational overhead. Srikar Yellapragada, Kowshik Thopalli, Vivek Sivaraman Narayanaswamy, Wesam A. Sakla, Yamen Mubarka, Dimitris Samaras, Jayaraman J. Thiagarajan |
ICASSP | 8 |
| 2025 | Conformal Edge-Weight Prediction in Latent SpaceabstractPredicting the edge weights of a graph is a critical task across many domains. Some examples include predicting traffic flow in transportation networks, strength of interactions in protein-protein networks, and collaboration frequency in co-authorship networks. Graph Neural Networks have been very successful in edge-weight prediction tasks. However, these predictions lack rigorous statistical uncertainty quantification. Recent work has demonstrated the efficacy of conformal inference in quantifying the uncertainties of the predictions made by graph neural networks. However, there has been limited research in conformal inference for edge-weight prediction. Akash Choudhuri, Yongjian Zhong, Mehrdad Moharrami, Christine Klymko, Mark Heimann, Jayaraman J. Thiagarajan, Bijaya Adhikari |
SDM | 6 |
| 2025 | Structure-Aware Representation Learning for Effective Performance PredictionabstractABSTRACT Application performance is a function of several unknowns stemming from the interactions between the application, runtime, OS, and underlying hardware, making it challenging to model performance using deep learning techniques, especially without a large labeled dataset. Collecting such labeled longitudinal datasets can take weeks. Intuitively, developers could save analysis time during code development by taking a comparative approach between multiple applications. However, the unknown dynamic interactions between applications and execution environments make it difficult for deep learning‐based models to predict the performance of new applications. In this paper, we address these problems by presenting a labeled dataset for the community and taking a comparative analysis approach to explore the source code differences between different correct implementations of the same problem. This paper assesses the feasibility of using purely static information, for example, Abstract Syntax Tree (AST), of applications to predict performance change based on code structure. We evaluate several deep learning‐based representation learning techniques for source code and propose an architecture for the tree‐based Long Short‐Term Memory (LSTM) models to discover latent representations for a source code's hierarchical structure. We demonstrate that our proposed architecture enables feed‐forward predictive models to predict change in performance using source code with up to 84% accuracy. Tarek Ramadan, Nathan Pinnow, Chase Phelps, Jayaraman J. Thiagarajan, Tanzima Z. Islam |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
Rakshith Subramanyam, Kowshik Thopalli, Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan |
ECCV (79) | 4 |
| 2024 | The Double-Edged Sword Of Ai Safety: Balancing Anomaly Detection and OOD Generalization Via Model AnchoringabstractSafe deployment of AI systems requires models to accurately flag anomalous or semantically unrelated data, while also generalizing to unseen shifts in the data distribution. While both these problems have been extensively studied, there is a risk for undesirable trade-off when exclusively optimizing for one of the objectives. In this paper, we systematically study this trade-off under the lens of model anchoring. Anchoring is a recently proposed training methodology that involves reparameterizing input data into anchor-residual pairs (anchors are drawn from the training data itself), thus establishing a combinatorial relationship with other samples in the training data. We make a surprising finding that the dual objectives of generalization and anomaly detection can be controlled by independently regularizing the model’s dependency on the distribution of anchors and residuals respectively. This enables, for the first time, a finer control of the detection-generalization trade-off without requiring any additional data (e.g., outlier exposure) or computationally intensive modeling strategies (e.g., deep ensembling). Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Jayaraman J. Thiagarajan |
ICASSP | 3 |
| 2024 | Exploring the Utility of Clip Priors for Visual Relationship PredictionabstractThis work explores the challenges of leveraging large-scale vision language models, such as CLIP, for visual relationship prediction (VRP), a task vital in understanding the relations between objects in a scene based on both image features and text descriptors. Despite its potential, we find that CLIP’s language priors are restrictive in effectively differentiating between various predicates for VRP. Towards this, we present CREPE (CLIP Representation Enhanced Predicate Estimation), which utilizes learnable prompts and a unique contrastive training strategy to derive reliable CLIP representations suited for VRP. CREPE can be seamlessly integrated into any VRP method. Our evaluations on the Visual Genome benchmark illustrate that using representations from CREPE significantly enhances the performance of vanilla VRP methods, such as UVTransE and VCTree. This enhancement is notable as CREPE can be seamlessly integrated into any VRP method, even without the need for additional calibration techniques, showcasing its efficacy as a powerful solution to VRP. CREPE’s performance on the Unrel benchmark reveals strong generalization to diverse and previously unseen predicate occurrences, despite lacking explicit training on such examples. Rakshith Subramanyam, T. S. Jayram, Rushil Anirudh, Jayaraman J. Thiagarajan |
ICASSP | 4 |
| 2024 | On Estimating Link Prediction Uncertainty Using Stochastic CenteringabstractAccurate confidence estimates are crucial for safe graph neural network (GNN) deployment, yet link prediction (LP) calibration is understudied. We provide novel insights into LP calibration by highlighting the importance of meaningful node-level uncertainties. In response, we propose E-ΔUQ, an architecture-agnostic framework leveraging stochastic centering to incorporate epistemic uncertainty into GNNs. Our work provides principles and three E-ΔUQ variants to improve trust in LP models, while introducing minimal overhead. Key results demonstrate properly handling node-level uncertainty improves edge calibration. We evaluate E-ΔUQ variants on citation networks and find that intermediate stochastic layers outperform alternatives by producing better node uncertainties. E-ΔUQ reduces calibration error by 15-50% and maintains comparable prediction fidelity. Puja Trivedi, Danai Koutra, Jayaraman J. Thiagarajan |
ICASSP | 3 |
| 2024 | Accurate and Scalable Estimation of Epistemic Uncertainty for Graph Neural NetworksabstractWhile graph neural networks (GNNs) are widely used for node and graph representation learning tasks, the reliability of GNN uncertainty estimates under distribution shifts remains relatively under-explored. Indeed, while post-hoc calibration strategies can be used to improve in-distribution calibration, they need not also improve calibration under distribution shift. However, techniques which produce GNNs with better intrinsic uncertainty estimates are particularly valuable, as they can always be combined with post-hoc strategies later. Therefore, in this work, we propose G-$\Delta$UQ, a novel training framework designed to improve intrinsic GNN uncertainty estimates. Our framework adapts the principle of stochastic data centering to graph data through novel graph anchoring strategies, and is able to support partially stochastic GNNs. While, the prevalent wisdom is that fully stochastic networks are necessary to obtain reliable estimates, we find that the functional diversity induced by our anchoring strategies when sampling hypotheses renders this unnecessary and allows us to support G-$\Delta$UQ on pretrained models. Indeed, through extensive evaluation under covariate, concept and graph size shifts, we show that G-$\Delta$UQ leads to better calibrated GNNs for node and graph classification. Further, it also improves performance on the uncertainty-based tasks of out-of-distribution detection and generalization gap estimation. Overall, our work provides insights into uncertainty estimation for GNNs, and demonstrates the utility of G-$\Delta$UQ in obtaining reliable estimates. Puja Trivedi, Mark Heimann, Rushil Anirudh, Danai Koutra, Jayaraman J. Thiagarajan |
ICLR | 5 |
| 2024 | PAGER: Accurate Failure Characterization in Deep Regression ModelsabstractSafe deployment of AI models requires proactive detection of failures to prevent costly errors. To this end, we study the important problem of detecting failures in deep regression models. Existing approaches rely on epistemic uncertainty estimates or inconsistency w.r.t the training data to identify failure. Interestingly, we find that while uncertainties are necessary they are insufficient to accurately characterize failure in practice. Hence, we introduce PAGER (Principled Analysis of Generalization Errors in Regressors), a framework to systematically detect and characterize failures in deep regressors. Built upon the principle of anchored training in deep models, PAGER unifies both epistemic uncertainty and complementary manifold non-conformity scores to accurately organize samples into different risk regimes. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Puja Trivedi, Rushil Anirudh |
ICML | 1 |
| 2024 | On the Use of Anchoring for Training Vision ModelsabstractAnchoring is a recent, architecture-agnostic principle for training deep neural networks that has been shown to significantly improve uncertainty estimation, calibration, and extrapolation capabilities. In this paper, we systematically explore anchoring as a general protocol for training vision models, providing fundamental insights into its training and inference processes and their implications for generalization and safety. Despite its promise, we identify a critical problem in anchored training that can lead to an increased risk of learning undesirable shortcuts, thereby limiting its generalization capabilities. To address this, we introduce a new anchored training protocol that employs a simple regularizer to mitigate this issue and significantly enhances generalization. We empirically evaluate our proposed approach across datasets and architectures of varying scales and complexities, demonstrating substantial performance gains in generalization and safety metrics compared to the standard training protocol. The open-source code is available at https://software.llnl.gov/anchoring. Vivek Sivaraman Narayanaswamy, Kowshik Thopalli, Rushil Anirudh, Yamen Mubarka, Wesam A. Sakla, Jayaraman J. Thiagarajan |
NeurIPS | 6 |
| 2023 | Cross-GAN Auditing: Unsupervised Identification of Attribute Level Similarities and Differences Between Pretrained Generative ModelsabstractGenerative Adversarial Networks (GANs) are notoriously difficult to train especially for complex distributions and with limited data. This has driven the need for tools to audit trained networks in human intelligible format, for example, to identify biases or ensure fairness. Existing GAN audit tools are restricted to coarse-grained, modeldata comparisons based on summary statistics such as FID or recall. In this paper, we propose an alternative approach that compares a newly developed GAN against a prior baseline. To this end, we introduce Cross-GAN Auditing (xGA) that, given an established “reference” GAN and a newly proposed “client” GAN, jointly identifies intelligible attributes that are either common across both GANs, novel to the client GAN, or missing from the client GAN. This provides both users and model developers an intuitive assessment of similarity and differences between GANs. We introduce novel metrics to evaluate attribute-based GAN auditing approaches and use these metrics to demonstrate quantitatively that xGA outperforms baseline approaches. We also include qualitative results that illustrate the common, novel and missing attributes identified by xGA from GANs trained on a variety of image datasets1 Matthew L. Olson, Shusen Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Weng-Keen Wong |
CVPR | 4 |
| 2023 | Single-Shot Domain Adaptation via Target-Aware Generative AugmentationsabstractThe problem of adapting models from a source domain using data from any target domain of interest has gained prominence, thanks to the brittle generalization in deep neural networks. While several test-time adaptation techniques have emerged, they typically rely on synthetic data augmentations in cases of limited target data availability. In this paper, we consider the challenging setting of single-shot adaptation and explore the design of augmentation strategies. We argue that augmentations utilized by existing methods are insufficient to handle large distribution shifts, and hence propose a new approach SiSTA (Single-Shot Target Augmentations), which first fine-tunes a generative model from the source domain using a single-shot target, and then employs novel sampling strategies for curating synthetic target data. Using experiments with a state-of-the-art domain adaptation method, we find that SiSTA produces improvements as high as 20% over existing baselines under challenging shifts in face attribute detection, and that it performs competitively to oracle models obtained by training on a larger target dataset. Our codes can be accessed at github.com/kowshikthopalli/SISTA. Rakshith Subramanyam, Kowshik Thopalli, Spring Berman, Pavan Turaga, Jayaraman J. Thiagarajan |
ICASSP | 5 |
| 2023 | A Closer Look At Scoring Functions And Generalization PredictionabstractGeneralization error predictors (GEPs) aim to predict model performance on unseen distributions by deriving dataset-level error estimates from sample-level scores. However, GEPs often utilize disparate mechanisms (e.g., regressors, thresholding functions, calibration datasets, etc), to derive such error estimates, which can obfuscate the benefits of a particular scoring function. Therefore, in this work, we rigorously study the effectiveness of popular scoring functions (confidence, local manifold smoothness, model agreement), independent of mechanism choice. We find, absent complex mechanisms, that state-of-the-art confidence- and smoothness- based scores fail to outperform simple model-agreement scores when estimating error under distribution shifts and corruptions. Furthermore, on realistic settings where the training data has been compromised (e.g., label noise, measurement noise, under-sampling), we find that model-agreement scores continue to perform well and that ensemble diversity is important for improving its performance. Finally, to better understand the limitations of scoring functions, we demonstrate that simplicity bias, or the propensity of deep neural networks to rely upon simple but brittle features, can adversely affect GEP performance. Overall, our work carefully studies the effectiveness of popular scoring functions in realistic settings and helps to better understand their limitations. Puja Trivedi, Danai Koutra, Jayaraman J. Thiagarajan |
ICASSP | 3 |
| 2023 | DOLCE: A Model-Based Probabilistic Diffusion Framework for Limited-Angle CT ReconstructionabstractLimited-Angle Computed Tomography (LACT) is a nondestructive 3D imaging technique used in a variety of applications ranging from security to medicine. The limited angle coverage in LACT is often a dominant source of severe artifacts in the reconstructed images, making it a challenging imaging inverse problem. Diffusion models are a recent class of deep generative models for synthesizing realistic images using image denoisers. In this work, we present DOLCE as the first framework for integrating conditionally-trained diffusion models and explicit physical measurement models for solving imaging inverse problems. DOLCE achieves the SOTA performance in highly ill-posed LACT by alternating between the data-fidelity and sampling updates of a diffusion model conditioned on the transformed sinogram. We show through extensive experimentation that unlike existing methods, DOLCE can synthesize high-quality and structurally coherent 3D volumes by using only 2D conditionally pre-trained diffusion models. We further show on several challenging real LACT datasets that the same pretrained DOLCE model achieves the SOTA performance on drastically different types of images. Jiaming Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Stewart He, K. Aditya Mohan, Ulugbek Kamilov, Hyojin Kim 0001 |
ICCV | 3 |
| 2023 | A Closer Look at Model Adaptation using Feature Distortion and Simplicity Bias
Puja Trivedi, Danai Koutra, Jayaraman J. Thiagarajan |
ICLR | 3 |
| 2023 | Target-Aware Generative Augmentations for Single-Shot AdaptationabstractIn this paper, we address the problem of adapting models from a source domain to a target domain, a task that has become increasingly important due to the brittle generalization of deep neural networks. While several test-time adaptation techniques have emerged, they typically rely on synthetic toolbox data augmentations in cases of limited target data availability. We consider the challenging setting of single-shot adaptation and explore the design of augmentation strategies. We argue that augmentations utilized by existing methods are insufficient to handle large distribution shifts, and hence propose a new approach SiSTA, which first fine-tunes a generative model from the source domain using a single-shot target, and then employs novel sampling strategies for curating synthetic target data. Using experiments on a variety of benchmarks, distribution shifts and image corruptions, we find that SiSTA produces significantly improved generalization over existing baselines in face attribute detection and multi-class object recognition. Furthermore, SiSTA performs competitively to models obtained by training on larger target datasets. Our codes can be accessed at https://github.com/Rakshith-2905/SiSTA Kowshik Thopalli, Rakshith Subramanyam, Pavan Turaga, Jayaraman J. Thiagarajan |
ICML | 4 |
| 2023 | Improving Diversity with Adversarially Learned Transformations for Domain GeneralizationabstractTo be successful in single source domain generalization (SSDG), maximizing diversity of synthesized domains has emerged as one of the most effective strategies. Recent success in SSDG comes from methods that pre-specify diversity inducing image augmentations during training, so that it may lead to better generalization on new domains. However, naïve pre-specified augmentations are not always effective, either because they cannot model large domain shift, or be-cause the specific choice of transforms may not cover the types of shift commonly occurring in domain generalization. To address this issue, we present a novel framework called ALT: adversarially learned transformations, that uses an adversary neural network to model plausible, yet hard image transformations that fool the classifier. ALT learns image transformations by randomly initializing the adversary net-work for each batch and optimizing it for a fixed number of steps to maximize classification error. The classifier is trained by enforcing a consistency between its predictions on the clean and transformed images. With extensive empirical analysis, we find that this new form of adversarial transformations achieves both objectives of diversity and hardness simultaneously, outperforming all existing techniques on competitive benchmarks for SSDG. We also show that ALT can seamlessly work with existing diversity modules to produce highly distinct, and large transformations of the source domain leading to state-of-the-art performance. Code: https://github.com/tejas-gokhale/ALT Tejas Gokhale, Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Chitta Baral, Yezhou Yang |
WACV | 3 |
| 2023 | Contrastive Knowledge-Augmented Meta-Learning for Few-Shot ClassificationabstractModel agnostic meta-learning algorithms aim to infer priors from several observed tasks that can then be used to adapt to a new task with few examples. Given the inherent diversity of tasks arising in existing benchmarks, recent methods have resorted to task-specific adaptation of the prior. Our goal is to improve generalization of meta learners when the task distribution contains challenging distribution shifts and semantic disparities. To this end, we introduce CAML (Contrastive Knowledge-Augmented Meta Learning), a knowledge-enhanced few-shot learning approach that evolves a knowledge graph to encode historical experience, and employs a contrastive distillation strategy to leverage the encoded knowledge for task-aware modulation of the base learner. In addition to the standard few-shot task adaptation, we also consider the more challenging multi-domain task adaptation and few-shot dataset generalization settings in our evaluation with standard benchmarks. Our empirical study shows that CAML (i) enables simple task encoding schemes; (ii) eliminates the need for knowledge extraction at inference time; and most importantly, (iii) effectively aggregates historical experience thus leading to improved performance in both multi-domain adaptation and dataset generalization. Rakshith Subramanyam, Mark Heimann, T. S. Jayram, Rushil Anirudh, Jayaraman J. Thiagarajan |
WACV | 5 |
| 2022 | Out of Distribution Detection via Neural Network Anchoring
Rushil Anirudh, Jayaraman J. Thiagarajan |
ACML | 2 |
| 2022 | Domain Alignment Meets Fully Test-Time Adaptation
Kowshik Thopalli, Pavan Turaga, Jayaraman J. Thiagarajan |
ACML | 3 |
| 2022 | Sparsity Improves Unsupervised Attribute Discovery in StyleganabstractRich semantics exist in latent spaces inferred using deep generative models. The ability to extract and interpret them is not only essential for understanding the underlying factors of variation in the data distribution, but also crucial for con-trolled image generation. Several methods have been proposed to identify semantically meaningful linear directions, either through existing annotations, or relying on identifying directions of large variation that arise from the data representation of the network. In this paper, we identify a new criterion, representation sparsity, that allows us to produce extremely efficient yet diverse semantic directions in GAN (generative adversarial network) latent spaces. The observation also reveals a potential deeper connection between representation sparsity and semantics in deep neural networks that worth further exploration. Shusen Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICASSP | 3 |
| 2022 | Predicting the Generalization Gap in Deep Models using AnchoringabstractWe address the problem of predicting the generalization gap of deep neural networks under large, natural, and synthetic distribution shifts between source and target domains. This is crucial in understanding how models behave in uncontrollable ‘in-the-wild’ scenarios, but existing techniques fail when target domain becomes very different from the source. Accurately capturing the relationship and distance between the source and target domains is critical for a reliable post-hoc estimation of generalization. In this paper, we propose a novel strategy for directly predicting accuracy on unseen target data with the help of anchoring and pre-text encoding in predictive models. Anchoring has been shown previously to perform effectively in characterizing domain shifts, which we exploit for predicting the generalization gap. Our experiments on the PACS dataset along with synthetic ablations indicate that our approach produces well calibrated accuracy estimates outperforming existing baselines. Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Irene Kim, Yamen Mubarka, Andreas Spanias, Jayaraman J. Thiagarajan |
ICASSP | 6 |
| 2022 | Improved StyleGAN-v2 based Inversion for Out-of-Distribution ImagesabstractInverting an image onto the latent space of pre-trained generators, e.g., StyleGAN-v2, has emerged as a popular strategy to leverage strong image priors for ill-posed restoration. Several studies have showed that this approach is effective at inverting images similar to the data used for training. However, with out-of-distribution (OOD) data that the generator has not been exposed to, existing inversion techniques produce sub-optimal results. In this paper, we propose SPHInX (StyleGAN with Projection Heads for Inverting X), an approach for accurately embedding OOD images onto the StyleGAN latent space. SPHInX optimizes a style projection head using a novel training strategy that imposes a vicinal regularization in the StyleGAN latent space. To further enhance OOD inversion, SPHInX can additionally optimize a content projection head and noise variables in every layer. Our empirical studies on a suite of OOD data show that, in addition to producing higher quality reconstructions over the state-of-the-art inversion techniques, SPHInX is effective for ill-posed restoration tasks while offering semantic editing capabilities. Rakshith Subramanyam, Vivek Sivaraman Narayanaswamy, Mark Naufel, Andreas Spanias, Jayaraman J. Thiagarajan |
ICML | 5 |
| 2022 | Single Model Uncertainty Estimation via Stochastic Data CenteringabstractWe are interested in estimating the uncertainties of deep neural networks, which play an important role in many scientific and engineering problems. In this paper, we present a striking new finding that an ensemble of neural networks with the same weight initialization, trained on datasets that are shifted by a constant bias gives rise to slightly inconsistent trained models, where the differences in predictions are a strong indicator of epistemic uncertainties. Using the neural tangent kernel (NTK), we demonstrate that this phenomena occurs in part because the NTK is not shift-invariant. Since this is achieved via a trivial input transformation, we show that this behavior can therefore be approximated by training a single neural network -- using a technique that we call $\Delta-$UQ -- that estimates uncertainty around prediction by marginalizing out the effect of the biases during inference. We show that $\Delta-$UQ's uncertainty estimates are superior to many of the current methods on a variety of benchmarks-- outlier rejection, calibration under distribution shift, and sequential design optimization of black box functions. Code for $\Delta-$UQ can be accessed at github.com/LLNL/DeltaUQ Jayaraman J. Thiagarajan, Rushil Anirudh, Vivek Sivaraman Narayanaswamy, Peer-Timo Bremer |
NeurIPS | 1 |
| 2022 | Analyzing Data-Centric Properties for Graph Contrastive LearningabstractRecent analyses of self-supervised learning (SSL) find the following data-centric properties to be critical for learning good representations: invariance to task-irrelevant semantics, separability of classes in some latent space, and recoverability of labels from augmented samples. However, given their discrete, non-Euclidean nature, graph datasets and graph SSL methods are unlikely to satisfy these properties. This raises the question: how do graph SSL methods, such as contrastive learning (CL), work well? To systematically probe this question, we perform a generalization analysis for CL when using generic graph augmentations (GGAs), with a focus on data-centric properties. Our analysis yields formal insights into the limitations of GGAs and the necessity of task-relevant augmentations. As we empirically show, GGAs do not induce task-relevant invariances on common benchmark datasets, leading to only marginal gains over naive, untrained baselines. Our theory motivates a synthetic data generation process that enables control over task-relevant information and boasts pre-defined optimal augmentations. This flexible benchmark helps us identify yet unrecognized limitations in advanced augmentation techniques (e.g., automated methods). Overall, our work rigorously contextualizes, both empirically and theoretically, the effects of data-centric properties on augmentation strategies and learning paradigms for graph SSL. Puja Trivedi, Ekdeep Singh Lubana, Mark Heimann, Danai Koutra, Jayaraman J. Thiagarajan |
NeurIPS | 5 |
| 2022 | Enabling machine learning-ready HPC ensembles with Merlin
Jayson Luc Peterson, Benjamin Bay, Joe Koning, Peter B. Robinson, Jessica Semler, Jeremy White, Rushil Anirudh, Kevin Athey, Peer-Timo Bremer, Francesco Di Natale, Jim Gaffney, Sam Ade Jacobs, Bhavya Kailkhura, Bogdan Kustowski, Steve H. Langer, Brian K. Spears, Jayaraman J. Thiagarajan, Brian Van Essen, Jae-Seung Yeom |
Future Gener. Comput. Syst. | 18 |
| 2022 | Improving Single-Stage Object Detectors for Nighttime Pedestrian DetectionabstractImproving the reliability of nighttime pedestrian detection is a crucial challenge towards the design of robust autonomous systems. Not surprisingly, most pedestrian fatalities occur in low-illumination settings, thus emphasizing the need for new algorithmic advances. This work presents a novel pedestrian detection approach that makes a number of crucial modifications to the state-of-the-art YOLOV5-PANet architecture, in order to improve the reliability of features extracted from nighttime images. More specifically, the proposed architecture systematically incorporates powerful shuffle attention mechanisms and a transformer module to improve the feature learning pipeline. Instead of advocating the use of other sensing modalities that are better suited for nighttime detection, our approach relies only on conventional RGB cameras and is hence broadly applicable. Our empirical studies with nighttime pedestrian detection benchmarks show that with only minimal increase in model complexity, our approach provides significant improvements in detection efficacy over existing solutions. Finally, we explore the impact of post-hoc network pruning on the speed-accuracy trade-off of our approach and demonstrate that it is well suited for reduced memory/compute requirements. Kowshik Thopalli, Jayaraman J. Thiagarajan |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsabstractWhile existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbations (such as an unknown degree of rotation) may be known. We consider a setup where robustness is expected over an unseen test domain that is not i.i.d. but deviates from the training domain. While this deviation may not be exactly known, its broad characterization is specified a priori, in terms of attributes. We propose an adversarial training approach which learns to generate new samples so as to maximize exposure of the classifier to the attributes-space, without having access to the data from the test domain. Our adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial perturbations, and the outer minimization finding model parameters by optimizing the loss on adversarial perturbations generated from the inner maximization. We demonstrate the applicability of our approach on three types of naturally occurring perturbations --- object-related shifts, geometric transformations, and common image corruptions. Our approach enables deep neural networks to be robust against a wide range of naturally occurring perturbations. We demonstrate the usefulness of the proposed approach by showing the robustness gains of deep neural networks trained using our adversarial training on MNIST, CIFAR-10, and a new variant of the CLEVR dataset. Tejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Chitta Baral, Yezhou Yang |
AAAI | 4 |
| 2021 | Uncertainty-Matching Graph Neural Networks to Defend Against Poisoning AttacksabstractGraph Neural Networks (GNNs), a generalization of neural networks to graph-structured data, are often implemented using message passes between entities of a graph. While GNNs are effective for node classification, link prediction and graph classification, they are vulnerable to adversarial attacks, i.e., a small perturbation to the structure can lead to a non-trivial performance degradation. In this work, we propose Uncertainty Matching GNN (UM-GNN), that is aimed at improving the robustness of GNN models, particularly against poisoning attacks to the graph structure, by leveraging epistemic uncertainties from the message passing framework. More specifically, we propose to build a surrogate predictor that does not directly access the graph structure, but systematically extracts reliable knowledge from a standard GNN through a novel uncertainty-matching strategy. Interestingly, this uncoupling makes UM-GNN immune to evasion attacks by design, and achieves significantly improved robustness against poisoning attacks. Using empirical studies with standard benchmarks and a suite of global and target attacks, we demonstrate the effectiveness of UM-GNN, when compared to existing baselines including the state-of-the-art robust GCN. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Andreas Spanias |
AAAI | 2 |
| 2021 | Accurate and Robust Feature Importance Estimation under Distribution ShiftsabstractWith increasing reliance on the outcomes of black-box models in critical applications, post-hoc explainability tools that do not require access to the model internals are often used to enable humans understand and trust these models. In particular, we focus on the class of methods that can reveal the influence of input features on the predicted outputs. Despite their wide-spread adoption, existing methods are known to suffer from one or more of the following challenges: computational complexities, large uncertainties and most importantly, inability to handle real-world domain shifts. In this paper, we propose PRoFILE (Producing Robust Feature Importances using Loss Estimates), a novel feature importance estimation method that addresses all these challenges. Through the use of a loss estimator jointly trained with the predictive model and a causal objective, PRoFILE can accurately estimate the feature importance scores even under complex distribution shifts, without any additional re-training. To this end, we also develop learning strategies for training the loss estimator, namely contrastive and dropout calibration, and find that it can effectively detect distribution shifts. Using empirical studies on several benchmark image and non-image data, we show significant improvements over state-of-the-art approaches, both in terms of fidelity and robustness. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer, Andreas Spanias |
AAAI | 1 |
| 2021 | College Life is Hard! - Shedding Light on Stress Prediction for Autistic College Students using Data-Driven AnalysisabstractAutistic college students face significant challenges in college settings and have a higher dropout rate than neurotypical college students. High physiological distress, depression, and anxiety are identified as critical challenges that contribute to this less than optimal college experience. In this paper, we leverage affordable mobile and wearable devices to collect large amounts of physiological and contextual data (biomarkers) and leverage a data-driven analysis approach for building stress prediction models. Such models can be used to provide real-time intervention for better stress management. We conducted a mixed-method study where we collected physiological and contextual data from 20 college students (10 neurotypical and 10 autistic). Our proposed data-driven analysis pipeline leverages an unsupervised representation learning technique with a semi-supervised label approximation method to predict the onset of stress based on biomarkers for autistic students, neurotypical students, and both populations with accuracies 69%, 72%, and 70%, respectively. Tanzima Z. Islam, Philip Wu Liang, Forest Sweeney, Cody Pranger, Jayaraman J. Thiagarajan, Moushumi Sharmin, Shameem Ahmed |
COMPSAC | 5 |
| 2021 | Using Deep Image Priors to Generate Counterfactual ExplanationsabstractThrough the use of carefully tailored convolutional neural network architectures, a deep image prior (DIP) can be used to obtain pre-images from latent representation encodings. Though DIP inversion has been known to be superior to conventional regularized inversion strategies such as total variation, such an over-parameterized generator is able to effectively reconstruct even images that are not in the original data distribution. This limitation makes it challenging to utilize such priors for tasks such as counterfactual reasoning, wherein the goal is to generate small, interpretable changes to an image that systematically leads to changes in the model prediction. To this end, we propose a novel regularization strategy based on an auxiliary loss estimator jointly trained with the predictor, which efficiently guides the prior to re-cover natural pre-images. Our empirical studies with a real-world ISIC skin lesion detection problem clearly evidence the effectiveness of the proposed approach in synthesizing meaningful counterfactuals. In comparison, we find that the standard DIP inversion often proposes visually imperceptible perturbations to irrelevant parts of the image, thus providing no additional insights into the model behavior. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICASSP | 2 |
| 2021 | On the Design of Deep Priors for Unsupervised Audio RestorationabstractUnsupervised deep learning methods for solving audio restoration problems extensively rely on carefully tailored neural architectures that carry strong inductive biases for defining priors in the time or spectral domain. In this context, lot of recent success has been achieved with sophisticated convolutional network constructions that recover audio signals in the spectral domain. However, in practice, audio priors require careful engineering of the convolutional kernels to be effective at solving ill-posed restoration tasks, while also being easy to train. To this end, in this paper, we propose a new U-Net based prior that does not impact either the network complexity or convergence behavior of existing convolutional architectures, yet leads to significantly improved restoration. In particular, we advocate the use of carefully designed dilation schedules and dense connections in the U-Net architecture to obtain powerful audio priors. Using empirical studies on standard benchmarks and a variety of ill-posed restoration tasks, such as audio denoising, in-painting and source separation, we demonstrate that our proposed approach consistently outperforms widely adopted audio prior architectures. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Andreas Spanias |
Interspeech | 2 |
| 2021 | Comparative Code Structure Analysis using Deep Learning for Performance PredictionabstractPerformance analysis has always been an afterthought during the application development process, focusing on application correctness first. The learning curve of the existing static and dynamic analysis tools are steep, which requires understanding low-level details to interpret the findings for actionable optimizations. Additionally, application performance is a function of a number of unknowns stemming from the application-, runtime-, and interactions between the OS and underlying hardware, making it difficult to model using any deep learning technique, especially without a large labeled dataset. In this paper, we address both of these problems by presenting a large corpus of a labeled dataset for the community and take a comparative analysis approach to mitigate all unknowns except their source code differences between different correct implementations of the same problem. We put the power of deep learning to the test for automatically extracting information from the hierarchical structure of abstract syntax trees to represent source code. This paper aims to assess the feasibility of using purely static information (e.g., abstract syntax tree or AST) of applications to predict performance change based on the change in code structure. This research will enable performance-aware application development since every version of the application will continue to contribute to the corpora, which will enhance the performance of the model. We evaluate several deep learning-based representation learning techniques for source code. Our results show that tree-based Long Short-Term Memory (LSTM) models can leverage source code's hierarchical structure to discover latent representations. Specifically, LSTM-based predictive models built using a single problem and a combination of multiple problems can correctly predict if a source code will perform better or worse up to 84% and 73% of the time, respectively. Tarek Ramadan, Tanzima Z. Islam, Chase Phelps, Nathan Pinnow, Jayaraman J. Thiagarajan |
ISPASS | 5 |
| 2021 | Designing Counterfactual Generators using Deep Model InversionabstractExplanation techniques that synthesize small, interpretable changes to a given image while producing desired changes in the model prediction have become popular for introspecting black-box models. Commonly referred to as counterfactuals, the synthesized explanations are required to contain discernible changes (for easy interpretability) while also being realistic (consistency to the data manifold). In this paper, we focus on the case where we have access only to the trained deep classifier and not the actual training data. While the problem of inverting deep models to synthesize images from the training distribution has been explored, our goal is to develop a deep inversion approach to generate counterfactual explanations for a given query image. Despite their effectiveness in conditional image synthesis, we show that existing deep inversion methods are insufficient for producing meaningful counterfactuals. We propose DISC (Deep Inversion for Synthesizing Counterfactuals) that improves upon deep inversion by utilizing (a) stronger image priors, (b) incorporating a novel manifold consistency objective and (c) adopting a progressive optimization strategy. We find that, in addition to producing visually meaningful explanations, the counterfactuals from DISC are effective at learning classifier decision boundaries and are robust to unknown test-time corruptions. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang, Akshay Chaudhari, Andreas Spanias |
NeurIPS | 1 |
| 2021 | Coverage-Based Designs Improve Sample Mining and Hyperparameter Optimization
Gowtham Muniraju, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Cihan Tepedelenlioglu, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Building Calibrated Deep Models via Uncertainty Matching with Auxiliary Interval PredictorsabstractWith rapid adoption of deep learning in critical applications, the question of when and how much to trust these models often arises, which drives the need to quantify the inherent uncertainties. While identifying all sources that account for the stochasticity of models is challenging, it is common to augment predictions with confidence intervals to convey the expected variations in a model's behavior. We require prediction intervals to be well-calibrated, reflect the true uncertainties, and to be sharp. However, existing techniques for obtaining prediction intervals are known to produce unsatisfactory results in at least one of these criteria. To address this challenge, we develop a novel approach for building calibrated estimators. More specifically, we use separate models for prediction and interval estimation, and pose a bi-level optimization problem that allows the former to leverage estimates from the latter through an uncertainty matching strategy. Using experiments in regression, time-series forecasting, and object localization, we show that our approach achieves significant improvements over existing uncertainty quantification methods, both in terms of model fidelity and calibration error. Jayaraman J. Thiagarajan, Bindya Venkatesh, Prasanna Sattigeri, Peer-Timo Bremer |
AAAI | 1 |
| 2020 | A Regularized Attention Mechanism for Graph Attention NetworksabstractMachine learning models that can exploit the inherent structure in data have gained prominence. In particular, there is a surge in deep learning solutions for graph-structured data, due to its wide-spread applicability in several fields. Graph attention networks (GAT), a recent addition to the broad class of feature learning models in graphs, utilizes the attention mechanism to efficiently learn continuous vector representations for semi-supervised learning problems. In this paper, we perform a detailed analysis of GAT models, and present interesting insights into their behavior. In particular, we show that the models are vulnerable to heterogeneous rogue nodes and hence propose novel regularization strategies to improve the robustness of GAT models. Using benchmark datasets, we demonstrate performance improvements on semi-supervised learning, using the proposed robust variant of GAT. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Andreas Spanias |
ICASSP | 2 |
| 2020 | Learn-By-Calibrating: Using Calibration As A Training ObjectiveabstractCalibration error is commonly adopted for evaluating the quality of uncertainty estimators in deep neural networks. In this paper, we argue that such a metric is highly beneficial for training predictive models, even when we do not explicitly measure the uncertainties. This is conceptually similar to heteroscedastic neural networks that produce variance estimates for each prediction, with the key difference that we do not place a Gaussian prior on the predictions. We propose a novel algorithm that performs simultaneous interval estimation for different calibration levels and effectively leverages the intervals to refine the mean estimates. Our results show that, our approach is consistently superior to existing regularization strategies in deep regression models. Finally, we propose to augment partial dependence plots, a model-agnostic interpretability tool, with expected prediction intervals to reveal interesting dependencies between data and the target. Jayaraman J. Thiagarajan, Bindya Venkatesh, Deepta Rajan |
ICASSP | 1 |
| 2020 | Unsupervised Audio Source Separation Using Generative PriorsabstractState-of-the-art under-determined audio source separation systems rely on supervised end-end training of carefully tailored neural network architectures operating either in the time or the spectral domain. However, these methods are severely challenged in terms of requiring access to expensive source level labeled data and being specific to a given set of sources and the mixing process, which demands complete re-training when those assumptions change. This strongly emphasizes the need for unsupervised methods that can leverage the recent advances in data-driven modeling, and compensate for the lack of labeled data through meaningful priors. To this end, we propose a novel approach for audio source separation based on generative priors trained on individual sources. Through the use of projected gradient descent optimization, our approach simultaneously searches in the source-specific latent spaces to effectively recover the constituent sources. Though the generative priors can be defined in the time domain directly, e.g. WaveGAN, we find that using spectral domain loss functions for our optimization leads to good-quality source estimates. Our empirical studies on standard spoken digit and instrument datasets clearly demonstrate the effectiveness of our approach over classical as well as state-of-the-art unsupervised baselines. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Rushil Anirudh, Andreas Spanias |
INTERSPEECH | 2 |
| 2020 | The Case of Performance Variability on Dragonfly-based SystemsabstractPerformance of a parallel code running on a large supercomputer can vary significantly from one run to another even when the executable and its input parameters are left unchanged. Such variability can occur due to perturbation of the computation and/or communication in the code. In this paper, we investigate the case of performance variability arising due to network effects on supercomputers that use a dragonfly topology - specifically, Cray XC systems equipped with the Aries interconnect. We perform post-mortem analysis of network hardware counters, profiling output, job queue logs, and placement information, all gathered from periodic representative application runs. We investigate the causes of performance variability using deviation prediction and recursive feature elimination. Additionally, using time-stepped performance data of individual applications, we train machine learning models that can forecast the execution time of future time steps. Abhinav Bhatele, Jayaraman J. Thiagarajan, Taylor L. Groves, Rushil Anirudh, Staci A. Smith, Brandon Cook 0001, David K. Lowenthal |
IPDPS | 2 |
| 2020 | A Statistical Mechanics Framework for Task-Agnostic Sample Design in Machine LearningabstractIn this paper, we present a statistical mechanics framework to understand the effect of sampling properties of training data on the generalization gap of machine learning (ML) algorithms. We connect the generalization gap to the spatial properties of a sample design characterized by the pair correlation function (PCF). In particular, we express generalization gap in terms of the power spectra of the sample design and that of the function to be learned. Using this framework, we show that space-filling sample designs, such as blue noise and Poisson disk sampling, which optimize spectral properties, outperform random designs in terms of the generalization gap and characterize this gain in a closed-form. Our analysis also sheds light on design principles for constructing optimal task-agnostic sample designs that minimize the generalization gap. We corroborate our findings using regression experiments with neural networks on: a) synthetic functions, and b) a complex scientific simulator for inertial confinement fusion (ICF). Bhavya Kailkhura, Jayaraman J. Thiagarajan, Qunwei Li, Jize Zhang, Yi Zhou 0017, Peer-Timo Bremer |
NeurIPS | 2 |
| 2020 | MimicGAN: Robust Projection onto Image Manifolds with Corruption Mimicking
Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Peer-Timo Bremer |
Int. J. Comput. Vis. | 2 |
| 2020 | GrAMME: Semisupervised Learning Using Multilayered Graph Attention ModelsabstractModern data analysis pipelines are becoming increasingly complex due to the presence of multiview information sources. While graphs are effective in modeling complex relationships, in many scenarios, a single graph is rarely sufficient to succinctly represent all interactions, and hence, multilayered graphs have become popular. Though this leads to richer representations, extending solutions from the single-graph case is not straightforward. Consequently, there is a strong need for novel solutions to solve classical problems, such as node classification, in the multilayered case. In this article, we consider the problem of semisupervised learning with multilayered graphs. Though deep network embeddings, e.g., DeepWalk, are widely adopted for community discovery, we argue that feature learning with random node attributes, using graph neural networks, can be more effective. To this end, we propose to use attention models for effective feature learning and develop two novel architectures, GrAMME-SG and GrAMME-Fusion, that exploit the interlayer dependences for building multilayered graph embeddings. Using empirical studies on several benchmark data sets, we evaluate the proposed approaches and demonstrate significant performance improvements in comparison with the state-of-the-art network embedding strategies. The results also show that using simple random features is an effective choice, even in cases where explicit node attributes are not available. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Scalable Topological Data Analysis and Visualization for Evaluating Data-Driven Models in Scientific ApplicationsabstractWith the rapid adoption of machine learning techniques for large-scale applications in science and engineering comes the convergence of two grand challenges in visualization. First, the utilization of black box models (e.g., deep neural networks) calls for advanced techniques in exploring and interpreting model behaviors. Second, the rapid growth in computing has produced enormous datasets that require techniques that can handle millions or more samples. Although some solutions to these interpretability challenges have been proposed, they typically do not scale beyond thousands of samples, nor do they provide the high-level intuition scientists are looking for. Here, we present the first scalable solution to explore and analyze high-dimensional functions often encountered in the scientific data analysis pipeline. By combining a new streaming neighborhood graph construction, the corresponding topology computation, and a novel data aggregation scheme, namely topology aware datacubes, we enable interactive exploration of both the topological and the geometric aspect of high-dimensional data. Following two use cases from high-energy-density (HED) physics and computational biology, we demonstrate how these capabilities have led to crucial new insights in both applications. Shusen Liu 0001, Jim Gaffney, Jayson Luc Peterson, Peter B. Robinson, Harsh Bhatia, Valerio Pascucci, Brian K. Spears, Peer-Timo Bremer, Dan Maljovec, Rushil Anirudh, Jayaraman J. Thiagarajan, Sam Ade Jacobs, Brian Van Essen, David Hysom, Jae-Seung Yeom |
IEEE Trans. Vis. Comput. Graph. | 12 |
| 2019 | Improved Deep Embeddings for Inferencing with Multi-Layered GraphsabstractInferencing with graph data necessitates the mapping of its nodes into a vector space, where the relationships are preserved. However, with multi-layered graphs, where multiple types of relationships exist for the same set of nodes, it is crucial to exploit the information shared between layers, in addition to the distinct aspects of each layer. In this paper, we propose a novel approach that first obtains node embeddings in all layers jointly via DeepWalk on a supra graph, which allows interactions between layers, and then fine-tunes the embeddings to encourage cohesive structure in the latent space. With empirical studies in node classification, link prediction and multi-layered community detection, we show that the proposed approach outperforms existing single-and multi-layered graph embedding algorithms on several benchmarks. In addition to effectively scaling to a large number of layers (tested up to 37), our approach consistently produces highly modular community structure, even when compared to methods that directly optimize for the modularity function. Huan Song, Jayaraman J. Thiagarajan |
IEEE BigData | 2 |
| 2019 | Parallelizing Training of Deep Generative Models on Massive Scientific DatasetsabstractTraining deep neural networks on large scientific data is a challenging task that requires enormous compute power, especially if no pre-trained models exist to initialize the process. We present a novel tournament method to train traditional as well as generative adversarial networks built on LBANN, a scalable deep learning framework optimized for HPC systems. LBANN combines multiple levels of parallelism and exploits some of the worlds largest supercomputers.We demonstrate our framework by creating a complex predictive model based on multi-variate data from high-energy-density physics containing hundreds of millions of images and hundreds of millions of scalar values derived from tens of millions of simulations of inertial confinement fusion. Our approach combines an HPC workflow and extends LBANN with optimized data ingestion and the new tournament-style training algorithm to produce a scalable neural network architecture using a CORAL-class supercomputer. Experimental results show that 64 trainers (1024 GPUs) achieve a speedup of 70.2× over a single trainer (16 GPUs) baseline, and an effective 109% parallel efficiency. Sam Ade Jacobs, Jim Gaffney, Tom Benson, Peter B. Robinson, Jayson Luc Peterson, Brian K. Spears, Brian Van Essen, David Hysom, Jae-Seung Yeom, Tim Moon, Rushil Anirudh, Jayaraman J. Thiagarajan, Shusen Liu 0001, Peer-Timo Bremer |
CLUSTER | 12 |
| 2019 | Bootstrapping Graph Convolutional Neural Networks for Autism Spectrum Disorder ClassificationabstractUsing predictive models to identify patterns that can act as biomarkers for different neuropathoglogical conditions is becoming highly prevalent. In this paper, we consider the problem of Autism Spectrum Disorder (ASD) classification where previous work has shown that it can be beneficial to incorporate a wide variety of meta features, such as socio-cultural traits, into predictive modeling. A graph-based approach naturally suits these scenarios, where a contextual graph captures traits that characterize a population, while the specific brain activity patterns are utilized as a multivariate signal at the nodes. Graph neural networks have shown improvements in inferencing with graph-structured data. Though the underlying graph strongly dictates the overall performance, there exists no systematic way of choosing an appropriate graph in practice, thus making predictive models non-robust. To address this, we propose a bootstrapped version of graph convolutional neural networks (G-CNNs) that utilizes an ensemble of weakly trained G-CNNs, and reduce the sensitivity of models on the choice of graph construction. We demonstrate its effectiveness on the challenging Autism Brain Imaging Data Exchange (ABIDE) dataset and show that our approach improves upon recently proposed graph-based neural networks. We also show that our method remains more robust to noisy graphs. Rushil Anirudh, Jayaraman J. Thiagarajan |
ICASSP | 2 |
| 2019 | Designing an Effective Metric Learning Pipeline for Speaker DiarizationabstractState-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on choosing the appropriate feature extractor, ranging from pre-trained i-vectors to representations learned via different sequence modeling architectures (e.g. 1D-CNNs, LSTMs, attention models), while adopting off-the-shelf metric learning solutions. In this paper, we argue that, regardless of the feature extractor, it is crucial to carefully design a metric learning pipeline, namely the loss function, the sampling strategy and the discriminative margin parameter, for building robust diarization systems. Furthermore, we propose to adopt a fine-grained validation process to obtain a comprehensive evaluation of the generalization power of metric learning pipelines. To this end, we measure diarization performance across different language speakers, and variations in the number of speakers in a recording. Using empirical studies, we provide interesting insights into the effectiveness of different design choices and make recommendations. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias |
ICASSP | 2 |
| 2019 | Unsupervised Dimension Selection Using a Blue Noise Graph SpectrumabstractUnsupervised dimension selection is an important problem that seeks to reduce dimensionality of data, while preserving the most useful characteristics. While dimensionality reduction is commonly utilized to construct low-dimensional embeddings, they produce feature spaces that are hard to interpret. Further, in applications such as sensor design, one needs to perform reduction directly in the input domain, instead of constructing transformed spaces. Consequently, dimension selection (DS) aims to solve the combinatorial problem of identifying the top-k dimensions, which is required for effective experiment design, reducing data while keeping it interpretable, and designing better sensing mechanisms. In this paper, we develop a novel approach for DS based on graph signal analysis to measure feature influence. By analyzing synthetic graph signals with a blue noise spectrum, we show that we can measure the importance of each dimension. Using experiments in supervised learning and image masking, we demonstrate the superiority of the proposed approach over existing techniques in capturing crucial characteristics of high dimensional spaces, using only a small subset of the original features. Jayaraman J. Thiagarajan, Rushil Anirudh, Rahul Sridhar, Peer-Timo Bremer |
ICASSP | 1 |
| 2019 | Understanding Deep Neural Networks through Input UncertaintiesabstractTechniques for understanding the functioning of complex machine learning models are becoming increasingly popular, not only to improve the validation process, but also to extract new insights about the data via exploratory analysis. Though a large class of such tools currently exists, most assume that predictions are point estimates and use a sensitivity analysis of these estimates to interpret the model. Using lightweight probabilistic networks we show how including prediction uncertainties in the sensitivity analysis leads to: (i) more robust and generalizable models; and (ii) a new approach for model interpretation through uncertainty decomposition. In particular, we introduce a new regularization that takes both the mean and variance of a prediction into account and demonstrate that the resulting networks provide improved generalization to unseen data. Furthermore, we propose a new technique to explain prediction uncertainties through uncertainties in the input domain, thus providing new ways to validate and interpret deep learning models. Jayaraman J. Thiagarajan, Irene Kim, Rushil Anirudh, Peer-Timo Bremer |
ICASSP | 1 |
| 2019 | Multiple Subspace Alignment Improves Domain AdaptationabstractWe present a novel unsupervised domain adaptation (DA) method for cross-domain visual recognition. Though subspace methods have found success in DA, their performance is often limited due to the assumption of approximating an entire dataset using a single low-dimensional subspace. Instead, we develop a method to effectively represent the source and target datasets via a collection of low-dimensional subspaces, and subsequently align them by exploiting the natural geometry of the space of subspaces, on the Grassmann manifold. We demonstrate the effectiveness of this approach, using empirical studies on two widely used benchmarks,with performance on par or better than the performance of the state of the art domain adaptation methods. Kowshik Thopalli, Rushil Anirudh, Jayaraman J. Thiagarajan, Pavan Turaga |
ICASSP | 3 |
| 2019 | Distill-to-Label: Weakly Supervised Instance Labeling Using Knowledge DistillationabstractWeakly supervised instance labeling using only image-level labels, in lieu of expensive fine-grained pixel annotations, is crucial in several applications including medical image analysis. In contrast to conventional instance segmentation in computer vision, the problems that we consider are characterized by a small number of training images and non-local patterns that lead to the diagnosis. In this paper, we explore the use of multiple instance learning (MIL) to design an instance label generator under this weakly supervised setting. Motivated by the observation that an MIL model can handle bags of varying sizes, we propose to repurpose an MIL model originally trained for bag-level classification to produce reliable predictions for single instances. To this end, we introduce a novel regularization strategy based on virtual adversarial training for improving MIL training, and subsequently develop a knowledge distillation technique for repurposing the trained MIL model. Using empirical studies on colon cancer and breast cancer detection from histopathological images, we show that the proposed approach produces high-quality instance-level prediction and significantly outperforms state-of-the MIL methods. Jayaraman J. Thiagarajan, Satyananda Kashyap, Alexandros Karargyris |
ICMLA | 1 |
| 2019 | Performance optimality or reproducibility: that is the questionabstractThe era of extremely heterogeneous supercomputing brings with itself the devil of increased performance variation and reduced reproducibility. There is a lack of understanding in the HPC community on how the simultaneous consideration of network traffic, power limits, concurrency tuning, and interference from other jobs impacts application performance. Tapasya Patki, Jayaraman J. Thiagarajan, Alexis Ayala, Tanzima Z. Islam |
SC | 2 |
| 2018 | Attend and Diagnose: Clinical Time Series Analysis Using Attention ModelsabstractWith widespread adoption of electronic health records, there is an increased emphasis for predictive models that can effectively deal with clinical time-series data. Powered by Recurrent Neural Network (RNN) architectures with Long Short-Term Memory (LSTM) units, deep neural networks have achieved state-of-the-art results in several clinical prediction tasks. Despite the success of RNN, its sequential nature prohibits parallelized computing, thus making it inefficient particularly when processing long sequences. Recently, architectures which are based solely on attention mechanisms have shown remarkable success in transduction tasks in NLP, while being computationally superior. In this paper, for the first time, we utilize attention models for clinical time-series modeling, thereby dispensing recurrence entirely. We develop the SAnD (Simply Attend and Diagnose) architecture, which employs a masked, self-attention mechanism, and uses positional encoding and dense interpolation strategies for incorporating temporal order. Furthermore, we develop a multi-task variant of SAnD to jointly infer models with multiple diagnosis tasks. Using the recent MIMIC-III benchmark datasets, we demonstrate that the proposed approach achieves state-of-the-art performance in all tasks, outperforming LSTM models and classical baselines with hand-engineered features. Huan Song, Deepta Rajan, Jayaraman J. Thiagarajan, Andreas Spanias |
AAAI | 3 |
| 2018 | Lose the Views: Limited Angle CT Reconstruction via Implicit Sinogram CompletionabstractComputed Tomography (CT) reconstruction is a fundamental component to a wide variety of applications ranging from security, to healthcare. The classical techniques require measuring projections, called sinograms, from a full 180° view of the object. However, obtaining a full-view is not always feasible, such as when scanning irregular objects that limit flexibility of scanner rotation. The resulting limited angle sinograms are known to produce highly artifact-laden reconstructions with existing techniques. In this paper, we propose to address this problem using CTNet - a system of 1D and 2D convolutional neural networks, that operates directly on a limited angle sinogram to predict the reconstruction. We use the x-ray transform on this prediction to obtain a "completed" sinogram, as if it came from a full 180°view. We feed this to standard analytical and iterative reconstruction techniques to obtain the final reconstruction. We show with extensive experimentation on a challenging real world dataset that this combined strategy outperforms many competitive baselines. We also propose a measure of confidence for the reconstruction that enables a practitioner to gauge the reliability of a prediction made by CTNet. We show that this measure is a strong indicator of quality as measured by the PSNR, while not requiring ground truth at test time. Finally, using a segmentation experiment, we show that our reconstruction also preserves the 3D structure of objects better than existing solutions. Rushil Anirudh, Hyojin Kim 0001, Jayaraman J. Thiagarajan, K. Aditya Mohan, Kyle Champley, Peer-Timo Bremer |
CVPR | 3 |
| 2018 | Bootstrapping Parameter Space Exploration for Fast TuningabstractThe task of tuning parameters for optimizing performance or other metrics of interest such as energy, variability, etc. can be resource and time consuming. Presence of a large parameter space makes a comprehensive exploration infeasible. In this paper, we propose a novel bootstrap scheme, called GEIST, for parameter space exploration to find performance-optimizing configurations quickly. Our scheme represents the parameter space as a graph whose connectivity guides information propagation from known configurations. Guided by the predictions of a semi-supervised learning method over the parameter graph, GEIST is able to adaptively sample and find desirable configurations using limited results from experiments. We show the effectiveness of GEIST for selecting application input options, compiler flags, and runtime/system settings for several parallel codes including LULESH, Kripke, Hypre, and OpenAtom. Jayaraman J. Thiagarajan, Rushil Anirudh, Alfredo Giménez, Rahul Sridhar, Aniruddha Marathe, Tao Wang 0077, Murali Emani, Abhinav Bhatele, Todd Gamblin |
ICS | 1 |
| 2018 | Triplet Network with Attention for Speaker DiarizationabstractIn automatic speech processing systems, speaker diarization is a crucial front-end component to separate segments from different speakers.Inspired by the recent success of deep neural networks (DNNs) in semantic inferencing, triplet loss-based architectures have been successfully used for this problem.However, existing work utilizes conventional i-vectors as the input representation and builds simple fully connected networks for metric learning, thus not fully leveraging the modeling power of DNN architectures.This paper investigates the importance of learning effective representations from the sequences directly in metric learning pipelines for speaker diarization.More specifically, we propose to employ attention models to learn embeddings and the metric jointly in an end-to-end fashion.Experiments are conducted on the CALLHOME conversational speech corpus.The diarization results demonstrate that, besides providing a unified model, the proposed approach achieves improved performance when compared against existing approaches. Huan Song, Megan M. Willi, Jayaraman J. Thiagarajan, Visar Berisha, Andreas Spanias |
INTERSPEECH | 3 |
| 2018 | PADDLE: Performance Analysis Using a Data-Driven Learning EnvironmentabstractThe use of machine learning techniques to model execution time and power consumption, and, more generally, to characterize performance data is gaining traction in the HPC community. Although this signifies huge potential for automating complex inference tasks, a typical analytics pipeline requires selecting and extensively tuning multiple components ranging from feature learning to statistical inferencing to visualization. Further, the algorithmic solutions often do not generalize between problems, thereby making it cumbersome to design and validate machine learning techniques in practice. In order to address these challenges, we propose a unified machine learning framework, PADDLE, which is specifically designed for problems encountered during analysis of HPC data. The proposed framework uses an information-theoretic approach for hierarchical feature learning and can produce highly robust and interpretable models. We present user-centric workflows for using PADDLE and demonstrate its effectiveness in different scenarios: (a) identifying causes of network congestion; (b) determining the best performing linear solver for sparse matrices; and (c) comparing performance characteristics of parent and proxy application pairs. Jayaraman J. Thiagarajan, Rushil Anirudh, Bhavya Kailkhura, Tanzima Z. Islam, Abhinav Bhatele, Jae-Seung Yeom, Todd Gamblin |
IPDPS | 1 |
| 2018 | Mitigating inter-job interference using adaptive flow-aware routing
Staci A. Smith, Clara E. Cromey, David K. Lowenthal, Jens Domke, Jayaraman J. Thiagarajan, Abhinav Bhatele |
SC | 6 |
| 2018 | Exploring High-Dimensional Structure via Axis-Aligned Decomposition of Linear ProjectionsabstractAbstract Two‐dimensional embeddings remain the dominant approach to visualize high dimensional data. The choice of embeddings ranges from highly non‐linear ones, which can capture complex relationships but are difficult to interpret quantitatively, to axis‐aligned projections, which are easy to interpret but are limited to bivariate relationships. Linear project can be considered as a compromise between complexity and interpretability, as they allow explicit axes labels, yet provide significantly more degrees of freedom compared to axis‐aligned projections. Nevertheless, interpreting the axes directions, which are often linear combinations of many non‐trivial components, remains difficult. To address this problem we introduce a structure aware decomposition of (multiple) linear projections into sparse sets of axis‐aligned projections, which jointly capture all information of the original linear ones. In particular, we use tools from Dempster‐Shafer theory to formally define how relevant a given axis‐aligned project is to explain the neighborhood relations displayed in some linear projection. Furthermore, we introduce a new approach to discover a diverse set of high quality linear projections and show that in practice the information of k linear projections is often jointly encoded in ∼ k axis‐aligned plots. We have integrated these ideas into an interactive visualization system that allows users to jointly browse both linear projections and their axis‐aligned representatives. Using a number of case studies we show how the resulting plots lead to more intuitive visualizations and new insights. Jayaraman J. Thiagarajan, Shusen Liu 0001, Karthikeyan Natesan Ramamurthy, Peer-Timo Bremer |
Comput. Graph. Forum | 1 |
| 2018 | A Spectral Approach for the Design of Experiments: Design, Analysis and AlgorithmsabstractThis paper proposes a new approach to construct high quality space-filling sample designs. First, we propose a novel technique to quantify the space-filling property and optimally trade-off uniformity and randomness in sample designs in arbitrary dimensions. Second, we connect the proposed metric (defined in the spatial domain) to the quality metric of the design performance (defined in the spectral domain). This connection serves as an analytic framework for evaluating the qualitative properties of space-filling designs in general. Using the theoretical insights provided by this spatial-spectral analysis, we derive the notion of optimal space-filling designs, which we refer to as space-filling spectral designs. Third, we propose an efficient estimator to evaluate the space-filling properties of sample designs in arbitrary dimensions and use it to develop an optimization framework for generating high quality space-filling designs. Finally, we carry out a detailed performance comparison on two different applications in varying dimensions: a) image reconstruction and b) surrogate modeling for several benchmark optimization functions and a physics simulation code for inertial confinement fusion (ICF). Our results clearly evidence the superiority of the proposed space-filling designs over existing approaches, particularly in high dimensions. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Charvi Rastogi, Pramod K. Varshney, Peer-Timo Bremer |
J. Mach. Learn. Res. | 2 |
| 2018 | Optimizing Kernel Machines Using Deep LearningabstractBuilding highly nonlinear and nonparametric models is central to several state-of-the-art machine learning systems. Kernel methods form an important class of techniques that induce a reproducing kernel Hilbert space (RKHS) for inferring non-linear models through the construction of similarity functions from data. These methods are particularly preferred in cases where the training data sizes are limited and when prior knowledge of the data similarities is available. Despite their usefulness, they are limited by the computational complexity and their inability to support end-to-end learning with a task-specific objective. On the other hand, deep neural networks have become the de facto solution for end-to-end inference in several learning paradigms. In this paper, we explore the idea of using deep architectures to perform kernel machine optimization, for both computational efficiency and end-to-end inferencing. To this end, we develop the deep kernel machine optimization framework, that creates an ensemble of dense embeddings using Nyström kernel approximations and utilizes deep learning to generate task-specific representations through the fusion of the embeddings. Intuitively, the filters of the network are trained to fuse information from an ensemble of linear subspaces in the RKHS. Furthermore, we introduce the kernel dropout regularization to enable improved training convergence. Finally, we extend this framework to the multiple kernel case, by coupling a global fusion layer with pretrained deep kernel machines for each of the constituent kernels. Using case studies with limited training data, and lack of explicit feature sources, we demonstrate the effectiveness of our framework over conventional model inferencing techniques. Huan Song, Jayaraman J. Thiagarajan, Prasanna Sattigeri, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Visual Exploration of Semantic Relationships in Neural Word EmbeddingsabstractConstructing distributed representations for words through neural language models and using the resulting vector spaces for analysis has become a crucial component of natural language processing (NLP). However, despite their widespread application, little is known about the structure and properties of these spaces. To gain insights into the relationship between words, the NLP community has begun to adapt high-dimensional visualization techniques. In particular, researchers commonly use t-distributed stochastic neighbor embeddings (t-SNE) and principal component analysis (PCA) to create two-dimensional embeddings for assessing the overall structure and exploring linear relationships (e.g., word analogies), respectively. Unfortunately, these techniques often produce mediocre or even misleading results and cannot address domain-specific visualization challenges that are crucial for understanding semantic relationships in word embeddings. Here, we introduce new embedding techniques for visualizing semantic and syntactic analogies, and the corresponding tests to determine whether the resulting views capture salient structures. Additionally, we introduce two novel views for a comprehensive study of analogy relationships. Finally, we augment t-SNE embeddings to convey uncertainty information in order to allow a reliable interpretation. Combined, the different views address a number of domain-specific tasks difficult to solve with existing tools. Shusen Liu 0001, Peer-Timo Bremer, Jayaraman J. Thiagarajan, Vivek Srikumar, Bei Wang 0001, Yarden Livnat, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | A deep learning approach to multiple kernel fusionabstractKernel fusion is a popular and effective approach for combining multiple features that characterize different aspects of data. Traditional approaches for Multiple Kernel Learning (MKL) attempt to learn the parameters for combining the kernels through sophisticated optimization procedures. In this paper, we propose an alternative approach that creates dense embeddings for data using the kernel similarities and adopts a deep neural network architecture for fusing the embeddings. In order to improve the effectiveness of this network, we introduce the kernel dropout regularization strategy coupled with the use of an expanded set of composition kernels. Experiment results on a real-world activity recognition dataset show that the proposed architecture is effective in fusing kernels and achieves state-of-the-art performance. Huan Song, Jayaraman J. Thiagarajan, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
ICASSP | 2 |
| 2017 | Performance modeling under resource constraints using deep transfer learningabstractTuning application parameters for optimal performance is a challenging combinatorial problem. Hence, techniques for modeling the functional relationships between various input features in the parameter space and application performance are important. We show that simple statistical inference techniques are inadequate to capture these relationships. Even with more complex ensembles of models, the minimum coverage of the parameter space required via experimental observations is still quite large. We propose a deep learning based approach that can combine information from exhaustive observations collected at a smaller scale with limited observations collected at a larger target scale. The proposed approach is able to accurately predict performance in the regimes of interest to performance analysts while outperforming many traditional techniques. In particular, our approach can identify the best performing configurations even when trained using as few as 1% of observations at the target scale. Aniruddha Marathe, Rushil Anirudh, Abhinav Bhatele, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Jae-Seung Yeom, Barry Rountree, Todd Gamblin |
SC | 5 |
| 2016 | Theoretical guarantees for poisson disk sampling using pair correlation functionabstractIn this paper, we study the problem of generating uniform random point samples on a domain of d dimensional space based on a minimum distance criterion between point samples (Poisson-disk sampling or PDS). First, we formally define PDS via the pair correlation function (PCF) to quantitatively evaluate properties of the sampling process. Surprisingly, none of the existing PDS techniques satisfy both uniformity and minimum distance criterion, simultaneously. These approaches typically create an approximate PDS with high regularity, and inherently present high risk for sample aliasing. Our new formulation based on PCF introduces a new approach to evaluate PDS properties which leads to theoretical bounds on the size of a PDS in arbitrary dimensions as well as a faster algorithm to create better quality samplings than the current PDS approaches. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Pramod K. Varshney |
ICASSP | 2 |
| 2016 | Beyond L2-loss functions for learning sparse modelsabstractIn sparse learning, the squared Euclidean distance is a popular choice for measuring the approximation quality. However, the use of other forms of parametrized loss functions, including asymmetric losses, has generated research interest. In this paper, we perform sparse learning using a broad class of smooth piecewise linear quadratic (PLQ) loss functions, including robust and asymmetric losses that are adaptable to many real-world scenarios. The proposed framework also supports heterogeneous data modeling by allowing different PLQ penalties for different blocks of residual vectors (split-PLQ). We demonstrate the impact of the proposed sparse learning in image recovery, and apply the proposed split-PLQ loss approach to tag refinement for image annotation and retrieval. Karthikeyan Natesan Ramamurthy, Aleksandr Y. Aravkin, Jayaraman J. Thiagarajan |
ICASSP | 3 |
| 2016 | Consensus inference on mobile phone sensors for activity recognitionabstractThe pervasive use of wearable sensors in activity and health monitoring presents a huge potential for building novel data analysis and prediction frameworks. In particular, approaches that can harness data from a diverse set of low-cost sensors for recognition are needed. Many of the existing approaches rely heavily on elaborate feature engineering to build robust recognition systems, and their performance is often limited by the inaccuracies in the data. In this paper, we develop a novel two-stage recognition system that enables a systematic fusion of complementary information from multiple sensors in a linear graph embedding setting, while employing an ensemble classifier phase that leverages the discriminative power of different feature extraction strategies. Experimental results on a challenging dataset show that our framework greatly improves the recognition performance when compared to using any single sensor. Huan Song, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Pavan Turaga |
ICASSP | 2 |
| 2016 | Auto-context modeling using multiple Kernel learningabstractIn complex visual recognition systems, feature fusion has become crucial to discriminate between a large number of classes. In particular, fusing high-level context information with image appearance models can be effective in object/scene recognition. To this end, we develop an auto-context modeling approach under the RKHS (Reproducing Kernel Hilbert Space) setting, wherein a series of supervised learners are used to approximate the context model. By posing the problem of fusing the context and appearance models using multiple kernel learning, we develop a computationally tractable solution to this challenging problem. Furthermore, we propose to use the marginal probabilities from a kernel SVM classifier to construct the auto-context kernel. In addition to providing better regularization to the learning problem, our approach leads to improved recognition performance in comparison to using only the image features. Huan Song, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
ICIP | 2 |
| 2016 | A machine learning framework for performance coverage analysis of proxy applicationsabstractProxy applications are written to represent subsets of performance behaviors of larger, and more complex applications that often have distribution restrictions. They enable easy evaluation of these behaviors across systems, e.g., for procurement or co-design purposes. However, the intended correlation between the performance behaviors of proxy applications and their parent codes is often based solely on the developer's intuition. In this paper, we present novel machine learning techniques to methodically quantify the coverage of performance behaviors of parent codes by their proxy applications. We have developed a framework, VERITAS, to answer these questions in the context of on-node performance: (a) which hardware resources are covered by a proxy application and how well, and (b) which resources are important, but not covered. We present our techniques in the context of two benchmarks, STREAM and DGEMM, and two production applications, OpenMC and CMTnek, and their respective proxy applications. Tanzima Z. Islam, Jayaraman J. Thiagarajan, Abhinav Bhatele, Martin Schulz 0001, Todd Gamblin |
SC | 2 |
| 2016 | Universal Collaboration Strategies for Signal Detection: A Sparse Learning ApproachabstractThis paper considers the problem of high-dimensional signal detection in a large distributed network whose nodes can collaborate with their one-hop neighboring nodes (spatial collaboration). We assume that only a small subset of nodes communicate with the fusion center (FC). We design optimal collaboration strategies which are universal for a class of deterministic signals. By establishing the equivalence between the collaboration strategy design problem and sparse principal component analysis (PCA), we solve the problem efficiently and evaluate the impact of collaboration on detection performance. Prashant Khanduri, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Pramod K. Varshney |
IEEE Signal Process. Lett. | 3 |
| 2016 | Stair blue noise samplingabstractA common solution to reducing visible aliasing artifacts in image reconstruction is to employ sampling patterns with a blue noise power spectrum. These sampling patterns can prevent discernible artifacts by replacing them with incoherent noise. Here, we propose a new family of blue noise distributions, Stair blue noise , which is mathematically tractable and enables parameter optimization to obtain the optimal sampling distribution. Furthermore, for a given sample budget, the proposed blue noise distribution achieves a significantly larger alias-free low-frequency region compared to existing approaches, without introducing visible artifacts in the mid-frequencies. We also develop a new sample synthesis algorithm that benefits from the use of an unbiased spatial statistics estimator and efficient optimization strategies. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Pramod K. Varshney |
ACM Trans. Graph. | 2 |
| 2015 | Subspace learning using consensus on the grassmannian manifoldabstractHigh-dimensional structure of data can be explored and task-specific representations can be obtained using manifold learning and low-dimensional embedding approaches. However, the uncertainties in data and the sensitivity of the algorithms to parameter settings, reduce the reliability of such representations, and make visualization and interpretation of data very challenging. A natural approach to combat challenges pertinent to data visualization is to use linearized embedding approaches. In this paper, we explore approaches to improve the reliability of linearized, subspace embedding frameworks by learning a plurality of subspaces and computing a geometric mean on the Grassmannian manifold. Using the proposed algorithm, we build variants of popular unsupervised and supervised graph embedding algorithms, and show that we can infer high-quality embeddings, thereby significantly improving their usability in visualization and classification. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy |
ICASSP | 1 |
| 2015 | A Randomized Ensemble Approach to Industrial CT SegmentationabstractTuning the models and parameters of common segmentation approaches is challenging especially in the presence of noise and artifacts. Ensemble-based techniques attempt to compensate by randomly varying models and/or parameters to create a diverse set of hypotheses, which are subsequently ranked to arrive at the best solution. However, these methods have been restricted to cases where the underlying models are well-established, e.g. natural images. In practice, it is difficult to determine a suitable base-model and the amount of randomization required. Furthermore, for multi-object scenes no single hypothesis may perform well for all objects, reducing the overall quality of the results. This paper presents a new ensemble-based segmentation framework for industrial CT images demonstrating that comparatively simple models and randomization strategies can significantly improve the result over existing techniques. Furthermore, we introduce a per-object based ranking, followed by a consensus inference that can outperform even the best case scenario of existing hypothesis ranking approaches. We demonstrate the effectiveness of our approach using a set of noise and artifact rich CT images from baggage security and show that it significantly outperforms existing solutions in this area. Hyojin Kim 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICCV | 2 |
| 2015 | Identifying the Culprits Behind Network CongestionabstractNetwork congestion is one of the primary causes of performance degradation, performance variability and poor scaling in communication-heavy parallel applications. However, the causes and mechanisms of network congestion on modern interconnection networks are not well understood. We need new approaches to analyze, model and predict this critical behaviour in order to improve the performance of large-scale parallel applications. This paper applies supervised learning algorithms, such as forests of extremely randomized trees and gradient boosted regression trees, to perform regression analysis on communication data and application execution time. Using data derived from multiple executions, we create models to predict the execution time of communication-heavy parallel applications. This analysis also identifies the features and associated hardware components that have the most impact on network congestion and intern, on execution time. The ideas presented in this paper have wide applicability: predicting the execution time on a different number of nodes, or different input datasets, or even for an unknown code, identifying the best configuration parameters for an application, and finding the root causes of network congestion on different architectures. Abhinav Bhatele, Andrew R. Titus, Jayaraman J. Thiagarajan, Todd Gamblin, Peer-Timo Bremer, Martin Schulz 0001, Laxmikant V. Kalé |
IPDPS | 3 |
| 2015 | Visual Exploration of High-Dimensional Data through Subspace Analysis and Dynamic ProjectionsabstractAbstract We introduce a novel interactive framework for visualizing and exploring high‐dimensional datasets based on subspace analysis and dynamic projections. We assume the high‐dimensional dataset can be represented by a mixture of low‐dimensional linear subspaces with mixed dimensions, and provide a method to reliably estimate the intrinsic dimension and linear basis of each subspace extracted from the subspace clustering. Subsequently, we use these bases to define unique 2D linear projections as viewpoints from which to visualize the data. To understand the relationships among the different projections and to discover hidden patterns, we connect these projections through dynamic projections that create smooth animated transitions between pairs of projections. We introduce the view transition graph, which provides flexible navigation among these projections to facilitate an intuitive exploration. Finally, we provide detailed comparisons with related systems, and use real‐world examples to demonstrate the novelty and usability of our proposed framework. Shusen Liu 0001, Bei Wang 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Valerio Pascucci |
Comput. Graph. Forum | 3 |
| 2015 | Learning Stable Multilevel Dictionaries for Sparse RepresentationsabstractSparse representations using learned dictionaries are being increasingly used with success in several data processing and machine learning applications. The increasing need for learning sparse models in large-scale applications motivates the development of efficient, robust, and provably good dictionary learning algorithms. Algorithmic stability and generalizability are desirable characteristics for dictionary learning algorithms that aim to build global dictionaries, which can efficiently model any test data similar to the training samples. In this paper, we propose an algorithm to learn dictionaries for sparse representations from large scale data, and prove that the proposed learning algorithm is stable and generalizable asymptotically. The algorithm employs a 1-D subspace clustering procedure, the K-hyperline clustering, to learn a hierarchical dictionary with multiple levels. We also propose an information-theoretic scheme to estimate the number of atoms needed in each level of learning and develop an ensemble approach to learn robust dictionaries. Using the proposed dictionaries, the sparse code for novel test data can be computed using a low-complexity pursuit procedure. We demonstrate the stability and generalization characteristics of the proposed algorithm using simulations. We also evaluate the utility of the multilevel dictionaries in compressed recovery and subspace learning applications. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Multiple kernel interpolation for inverting non-linear dimensionality reduction and dimension estimationabstractThe problem of stably inverting a non-linear dimensionality reduction map has applications in data visualization and machine learning, besides being of theoretical interest. In this paper, we propose a meshfree interpolation method for obtaining such inverse maps using a non-negative linear combination of multiple interpolants. We show that the proposed scheme can improve upon the approximation power of its individual constituent kernels, and discuss the conditions under which its parameters can be uniquely estimated. We also provide an approach for estimating the intrinsic dimensionality (ID) of manifolds using the proposed inverse map. Experiments using multiple kernel interpolation for reconstruction of novel test data and ID estimation show an improved or similar performance compared to existing techniques. Jayaraman J. Thiagarajan, Peer-Timo Bremer, Karthikeyan Natesan Ramamurthy |
ICASSP | 1 |
| 2014 | Image segmentation using consensus from hierarchical segmentation ensemblesabstractUnsupervised, automatic image segmentation without contextual knowledge, or user intervention is a challenging problem. The key to robust segmentation is an appropriate selection of local features and metrics. However, a single aggregation of the local features using a greedy merging order often results in incorrect segmentation. This paper presents an unsupervised approach, which uses the consensus inferred from hierarchical segmentation ensembles, for partitioning images into foreground and background regions. By exploring an expanded set of possible aggregations of the local features, the proposed method generates meaningful segmentations that are not often revealed when only the optimal hierarchy is considered. A graph cuts-based approach is employed to combine the consensus along with a foreground-background model estimate, obtained using the ensemble, for effective segmentation. Experiments with a standard dataset show promising results when compared to several existing methods including the state-of-the-art weak supervised techniques that use co-segmentation. Hyojin Kim 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICIP | 2 |
| 2014 | Automatic image annotation using inverse maps from semantic embeddingsabstractHuman annotation in large scale image databases is time-consuming and error-prone. Since it is very hard to mine image databases using just visual features or textual descriptors, it is common to transform the image features into a semantically meaningful space. In this paper, we propose to perform image annotation in a semantic space inferred based on sparse representations. By constructing a semantic embedding for the visual features, that is constrained to be close to the tag embedding, we show that a robust inverse map can be used to predict the tags. Experiments using standard datasets show the effectiveness of the proposed approach in automatic image annotation when compared to existing methods. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Peer-Timo Bremer, Andreas Spanias |
ICIP | 1 |
| 2014 | Multiple Kernel Sparse Representations for Supervised and Unsupervised LearningabstractIn complex visual recognition tasks, it is typical to adopt multiple descriptors, which describe different aspects of the images, for obtaining an improved recognition performance. Descriptors that have diverse forms can be fused into a unified feature space in a principled manner using kernel methods. Sparse models that generalize well to the test data can be learned in the unified kernel space, and appropriate constraints can be incorporated for application in supervised and unsupervised learning. In this paper, we propose to perform sparse coding and dictionary learning in the multiple kernel space, where the weights of the ensemble kernel are tuned based on graph-embedding principles such that class discrimination is maximized. In our proposed algorithm, dictionaries are inferred using multiple levels of 1D subspace clustering in the kernel space, and the sparse codes are obtained using a simple levelwise pursuit scheme. Empirical results for object recognition and image clustering show that our algorithm outperforms existing sparse coding based approaches, and compares favorably to other state-of-the-art methods. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
IEEE Trans. Image Process. | 1 |
| 2013 | A heterogeneous dictionary model for representation and recognition of human actionsabstractIn this paper, we consider low-dimensional and sparse representation models for human actions, that are consistent with how actions evolve in high-dimensional feature spaces. We first show that human actions can be well approximated by piecewise linear structures in the feature space. Based on this, we propose a new dictionary model that considers each atom in the dictionary to be an affine subspace defined by a point and a corresponding line. When compared to centered clustering approaches such as K-means, we show that the proposed dictionary is a better generative model for human actions. Furthermore, we demonstrate the utility of this model in efficient representation and recognition of human activities that are not available in the training set. Rushil Anirudh, Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Pavan Turaga, Andreas Spanias |
ICASSP | 3 |
| 2013 | Boosted dictionaries for image restoration based on sparse representationsabstractSparse representations using learned dictionaries have been successful in several image processing applications. However, using a single dictionary model in inverse problems may lead to instability in estimation. In this paper, we propose to perform image restoration using an ensemble of weak dictionaries that incorporate prior knowledge about the form of linear corruption. The dictionary learned in each round of the training procedure is optimized for the training examples having high reconstruction error in the previous round. The weak dictionaries are either obtained using a weighted K-Means or an example-selection approach. The final restored data is computed as a convex combination of data restored in individual rounds. Results with compressed recovery of standard images show that the proposed dictionaries result in a better performance compared to using a single dictionary obtained with a traditional alternating minimization approach. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias, Prasanna Sattigeri |
ICASSP | 2 |
| 2012 | Automated tumor segmentation using kernel sparse representationsabstractIn this paper, we describe a pixel based approach for automated segmentation of tumor components from MR images. Sparse coding with data-adapted dictionaries has been successfully employed in several image recovery and vision problems. Since it is trivial to obtain sparse codes for pixel values, we propose to consider their non-linear similarities to perform kernel sparse coding in a high dimensional feature space. We develop the kernel K-lines clustering procedure for inferring kernel dictionaries and use the kernel sparse codes to determine if a pixel belongs to a tumorous region. By incorporating spatial locality information of the pixels, contiguous tumor regions can be efficiently identified. A low complexity segmentation approach, which allows the user to initialize the tumor region, is also presented. Results show that both of the proposed approaches lead to accurate tumor identification with a low false positive rate, when compared to manual segmentation by an expert. Jayaraman J. Thiagarajan, Deepta Rajan, Karthikeyan Natesan Ramamurthy, David H. Frakes, Andreas Spanias |
BIBE | 1 |
| 2012 | Work in progress: Performing signal analysis laboratories using Android devicesabstractIn this paper, we present a graphical-programming application to support signal processing education on the Android operating system. This application features a simulation environment and a palette of DSP functions, which will allow students to perform laboratories using Android smartphones and tablets. In order to demonstrate the application of the software in a classroom setting, a number of laboratories which incorporate the proposed functionalities have been developed. A set of assessments designed to evaluate the effectiveness of the software is also presented. Suhas Ranganath, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Mahesh K. Banavar, Andreas Spanias |
FIE | 2 |
| 2012 | Interactive DSP laboratories on mobile phones and tabletsabstractThe use of mobile devices and tablets in engineering education has been gaining lot of interest, due to its interactive capabilities and its ability to stimulate student interest. On the other hand, this technology can also enable instructors to broaden the scope of their curriculum and increase student participation. In this paper, we describe an interactive application to perform signal processing simulations on iOS devices such as the iPhone and the iPad. Furthermore, we describe two laboratory exercises to introduce continuous/discrete convolution and filter design. The exercises and the proposed application will be evaluated by students of an undergraduate DSP course at Arizona State University during Fall 2011. Finally, we describe the planned assessment methodology which will enable us to provide prescriptive recommendations for using i-JDSP in DSP courses. Jinru Liu, Jayaraman J. Thiagarajan, Xue Zhang 0002, Suhas Ranganath, Mahesh K. Banavar, Andreas Spanias |
ICASSP | 3 |
| 2012 | Supervised local sparse coding of sub-image features for image retrievalabstractThe success of sparse representations in image modeling and recovery has motivated its use in computer vision applications. Image retrieval and classification tasks require extracting features that discriminate different image classes. State-of-the-art object recognition methods based on sparse coding use spatial pyramid features obtained from dense descriptors. In this paper, we develop a feature extraction method that uses multiple global/local features extracted from large overlapping regions of an image, which we refer to as sub-images. We propose a procedure for dictionary design and supervised local sparse coding of sub-image heterogeneous features. We perform image retrieval on the Microsoft Research Cambridge image dataset and show that the proposed features outperform the spatial pyramid features obtained using dense descriptors. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Andreas Spanias |
ICIP | 1 |
| 2011 | Work in progress - Interactive signal-processing labs and simulations on iOS devicesabstractHandheld devices are increasingly finding more applications in STEM education. In this paper, we present the design of an interactive signal processing simulation software operating on both the iPhone OS (iOS) and Android platforms. This object-oriented application is called i-JDSP and is conceptually based on the award-winning Java-DSP (J-DSP) simulation environment. The i-JDSP app offers a user-friendly visual programming interface and provides users with a compelling multi-touch programming experience. It supports basic signal processing simulation functions such as the FFT, filtering, frequency response, pole-zero plots, and sound recording and playback. Initial assessments have been promising and we believe that this new attractive smartphone interface will make signal processing education among undergraduate students more appealing. Jinru Liu, Andreas Spanias, Mahesh K. Banavar, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Xue Zhang 0002 |
FIE | 4 |
| 2011 | Work in progress - Modules and laboratories for a pathways course in signals and systemsabstractA gap between theory and practice in signals and systems courses is often reported at many universities as a key problem in recruiting signals and systems students. On the other hand, instructors often cite a lack of fundamental understanding in mathematics as an issue in this course. Students seem to be discontent with some of the abstraction of the signals and systems courses. In this work-in-progress paper, we describe a new pathways concept we introduced to address these problems by introducing in-depth discussions, several applications and hands-on exercises. Kostas Tsakalis, Jayaraman J. Thiagarajan, Tolga M. Duman, Martin Reisslein, G. Tong Zhou, Xiaoli Ma, Photini Spanias |
FIE | 2 |
| 2011 | Improved sparse coding using manifold projectionsabstractSparse representations using predefined and learned dictionaries have widespread applications in signal and image processing. Sparse approximation techniques can be used to recover data from its low dimensional corrupted observations, based on the knowledge that the data is sparsely representable using a known dictionary. In this paper, we propose a method to improve data recovery by ensuring that the data recovered using sparse approximation is close its manifold. This is achieved by performing regularization using examples from the data manifold. This technique is particularly useful when the observations are highly reduced in dimensions when compared to the data and corrupted with high noise. Using an example application of image inpainting, we demonstrate that the proposed algorithm achieves a reduction in reconstruction error in comparison to using only sparse coding with predefined and learned dictionaries, when the percentage of missing pixels is high. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICIP | 2 |
| 2011 | Optimality and stability of the K-hyperline clustering algorithm
Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
Pattern Recognit. Lett. | 1 |
| 2009 | Fast image registration with non-stationary Gauss-Markov random field templatesabstractNon-stationary Gauss-Markov random fields are required in modeling images with complex patterns. In this paper, we propose a framework for registering images to a non-stationary Gauss-Markov random field template in an M×M lattice, with a complexity of order M2log M, considering only global translations. We simplify the likelihood computation by expressing it as a scalar product and we estimate the maximal likelihood translation using 2-D FFTs. We demonstrate the utility of this framework by applying it to image registration in a wavelet-domain template learning application. Results reveal that significant complexity reduction is achieved in image registration compared to straightforward registration in the wavelet domain. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICIP | 2 |