Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chandramouli Shama Sastry

dblp:223/6317 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 34% Trustworthy machine learning · 18% Vision and language · 17%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522024
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers · NeurIPS 2024
Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion · ICML 2024
Machine learning › Trustworthy machine learning
robustness
1.222024
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers · NeurIPS 2024
Detecting Out-of-Distribution Examples with Gram Matrices · ICML 2020
Machine learning › Optimization for machine learning
neural network training acceleration
0.912025
Accelerating neural network training: An analysis of the AlgoPerf competition · ICLR 2025
Computer vision › Vision and language
compositionality
0.812024
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations · NeurIPS 2024
Machine learning › Deep learning architectures and training
data augmentation
0.812024
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation
0.812024
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
guided diffusion
0.812024
Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion · ICML 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations · NeurIPS 2024
Audio and music processing › music generation
symbolic music generation
0.812024
Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion · ICML 2024
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.412020
Detecting Out-of-Distribution Examples with Gram Matrices · ICML 2020
Machine learning › Optimization for machine learning
preconditioning
0.312025
Accelerating neural network training: An analysis of the AlgoPerf competition · ICLR 2025
Computer vision › Image recognition and object detection
image classification
0.212024
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

stochastic control guidance · 1.5latent diffusion · 1.5schedule-free adamw · 0.9hyperparameter tuning · 0.9distributed shampoo · 0.9reverse diffusion · 0.8forward diffusion · 0.8deepaugment · 0.8benchmark dataset construction · 0.8augmix · 0.8
YearPublicationVenuePosition
2025 Accelerating neural network training: An analysis of the AlgoPerf competition
abstract
The goal of the AlgoPerf: Training Algorithms competition is to evaluate practical speed-ups in neural network training achieved solely by improving the underlying training algorithms. In the external tuning ruleset, submissions must provide workload-agnostic hyperparameter search spaces, while in the self-tuning ruleset they must be completely hyperparameter-free. In both rulesets, submissions are compared on time-to-result across multiple deep learning workloads, training on fixed hardware. This paper presents the inaugural AlgoPerf competition's results, which drew 18 diverse submissions from 10 teams. Our investigation reveals several key findings: (1) The winning submission in the external tuning ruleset, using Distributed Shampoo, demonstrates the effectiveness of non-diagonal preconditioning over popular methods like Adam, even when compared on wall-clock runtime. (2) The winning submission in the self-tuning ruleset, based on the Schedule Free AdamW algorithm, demonstrates a new level of effectiveness for completely hyperparameter-free training algorithms. (3) The top-scoring submissions were surprisingly robust to workload changes. We also discuss the engineering challenges encountered in ensuring a fair comparison between different training algorithms. These results highlight both the significant progress so far, and the considerable room for further improvements.
Priya Kasimbeg, Frank Schneider 0001, Runa Eschenhagen, Juhan Bae, Chandramouli Shama Sastry, Mark Saroufim, Boyuan Feng, Less Wright, Edward Z. Yang, Zachary Nado, Sourabh Medapati, Philipp Hennig, Michael G. Rabbat, George E. Dahl
ICLR5
2025 Test-Time Training for Speech-based Depression Detection
Sri Harsha Dumpala, Chandramouli Shama Sastry, Rudolf Uher, Sageev Oore
INTERSPEECH2
2024 Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
abstract
We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density or chord progression, many of which are non-differentiable which pose a challenge when using them for guided diffusion. We propose Stochastic Control Guidance (SCG), a novel guidance method that only requires forward evaluation of rule functions that can work with pre-trained diffusion models in a plug-and-play way, thus achieving training-free guidance for non-differentiable rules for the first time. Additionally, we introduce a latent diffusion architecture for symbolic music generation with high time resolution, which can be composed with SCG in a plug-and-play fashion. Compared to standard strong baselines in symbolic music generation, this framework demonstrates marked advancements in music quality and rule-based controllability, outperforming current state-of-the-art generators in a variety of settings. For detailed demonstrations, code and model checkpoints, please visit our [project website](https://scg-rule-guided-music.github.io/).
Yujia Huang, Adishree Ghatare, Yuanzhe Liu 0001, Ziniu Hu, Qinsheng Zhang, Chandramouli Shama Sastry, Siddharth Gururani, Sageev Oore, Yisong Yue
ICML6
2024 XANE: eXplainable Acoustic Neural Embeddings
Sri Harsha Dumpala, Dushyant Sharma, Chandramouli Shama Sastry, Stanislav Yu. Kruchinin, James Fosburgh, Patrick A. Naylor
INTERSPEECH3
2024 SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
abstract
Despite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understand precise semantics. For example, semantically equivalent sentences expressed using different lexical compositions elicit diverging representations. The degree of this divergence and its impact on encoded semantics is not very well understood. In this paper, we introduce the SUGARCREPE++ dataset to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. Each sample in SUGARCREPE++ dataset consists of an image and a corresponding triplet of captions: a pair of semantically equivalent but lexically different positive captions and one hard negative caption. This poses a 3-way semantic (in)equivalence problem to the language models. We comprehensively evaluate VLMs and ULMs that differ in architecture, pre-training objectives and datasets to benchmark the performance of SUGARCREPE++ dataset. Experimental results highlight the difficulties of VLMs in distinguishing between lexical and semantic variations, particularly to object attributes and spatial relations. Although VLMs with larger pre-training datasets, model sizes, and multiple pre-training objectives achieve better performance on SUGARCREPE++, there is a significant opportunity for improvement. We demonstrate that models excelling on compositionality datasets may not perform equally well on SUGARCREPE++. This indicates that compositionality alone might not be sufficient to fully understand semantic and lexical alterations. Given the importance of the property that the SUGARCREPE++ dataset targets, it serves as a new challenge to the vision-and-language community. Data and code is available at https://github.com/Sri-Harsha/scpp.
Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Shama Sastry, Evangelos E. Milios, Sageev Oore, Hassan Sajjad 0001
NeurIPS3
2024 DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers
abstract
We introduce DiffAug, a simple and efficient diffusion-based augmentation technique to train image classifiers for the crucial yet challenging goal of improved classifier robustness. Applying DiffAug to a given example consists of one forward-diffusion step followed by one reverse-diffusion step. Using both ResNet-50 and Vision Transformer architectures, we comprehensively evaluate classifiers trained with DiffAug and demonstrate the surprising effectiveness of single-step reverse diffusion in improving robustness to covariate shifts, certified adversarial accuracy and out of distribution detection. When we combine DiffAug with other augmentations such as AugMix and DeepAugment we demonstrate further improved robustness. Finally, building on this approach, we also improve classifier-guided diffusion wherein we observe improvements in: (i) classifier-generalization, (ii) gradient quality (i.e., improved perceptual alignment) and (iii) image generation performance. We thus introduce a computationally efficient technique for training with improved robustness that does not require any additional data, and effectively complements existing augmentation approaches.
Chandramouli Shama Sastry, Sri Harsha Dumpala, Sageev Oore
NeurIPS1
2022 On Combining Global and Localized Self-Supervised Models of Speech
Sri Harsha Dumpala, Chandramouli Shama Sastry, Rudolf Uher, Sageev Oore
INTERSPEECH2
2021 Controlling BigGAN Image Generation with a Segmentation Network
Aman Jaiswal, Harpreet Singh Sodhi, Mohamed Muzamil H, Rajveen Singh Chandhok, Sageev Oore, Chandramouli Shama Sastry
DS6
2020 Detecting Out-of-Distribution Examples with Gram Matrices
abstract
When presented with Out-of-Distribution (OOD) examples, deep neural networks yield confident, incorrect predictions; detecting OOD examples is challenging, and the potential risks are high. In this paper, we propose to detect OOD examples by identifying inconsistencies between activity patterns and predicted class. We find that characterizing activity patterns by Gram matrices and identifying anomalies in Gram matrix values can yield high OOD detection rates. We identify anomalies in the Gram matrices by simply comparing each value with its respective range observed over the training data. Unlike many approaches, this can be used with any pre-trained softmax classifier and neither requires access to OOD data for fine-tuning hyperparameters, nor does it require OOD access for inferring parameters. We empirically demonstrate applicability across a variety of architectures and vision datasets and, for the important and surprisingly hard task of detecting far out-of-distribution examples, it generally performs better than or equal to state-of-the-art OOD detection methods (including those that do assume access to OOD examples).
Chandramouli Shama Sastry, Sageev Oore
ICML1
2020 Active neural learners for text with dual supervision
Chandramouli Shama Sastry, Evangelos E. Milios
Neural Comput. Appl.1
2017 Visualizing Textbook Concepts: Beyond Word Co-occurrences
Chandramouli Shama Sastry, Darshan Siddesh Jagaluru, Kavi Mahesh
CICLing (1)1