VLDB 2026 Research / reviewers in the wild / expert
Chandra Kiran Reddy Evuru
dblp:355/1221
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 24% Vision and language · 22% Speech recognition and synthesis · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.6 | 2 | 2025 | Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs · ICLR 2025 GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities · EMNLP 2024 |
Natural language and speech › Speech recognition and synthesis
audio-language model |
1.5 | 2 | 2024 | CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models · ICLR 2024 GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities · EMNLP 2024 |
Machine learning › Deep learning architectures and training
data augmentation |
1.4 | 2 | 2024 | ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions · ACL (1) 2024 DALE: Generative Data Augmentation for Low-Resource Legal NLP · EMNLP 2023 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.9 | 1 | 2025 | Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › audio-language model
audio reasoning |
0.8 | 1 | 2024 | GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities · EMNLP 2024 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.8 | 1 | 2024 | CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models · ICLR 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models · ICLR 2024 |
Machine learning › Trustworthy machine learning
hallucination |
0.8 | 1 | 2024 | A Closer Look at the Limitations of Instruction Tuning · ICML 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.8 | 1 | 2024 | A Closer Look at the Limitations of Instruction Tuning · ICML 2024 |
Natural language and speech › Language models and text generation › natural language understanding
low-resource natural language understanding |
0.8 | 1 | 2024 | ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions · ACL (1) 2024 |
Machine learning › Deep learning architectures and training › data augmentation
generative data augmentation |
0.7 | 1 | 2023 | DALE: Generative Data Augmentation for Low-Resource Legal NLP · EMNLP 2023 |
Natural language and speech › Language models and text generation
decoding |
0.3 | 1 | 2025 | Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
visual description grounded decoding · 0.9KL divergence · 0.9modular contrastive loss · 0.8large language model · 0.8instruction tuning · 0.8hard negatives · 0.8full-parameter fine-tuning · 0.8contrastive learning · 0.8audio language model · 0.8LoRA fine-tuning · 0.8generative model · 0.7LLM-based augmentation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReCLAP: Improving Zero Shot Audio Classification by Describing SoundsabstractOpen-vocabulary audio-language models, like CLAP [1], offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC with CLAP. Specifically, we shift from the conventional method of using prompts with abstract category labels (e.g., Sound of an organ) to prompts that describe sounds using their inherent descriptive features in a diverse context (e.g., The organ’s deep and resonant tones filled the cathedral.). To achieve this, we first propose ReCLAP, a CLAP model trained with rewritten audio captions for improved understanding of sounds in the wild. These rewritten captions describe each sound event in the original caption using their unique discriminative characteristics. ReCLAP outperforms all baselines on both multi-modal audio-text retrieval and ZSAC. Next, to improve zero-shot audio classification with ReCLAP, we propose prompt augmentation. In contrast to the traditional method of employing hand-written template prompts, we generate custom prompts for each unique label in the dataset. These custom prompts first describe the sound event in the label and then employ them in diverse scenes. Our proposed method improves ReCLAP’s performance on ZSAC by 1%-18% and outperforms 1 all baselines by 1% - 55%1. Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha |
ICASSP | 3 |
| 2025 | Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMsabstractLarge Vision-Language Models (LVLMs) often produce responses that misalign with factual information, a phenomenon known as hallucinations. While hallucinations are well-studied, the exact causes behind them remain underexplored. In this paper, we first investigate the root causes of hallucinations in LVLMs. Our findings reveal that existing mitigation techniques primarily reduce hallucinations for visual recognition prompts—those that require simple descriptions of visual elements—but fail for cognitive prompts that demand deliberate reasoning. We identify the core issue as a lack of true visual perception in LVLMs: although they can accurately recognize visual elements, they struggle to fully interpret these elements in the context of the input prompt and effectively link this recognition to their internal knowledge, which is critical for reasoning. To address this gap, we introduce Visual Description Grounded Decoding (VDGD), a simple, robust, and training-free method designed to enhance visual perception and improve reasoning capabilities in LVLMs. VDGD works by first generating a detailed description of the image and appending it as a prefix to the instruction. During response generation, tokens are sampled based on their KL divergence to the description, favoring candidates with lower divergence. Experimental results on multiple visual reasoning benchmarks and LVLMs demonstrate that VDGD consistently outperforms existing baselines 2% - 33%. Finally, we introduce VaLLu, a benchmark designed for comprehensive evaluation of the cognitive capabilities of LVLMs. Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Utkarsh Tyagi, Oriol Nieto, Zeyu Jin, Dinesh Manocha |
ICLR | 2 |
| 2025 | ProSE: Diffusion Priors for Speech EnhancementabstractSonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha |
NAACL (Long Papers) | 5 |
| 2024 | ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsabstractSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Evuru, Ramaneswaran S, S Sakshi, Dinesh Manocha. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramaneswaran S., S. Sakshi, Dinesh Manocha |
ACL (1) | 4 |
| 2024 | GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning AbilitiesabstractSreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S. Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha |
EMNLP | 4 |
| 2024 | Recap: Retrieval-Augmented Audio CaptioningabstractWe present RECAP (REtrieval-Augmented Audio CAPtioning), a novel and effective audio captioning system that generates captions conditioned on an input audio and other captions similar to the audio retrieved from a datastore. Additionally, our proposed method can transfer to any domain without the need for any additional fine-tuning. To generate a caption for an audio sample, we leverage an audio-text model CLAP [1] to retrieve captions similar to it from a replaceable datastore, which are then used to construct a prompt. Next, we feed this prompt to a GPT-2 decoder and introduce cross-attention layers between the CLAP encoder and GPT-2 to condition the audio for caption generation. Experiments on two benchmark datasets, Clotho and AudioCaps, show that RECAP achieves competitive performance in in-domain settings and significant improvements in out-of-domain settings. Additionally, due to its capability to exploit a large text-captions-only datastore in a training-free fashion, RECAP shows unique capabilities of captioning novel audio events never seen during training and compositional audios with multiple events. To promote research in this space, we also release 150,000+ new weakly labeled captions for AudioSet, AudioCaps, and Clotho1. Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha |
ICASSP | 3 |
| 2024 | CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsabstractA fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot audio classification, audio retrieval, etc. However, the ability of these models to effectively perform compositional reasoning remains largely unexplored and necessitates additional research. In this paper, we propose CompA, a collection of two expert-annotated benchmarks with a majority of real-world audio samples, to evaluate compositional reasoning in ALMs. Our proposed CompA-order evaluates how well an ALM understands the order or occurrence of acoustic events in audio, and CompA-attribute evaluates attribute-binding of acoustic events. An instance from either benchmark consists of two audio-caption pairs, where both audios have the same acoustic events but with different compositions. An ALM is evaluated on how well it matches the right audio to the right caption. Using this benchmark, we first show that current ALMs perform only marginally better than random chance, thereby struggling with compositional reasoning. Next, we propose CompA-CLAP, where we fine-tune CLAP using a novel learning method to improve its compositional reasoning abilities. To train CompA-CLAP, we first propose improvements to contrastive training with composition-aware hard negatives, allowing for more focused training. Next, we propose a novel modular contrastive loss that helps the model learn fine-grained compositional understanding and overcomes the acute scarcity of openly available compositional audios. CompA-CLAP significantly improves over all our baseline models on the CompA benchmark, indicating its superior compositional reasoning capabilities. Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S., Sakshi Singh, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha |
ICLR | 5 |
| 2024 | A Closer Look at the Limitations of Instruction TuningabstractInstruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limitations and shortcomings remain underexplored. In this paper, through rigorous experiments and an in-depth analysis of the changes LLMs undergo through IT, we reveal various limitations of IT. In particular, we show that (1) IT fails to enhance knowledge or skills in LLMs. LoRA fine-tuning is limited to learning response initiation and style tokens, and full-parameter fine-tuning leads to knowledge degradation. (2) Copying response patterns from IT datasets derived from knowledgeable sources leads to a decline in response quality. (3) Full-parameter fine-tuning increases hallucination by inaccurately borrowing tokens from conceptually similar instances in the IT dataset for generating responses. (4) Popular methods to improve IT do not lead to performance improvements over a simple LoRA fine-tuned model. Our findings reveal that responses generated solely from pre-trained knowledge consistently outperform responses by models that learn any form of new knowledge from IT on open-source datasets. We hope the insights and challenges revealed in this paper inspire future work in related directions. Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Deepali Aneja, Zeyu Jin, Ramani Duraiswami, Dinesh Manocha |
ICML | 2 |
| 2023 | DALE: Generative Data Augmentation for Low-Resource Legal NLPabstractSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, Dinesh Manocha. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Sakshi Singh, Utkarsh Tyagi, Dinesh Manocha |
EMNLP | 2 |