VLDB 2026 Research / reviewers in the wild / expert
Arijit Ray
dblp:164/9384
· DBLP profile ↗
6ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Vision and language · 57% Reinforcement learning · 15% Language models and text generation · 15% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
0.8 | 1 | 2024 | Feedback-Guided Autonomous Driving · CVPR 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Feedback-Guided Autonomous Driving · CVPR 2024 |
Computer vision › Vision and language
compositionality |
0.7 | 1 | 2023 | Cola: A Benchmark for Compositional Text-to-image Retrieval · NeurIPS 2023 |
Computer vision › Vision and language › image-text retrieval
text-to-image retrieval |
0.7 | 1 | 2023 | Cola: A Benchmark for Compositional Text-to-image Retrieval · NeurIPS 2023 |
Audio and music processing
source separation |
0.7 | 1 | 2023 | Language-Guided Audio-Visual Source Separation via Trimodal Consistency · CVPR 2023 |
Computer vision › Vision and language
visual question answering |
0.6 | 2 | 2019 | Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation · EMNLP/IJCNLP (1) 2019 Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions · EMNLP 2016 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.4 | 1 | 2019 | Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems
question understanding |
0.2 | 1 | 2016 | Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions · EMNLP 2016 |
Computer vision › Vision and language › vision-language model
vision-language model adaptation |
0.2 | 1 | 2023 | Cola: A Benchmark for Compositional Text-to-image Retrieval · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
vision-language foundation model · 1.3self-supervised learning · 1.3pseudo-target supervision · 1.3large language model feedback · 0.8corrective feedback · 0.8multimodal attention layer · 0.7fine-tuning · 0.7contrastive learning · 0.7question generation · 0.4entailment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Feedback-Guided Autonomous DrivingabstractWhile behavior cloning has recently emerged as a highly successful paradigm for autonomous driving, humans rarely learn to perform complex tasks, such as driving, via imitation or behavior cloning alone. In contrast, learning in humans often involves additional detailed guidance throughout the interactive learning process, i.e., where feedback, often via language, provides detailed information as to which part of their trial was performed incorrectly or suboptimally and why. Motivated by this observation, we introduce an efficient feedback-based framework for improving behavior-cloning-based training of sensorimotor driving agents. Our key insight is to leverage recent advances in Large Language Models (LLMs) to provide corrective fine-grained feedback regarding the underlying reason behind driving prediction failures. Moreover, our introduced network architecture is efficient, enabling the first sensorimotor end-to-end training and evaluation of LLM-based driving models. The resulting agent achieves state-of-the-art performance in open-loop evaluation on nuScenes, outperforming prior state-of-the-art by over 8.1% and 57.1% in accuracy and collision rate, respectively. In CARLA, our camera-based agent improves by 16.6% in driving score over prior LIDAR-based approaches. Jimuyang Zhang, Zanming Huang, Arijit Ray, Eshed Ohn-Bar |
CVPR | 3 |
| 2023 | Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyabstractWe propose a self-supervised approach for learning to perform audio source separation in videos based on natu-ral language queries, using only unlabeled video and au-dio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object to its visual features and the corresponding components of the audio waveform, all without access to annotations during training. To overcome this challenge, we adapt off-the-shelf vision-language foundation models to provide pseudo-target supervision via two novel loss functions and encourage a stronger alignment between the audio, visual and natural language modalities. During inference, our approach can separate sounds given text, video and audio input, or given text and audio input alone. We demonstrate the effectiveness of our self-supervised approach on three audio-visual separation datasets, including MUSIC, SOLOS and AudioSet, where we outperform state-of-the-art strongly supervised approaches despite not using object detectors or text labels during training. Our project page including publicly available code can be found at https://cs-people.bu.edu/rxtan/projectsNAST. Reuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer, Justin Salamon, Oriol Nieto, Bryan C. Russell, Kate Saenko |
CVPR | 2 |
| 2023 | Cola: A Benchmark for Compositional Text-to-image RetrievalabstractCompositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To measure this lack of compositional capability, we design Cola, a text-to-image retrieval benchmark to Compose Objects Localized with Attributes. To solve Cola, a model must retrieve images with the correct configuration of attributes and objects and avoid choosing a distractor image with the same objects and attributes but in the wrong configuration. Cola contains about 1.2k composed queries of 168 objects and 197 attributes on around 30K images. Our human evaluation finds that Cola is 83.33% accurate, similar to contemporary compositionality benchmarks. Using Cola as a testbed, we explore empirical modeling designs to adapt pre-trained vision-language models to reason compositionally. We explore 6 adaptation strategies on 2 seminal vision-language models, using compositionality-centric test benchmarks - Cola and CREPE. We find the optimal adaptation strategy is to train a multi-modal attention layer that jointly attends over the frozen pre-trained image and language features. Surprisingly, training multimodal layers on CLIP performs better than tuning a larger FLAVA model with already pre-trained multimodal layers. Furthermore, our adaptation strategy improves CLIP and FLAVA to comparable levels, suggesting that training multimodal layers using contrastive attribute-object data is key, as opposed to using them pre-trained. Lastly, we show that Cola is harder than a closely related contemporary benchmark, CREPE, since simpler fine-tuning strategies without multimodal layers suffice on CREPE, but not on Cola. However, we still see a significant gap between our best adaptation and human accuracy, suggesting considerable room for further research. Project page: https://cs-people.bu.edu/array/research/cola/ Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan A. Plummer, Ranjay Krishna, Kate Saenko |
NeurIPS | 1 |
| 2019 | Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question GenerationabstractArijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image RetrievalabstractWhile there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting human-AI collaborative tasks. To bridge the gap, we propose a Twenty-Questions style collaborative image retrieval game, Explanation-assisted Guess Which (ExAG), as a method of evaluating the efficacy of explanations (visual evidence or textual justification) in the context of Visual Question Answering (VQA). In our proposed ExAG, a human user needs to guess a secret image picked by the VQA agent by asking natural language questions to it. We show that overall, when AI explains its answers, users succeed more often in guessing the secret image correctly. Notably, a few correct explanations can readily improve human performance when VQA answers are mostly incorrect as compared to no-explanation games. Furthermore, we also show that while explanations rated as “helpful” significantly improve human performance, “incorrect” and “unhelpful” explanations can degrade performance as compared to no-explanation games. Our experiments, therefore, demonstrate that ExAG is an effective means to evaluate the efficacy of AI-generated explanation on a human-AI collaborative task. Arijit Ray, Rakesh Kumar 0001, Ajay Divakaran, Giedrius Burachas |
HCOMP | 1 |
| 2016 | Question Relevance in VQA: Identifying Non-Visual And False-Premise QuestionsabstractVisual Question Answering (VQA) is the task of answering natural-language questions about images.We introduce the novel problem of determining the relevance of questions to images in VQA.Current VQA models do not reason about whether a question is even related to the given image (e.g., What is the capital of Argentina?) or if it requires information from external resources to answer correctly.This can break the continuity of a dialogue in human-machine interaction.Our approaches for determining relevance are composed of two stages.Given an image and a question, (1) we first determine whether the question is visual or not, (2) if visual, we determine whether the question is relevant to the given image or not.Our approaches, based on LSTM-RNNs, VQA model uncertainty, and caption-question similarity, are able to outperform strong baselines on both relevance tasks.We also present human studies showing that VQA models augmented with such question relevance reasoning are perceived as more intelligent, reasonable, and human-like. Arijit Ray, Gordon A. Christie, Mohit Bansal, Dhruv Batra, Devi Parikh |
EMNLP | 1 |