VLDB 2026 Research / reviewers in the wild / expert
Raghav Kapoor
dblp:200/2387
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-1030-7423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 26% Multi-agent systems · 26% Reinforcement learning · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 89% Data integration and cleaning · 11% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents |
0.8 | 1 | 2024 | OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web · ECCV (68) 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | SkillCLIP: Skill Aware Modality Fusion Visual Question Answering (Student Abstract) · AAAI 2024 |
Machine learning › Reinforcement learning › exploration
embodied exploration |
0.7 | 1 | 2023 | EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023 |
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.7 | 1 | 2023 | EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023 |
Natural language and speech › Information extraction and text analysis › abusive language detection
hate speech detection |
0.4 | 1 | 2019 | Mind Your Language: Abuse and Offense Detection for Code-Switched Languages · AAAI 2019 |
Query processing and optimization › query optimization
cost-based optimization |
0.3 | 1 | 2017 | Provenance-Aware Query Optimization · ICDE 2017 |
Computer vision › Vision and language
cross-modal interaction |
0.2 | 1 | 2024 | OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web · ECCV (68) 2024 |
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
interactive agents |
0.2 | 1 | 2023 | EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 1 | 2019 | Mind Your Language: Abuse and Offense Detection for Code-Switched Languages · AAAI 2019 |
Data integration and cleaning
data provenance |
0.1 | 1 | 2019 | Heuristic and Cost-Based Optimization for Diverse Provenance Tasks · IEEE Trans. Knowl. Data Eng. 2019 |
Methods — techniques the papers use, named apart from their topics
skill embedding · 0.8modality fusion · 0.8expert modules · 0.8desktop and web automation benchmark · 0.8heuristic optimization · 0.7algebraic equivalences · 0.7virtual reality interface · 0.7transfer learning · 0.4cost-based optimization · 0.4LSTM · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SkillCLIP: Skill Aware Modality Fusion Visual Question Answering (Student Abstract)abstractWhen humans are posed with a difficult problem, they often approach it by identifying key skills, honing them, and finally effectively combining them. We propose a novel method and apply it for the VizWiz VQA task to predict the visual skills needed to answer a question, and leverage expert modules to produce intermediary outputs and fuse them in a skill-aware manner. Unlike prior works in visual question-answering (VQA) that use intermediate outputs such as detected objects and Optical Character Recognition (OCR), our approach explicitly guides the model with a skill embedding on what to focus on. While our results show that using skill-aware fusion outperforms skill-unaware models for only a subset of questions, we believe our results provide interesting directions for future work. We also release our code, model, and illustrative demonstrations for future research purposes. Atharva Naik, Yash Butala, Navaneethan Vaikunthan, Raghav Kapoor |
AAAI | 4 |
| 2024 | OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Raghav Kapoor, Yash Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, Ruslan Salakhutdinov |
ECCV (68) | 1 |
| 2023 | EXCALIBUR: Encouraging and Evaluating Embodied ExplorationabstractExperience precedes understanding. Humans constantly explore and learn about their environment out of curiosity, gather information, and update their models of the world. On the other hand, machines are either trained to learn passively from static and fixed datasets, or taught to complete specific goal-conditioned tasks. To encourage the development of exploratory interactive agents, we present the EXCALIBUR benchmark. EXCALIBUR allows agents to explore their environment for long durations and then query their understanding of the physical world via inquiries like: “is the small heavy red bowl made from glass?” or “is there a silver spoon heavier than the egg?”. This design encourages agents to perform free-form home exploration without myopia induced by goal conditioning. Once the agents have answered a series of questions, they can renter the scene to refine their knowledge, update their beliefs, and improve their performance on the questions. Our experiments demonstrate the challenges posed by this dataset for the present-day state-of-the-art embodied systems and the headroom afforded to develop new innovative methods. Finally, we present a virtual reality interface that enables humans to seamlessly interact within the simulated world and use it to gather human performance measures. EXCALIBUR affords unique challenges in comparison to presentday benchmarks and represents the next frontier for embodied AI research. Hao Zhu 0011, Raghav Kapoor, So Yeon Min, Winson Han, Jiatai Li, Kaiwen Geng, Graham Neubig, Yonatan Bisk, Aniruddha Kembhavi, Luca Weihs |
CVPR | 2 |
| 2023 | MoEmo Vision Transformer: Integrating Cross-Attention and Movement Vectors in 3D Pose Estimation for HRI Emotion DetectionabstractEmotion detection presents challenges to intelligent human-robot interaction (URI). Foundational deep learning techniques used in emotion detection are limited by information-constrained datasets or models that lack the necessary complexity to learn interactions between input data elements, such as the the variance of human emotions across different contexts. In the current effort, we introduce 1) MoEmo (Motion to Emotion), a cross-attention vision transformer (ViT) for human emotion detection within robotics systems based on 3D human pose estimations across various contexts, and 2) a data set that offers full-body videos of human movement and corresponding emotion labels based on human gestures and environmental contexts. Compared to existing approaches, our method effectively leverages the subtle connections between movement vectors of gestures and environmental contexts through the use of cross-attention on the extracted movement vectors of full-body human gestures/poses and feature maps of environmental contexts. We implement a cross-attention fusion model to combine movement vectors and environment contexts into a joint representation to derive emotion estimation. Leveraging our Naturalistic Motion Database, we train the MoEmo system to jointly analyze motion and context, yielding emotion detection that outperforms the current state-of-the-art. David C. Jeong, Tianma Shen, Hongji Liu, Raghav Kapoor, Casey Nguyen, Song Liu 0003, Christopher Kitts |
IROS | 4 |
| 2019 | Mind Your Language: Abuse and Offense Detection for Code-Switched LanguagesabstractIn multilingual societies like the Indian subcontinent, use of code-switched languages is much popular and convenient for the users. In this paper, we study offense and abuse detection in the code-switched pair of Hindi and English (i.e, Hinglish), the pair that is the most spoken. The task is made difficult due to non-fixed grammar, vocabulary, semantics and spellings of Hinglish language. We apply transfer learning and make a LSTM based model for hate speech classification. This model surpasses the performance shown by the current best models to establish itself as the state-of-the-art in the unexplored domain of Hinglish offensive text classification. We also release our model and the embeddings trained for research purposes. Raghav Kapoor, Yaman Singla, Kshitij Rajput, Rajiv Ratn Shah, Ponnurangam Kumaraguru, Roger Zimmermann |
AAAI | 1 |
| 2019 | Heuristic and Cost-Based Optimization for Diverse Provenance TasksabstractA well-established technique for capturing database provenance as annotations on data is to instrument queries to propagate such annotations. However, even sophisticated query optimizers often fail to produce efficient execution plans for instrumented queries. We develop provenance-aware optimization techniques to address this problem. Specifically, we study algebraic equivalences targeted at instrumented queries and alternative ways of instrumenting queries for provenance capture. Furthermore, we present an extensible heuristic and cost-based optimization framework utilizing these optimizations. Our experiments confirm that these optimizations are highly effective, improving performance by several orders of magnitude for diverse provenance tasks. Xing Niu 0002, Raghav Kapoor, Boris Glavic, Dieter Gawlick, Zhen Hua Liu, Vasudha Krishnaswamy, Venkatesh Radhakrishnan |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Provenance-Aware Query OptimizationabstractData provenance is essential for debugging query results, auditing data in cloud environments, and explaining outputs of Big Data analytics. A well-established technique is to represent provenance as annotations on data and to instrument queries to propagate these annotations to produce results annotated with provenance. However, even sophisticated optimizers are often incapable of producing efficient execution plans for instrumented queries, because of their inherent complexity and unusual structure. Thus, while instrumentation enables provenance support for databases without requiring any modification to the DBMS, the performance of this approach is far from optimal. In this work, we develop provenancespecific optimizations to address this problem. Specifically, we introduce algebraic equivalences targeted at instrumented queries and discuss alternative, equivalent ways of instrumenting a query for provenance capture. Furthermore, we present an extensible heuristic and cost-based optimization (CBO) framework that governs the application of these optimizations and implement this framework in our GProM provenance system. Our CBO is agnostic to the plan space shape, uses a DBMS for cost estimation, and enables retrofitting of optimization choices into existing code by adding a few LOC. Our experiments confirm that these optimizations are highly effective, often improving performance by several orders of magnitude for diverse provenance tasks. Xing Niu 0002, Raghav Kapoor, Boris Glavic, Dieter Gawlick, Zhen Hua Liu, Venkatesh Radhakrishnan |
ICDE | 2 |