VLDB 2026 Research / reviewers in the wild / expert
Chancharik Mitra
dblp:360/6323
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0008-9826-7534ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational AssistantsabstractWith the growing use of Large Language Model (LLM)-based Question-Answering (QA) systems in education, it is critical to evaluate their performance across individual pipeline components. In this work, we introduce EduMod-LLM, a modular function-calling LLM pipeline, and present a comprehensive evaluation along three key axes: function calling strategies, retrieval methods, and generative language models. Our framework enables fine-grained analysis by isolating and assessing each component. We benchmark function-calling performance across LLMs, compare our novel structure-aware retrieval method to vector-based and LLM-scoring baselines, and evaluate various LLMs for response synthesis. This modular approach reveals specific failure modes and performance patterns, supporting the development of interpretable and effective educational QA systems. Our findings demonstrate the value of modular function calling in improving system transparency and pedagogical alignment. Meenakshi Mittal, Rishi Khare, Mihran Miroyan, Chancharik Mitra, Narges Norouzi |
AAAI | 4 |
| 2026 | Improving Online Learning: Using Utterance Distribution to Improve Student-Facing Assistants in Discussion ForumsabstractRecent advancements in large language models (LLMs) have paved the way for AI educational assistants in academic settings. However, AI assistants often respond differently than TAs, providing extensive explanations that may overwhelm students or inadvertently reveal more than intended. This study identifies the main differences between TA and LLM responses to students by using a four-class utterance classification system to compare the utterance distributions found in TA replies and in responses generated by Edison, a state-of-the-art AI educational assistant. Using this classification, striking distributional differences are observed: Edison produces far more Advance utterances, whereas TAs use many more React and Social Convention utterances. This research examines how differences in these distributions relate to response quality in student–TA interactions. Through prompt engineering, we align Edison's utterance distribution with TA patterns, producing responses that are more concise, directly address student questions, and avoid unnecessary elaboration. Wolfgang Edholm, Justin Park, Mihran Miroyan, Chancharik Mitra, Narges Norouzi |
SIGCSE (2) | 4 |
| 2026 | Edison 3.0: A Multimodal RAG System for Large-Scale Educational Q&A with Human-in-the-Loop Oversight
Meenakshi Mittal, Rishi Khare, Mihran Miroyan, Chancharik Mitra, Narges Norouzi |
SIGCSE (2) | 4 |
| 2025 | Enhancing Few-Shot Vision-Language Classification With Large Multimodal Model Features
Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin, Assaf Arbelle, Rogério Feris, Leonid Karlinsky, Trevor Darrell, Deva Ramanan, Roei Herzig |
ICCV | 1 |
| 2025 | Towards Understanding Camera Motions in Any VideoabstractWe introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality control process. One of our core contributions is a taxonomy or "language" of camera motion primitives, designed in collaboration with cinematographers. We find, for example, that some motions like "follow" (or tracking) require understanding scene content like moving subjects. We conduct a large-scale human study to quantify human performance, revealing that domain expertise and tutorial-based training can significantly enhance accuracy. For example, a novice may confuse zoom-in (a change of intrinsics) with translating forward (a change of extrinsics), but can be trained to differentiate the two. Using CameraBench, we evaluate Structure-from-Motion (SfM) and Video-Language Models (VLMs), finding that SfM models struggle to capture semantic primitives that depend on scene content, while generative VLMs struggle to capture geometric primitives that require precise estimation of trajectories. We then fine-tune a generative VLM on CameraBench to achieve the best of both worlds and showcase its applications, including motion-augmented captioning, video question answering, and video-text retrieval. We hope our taxonomy, benchmark, and tutorials will drive future efforts towards the ultimate goal of understanding camera motions in any video. Zhiqiu Lin, Siyuan Cen, Jay Karhade, Hewei Wang 0001, Chancharik Mitra, Yu Tong Tiffany Ling, Rushikesh Zawar, Yilun Du, Chuang Gan 0001, Deva Ramanan |
NeurIPS | 6 |
| 2025 | Analyzing Pedagogical Quality and Efficiency of LLM Responses with TA Feedback to Live Student QuestionsabstractWhile Large Language Models (LLMs) have emerged as promising methods for automated student question-answering, guaranteeing consistent instructional effectiveness of the response remains a key challenge. Therefore, there is a need for fine-grained analysis of State-Of-The-Art (SOTA) LLM-powered educational assistants. Mihran Miroyan, Chancharik Mitra, Gireeja Ranade, Narges Norouzi |
SIGCSE (1) | 2 |
| 2025 | Raising the Bar: Automating Consistent and Equitable Student Support with LLMsabstractLarge Language Models (LLMs) can be used to automate many aspects of the educational field. In this paper, we look into the benefits of automating responses to student questions in course discussion forums using our Retrieval-Augmented Generation (RAG)-based LLM pipeline (Edison). Our research questions are: Meenakshi Mittal, Azalea Bailey, Victoria Phelps, Mihran Miroyan, Chancharik Mitra, Rose Niousha, Gireeja Ranade, Narges Norouzi |
SIGCSE (2) | 5 |
| 2024 | RetLLM-E: Retrieval-Prompt Strategy for Question-Answering on Student Discussion ForumsabstractThis paper focuses on using Large Language Models to support teaching assistants in answering questions on large student forums such as Piazza and EdSTEM. Since student questions on these forums are often closely tied to specific aspects of the institution, instructor, and course delivery, general-purpose LLMs do not directly do well on this task. We introduce RetLLM-E, a method that combines text-retrieval and prompting approaches to enable LLMs to provide precise and high-quality answers to student questions. When presented with a student question, our system initiates a two-step process. First, it retrieves relevant context from (i) a dataset of student questions addressed by course instructors (Q&A Retrieval) and (ii) relevant segments of course materials (Document Retrieval). RetLLM-E then prompts LLM using the retrieved text and an engineered prompt structure to yield an answer optimized for the student question. We present a set of quantitative and human evaluation experiments, comparing our method to ground truth answers to questions in a test set of actual student questions. Our results demonstrate that our approach provides higher-quality responses to course-related questions than an LLM operating without context or relying solely on retrieval-based context. RetLLM-E can easily be adopted in different courses, providing instructors and students with context-aware automatic responses. Chancharik Mitra, Mihran Miroyan, Vedant Kumud, Gireeja Ranade, Narges Norouzi |
AAAI | 1 |
| 2024 | Compositional Chain-of-Thought Prompting for Large Multimodal ModelsabstractThe combination of strong visual backbones and Large Language Model (LLM) reasoning has led to Large Multimodal Models (LMMs) becoming the current standard for a wide range of vision and language (VL) tasks. However, recent research has shown that even the most advanced LMMs still struggle to capture aspects of compositional visual reasoning, such as attributes and relationships between objects. One solution is to utilize scene graphs (SGs)-a formalization of objects and their relations and attributes that has been extensively used as a bridge between the visual and textual domains. Yet, scene graph data requires scene graph annotations, which are expensive to collect and thus not easily scalable. Moreover, finetuning an LMM based on SG data can lead to catastrophic forgetting of the pretraining objective. To overcome this, inspired by chain-of-thought methods, we propose Compositional Chain-of-Thought (CCoT), a novel zero-shot Chain-of-Thought prompting method that utilizes SG representations in order to extract compositional knowledge from an LMM. Specifically, we first generate an SG using the LMM, and then use that SG in the prompt to produce a response. Through extensive experiments, we find that the proposed CCoT approach not only improves LMM performance on several vision and language (VL) compositional benchmarks but also improves the performance of several popular LMMs on general multimodal benchmarks, without the need for fine-tuning or annotated ground-truth SGs. Code: https://github.com/chancharikmitra/CCoT. Chancharik Mitra, Brandon Huang, Trevor Darrell, Roei Herzig |
CVPR | 1 |
| 2024 | Which One? Leveraging Context Between Objects and Multiple Views for Language GroundingabstractChancharik Mitra, Abrar Anwar, Rodolfo Corona, Dan Klein, Trevor Darrell, Jesse Thomason. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chancharik Mitra, Abrar Anwar, Rodolfo Corona, Daniel Klein 0001, Trevor Darrell, Jesse Thomason |
NAACL-HLT | 1 |
| 2024 | Multimodal Task Vectors Enable Many-Shot Multimodal In-Context LearningabstractThe recent success of interleaved Large Multimodal Models (LMMs) in few-shot learning suggests that in-context learning (ICL) with many examples can be promising for learning new tasks. However, this many-shot multimodal ICL setting has one crucial problem: it is fundamentally limited by the model's context length set at pretraining. The problem is especially prominent in the multimodal domain, which processes both text and images, requiring additional tokens. This motivates the need for a multimodal method to compress many shots into fewer tokens without finetuning. In this work, we enable LMMs to perform multimodal, many-shot in-context learning by leveraging Multimodal Task Vectors (MTV)---compact implicit representations of in-context examples compressed in the model's attention heads. Specifically, we first demonstrate the existence of such MTV in LMMs and then leverage these extracted MTV to enable many-shot in-context learning for various vision-and-language tasks. Our experiments suggest that MTV can scale in performance with the number of compressed shots and generalize to similar out-of-domain tasks without additional context length for inference. Code: https://github.com/Brandon3964/MultiModal-Task-Vector Brandon Huang, Chancharik Mitra, Leonid Karlinsky, Assaf Arbelle, Trevor Darrell, Roei Herzig |
NeurIPS | 2 |
| 2024 | Elevating Learning Experiences: Leveraging Large Language Models as Student-Facing Assistants in Discussion ForumsabstractRecent advancements in instruction-tuned large language models offer new potential for enhancing students' experiences in large-scale classes. Deploying LLMs as student-facing assistants, however, presents challenges. Key issues include integrating class-specific content into responses and applying effective pedagogical techniques. This study addresses these challenges through retrieval and prompting techniques, focusing on mitigating hallucinations in LLM-generated responses, a crucial concern in education. Furthermore, practical deployment brings further challenges related to student data privacy and computational constraints. This research strives to enhance the quality and relevance of LLM responses while addressing practical deployment issues, with an emphasis on creating a versatile system for diverse domains and teaching styles. Chancharik Mitra, Mihran Miroyan, Vedant Kumud, Gireeja Ranade, Narges Norouzi |
SIGCSE (2) | 1 |