VLDB 2026 Research / reviewers in the wild / expert
Nimisha Roy
dblp:356/8683
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-2480-9974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | To Tell or to Ask? Comparing the Effects of Targeted vs. Socratic AI HintsabstractAs enrollment in CS1 courses continues to increase, extensive research has focused on autonomous support to offer personalized assistance for struggling students at scale. However, it is crucial that these intervention techniques do not inadvertently hinder the development of higher-order, computational thinking skills for novice programmers. This poster extends upon the research on LLM-based support by assessing the short- and long-term student outcomes from two carefully prompt-engineered, LLM-generated hint styles: Targeted and Socratic hints. A randomized controlled trial with 178 students was conducted over two semesters in a CS1 course at a large university, allowing students to interact with a hint generation AI agent while attempting course coding assignments. In the short-term, students receiving Socratic hints spent more time, took more attempts, and used more keystrokes to solve coding questions, while committing more repeat errors. Furthermore, this short-term loss in debugging efficiency is not counteracted by any evidence of an improvement in long-term student outcomes. Further research is being conducted to quantify the tradeoff between short-term performance and long-term, higher-order coding skill improvement in the development of educational AI agents. Zhixian Christopher Liding, Michael Osmolovskiy, Harshith Lanka, Ronnie Howard, Nimisha Roy, Rodrigo Borela |
SIGCSE (2) | 5 |
| 2026 | AI-Augmented Instruction: Real-Time Misconception DetectionabstractEnrollments in introductory computer science (CS1) courses continue to rise, making it difficult for instructors to deliver rapid, individualized feedback that addresses students' misconceptions at scale. We present an analysis framework and instructor tool that leverage large language models (LLMs) to classify, cluster, and present students' coding errors in real time. Our approach comprises two main contributions: (1) a prompt-engineered workflow for automatic error detection and a clustering pipeline using universal sentence encoders, KMeans, and t-SNE to group errors into thematic clusters; and (2) a dashboard that enables instructors to review class-wide, LLM-identified errors and dynamically tailor instruction toward current student misunderstandings. Our automated thematic clustering system is able to surface conceptual and strategic pitfalls that often persist beneath superficial debugging. A pilot study is being conducted to evaluate the effectiveness of the dashboard tool in large-scale CS1 instructional settings to enhance active learning at scale. Zhixian Christopher Liding, Michael Osmolovskiy, Harshith Lanka, Nimisha Roy, Rodrigo Borela |
SIGCSE (2) | 4 |
| 2026 | Benchmarking AI Tools for Software Engineering Education: Insights into Design, Implementation, and TestingabstractAs generative AI (Gen AI) tools reshape software engineering (SE) workflows, educators are exploring how to meaningfully integrate them into computing education. This experience report presents a structured benchmarking of widely used AI tools -- such as GitHub Copilot, GPT-4, Codeium, Claude 3.5, Gemini 1.5, Supermaven, TabNine, Testim, Postman, Eraser.io, and Lucidchart AI -- across key SE phases: design, implementation, debugging, and testing. Tools were selected based on industry relevance, accessibility for students, and alignment with common SE tasks. Through controlled experiments conducted by five AI-experienced evaluators with matched exposure levels, we assessed tool performance using standardized prompts, counterbalanced task roles, and a range of proxy metrics -- including prompt iterations, task completion time, human correction burden, hallucination frequency, output accuracy, and cross-file consistency -- to capture both cognitive load and tool limitations. While AI tools accelerated tasks such as boilerplate generation and UML sketching, they exhibited challenges in test coverage quality, cross-file coherence, and reliability under complex prompts. We discuss educational implications, including managing cognitive load, aligning tools with task types, and explicitly teaching prompt refinement and verification strategies. The paper offers actionable guidance for instructors, curriculum-ready artifacts, and a roadmap for scaling AI integration in SE classrooms, while also noting key limitations to support replication and contextual adoption. Nimisha Roy, Oleksandr Horielko, Olufisayo Omojokun |
SIGCSE (1) | 1 |
| 2025 | What Computing Faculty Want: Designing AI Tools for High-Enrollment Courses Beyond CS1abstractDespite the rapid adoption of GenAI assistants in computing education, we still lack insight into the AI designs that instructors consider essential for effective teaching in large computing courses beyond CS1.Prior work has analyzed instructors attitudes towards Rodrigo Borela, Meryem Yilmaz Soylu, Nimisha Roy |
ICER (2) | 4 |
| 2025 | Benchmarking of Generative AI Tools in Software Engineering Education: Formative Insights for Curriculum IntegrationabstractGenerative Artificial Intelligence (Gen-AI) has revolutionized software engineering (SE) by automating tasks across design, coding, and testing [1] [2].Tools like ChatGPT and GitHub Copilot streamline code generation, architectural modeling, debugging, and testcase creation [3] [4].Despite their rapid adoption in industry, the pedagogical implications of these tools in computing education have not been systematically examined.This study solves the existing gap by conducting a comprehensive benchmarking study of Gen-AI tools across four core SE phases-design documentation, feature implementation, debugging support, and testing -to address two research questions:RQ1: What strengths and limitations do Gen-AI tools exhibit in each phase?RQ2: How can insights from benchmarking inform effective integration of Gen-AI into SE curricula?To answer these questions, a diverse set of Gen-AI tools is evaluated, ranging from design-focused assistants such as Lucidchart, Mermaid.js and UIzard; implementation-oriented systems including GitHub Copilot, TabNine, Codeium and Supermaven; debugging supports like GPT-4 and Claude 3.5 Sonnet; and testing frameworks such as Testim, Mabl and Applitools-while also surveying emerging platforms (as of summer 2024) like Replit, Postman, Visily, Gemini, Eraser.io and others.For each tool and development phase, we applied phase-specific metrics: in design documentation, we assessed diagram accuracy, completeness, user effort, and IDE integration; in feature implementation, we measured pattern-based code generation quality, code-completion effectiveness, refactoring robustness, and UI/UX scaffolding; in debugging, we evaluated error-detection accuracy, hallucination rates, and clarity of explanatory feedback; and in testing, we examined test-case relevance and defect-detection coverage.Across all phases, we tracked prompt engineering complexity as a key mediating factor influencing tool performance.Our evaluation reveals speed-fidelity trade-offs: Code-completion assistants accelerate boilerplate generation but demand manual oversight to ensure cross-file consistency and manage higher-order abstractions; diagramming tools can produce precise UML models with minimal effort-but at the cost of iterative prompt refinement for complex cases; LLM debuggers deliver context-sensitive fixes Nimisha Roy, Oleksandr Horielko, Olufisayo Omojokun |
ICER (2) | 1 |
| 2025 | Beyond Buzzwords: Making Sustainability a Pillar of the Computing CurriculumabstractThe rapid digitalization of the global economy, driven by big data and artificial intelligence, has significantly increased energy consumption, reshaped labor markets, and impacted politics and communities to an extraordinary extent. Addressing these challenges involves educating future computer scientists about the carbon emissions associated with their code and the broader societal consequences of the technologies they design. Traditionally, computing education has focused on optimizing runtime and memory efficiency, frequently overlooking the links to energy efficiency and carbon footprint considerations. Additionally, the integration of ethics into the curriculum has not been comprehensive. This paper proposes a framework for integrating sustainability into the computing curriculum, prioritizing it as a critical consideration for students. It outlines the competencies required for sustainability education and identifies topics directly related to the UN SDGs as a natural entry point for sustainability concepts. Additionally, it reports on a pilot framework implementation at a major US public university, where over 3,200 students from over 30 disciplines were exposed to sustainable coding practices and the ethics of AI and machine learning. The curriculum incorporated transformative teaching and learning methodologies with lectures, supplemental materials, and interactive projects highlighting these themes. Challenges to implementation, which may be encountered by other institutions, are also discussed. Survey results demonstrate that sustainability can be seamlessly integrated into early coding education, encouraging ongoing effort to accentuate energy-efficient coding in computer science courses. Nikhila Alavala, Nimisha Roy, Melinda McDaniel, Max Mahdi Roozbahani, Rodrigo Borela, Parisa Babolhavaeji |
ITiCSE (1) | 2 |
| 2025 | Tracking the Progression of Errors Across Successive CS1 Code SubmissionsabstractUnderstanding the debugging process of novice programmers as they iteratively solve coding challenges is essential for developing intelligent tutoring systems that address gaps in comprehension and procedural coding skills. This poster presents a framework for systematically analyzing student coding attempts using large language models (LLMs) to identify syntactical, conceptual, and strategic errors. This study investigates 346 coding attempts for three live-coding challenges in a CS1 course, tracking the progression of errors over successive submissions. Preliminary results indicate that among students who attempted the challenges at least ten times, syntactical errors decrease more rapidly within the first ten attempts compared to conceptual or strategic errors. Although students effectively resolve syntax issues early in the debugging process, higher-level conceptual and strategic errors persist, suggesting the need for targeted instructional support at this stage. Zhixian Christopher Liding, Nimisha Roy, Rodrigo Borela |
ITiCSE (2) | 2 |
| 2025 | Empowering Future Software Engineers: Integrating AI Tools into Advanced CS CurriculumabstractArtificial Intelligence (AI) tools have transformed software development, making it crucial to equip computer science (CS) students with the skills to leverage these technologies. This talk presents an innovative curriculum approach, integrating AI tools into an advanced CS capstone course at a stage where students possess foundational skills in software engineering. This strategic timing ensures students can critically engage with AI, recognizing biases and managing challenges like hallucinations in AI-generated outputs. Nimisha Roy, Olufisayo Omojokun, Oleksandr Horielko |
SIGCSE (2) | 1 |
| 2025 | Scaling Academic Decision-Making with NLP: Automating Transfer Credit EvaluationsabstractManual processes for evaluating external course syllabi for transfer credit in higher education are time-consuming, inconsistent, and prone to bias. This project leverages Natural Language Processing (NLP) and large language models (LLMs) to automate the transfer credit evaluation process. The system processes external syllabi by embedding course content, conducting similarity searches, and providing structured reasoning for each match. Using techniques such as chain-of-thought reasoning and reflection agents, the system generates similarity scores and detailed explanations to support informed, data-driven decision-making by faculty. Validated against faculty decisions, the system promises to significantly improve the efficiency, consistency, and fairness of transfer credit evaluations. Future directions include expanding the system for advanced standing test evaluations and allowing faculty to query specific course components for more targeted analysis. Nimisha Roy, Olufisayo Omojokun, Huaijin Tu |
SIGCSE (2) | 1 |
| 2024 | Learning by Teaching: Insights on Student-Created Instructional Videos for Large CS ClassesabstractPromoting active learning is challenging in large computing courses, often with hundreds of students. We present insights from a pedagogical strategy we designed to foster college Computer Science (CS) students' learning in a large introductory course on software design and engineering (SWE). Students first created an instructional video explaining SWE concepts. Next, they peer-reviewed their classmates' explanations. Our findings suggest that using student-created instructional videos can support students' learning processes and help promote active learning in large computing courses. Pedro Guillermo Feijóo García, Nimisha Roy, Olufisayo Omojokun |
ITiCSE (2) | 2 |
| 2024 | Active Learning at Large-Scale: Using Video Tutorials to Learn by TeachingabstractIn an era where digital platforms like TikTok and Instagram redefine interaction, integrating active learning in large-scale computer science (CS) courses presents both a unique challenge and opportunity. This lightning talk intends to discuss a two-step strategy piloted within a second-year large-scale (i.e., over 500 students) introductory course to software engineering at a Southeast North American university. First, we tasked CS students to create a video tutorial as their midterm exam deliverable. The exam consisted of four questions: three open-ended questions for concepts on software engineering and one diagramming question that had them transform a descriptive context into a UML domain model diagram. Next, students took part in a double-masked peer review process that had them evaluate their peers' deliverables and explanations. Overall, we intended to assess students' understanding of software engineering concepts while enabling them to reflect upon their learning processes as they taught what they learned. Also, by having them peer-review among themselves, we aimed to foster an experience to enrich students' engagement and develop feedback skills. This lightning talk is to gather feedback from the CS Education community on this instructional strategy and possibly collaborate for a more extensive study on student-created artifacts in large-scale CS courses. Pedro Guillermo Feijóo García, Nimisha Roy |
SIGCSE (2) | 2 |
| 2024 | VG: Automatic Grading of D3 VisualizationsabstractManually grading D3 data visualizations is a challenging endeavor, and is especially difficult for large classes with hundreds of students. Grading an interactive visualization requires a combination of interactive, quantitative, and qualitative evaluation that are conventionally done manually and are difficult to scale up as the visualization complexity, data size, and number of students increase. We present VISGRADER, a first-of-its kind automatic grading method for D3 visualizations that scalably and precisely evaluates the data bindings, visual encodings, interactions, and design specifications used in a visualization. Our method enhances students' learning experience, enabling them to submit their code frequently and receive rapid feedback to better inform iteration and improvement to their code and visualization design. We have successfully deployed our method and auto-graded D3 submissions from more than 4000 students in a visualization course at Georgia Tech, and received positive feedback for expanding its adoption. Matthew Hull, Vivian Pednekar, Hannah Murray, Nimisha Roy, Emmanuel Tung, Susanta Routray, Connor Guerin, Zijie J. Wang, Seongmin Lee 0007, Max Mahdi Roozbahani, Polo Chau |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Creating Equitable Grading Practices with Rubrics: A Teaching Assistant Training ActivityabstractManually grading coding assignments in large computer science (CS) classes is a challenging logistical task. The evaluation of code correctness is subjective, leading to grading bias and inconsistencies. This problem is exacerbated when multiple teaching assistants (TAs) grade different submissions of the same problem. Rodrigo Borela, Nimisha Roy |
ICER (2) | 2 |