Parshiv Kapoor

dblp:429/6897 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0004-5804-0547ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 44% Video understanding and tracking · 22% Image recognition and object detection · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › visual reasoning
compositional visual reasoning
1.012026
Zero-Shot Vision Language Reasoning via Dual-layer Scene Graph Chain of Thoughts (Student Abstract) · AAAI 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
1.012026
Zero-Shot Vision Language Reasoning via Dual-layer Scene Graph Chain of Thoughts (Student Abstract) · AAAI 2026
Computer vision › Video understanding and tracking › video analytics
surgical video analysis
1.012026
TWiST: Temporal Weakly-Supervised Triplets Recognition in Surgical Videos (Student Abstract) · AAAI 2026
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization
1.012026
TWiST: Temporal Weakly-Supervised Triplets Recognition in Surgical Videos (Student Abstract) · AAAI 2026
Machine learning › Graph learning › graph representation
scene graph representation
0.312026
Zero-Shot Vision Language Reasoning via Dual-layer Scene Graph Chain of Thoughts (Student Abstract) · AAAI 2026
Computer vision › Segmentation and scene understanding
scene understanding
0.312026
Zero-Shot Vision Language Reasoning via Dual-layer Scene Graph Chain of Thoughts (Student Abstract) · AAAI 2026

Methods — techniques the papers use, named apart from their topics

weakly supervised learning · 1.0temporal attention · 1.0scene graph prompting · 1.0chain-of-thought prompting · 1.0
YearPublicationVenuePosition
2026 Zero-Shot Vision Language Reasoning via Dual-layer Scene Graph Chain of Thoughts (Student Abstract)
abstract
Large Multimodal Models (LMMs) often hallucinate objects and struggle with compositional reasoning in complex visual scenes. Structured Scene Graph (SG) representations explicitly encoding objects, attributes, and relations can mitigate these issues, however finetuning risks catastrophic forgetting. Recent zero-shot approaches prompt LMMs with scene graphs, yet typically rely on a single SG generated in one step, limiting capture of holistic context and question-specific details. We introduce a Dual-Layer Scene Graph Chain-of-Thought DLSG-CoT framework that enriches reasoning by combining two structured SGs: a Global Scene Graph (G-SG) that offers comprehensive image context, and a Query-Specific Scene Graph (Q-SG) produced through a two-step process targeting information relevant to the input query. Extensive experiments demonstrate that DLSG-CoT substantially improves LMM performance on compositional and context-sensitive tasks.
Yash Bansal, Parshiv Kapoor, Agam Pandey
AAAI2
2026 TWiST: Temporal Weakly-Supervised Triplets Recognition in Surgical Videos (Student Abstract)
abstract
Deep learning is increasingly applied to intraoperative and surgical video analysis to enable real-time workflow recognition, and decision support for improved surgical precision. A key direction is modeling surgical activity as triplets of instrument, action, and target, which provide a richer representation of procedures. However, existing approaches often depend on bounding-box annotations or lack temporal context. We propose TWiST (Temporal Weakly Supervised Triplet detection), a framework that combines weakly supervised instrument localization, temporal attention for triplet prediction, and grounding of triplets with detected instruments. Our experiments show that TWiST outperforms prior weakly supervised baselines.
Pranshu Danani, Yash Bansal, Parshiv Kapoor
AAAI3