Anurag Ghosh

dblp:02/7988 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-1617-0851ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Image recognition and object detection · 40% Autonomous driving · 35% Efficient and distributed learning · 17%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 50% Rendering · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 50% Cloud and datacenter computing · 50%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Human-computer interaction and pervasive computing
3 papers
Ubiquitous computing and smart environments · 33% Wearable and physiological sensing · 33% Collaborative and social computing · 33%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object recognition
0.912025
ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones · ICCV 2025
Robotics › Autonomous driving
perception and planning
0.912025
ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones · ICCV 2025
Geometric modeling and processing
3d reconstruction
0.912025
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis · CVPR 2025
Rendering
novel view synthesis
0.912025
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis · CVPR 2025
Computer vision › Image recognition and object detection › object detection
efficient object detection
0.712023
Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection · CVPR 2023
Cloud and datacenter computing › resource allocation › dynamic resource allocation
adaptive resource allocation
0.712023
Chanakya: Learning Runtime Decisions for Adaptive Real-Time Perception · NeurIPS 2023
Embedded and real-time systems › autonomous systems
real-time perception
0.712023
Chanakya: Learning Runtime Decisions for Adaptive Real-Time Perception · NeurIPS 2023
Empirical software engineering
developer studies
0.412019
Signals Matter: Understanding Popularity and Impact of Users on Stack Overflow · WWW 2019
Computer vision › 3D vision
camera pose estimation
0.312025
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis · CVPR 2025
Robotics › Autonomous driving
driving scene understanding
0.312025
ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones · ICCV 2025
Robotics › Autonomous driving
perception
0.212023
Chanakya: Learning Runtime Decisions for Adaptive Real-Time Perception · NeurIPS 2023
Wearable and physiological sensing
eye tracking
0.112019
ALT: towards automating driver license testing using smartphones · SenSys 2019
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing
0.112019
Smartphone-based driver license testing: demo abstract · SenSys 2019
Empirical software engineering
mining software repositories
0.112019
Signals Matter: Understanding Popularity and Impact of Users on Stack Overflow · WWW 2019
Empirical software engineering
user behavior analysis
0.112019
Signals Matter: Understanding Popularity and Impact of Users on Stack Overflow · WWW 2019

Methods — techniques the papers use, named apart from their topics

pseudo-synthetic rendering · 1.7fine-tuning · 1.7reinforcement learning · 1.3learned execution policy · 1.3statistical modeling · 0.8smartphone-based sensing · 0.8smartphone camera · 0.8inertial sensors · 0.8gaze estimation · 0.8two-plane perspective prior · 0.7adaptive sampling · 0.7
YearPublicationVenuePosition
2025 AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
abstract
We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-registered aerial-ground datasets for training is a key reason for this failure. Such data is difficult to assemble precisely because it is difficult to reconstruct in a scalable way. To overcome this challenge, we propose a scalable framework combining pseudo-synthetic renderings from 3D city-wide meshes (e.g., Google Earth) with real, ground-level crowd-sourced images (e.g., MegaDepth [29]). The pseudo-synthetic data simulates a wide range of aerial viewpoints, while the real, crowd-sourced images help improve visual fidelity for ground-level images where mesh-based renderings lack sufficient detail, effectively bridging the domain gap between real images and pseudo-synthetic renderings. Using this hybrid dataset, we fine-tune several state-of-the-art algorithms and achieve significant improvements on real-world, zero-shot aerial-ground tasks. For example, we observe that baseline DUSt3R [64] localizes fewer than 5% of aerial-ground pairs within 5 degrees of camera rotation error, while fine-tuning with our data raises accuracy to nearly 56%, addressing a major failure point in handling large viewpoint changes. Beyond camera estimation and scene reconstruction, our dataset also improves performance on downstream tasks like novel-view synthesis in challenging aerial-ground scenarios, demonstrating the practical value of our approach in real-world applications.
Khiem Vuong, Anurag Ghosh, Deva Ramanan, Srinivasa G. Narasimhan, Shubham Tulsiani
CVPR2
2025 ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Anurag Ghosh, Robert Tamburo, Khiem Vuong, Juan R. Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan
ICCV1
2025 Instance-Warp: Saliency Guided Image Warping for Unsupervised Domain Adaptation
Anurag Ghosh, Srinivasa G. Narasimhan
WACV2
2024 Holistic Energy Awareness and Robustness for Intelligent Drones
abstract
Drones represent a significant technological shift at the convergence of on-demand cyber-physical systems and edge intelligence. However, realizing their full potential necessitates managing the limited energy resources carefully. Prior work looks at factors such as battery characteristics, intelligent edge sensing considerations, planning, and robustness in isolation. But a global view of energy awareness that considers these factors and looks at various tradeoffs is essential. To this end, we present results from our detailed empirical study of battery charge-discharge characteristics and the impact of altitude and lighting on edge inference accuracy. Our energy models, derived from these observations, predict energy usage while performing various manoeuvres with an error of 5.6%, a 2.5X improvement over the state-of-the-art. Furthermore, we propose a holistic energy-aware multi-drone scheduling system that decreases the energy consumed by 21.14% and the mission times by 46.91% over state-of-the-art baselines. To achieve system robustness in the event of link or drone failure, we observe trends in Packet Delivery Ratio to propose a methodology to establish reliable communication between nodes. We release an open-source implementation of our system. Finally, we tie all of these pieces together using a people-counting case study.
Ravi Raj Saxena, Joydeep Pal, Srinivasan Iyengar, Bhawana Chhaglani, Anurag Ghosh, Venkat N. Padmanabhan, Prabhakar Venkata Tamma
ACM Trans. Sens. Networks5
2023 Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection
abstract
Real-time efficient perception is critical for autonomous navigation and city scale sensing. Orthogonal to architectural improvements, streaming perception approaches have exploited adaptive sampling improving real-time detection performance. In this work, we propose a learnable geometry-guided prior that incorporates rough geometry of the 3D scene (a ground plane and a plane above) to resample images for efficient object detection. This significantly improves small and far-away object detection performance while also being more efficient both in terms of latency and memory. For autonomous navigation, using the same detector and scale, our approach improves detection rate by +4.1 APsor +39% and in real-time performance by +5.3 sAPs or +63% for small objects over state-of-the-art (SOTA). For fixed traffic cameras, our approach detects small objects at image scales other methods cannot. At the same scale, our approach improves detection of small objects by 195% (+12.5 APS) over naive-downsampling and 63% (+4.2 APS) over SOTA.
Anurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. Narasimhan
CVPR1
2023 Chanakya: Learning Runtime Decisions for Adaptive Real-Time Perception
abstract
Real-time perception requires planned resource utilization. Computational planning in real-time perception is governed by two considerations -- accuracy and latency. There exist run-time decisions (e.g. choice of input resolution) that induce tradeoffs affecting performance on a given hardware, arising from intrinsic (content, e.g. scene clutter) and extrinsic (system, e.g. resource contention) characteristics. Earlier runtime execution frameworks employed rule-based decision algorithms and operated with a fixed algorithm latency budget to balance these concerns, which is sub-optimal and inflexible. We propose Chanakya, a learned approximate execution framework that naturally derives from the streaming perception paradigm, to automatically learn decisions induced by these tradeoffs instead. Chanakya is trained via novel rewards balancing accuracy and latency implicitly, without approximating either objectives. Chanakya simultaneously considers intrinsic and extrinsic context, and predicts decisions in a flexible manner. Chanakya, designed with low overhead in mind, outperforms state-of-the-art static and dynamic execution policies on public datasets on both server GPUs and edge devices.
Anurag Ghosh, Vaibhav Balloli, Akshay Uttama Nambi, Tanuja Ganu
NeurIPS1
2019 Smartphone-based driver license testing: demo abstract
abstract
Road safety is compromised today by the inadequacies in driver license testing. Testing is typically still performed manually, and efforts aimed at automating testing are stymied by the cost of outfitting a testing track with sensors. We demonstrate a low-cost, smartphone-based system for automating key aspects of the driver license test. We have a pilot deployment of our system at an official testing track in India. We will present an analysis of license test results obtained from this pilot, comparing the smartphone-based testing results with manual evaluation.
Anurag Ghosh, Vijay Lingam, Ishit Mehta, Akshay Uttama Nambi, Venkat N. Padmanabhan, Satish Sangameswaran
SenSys1
2019 ALT: towards automating driver license testing using smartphones
abstract
Can a smartphone administer a driver license test? We ask this question because of the inadequacy of manual testing and the expense of outfitting an automated testing track with sensors such as cameras, leading to less-than-thorough testing and ultimately compromising road safety. We present ALT, a low-cost smartphone-based system for automating key aspects of the driver license test. A windshield-mounted smartphone serves as the sole sensing platform, with the front camera being used to monitor driver's gaze, and the rear camera, together with inertial sensors, being used to evaluate driving maneuvers such as parallel parking. The sensors are also used in tandem, for instance, to check that the driver scanned their mirror during a lane change.
Akshay Uttama Nambi, Ishit Mehta, Anurag Ghosh, Vijay Lingam, Venkat N. Padmanabhan
SenSys3
2019 Signals Matter: Understanding Popularity and Impact of Users on Stack Overflow
abstract
Stack Overflow, a Q&A site on programming, awards reputation points and badges (game elements) to users on performing various actions. Situating our work in Digital Signaling Theory, we investigate the role of these game elements in characterizing social qualities (specifically, popularity and impact) of its users. We operationalize these attributes using common metrics and apply statistical modeling to empirically quantify and validate the strength of these signals. Our results are based on a rich dataset of 3,831,147 users and their activities spanning nearly a decade since the site's inception in 2008. We present evidence that certain non-trivial badges, reputation scores and age of the user on the site positively correlate with popularity and impact. Further, we find that the presence of costly to earn and hard to observe signals qualitatively differentiates highly impactful users from highly popular users.
Arpit Merchant, Daksh Shah, Gurpreet Singh Bhatia 0001, Anurag Ghosh, Ponnurangam Kumaraguru
WWW4
2018 Towards Structured Analysis of Broadcast Badminton Videos
abstract
Sports video data is recorded for nearly every major tournament but remains archived and inaccessible to large scale data mining and analytics. It can only be viewed sequentially or manually tagged with higher-level labels which is time consuming and prone to errors. In this work, we propose an end-to-end framework for automatic attributes tagging and analysis of sport videos. We use commonly available broadcast videos of matches and, unlike previous approaches, does not rely on special camera setups or additional sensors. Our focus is on Badminton as the sport of interest. We propose a method to analyze a large corpus of badminton broadcast videos by segmenting the points played, tracking and recognizing the players in each point and annotating their respective badminton strokes. We evaluate the performance on 10 Olympic matches with 20 players and achieved 95.44% point segmentation accuracy, 97.38% player detection score ([email protected]), 97.98% player identification accuracy, and stroke segmentation edit scores of 80.48%. We further show that the automatically annotated videos alone could enable the gameplay analysis and inference by computing understandable metrics such as player's reaction time, speed, and footwork around the court, etc.
Anurag Ghosh, Suriya Singh, C. V. Jawahar
WACV1