Anupam Sobti

dblp:218/9523 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-7698-2432ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Time2Agri: Temporal Pretext Tasks for Agricultural Monitoring
abstract
Self Supervised Learning (SSL) has emerged as a prominent paradigm for label-efficient learning, and has been widely utilized by remote sensing foundation models (RSFMs). Recent RSFMs including SatMAE and DoFA primarily rely on masked autoencoding (MAE), contrastive learning or some combination of them. However, these pretext tasks often overlook the unique temporal characteristics of agricultural landscape, namely nature's cycle of sowing, growth, and harvest. Motivated by this gap, we propose three novel agriculture-specific pretext tasks, namely Time-Difference Prediction (TD), Temporal Frequency Prediction (FP), and Future-Frame Prediction (FF). Comprehensive evaluation on SICKLE dataset shows FF achieves 69.6% IoU on crop mapping and FP reduces yield prediction error to 30.7% MAPE, outperforming all baselines, and TD remains competitive on most tasks. Further, we also scale FF to the national scale of India, achieving 54.2% IoU outperforming all baselines on field boundary delineation on FTW India dataset.
Moti Rattan Gupta, Anupam Sobti
AAAI2
2025 Geospatial Active Learning for Efficient Data Annotation: A case study on cool roof detection
Yogendra Kumar, Anupam Sobti
COMPASS2
2024 Evaluation of computer vision pipeline for farm-level analytics: A case study in Sugarcane
abstract
Analyzing agricultural imagery for farm level insights has been an active area of research in the recent times. For providing the necessary information to stakeholders - be it farmers, financial institutions or governments, various computer vision tasks have to come together. For example, to provide information to a farmer about crop stress in their farm, accurate localization of the farm, identification of the crop type and a monitoring of the field’s micro-climate must be done together. In this work, we set performance benchmarks for three computer vision tasks - farm boundary detection, crop classification and sub-field stress estimation with different modalities of images - Sentinel2, PlanetScope and Drone Imagery. We use public dataset benchmarks for farm boundaries and crop classification and do a controlled field study on a large sugarcane farm in Uttar Pradesh, India for the stress estimation.
Sambal Shikhar, Rajiv Ranjan 0004, Aman Sa, Anshika Srivastava, Yash Srivastava, Shashank Tamaskar, Anupam Sobti
COMPASS8
2022 Reliable Energy Consumption Modeling for an Electric Vehicle Fleet
abstract
Accurately predicting the energy consumption of an electric vehicle (EV) under real-world circumstances (such as varying road, traffic, weather conditions, etc.) is critical for a number of decisions like range estimation and route planning. A major concern for electric vehicle owners is the uncertain nature of the battery consumption. This results in the “range anxiety” and reluctance from users for mass adoption of EVs, since they are concerned about untimely drainage of battery. Even at the organizational level, a company running a fleet of electric vehicles must understand the battery consumption profiles accurately for tasks such as route and driver planning, battery sizing, maintenance planning, etc.
Millend Roy, Akshay Uttama Nambi, Anupam Sobti, Tanuja Ganu, Shivkumar Kalyanaraman, Shankar Akella, Jaya Subha Devi, S. A. Sundaresan
COMPASS3
2021 VmAP: A Fair Metric for Video Object Detection
abstract
Video object detection is the task of detecting objects in a sequence of frames, typically, with a significant overlap in content among consecutive frames. Mean Average Precision (mAP) was originally proposed for evaluating object detection techniques in independent frames, but has been used for evaluating video based object detectors as well. This is undesirable since the average precision over all frames masks the biases that a certain object detector might have against certain types of objects depending on the number of frames for which the object is present in a video sequence. In this paper we show several disadvantages of mAP as a metric for evaluating video based object detection. Specifically, we show that: (a) some object detectors could be severely biased against some specific kind of objects, such as small, blurred, or low contrast objects, and such differences may not reflect in mAP based evaluation, (b) operating a video based object detector at the best frame based precision/recall value (high F1 score) may lead to many false positives without a significant increase in the number of objects detected. (c) mAP does not take into account that tracking can be potentially used to recover missed detections in the temporal neighborhood while this can be account for while evaluating detectors. As an alternate, we suggest a novel evaluation metric (VmAP) which takes the focus away from evaluating detections on every frame. Unlike mAP, VmAP rewards a high recall of different object views throughout the video. We form sets of bounding boxes having similar views of an object in a temporal neighborhood and use a set-level recall for evaluation. We show that VmAP is able to address all the challenges with the mAP listed above. Our experiments demonstrate hidden biases in object detectors, shows upto 99% reduction in false positives while maintaining similar object recall and shows a 9% improvement in correlation with post-tracking performance.
Anupam Sobti, Vaibhav Mavi, M. Balakrishnan, Chetan Arora 0001
ACM Multimedia1
2019 Multi-sensor Energy Efficient Obstacle Detection
abstract
With the improvement in technology, both the cost and the power requirement of cameras, as well as other sensors have come down significantly. It has allowed these sensors to be integrated into portable as well as wearable systems. Such systems are usually operated in a hands-free and always-on manner where they need to function continuously in a variety of scenarios. In such situations, relying on a single sensor or a fixed sensor combination can be detrimental to both performance as well as energy requirements. Consider the case of an obstacle detection task. Here using an RGB camera helps in recognizing the obstacle type but takes much more energy than an ultrasonic sensor. Infrared cameras can perform better than RGB camera at night but consume twice the energy. Therefore, an efficient system must use a combination of sensors, with an adaptive control that ensures the use of the sensors appropriate to the context. In this adaptation, one needs to consider both performance and energy and their trade-off. In this paper, we explore the strengths of different sensors as well their trade-off for developing a deep neural network based wearable device. We choose a specific case study in the context of a mobility assistance device for the visually impaired. The device detects obstacles in the path of a visually impaired person and is required to operate both at day and night with minimal energy to increase the usage time on a single charge. The device employs multiple sensors: ultrasonic sensor, RGB Camera, and NIR Camera along with a deep neural network accelerator for speeding up computation. We show that by adaptively choosing the appropriate sensor for the context, we can achieve up to 90% reduction in energy while maintaining comparable performance to a single sensor system.
Anupam Sobti, M. Balakrishnan, Chetan Arora 0001
DSD1
2018 Object Detection in Real-Time Systems: Going Beyond Precision
abstract
Applications like autonomous driving, industrial robotics, surveillance, and wearable assistive technology rely on object detectors as an integral part of the system. Thus, an increase in performance of object detectors directly affects the quality of such systems. In the recent years, convolutional neural networks (CNNs) and its variants emerged as the state of art in object detection, where performance is usually measured either in terms of mean average precision (mAP) or number of frames processed per second (fps). Many applications which use object detectors are resource constrained in practice. Even though it is clear from the published results, that a frame-level analysis of the system in terms of mAP or fps proves the superiority of one algorithm over the other, we observe that such metrics do not necessarily apply to real time applications with resource constraints. A slower algorithm even though highly accurate may need to drop frames to maintain the necessary frame rate and lose on the accuracy. We propose a closer look at the metrics used for performance in real-time applications, and suggest some new evaluation criterion. Our comparison of state of the art detectors on these metrics has also thrown some surprises in terms of conventional wisdom, which we present in this paper. Our framework is available at https://www.github.com/anupamsobti/object-detectionreal-time-systems.
Anupam Sobti, Chetan Arora 0001, M. Balakrishnan
WACV1