VLDB 2026 Research / reviewers in the wild / expert
Motilal Agrawal
dblp:32/3704
· DBLP profile ↗
16ranked-venue papers
8as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Representation and self-supervised learning · 57% Robot navigation and mapping · 21% 3D vision · 17% | |
| Network and information security
1 paper |
Hardware security and side channels · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
adaptive masking |
0.7 | 1 | 2023 | AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.7 | 1 | 2023 | AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning
spatio-temporal representation learning |
0.7 | 1 | 2023 | AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders · CVPR 2023 |
Robotics › Robot navigation and mapping
SLAM |
0.2 | 3 | 2009 | Leaving Flatland: Toward real-time 3D navigation · ICRA 2009 FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 Frame-Frame Matching for Realtime Consistent Visual Mapping · ICRA 2007 |
Hardware security and side channels › hardware reverse engineering
integrated circuit reverse engineering |
0.2 | 1 | 2014 | Non-Invasive Recognition of Poorly Resolved Integrated Circuit Elements · IEEE Trans. Inf. Forensics Secur. 2014 |
Robotics › Robot navigation and mapping › SLAM
loop closure |
0.2 | 2 | 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 Frame-Frame Matching for Realtime Consistent Visual Mapping · ICRA 2007 |
Robotics › Robot navigation and mapping › robot mapping
visual mapping |
0.2 | 2 | 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 Frame-Frame Matching for Realtime Consistent Visual Mapping · ICRA 2007 |
Robotics › Robot navigation and mapping › mobile robot navigation
3d navigation |
0.1 | 1 | 2009 | Leaving Flatland: Toward real-time 3D navigation · ICRA 2009 |
Robotics › Motion planning and robot control › path planning
3d path planning |
0.1 | 1 | 2009 | Leaving Flatland: Toward real-time 3D navigation · ICRA 2009 |
Robotics › Motion planning and robot control › motion planning › mobile robot motion planning
terrain-aware planning |
0.1 | 1 | 2009 | Leaving Flatland: Toward real-time 3D navigation · ICRA 2009 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.1 | 2 | 2004 | Window-Based, Discontinuity Preserving Stereo · CVPR (1) 2004 Trinocular Stereo Using Shortest Paths and the Ordering Constraint · Int. J. Comput. Vis. 2002 |
Computer vision › 3D vision › structure from motion
bundle adjustment |
0.1 | 1 | 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 |
Computer vision › 3D vision
feature detection and matching |
0.1 | 1 | 2008 | CenSurE: Center Surround Extremas for Realtime Feature Detection and Matching · ECCV (4) 2008 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.1 | 1 | 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 |
Integrated circuit design › digital circuit design
CMOS circuit design |
0.1 | 1 | 2014 | Non-Invasive Recognition of Poorly Resolved Integrated Circuit Elements · IEEE Trans. Inf. Forensics Secur. 2014 |
Computer vision › 3D vision › stereo vision › stereo matching
discontinuity-preserving stereo |
0.0 | 1 | 2004 | Window-Based, Discontinuity Preserving Stereo · CVPR (1) 2004 |
Computer vision › 3D vision
camera calibration |
0.0 | 1 | 2003 | Camera calibration using spheres: A semi-definite programming approach · ICCV 2003 |
Computer vision › 3D vision › camera calibration
multi-camera calibration |
0.0 | 1 | 2003 | Camera calibration using spheres: A semi-definite programming approach · ICCV 2003 |
Computer vision › 3D vision › stereo vision
trinocular stereo |
0.0 | 1 | 2002 | Trinocular Stereo Using Shortest Paths and the Ordering Constraint · Int. J. Comput. Vis. 2002 |
Computer vision › 3D vision
3d reconstruction |
0.0 | 1 | 2001 | A Probabilistic Framework for Surface Reconstruction from Multiple Images · CVPR (2) 2001 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.0 | 1 | 2001 | A Probabilistic Framework for Surface Reconstruction from Multiple Images · CVPR (2) 2001 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.0 | 1 | 2001 | A Probabilistic Framework for Surface Reconstruction from Multiple Images · CVPR (2) 2001 |
Computer vision › 3D vision
structure from motion |
0.0 | 1 | 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual Mapping · IEEE Trans. Robotics 2008 |
Computer vision › 3D vision
depth estimation |
0.0 | 1 | 2004 | Window-Based, Discontinuity Preserving Stereo · CVPR (1) 2004 |
Graph algorithms and graph theory
shortest path |
0.0 | 1 | 2002 | Trinocular Stereo Using Shortest Paths and the Ordering Constraint · Int. J. Comput. Vis. 2002 |
Computer vision › 3D vision › geometric estimation
visibility estimation |
0.0 | 1 | 2001 | A Probabilistic Framework for Surface Reconstruction from Multiple Images · CVPR (2) 2001 |
Methods — techniques the papers use, named apart from their topics
policy gradient · 0.7auxiliary sampling network · 0.7machine learning · 0.4light-induced voltage alteration · 0.4confocal infrared laser scanning optical microscopy · 0.4binary representation · 0.4visual odometry · 0.1stereo mapping · 0.1nonlinear least squares · 0.1feature detection · 0.1bundle adjustment · 0.1nonlinear measurement · 0.1monocular vision · 0.1binocular vision · 0.1shortest path · 0.0ordering constraint · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked AutoencodersabstractMasked Autoencoders (MAEs) learn generalizable representations for image, text, audio, video, etc., by reconstructing masked input data from tokens of the visible data. Current MAE approaches for videos rely on random patch, tube, or frame based masking strategies to select these tokens. This paper proposes AdaMAE, an adaptive masking strategy for MAEs that is end-to-end trainable. Our adaptive masking strategy samples visible tokens based on the semantic context using an auxiliary sampling network. This network estimates a categorical distribution over spacetime-patch tokens. The tokens that increase the expected reconstruction error are rewarded and selected as visible tokens, motivated by the policy gradient algorithm in reinforcement learning. We show that AdaMAE samples more tokens from the high spatiotemporal information regions, thereby allowing us to mask 95% of tokens, resulting in lower memory requirements and faster pre-training. We conduct ablation studies on the Something-Something v2 (SSv2) dataset to demonstrate the efficacy of our adaptive sampling approach and report state-of-the-art results of 70.0% and 81.7% in top-1 accuracy on SSv2 and Kinetics-400 action classification datasets with a ViT-Base backbone and 800 pre-training epochs. Code and pre-trained models are available at: https://github.com/wgcban/adamae.git. Wele Gedara Chaminda Bandara, Naman Patel, Mehdi Nikkhah, Motilal Agrawal, Vishal M. Patel |
CVPR | 5 |
| 2014 | Non-Invasive Recognition of Poorly Resolved Integrated Circuit ElementsabstractWe present a non-invasive method for recognition of components in a digital CMOS integrated circuit (IC). We use a confocal infrared laser scanning optical microscope to collect multimodal images through the backside of the IC. Individual modes correspond to passive reflectivity measurements or active measurements, such as light-induced voltage alteration. The modes are registered and stored in a multidimensional data cube. We apply a machine learning algorithm using a binary representation to identify a variety of data structures from transistors to entire logic cells. Because of the compact representation, objects can be detected rapidly. We show that by increasing the number of imaging modes used to develop the descriptor, we can significantly increase recognition accuracy. The approach allows recognition of poorly resolved components, whose primary distinguishing features are below traditional optical resolution limits, and is general enough to be applied to multiple design processes. We believe this represents a significant step toward a fully non-invasive IC reverse engineering system. Erik Matlin, Motilal Agrawal, David Stoker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Accurate content-based video copy detection with efficient feature indexingabstractWe describe an accurate content-based copy detection system that uses both local and global visual features to ensure robustness. Our system advances state-of-the-art techniques in four key directions. (1) Multiple-codebook-based product quantization: conventional product quantization methods encode feature vectors using a single codebook, resulting in large quantization error. We propose a novel codebook generation method for an arbitrary number of codebooks. (2) Handling of temporal burstiness: for a stationary scene, once a query feature matches incorrectly, the match continues in successive frames, resulting in a high false-alarm rate. We present a temporal-burstiness-aware scoring method that reduces the impact from similar features, thereby reducing false alarms. (3) Densely sampled SIFT descriptors: conventional global features suffer from a lack of distinctiveness and invariance to non-photometric transformations. Our densely sampled global SIFT features are more discriminative and robust against logo or pattern insertions. (4) Bigram- and multiple-assignment-based indexing for global features: we extract two SIFT descriptors from each location, which makes them more distinctive. To improve recall, we propose multiple assignments on both the query and reference sides. Performance evaluation on the TRECVID 2009 dataset indicates that both local and global approaches outperform conventional schemes. Furthermore, the integration of these two approaches achieves a three-fold reduction in the error rate when compared with the best performance reported in the TRECVID 2009 workshop. Yusuke Uchida, Motilal Agrawal, Shigeyuki Sakazawa |
ICMR | 2 |
| 2009 | Leaving Flatland: Toward real-time 3D navigationabstractWe report our first experiences with Leaving Flatland, an exploratory project that studies the key challenges of closing the loop between autonomous perception and action on challenging terrain. We propose a comprehensive system for localization, mapping, and planning for the RHex mobile robot in fully 3D indoor and outdoor environments. This system integrates Visual Odometry-based localization with new techniques in real-time 3D mapping from stereo data. The motion planner uses a new decomposition approach to adapt existing 2D planning techniques to operate in 3D terrain. We test the map-building and motion-planning subsystems on real and synthetic data, and show that they have favorable computational performance for use in high-speed autonomous navigation. Benoit Morisset, Radu Bogdan Rusu, Aravind Sundaresan, Kris Hauser, Motilal Agrawal, Jean-Claude Latombe, Michael Beetz |
ICRA | 5 |
| 2008 | CenSurE: Center Surround Extremas for Realtime Feature Detection and Matching
Motilal Agrawal, Kurt Konolige, Morten Rufus Blas |
ECCV (4) | 1 |
| 2008 | Fast color/texture segmentation for outdoor robotsabstractWe present a fast integrated approach for online segmentation of images for outdoor robots. A compact color and texture descriptor has been developed to describe local color and texture variations in an image. This descriptor is then used in a two-stage fast clustering framework using K-means to perform online segmentation of natural images. We present results of applying our descriptor for segmenting a synthetic image and compare it against other state-of-the-art descriptors. We also apply our segmentation algorithm to the task of detecting natural paths in outdoor images. The whole system has been demonstrated to work online alongside localization, 3D obstacle detection, and planning. Morten Rufus Blas, Motilal Agrawal, Aravind Sundaresan, Kurt Konolige |
IROS | 2 |
| 2008 | FrameSLAM: From Bundle Adjustment to Real-Time Visual MappingabstractMany successful indoor mapping techniques employ frame-to-frame matching of laser scans to produce detailed local maps as well as the closing of large loops. In this paper, we propose a framework for applying the same techniques to visual imagery. We match visual frames with large numbers of point features, using classic bundle adjustment techniques from computational vision, but we keep only relative frame pose information (askeleton). The skeleton is a reduced nonlinear system that is a faithful approximation of the larger system and can be used to solve large loop closures quickly, as well as forming a backbone for data association and local registration. We illustrate the workings of the system with large outdoor datasets (10 km), showing large-scale loop closure and precise localization in real time. Kurt Konolige, Motilal Agrawal |
IEEE Trans. Robotics | 2 |
| 2007 | Frame-Frame Matching for Realtime Consistent Visual MappingabstractMany successful indoor mapping techniques employ frame-to-frame matching of laser scans to produce detailed local maps, as well as closing large loops. In this paper, we propose a framework for applying the same techniques to visual imagery, matching visual frames with large numbers of point features. The relationship between frames is kept as a nonlinear measurement, and can be used to solve large loop closures quickly. Both monocular (bearing-only) and binocular vision can be used to generate matches. Other advantages of our system are that no special landmark initialization is required, and large loops can be solved very quickly. Kurt Konolige, Motilal Agrawal |
ICRA | 2 |
| 2007 | Large-Scale Visual Odometry for Rough Terrain
Kurt Konolige, Motilal Agrawal, Joan Solà |
ISRR | 2 |
| 2007 | Localization and Mapping for Autonomous Navigation in Outdoor Terrains : A Stereo Vision ApproachabstractWe consider the problem of autonomous navigation in unstructured outdoor terrains using vision sensors. The goal is for a robot to come into a new environment, map it and move to a given goal at modest speeds (1 m/sec). The biggest challenges are in building good maps and keeping the robot well localized as it advances towards the goal. In this paper, we concentrate on showing how it is possible to build a consistent, globally correct map in real time, using efficient precise stereo algorithms for map making and visual odometry for localization. While we have made advances in both localization and mapping using stereo vision, it is the integration of the techniques that is the biggest contribution of the research. The validity of our approach is tested in blind experiments, where we submit our code to an independent testing group that runs and validates it on an outdoor robot Motilal Agrawal, Kurt Konolige, Robert C. Bolles |
WACV | 1 |
| 2006 | A Lie Algebraic Approach for Consistent Pose Registration for General Euclidean MotionabstractWe study the problem of registering local relative pose estimates to produce a global consistent trajectory of a moving robot. Traditionally, this problem has been studied with a flat world assumption wherein the robot motion has only three degrees of freedom. In this paper, we generalize this for the full six-degrees-of-freedom Euclidean motion. Given relative pose estimates and their covariances, our formulation uses the underlying Lie algebra of the Euclidean motion to compute the absolute poses. Ours is an iterative algorithm that minimizes the sum of Mahalanobis distances by linearizing around the current estimate at each iteration. Our algorithm is fast, does not depend on a good initialization, and can be applied to large sequences in complex outdoor terrains. It can also be applied to fuse uncertain pose information from different available sources including GPS, LADAR, wheel encoders and vision sensing to obtain more accurate odometry. Experimental results using both simulated and real data support our claim Motilal Agrawal |
IROS | 1 |
| 2004 | Window-Based, Discontinuity Preserving Stereo
Motilal Agrawal, Larry Davis 0001 |
CVPR (1) | 1 |
| 2004 | On automatic determination of varying focal lengths using semidefinite programmingabstractWe describe a novel approach to the determination of the focal length of a moving camera with rectangular pixels. The principal point of the camera assumed to be known and fixed, whereas the focal length is allowed to vary across the sequence. Given three or more such images and a projective reconstruction, we describe a novel auto-calibration technique to obtain a metric reconstruction. Our technique uses semidefinite programming to recover these focal lengths (and hence the metric reconstruction). Our approach is efficient, well behaved with guaranteed convergence, and can be applied to long sequences of video. We present results for our approach using both simulated and real video sequences. Motilal Agrawal |
ICIP | 1 |
| 2003 | Camera calibration using spheres: A semi-definite programming approachabstractVision algorithms utilizing camera networks with a common field of view are becoming increasingly feasible and important. Calibration of such camera networks is a challenging and cumbersome task. The current approaches for calibration using planes or a known 3D target may not be feasible as these objects may not be simultaneously visible in all the cameras. In this paper, we present a new algorithm to calibrate cameras using occluding contours of spheres. In general, an occluding contour of a sphere projects to an ellipse in the image. Our algorithm uses the projection of the occluding contours of three spheres and solves for the intrinsic parameters and the locations of the spheres. The problem is formulated in the dual space and the parameters are solved for optimally and efficiently using semidefinite programming. The technique is flexible, accurate and easy to use. In addition, since the contour of a sphere is simultaneously visible in all the cameras, our approach can greatly simplify calibration of multiple cameras with a common field of view. Experimental results from computer simulated data and real world data, both for a single camera and multiple cameras, are presented. Motilal Agrawal, Larry Davis 0001 |
ICCV | 1 |
| 2002 | Trinocular Stereo Using Shortest Paths and the Ordering Constraint
Motilal Agrawal |
Int. J. Comput. Vis. | 1 |
| 2001 | A Probabilistic Framework for Surface Reconstruction from Multiple ImagesabstractThe paper presents a novel probabilistic framework for 3D surface reconstruction from multiple stereo images. The method works on a discrete voxelized representation of the scene. An iterative scheme is used to estimate the probability that a scene point lies on the true 3D surface. The novelty of our approach lies in the ability to model and recover surfaces which may be occluded in some views. This is done by explicitly estimating the probabilities that a 3D scene point is visible in a particular view from the set of given images. This relies on the fact that for a point on a lambertian surface, if the pixel intensities of its projection along two views differ, then the point is necessarily occluded in one of the views. We present results of surface reconstruction from both real and synthetic image sets. Motilal Agrawal, Larry Davis 0001 |
CVPR (2) | 1 |