VLDB 2026 Research / reviewers in the wild / expert
Sarthak Pathak
dblp:191/2450
· DBLP profile ↗
11ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-5271-1782ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fisheye Stereo Camera Using Fisheye Vertical Stereo MethodabstractIn this paper, we propose a wide-range and high-accuracy fisheye stereo camera using the fisheye vertical stereo method. In stereo measurement with two fisheye cameras placed horizontally, increased mismatching occurs due to template matching along curved epipolar lines. Additionally, because the baseline is horizontal, there is a decrease in the distance measurement accuracy in the left and right areas of the image. Therefore, by placing fisheye cameras vertically and straightening the epipolar lines, we reduce mismatching during template matching in stereo measurement, achieving high-accuracy stereo measurement. Moreover, by making the baseline vertical, we improve the distance measurement accuracy in the left and right areas. Experiments demonstrate that the distance measurement accuracy of our proposed method is higher than that of conventional methods. Hikaru Chikugo, Kento Arai, Sarthak Pathak, Kazunori Umeda |
ICIP | 3 |
| 2024 | Visual Feedback Control of an Underactuated Hand for Grasping Brittle and Soft FoodsabstractThis paper presents a novel method to control an underactuated hand by using only a monocular camera, not using any internal sensors. In food factories, robots are required to handle a wide variety of foods without damaging them. To accomplish this, the use of underactuated hands is effective because they can adapt to various food shapes. However, if internal sensors such as tactile sensors and force sensors are used in the underactuated hands, it may cause a problem with hygiene and require complicated calibration. Moreover, if external sensors such as cameras are used, it is necessary to grasp foods without damaging them by using external information such as images. In our method, to tackle these problems, a camera is used as an external sensor. First, contact between the hand and the object is detected by using the contours of both, obtained from a camera image. Then, to avoid damaging the object, the following information is extracted from camera images and observed: the centroid of both the hand and object, the deformation of the object, and the occlusion rate of the hand. Furthermore, to prevent the object from dropping while the robotic arm is in motion, the distance between the centroid of the hand and the object is calculated. The experiments were conducted using twelve different food items. Ryogo Kai, Yuzuka Isobe, Sarthak Pathak, Kazunori Umeda |
ICRA | 3 |
| 2024 | Robust Gesture-based Appliance Control via Operator Identification and TrackingabstractIn this paper, we improve the robustness of a multi-camera gesture recognition system in multi-person situations by identifying and tracking the operator. This system is meant as an intelligent room to operate and interact with surrounding devices based on pointing gestures. In the method, we identify the operator by a hand-raising gesture, followed by tracking using the coordinates of the center of the operator’s head and extracting only the operator’s whole body. From the experimental results, we confirmed that highly accurate tracking could be performed in a multi-person situation of 2 to 5 persons, and that the success rate of extracting images of the operator’s whole body was more than 70%. We also clarified issues in the operator identification process and the extraction process. Masae Yokota, Sarthak Pathak, Kazunori Umeda |
RO-MAN | 2 |
| 2023 | Vision-Based In-Hand Manipulation of Variously Shaped Objects via Contact Point PredictionabstractIn-hand manipulation (IHM) is an important ability for robotic hands. This ability refers to changing the position and orientation of a grasped object without dropping it from the hand workspace. One major challenge of IHM is to achieve a large range of manipulation (especially rotation), regardless of the shape, size, and the orientation during manipulation of the grasped object. There are two main challenges - the manipulation range (due to the range of motion of the hand) and keeping the object grasped under all shapes and orientations. Specifically, even when the contact points between the hand and the object switch and the positions of these points change due to its shape and changing orientation, constant grasp of the object is required. This paper presents an IHM method for a robotic hand with belts, based on the prediction of the contact-point changes via image information. The focus is on a robotic hand that has a two-fingered parallel gripper with conveyor belts which can continuously manipulate an object through a large range. A stereo camera is attached to the hand. First, the contour of the grasped object is acquired from the camera. From the contour, the switching of the contact points between the surfaces of the belts and the object is predicted. Then, the positions of the contact points in the next frame are estimated by rotating the contour. The velocities of the belts are calculated based on the prediction of the switching. The fingers are controlled to follow the estimated positions of the contact points, via a feed-forward control. The effectiveness of the proposed method is verified through in-hand manipulation experiments for 22 objects of various shapes and sizes. Yuzuka Isobe, Sunhwi Kang, Takeshi Shimamoto, Yoshinari Matsuyama, Sarthak Pathak, Kazunori Umeda |
IROS | 5 |
| 2023 | Intuitive Arm-Pointing based Home-Appliance Control from Multiple Camera ViewsabstractThe purpose of this paper is to construct and evaluate a system to operate home appliances by pointing. In Human Machine Interface (HMI) design, a natural operating method is important. Pointing is a universal gesture for selecting an object. Arm-pointing to an appliance and selecting it to perform a simple operation is a very intuitive and easy-to-use method of operation. Many studies prepare data with locations of appliances and their sizes. In this paper, we a camera-based system where the user can simply point at an appliance to select and operate it is proposed. The user’s pointing direction and appliance locations are estimated automatically from image frames. This eliminates the need for any preparation beforehand and the appliances can be moved during operation. The proposed method was implemented and experimentally evaluated. It was found that the average recognition rates were about 87% and 57% when a humidifier and a TV were operated. Masae Yokota, Soichiro Majima, Sarthak Pathak, Kazunori Umeda |
RO-MAN | 3 |
| 2021 | Scale Optimization of Structure from Motion for Structured Light-based All-round 3D MeasurementabstractIn this paper, we propose a novel method for 3D measurement of large structures that have sparse features. The proposed method uses a structured-light approach based on a spherical camera and an omnidirectional ring laser. The spherical camera can capture omnidirectional images which enable it to view all sparse feature points existing in the target environment. The omnidirectional ring laser can enable dense 3D measurement of cross-sections of the environment via the structured light method. Structure from Motion (SfM) is used to measure the motion of the camera to integrate the laser cross sections to obtain a dense 3D model.However, the result of SfM does not contain real-world scale information. The novelty of this research lies in a new method to obtain the real-world scale. The real-world scale is determined by comparing a mesh generated from the resultant SfM point cloud and the integrated laser sections in terms of each scale.In a simulated environment, the proposed method was found to be accurate up to 1 mm. It was also able to accurately measure the 3D shape of a real environment. Momoko Kawata, Hiroshi Higuchi, Sarthak Pathak, Atsushi Yamashita, Hajime Asama |
SMC | 3 |
| 2019 | E-CNN: Accurate Spherical Camera Rotation Estimation via Uniformization of Distorted Optical Flow FieldsabstractSpherical cameras, which can acquire all-round information, are effective to estimate rotation for robotic applications. Recently, Convolutional Neural Networks have shown great robustness in solving such regression problems. However they are designed for planar images and cannot deal with the non-uniform distortion present in spherical images, when expressed in the planar equirectangular projection. This can lower the accuracy of motion estimation. In this research, we propose an Equirectangular-Convolutional Neural Network (E-CNN) to solve this issue. This novel network regresses 3D spherical camera rotation by uniformizing distorted optical flow patterns in the equirectangular projection. We experimentally show that this results in consistently lower error as opposed to learning from the distorted optical flow. Dabae Kim, Sarthak Pathak, Alessandro Moro, Ren Komatsu, Atsushi Yamashita, Hajime Asama |
ICASSP | 2 |
| 2019 | Accurate All-Round 3D Measurement Using Trinocular Spherical Stereo via Weighted Reprojection Error MinimizationabstractComparing to perspective cameras, the all-round 3D measurement of the environment can be done by spherical cameras in a more efficient way. However, the measurement using binocular spherical stereo has two singularity points at the epipoles of each spherical camera, where the measurement result gets extremely sensitive to the error when getting close to the epipoles and along the epipolar directions. This affects the accuracy of 3D reconstruction along with the epipolar directions. A three-way measurement method using three spherical cameras with trinocular spherical stereo setup is proposed in this paper to achieve accurate all-round 3D measurement. The improved accuracy of 3D measurement by the implementation of weighted reprojection error optimization was verified in experiments. Wanqi Yin, Sarthak Pathak, Alessandro Moro, Atsushi Yamashita, Hajime Asama |
ISM | 2 |
| 2018 | Distortion-Robust Spherical Camera Motion Estimation via Dense Optical FlowabstractConventional techniques for frame-to-frame camera motion estimation rely on tracking a set of sparse feature points. However, images taken from spherical cameras have high distortion which can induce mistakes in feature point tracking, offsetting the advantage of their large fields-of-view. Hence, in this research, we attempt a novel approach of using dense optical flow for distortion-robust spherical camera motion estimation. Dense optical flow incorporates smoothing terms and is free of local outliers. It encodes the camera motion as well as dense 3D information. Our approach decomposes dense optical flow into epipolar geometry and the dense disparity map, and reprojects this disparity map to estimate 6 DoF camera motion. The approach handles spherical image distortion in a natural way. We experimentally demonstrate its accuracy and robustness. Sarthak Pathak, Alessandro Moro, Hiromitsu Fujii, Atsushi Yamashita, Hajime Asama |
ICIP | 1 |
| 2018 | Line-Based Global Localization of a Spherical Camera in Manhattan WorldsabstractLocalization is an important task for mobile service robots in indoor spaces. In this research, we propose a novel technique for indoor localization using a spherical camera. Spherical cameras can obtain a complete view of the surroundings allowing the use of global environmental information. We take advantage of this in order to estimate camera position and the orientation with respect to a known 3D line map of an indoor environment, using a single image. We robustly extract 2D line information from the spherical image via spherical-gradient filtering and match it to 3D line information in the line map. Our method requires no information about the 3D-2D line correspondences. In order to avoid a complicated six degrees of freedom (6 DoF) search for position and orientation, we use a Manhattan world assumption to decompose the line information in the image. The 6 DoF localization process is divided into two phases. First, we estimate the orientation by extracting the three principle directions from the image. Then, the position is estimated by robustly matching the distribution of lines between the image and the 3D model via a spherical Hough representation. This decoupled search can robustly localize a spherical camera using a single image, as we demonstrate experimentally. Tsubasa Goto, Sarthak Pathak, Yonghoon Ji, Hiromitsu Fujii, Atsushi Yamashita, Hajime Asama |
ICRA | 2 |
| 2016 | A decoupled virtual camera using spherical optical flowabstractIn camera-equipped teleoperated robots, it is often tedious for the operator to manage both the viewpoint and the shaky/unstable navigation, leading to disorientation. Our proposal is to create a virtual, freely rotatable camera that is decoupled from the robot's rotation. It is implemented using a complete spherical camera and removing its rotation in-image with a novel algorithm based on aligning the dense spherical optical flow field along the epipolar direction. Finally, any area on the rotation-less image sequence can be undistorted, resulting in the desired decoupled camera. We illustrate the concept by showing the effect on some videos taken from a spherical camera under different robot motions. Sarthak Pathak, Alessandro Moro, Atsushi Yamashita, Hajime Asama |
ICIP | 1 |