Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yuichi Taguchi

dblp:99/5287 · DBLP profile ↗
← Back
47ranked-venue papers
14as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-authorSystems, architecture and hardware · 15 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
23 papers
3D vision · 54% Robot navigation and mapping · 26% Robot manipulation · 12%
Computer graphics and multimedia
12 papers
Computational photography and imaging · 39% Geometric modeling and processing · 28% Virtual and augmented reality · 19%

Topics — the 30 heaviest of 57, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
SLAM
0.952017
MonoRGBD-SLAM: Simultaneous localization and mapping using both monocular and RGBD cameras · ICRA 2017
Pinpoint SLAM: A hybrid of 2D and 3D simultaneous localization and mapping for RGB-D sensors · ICRA 2016
Point-plane SLAM for hand-held 3D sensors · ICRA 2013
Computer vision › 3D vision
3d reconstruction
0.432013
Point-plane SLAM for hand-held 3D sensors · ICRA 2013
SLAM using both points and planes for hand-held 3D sensors · ISMAR 2012
Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors · CVPR 2011
Computer vision › 3D vision
depth estimation
0.442014
Rainbow Flash Camera: Depth Edge Extraction Using Complementary Colors · Int. J. Comput. Vis. 2014
Joint Geodesic Upsampling of Depth Images · CVPR 2013
Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns · ECCV (5) 2012
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.422017
MonoRGBD-SLAM: Simultaneous localization and mapping using both monocular and RGBD cameras · ICRA 2017
ROS2D: Image feature detector using rank order statistics · ICRA 2017
Computer vision › 3D vision › range sensing
depth sensing
0.322016
Learning to remove multipath distortions in Time-of-Flight range images for a robotic arm setup · ICRA 2016
Pinpoint SLAM: A hybrid of 2D and 3D simultaneous localization and mapping for RGB-D sensors · ICRA 2016
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
planar surface reconstruction
0.322013
Point-plane SLAM for hand-held 3D sensors · ICRA 2013
SLAM using both points and planes for hand-held 3D sensors · ISMAR 2012
Computer vision › 3D vision › low-level vision
feature detection
0.312017
ROS2D: Image feature detector using rank order statistics · ICRA 2017
Robotics › Robot navigation and mapping › SLAM
multi-sensor SLAM
0.312017
MonoRGBD-SLAM: Simultaneous localization and mapping using both monocular and RGBD cameras · ICRA 2017
Geometric modeling and processing › curve fitting
spline fitting
0.312017
FasTFit: A fast T-spline fitting algorithm · Comput. Aided Des. 2017
Robotics › Robot manipulation
assembly
0.332012
Finding a needle in a specular haystack · ICRA 2011
Rao-Blackwellized particle filtering for probing-based 6-DOF localization in robotic assembly · ICRA 2010
Voting-based pose estimation for robotic assembly using a 3D sensor · ICRA 2012
Computer vision › 3D vision
point cloud registration
0.322013
A Theory of Minimal 3D Point to 3D Plane Registration and Its Generalization · Int. J. Comput. Vis. 2013
P2Pi: A Minimal Solution for Registration of 3D Points to 3D Planes · ECCV (5) 2010
Computer vision › 3D vision
camera calibration
0.322012
A theory of multi-layer flat refractive geometry · CVPR 2012
Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors · CVPR 2011
Computer vision › 3D vision
pose estimation
0.322012
Voting-based pose estimation for robotic assembly using a 3D sensor · ICRA 2012
Finding a needle in a specular haystack · ICRA 2011
Robotics › Robot navigation and mapping › SLAM › visual SLAM
RGB-D SLAM
0.212016
Pinpoint SLAM: A hybrid of 2D and 3D simultaneous localization and mapping for RGB-D sensors · ICRA 2016
Robotics › Robot manipulation › grasping › grasping in clutter
bin picking
0.222014
Fast graspability evaluation on single depth maps for bin picking with general grippers · ICRA 2014
Voting-based pose estimation for robotic assembly using a 3D sensor · ICRA 2012
Computer vision › 3D vision › geometric deep learning
3d feature learning
0.212014
Learning to Rank 3D Features · ECCV (1) 2014
Robotics › Robot manipulation › grasping › grasp analysis
graspability analysis
0.212014
Fast graspability evaluation on single depth maps for bin picking with general grippers · ICRA 2014
Robotics › Robot manipulation
grasping
0.212014
Fast graspability evaluation on single depth maps for bin picking with general grippers · ICRA 2014
Machine learning › Probabilistic and Bayesian machine learning › clustering
hierarchical agglomerative clustering
0.212014
Fast plane extraction in organized point clouds using agglomerative hierarchical clustering · ICRA 2014
Computer vision › 3D vision › 3d scene understanding › scene geometry
plane detection
0.212014
Fast plane extraction in organized point clouds using agglomerative hierarchical clustering · ICRA 2014
Computer vision › 3D vision
point cloud processing
0.212014
Fast plane extraction in organized point clouds using agglomerative hierarchical clustering · ICRA 2014
Computer vision › 3D vision › range sensing
structured light
0.222012
Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns · ECCV (5) 2012
Rainbow Flash Camera: Depth Edge Extraction Using Complementary Colors · ECCV (6) 2012
Computer vision › 3D vision › depth estimation
depth map upsampling
0.212013
Joint Geodesic Upsampling of Depth Images · CVPR 2013
Computer vision › 3D vision › 3d scene understanding
room layout estimation
0.212013
Manhattan Junction Catalogue for Spatial Reasoning of Indoor Scenes · CVPR 2013
Computer vision › Segmentation and scene understanding
scene understanding
0.212013
Manhattan Junction Catalogue for Spatial Reasoning of Indoor Scenes · CVPR 2013
Computational photography and imaging › depth imaging
depth edge detection
0.112012
Rainbow Flash Camera: Depth Edge Extraction Using Complementary Colors · ECCV (6) 2012
Geometric modeling and processing › 3d reconstruction › shape from silhouette
visual hull reconstruction
0.112012
Convex bricks: A new primitive for visual hull modeling and reconstruction · ICRA 2012
Computer vision › 3D vision › structure from motion
bundle adjustment
0.112011
Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors · CVPR 2011
Computer vision › 3D vision
object pose estimation
0.112010
Rao-Blackwellized particle filtering for probing-based 6-DOF localization in robotic assembly · ICRA 2010
Robotics › Robot navigation and mapping
state estimation
0.112010
Rao-Blackwellized particle filtering for probing-based 6-DOF localization in robotic assembly · ICRA 2010

Methods — techniques the papers use, named apart from their topics

RANSAC · 0.5structured light · 0.4triangulation · 0.4minimal solver · 0.3t-spline · 0.3rank order statistics · 0.3minimum spanning tree · 0.3bundle adjustment · 0.3RGBD-to-monocular registration · 0.3keypoint extraction · 0.2deep learning · 0.2complementary color flash · 0.2registration · 0.2minimal solvers · 0.2geodesic distance · 0.2all-pair-shortest-path approximation · 0.2rainbow flash · 0.1linear programming · 0.1
YearPublicationVenuePosition
2019 Unified Underwater Structure-from-Motion
abstract
This paper shows that accurate underwater 3D shape reconstruction is possible using a single camera, observing a target through a refractive interface. We provide unified reconstruction techniques for a variety of scenarios such as single static camera and moving refractive interface, single moving camera and static refractive interface, and single moving camera and moving refractive interface. In our basic setup, we assume that the refractive interface is planar, and simultaneously estimate the unknown transformations of the planar interface and the camera, and the unknown target shape using bundle adjustment. We also extend it to relax the planarity assumption, which enables us to use waves of the refractive interface for the reconstruction task. Experiments with real data show the superiority of our method to existing methods.
Kazuto Ichimaru, Yuichi Taguchi, Hiroshi Kawasaki
3DV2
2019 Visual odometry with a single-camera stereo omnidirectional system
Carlos Jaramillo, Juan Pablo Muñoz, Yuichi Taguchi, Jizhong Xiao
Mach. Vis. Appl.4
2018 System Resource Management to Control the Risk of Data-Loss in a Cloud-Based Disaster Recovery
abstract
A method for system resource management of a cloud-based disaster recovery is proposed. This method can determine whether system performance achieves the "recovery point objective" (RPO) or not. It is also necessary to resize resources for satisfying required performance at lower cost. Therefore, a formula for simulating system performance is established. The results of the simulation show it is possible to determine an appropriate system size that achieves RPO as well as minimizes amount of system resources. It is demonstrated that no error between the experimentally measured and simulated recovery points exists at peak workload. It is thus concluded that the proposed simulation is accurate enough. Also, it is estimated that 89% of system resource cost could be reduced in comparison to the case that the disaster recovery system is designed to fit a peak workload.
Yuichi Taguchi, Tsutomu Yoshinaga
COMPSAC (2)1
2018 VLASE: Vehicle Localization by Aggregating Semantic Edges
abstract
We propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pairs of distinct objects such as building-sky, road-sidewalk, and building-ground. While prior work has shown promising results by utilizing the boundary between prominent classes such as sky and building using skylines, we generalize this to consider 19 semantic classes. We extract semantic edge features using CASENet architecture and utilize VLAD framework to perform image retrieval. We achieve improvement over state-of-the-art localization algorithms such as SIFT-VLAD and its deep variant NetVLAD. Ablation study shows the importance of different semantic classes, and our unified approach achieves better performance compared to individual prominent features such as skylines. We also introduce SLC Marathon dataset, a challenging dataset covering most of Salt Lake City with sufficient lighting variations.
Xin Yu 0003, Sagar Chaturvedi, Chen Feng 0002, Yuichi Taguchi, Teng-Yok Lee, Clinton Fernandes, Srikumar Ramalingam
IROS4
2017 3D Object Discovery and Modeling Using Single RGB-D Images Containing Multiple Object Instances
abstract
Unsupervised object modeling is important in robotics, especially for handling a large set of objects. We present a method for unsupervised 3D object discovery, reconstruction, and localization that exploits multiple instances of an identical object contained in a single RGB-D image. The proposed method does not rely on segmentation, scene knowledge, or user input, and thus is easily scalable. Our method aims to find recurrent patterns in a single RGB-D image by utilizing appearance and geometry of the salient regions. We extract keypoints and match them in pairs based on their descriptors. We then generate triplets of the keypoints matching with each other using several geometric criteria to minimize false matches. The relative poses of the matched triplets are computed and clustered to discover sets of triplet pairs with similar relative poses. Triplets belonging to the same set are likely to belong to the same object and are used to construct an initial object model. Detection of remaining instances with the initial object model using RANSAC allows to further expand and refine the model. The automatically generated object models are both compact and descriptive. We show quantitative and qualitative results on RGB-D images with various objects including some from the Amazon Picking Challenge. We also demonstrate the use of our method in an object picking scenario with a robotic arm.
Wim Abbeloos, Esra Ataer Cansizoglu, Sergio Caccamo, Yuichi Taguchi, Yukiyasu Domae
3DV4
2017 Joint 3D Reconstruction of a Static Scene and Moving Objects
abstract
We present a technique for simultaneous 3D reconstruction of static regions and rigidly moving objects in a scene. An RGB-D frame is represented as a collection of features, which are points and planes. We classify the features into static and dynamic regions and grow separate maps, static and object maps, for each of them. To robustly classify the features in each frame, we fuse multiple RANSAC-based registration results obtained by registering different groups of the features to different maps, including (1) all the features to the static map, (2) all the features to each object map, and (3) subsets of the features, each forming a segment, to each object map. This multi-group registration approach is designed to overcome the following challenges: scenes can be dominated by static regions, making object tracking more difficult; and moving object might have larger pose variation between frames compared to the static regions. We show qualitative results from indoor scenes with objects in various shapes. The technique enables on-the-fly object model generation to be used for robotic manipulation.
Sergio Caccamo, Esra Ataer Cansizoglu, Yuichi Taguchi
3DV3
2017 Barcode: Global Binary Patterns for Fast Visual Inference
abstract
We present Barcode, a global binary descriptor for images captured from a vehicle-mounted camera with two applications: localization and turn classification. Barcode characterizes an image by encoding the distribution of vertical lines into a binary descriptor: in each vertical stripe of an image, if any vertical line exists the corresponding bit is set to 1, otherwise 0. For localization, our approach uses a database of geolocated images, each having its Barcode precomputed during a preprocessing stage. In the run time, we first generate the binary descriptor for each image and then use the descriptor to find the location in the database via Hamming distance metric. For turn classification, we train a deep neural network that uses a set of Barcodes from consecutive images to classify turns (left, right, straight, and stationary). We show that Barcode extraction can be done at 100-1000~Hz, localization at 10~kHz, and turn classification at 1~kHz. We show compelling experimental results on KITTI dataset and other sequences captured near Harvard and Purdue campuses.
Teng-Yok Lee, Sonali Patil, Srikumar Ramalingam, Yuichi Taguchi, Bedrich Benes
3DV4
2017 Compression of 3-D point clouds using hierarchical patch fitting
abstract
For applications such as virtual reality and mobile mapping, point clouds are an effective means for representing 3-D environments. The need for compressing such data is rapidly increasing, given the widespread use and precision of these systems. This paper presents a method for compressing organized point clouds. 3-D point cloud data is mapped to a 2-D organizational grid, where each element on the grid is associated with a point in 3-D space and its corresponding attributes. The data on the 2-D grid is hierarchically partitioned, and a Bezier patch is fit to the 3-D coordinates associated with each partition. Residual values are quantized and signaled along with data necessary to reconstruct the patch hierarchy in the decoder. We show how this method can be used to process point clouds captured by a mobile-mapping system, in which laser-scanned point locations are organized and compressed. The performance of the patch-fitting codec exceeds or is comparable to that of an octree-based codec.
Robert A. Cohen, Maja Krivokuca, Chen Feng 0002, Yuichi Taguchi, Hideaki Ochimizu, Dong Tian, Anthony Vetro
ICIP4
2017 Fusion of multi-angular aerial images based on epipolar geometry and matrix completion
abstract
We consider the problem of fusing multiple cloud-contaminated aerial images of a 3D scene to generate a cloud-free image, where the images are captured from multiple unknown view angles. In order to fuse these images, we propose an end-to-end framework incorporating epipolar geometry and low-rank matrix completion. In particular, we first warp the multi-angular images to single-angle ones based on the estimated fundamental matrices that relate the multi-angular images according to their projective relations to the 3D scene. Then we formulate the fusion process of the warpped images as a low-rank matrix completion problem where each column of the matrix corresponds to a vectorized image with missing entries corresponding to cloud or occluded areas. Results using DigitalGlobe high spatial resolution images demonstrate that our algorithm outperforms existing approaches.
Yanting Ma, Dehong Liu, Hassan Mansour, Ulugbek Kamilov, Yuichi Taguchi, Petros Boufounos, Anthony Vetro
ICIP5
2017 MonoRGBD-SLAM: Simultaneous localization and mapping using both monocular and RGBD cameras
abstract
RGBD SLAM systems have shown impressive results, but the limited field of view (FOV) and depth range of typical RGBD cameras still cause problems for registering distant frames. Monocular SLAM systems, in contrast, can exploit wide-angle cameras and do not have the depth range limitation, but are unstable for textureless scenes. We present a SLAM system that uses both an RGBD camera and a wide-angle monocular camera for combining the advantages of the two sensors. Our system extracts 3D point features from RGBD frames and 2D point features from monocular frames, which are used to perform both RGBD-to-RGBD and RGBD-to-monocular registration. To compensate for different FOV and resolution of the cameras, we generate multiple virtual images for each wide-angle monocular image and use the feature descriptors computed on the virtual images to perform the RGBD-to-monocular matches. To compute the poses of the frames, we construct a graph where nodes represent RGBD and monocular frames and edges denote the pairwise registration results between the nodes. We compute the global poses of the nodes by first finding the minimum spanning trees (MSTs) of the graph and then pruning edges that have inconsistent poses due to possible mismatches using the MST result. We finally run bundle adjustment on the graph using all the consistent edges. Experimental results show that our system registers a larger number of frames than using only an RGBD camera, leading to larger-scale 3D reconstruction.
Khalid Yousif, Yuichi Taguchi, Srikumar Ramalingam
ICRA2
2017 ROS2D: Image feature detector using rank order statistics
abstract
We present a new image feature detection method. Our method selects features based on segmenting points with high local intensity variations across different scales using a robust rank order statistics approach. Our method produces a large number of repeatable features that are invariant to several image transformations such as rotation, scaling, viewpoint, and lighting variations. We show the advantages of our feature in comparison to other existing features using the Oxford dataset. We also show that, when used in monocular and stereo SLAM systems, our feature outperforms SIFT in terms of the pose estimation accuracy using several public datasets including the KITTI dataset.
Khalid Yousif, Yuichi Taguchi, Srikumar Ramalingam, Alireza Bab-Hadiashar
ICRA2
2017 FasTFit: A fast T-spline fitting algorithm
Chen Feng 0002, Yuichi Taguchi
Comput. Aided Des.2
2016 Pinpoint SLAM: A hybrid of 2D and 3D simultaneous localization and mapping for RGB-D sensors
abstract
Conventional SLAM systems with an RGB-D sensor use depth measurements only in a limited depth range due to hardware limitation and noise of the sensor, ignoring regions that are too far or too close from the sensor. Such systems introduce registration errors especially in scenes with large depth variations. In this paper, we present a novel RGB-D SLAM system that makes use of both 2D and 3D measurements. Our system first extracts keypoints from RGB images and generates 2D and 3D point features from the keypoints with invalid and valid depth values, respectively. It then establishes 3D-to-3D, 2D-to-3D, and 2D-to-2D point correspondences among frames. For the 2D-to-3D point correspondences, we use the rays defined by the 2D point features to “pinpoint” the corresponding 3D point features, generating longer-range constraints than using only 3D-to-3D correspondences. For the 2D-to-2D point correspondences, we triangulate the rays to generate 3D points that are used as 3D point features in the subsequent process. We use the hybrid correspondences in both online SLAM and offline postprocessing: the online SLAM focuses more on the speed by computing correspondences among consecutive frames for real-time operations, while the offline postprocessing generates more correspondences among all the frames for higher accuracy. The results on RGB-D SLAM benchmarks show that the online SLAM provides higher accuracy than conventional SLAM systems, while the postprocessing further improves the accuracy.
Esra Ataer Cansizoglu, Yuichi Taguchi, Srikumar Ramalingam
ICRA2
2016 Learning to remove multipath distortions in Time-of-Flight range images for a robotic arm setup
abstract
Range images captured by Time-of-Flight (ToF) cameras are corrupted with multipath distortions due to interaction between modulated light signals and scenes. The interaction is often complicated, which makes a model-based solution elusive. We propose a learning-based approach for removing the multipath distortions for a ToF camera in a robotic arm setup. Our approach is based on deep learning. We use the robotic arm to automatically collect a large amount of ToF range images containing various multipath distortions. The training images are automatically labeled by leveraging a high precision structured light sensor available only in the training time. In the test time, we apply the learned model to remove the multipath distortions. This allows our robotic arm setup to enjoy the speed and compact form of the ToF camera without compromising with its range measurement errors. We conduct extensive experimental validations and compare the proposed method to several baseline algorithms. The experiment results show that our method achieves 55% error reduction in range estimation and largely outperforms the baseline algorithms.
Kilho Son, Ming-Yu Liu 0001, Yuichi Taguchi
ICRA3
2016 Object detection and tracking in RGB-D SLAM via hierarchical feature grouping
abstract
We present an object detection and tracking framework integrated into a simultaneous localization and mapping (SLAM) system using an RGB-D camera. We propose a compact representation of objects by grouping features hierarchically. Similar to a keyframe being a collection of features, an object is represented as a set of segments, where a segment is a subset of features in a frame. Just like keyframes, segments are registered with each other in a map, which we call an object map. We use the same SLAM procedure in both offline object scanning and online object detection modes. In the offline scanning mode, we scan an object using an RGB-D camera to generate an object map. In the online detection mode, a set of object maps for different objects is given, and the objects are detected via appearance-based matching between the segments in the current frame and in the object maps. In the case of a match, the object is localized with respect to the map being reconstructed by the SLAM system by a RANSAC registration. In the subsequent frames, the tracking is done by predicting the poses of the objects. We also incorporate constraints obtained from the objects into bundle adjustment to improve the object pose estimation accuracy as well as the SLAM reconstruction accuracy. We demonstrate our technique in an object picking scenario using a robot arm. Experimental results show that the system is able to detect and pick up objects successfully from different viewpoints and distances.
Esra Ataer Cansizoglu, Yuichi Taguchi
IROS2
2016 Parameter learning for improving binary descriptor matching
abstract
Binary descriptors allow fast detection and matching algorithms in computer vision problems. Though binary descriptors can be computed at almost two orders of magnitude faster than traditional gradient based descriptors, they suffer from poor matching accuracy in challenging conditions. In this paper we propose three improvements for binary descriptors in their computation and matching that enhance their performance in comparison to traditional binary and non-binary descriptors without compromising their speed. This is achieved by learning some weights and threshold parameters that allow customized matching under some variations such as lighting and viewpoint. Our suggested improvements can be easily applied to any binary descriptor. We demonstrate our approach on the ORB (Oriented FAST and Rotated BRIEF) descriptor and compare its performance with the traditional ORB and SIFT descriptors on a wide variety of datasets. In all instances, our enhancements outperform standard ORB and are comparable to SIFT.
Bharath Sankaran, Srikumar Ramalingam, Yuichi Taguchi
IROS3
2015 Estimating Drivable Collision-Free Space from Monocular Video
abstract
In this paper we propose a novel algorithm for estimating the drivable collision-free space for autonomous navigation of on-road and on-water vehicles. In contrast to previous approaches that use stereo cameras or LIDAR, we show a method to solve this problem using a single camera. Inspired by the success of many vision algorithms that employ dynamic programming for efficient inference, we reduce the free space estimation task to an inference problem on a 1D graph, where each node represents a column in the image and its label denotes a position that separates the free space from the obstacles. Our algorithm exploits several image and geometric features based on edges, color, and homography to define potential functions on the 1D graph, whose parameters are learned through structured SVM. We show promising results on the challenging KITTI dataset as well as video collected from boats.
Srikumar Ramalingam, Yuichi Taguchi, Yohei Miki 0003, Raquel Urtasun
WACV3
2014 Calibration of Non-overlapping Cameras Using an External SLAM System
abstract
We present a simple method for calibrating a set of cameras that may not have overlapping field of views. We reduce the problem of calibrating the non-overlapping cameras to the problem of localizing the cameras with respect to a global 3D model reconstructed with a simultaneous localization and mapping (SLAM) system. Specifically, we first reconstruct such a global 3D model by using a SLAM system using an RGB-D sensor. We then perform localization and intrinsic parameter estimation for each camera by using 2D-3D correspondences between the camera and the 3D model. Our method locates the cameras within the 3D model, which is useful for visually inspecting camera poses and provides a model-guided browsing interface of the images. We demonstrate the advantages of our method using several indoor scenes.
Esra Ataer Cansizoglu, Yuichi Taguchi, Srikumar Ramalingam, Yohei Miki 0003
3DV2
2014 Learning to Rank 3D Features
Oncel Tuzel, Ming-Yu Liu 0001, Yuichi Taguchi, Arvind Raghunathan
ECCV (1)3
2014 Fast graspability evaluation on single depth maps for bin picking with general grippers
abstract
We present a method that estimates graspability measures on a single depth map for grasping objects randomly placed in a bin. Our method represents a gripper model by using two mask images, one describing a contact region that should be filled by a target object for stable grasping, and the other describing a collision region that should not be filled by other objects to avoid collisions during grasping. The graspability measure is computed by convolving the mask images with binarized depth maps, which are thresholded differently in each region according to the minimum height of the 3D points in the region and the length of the gripper. Our method does not assume any 3-D model of objects, thus applicable to general objects. Our representation of the gripper model using the two mask images is also applicable to general grippers, such as multi-finger and vacuum grippers. We apply our method to bin picking of piled objects using a robot arm and demonstrate fast pick-and-place operations for various industrial objects.
Yukiyasu Domae, Haruhisa Okuda, Yuichi Taguchi, Kazuhiko Sumi, Takashi Hirai
ICRA3
2014 Fast plane extraction in organized point clouds using agglomerative hierarchical clustering
abstract
Real-time plane extraction in 3D point clouds is crucial to many robotics applications. We present a novel algorithm for reliably detecting multiple planes in real time in organized point clouds obtained from devices such as Kinect sensors. By uniformly dividing such a point cloud into non-overlapping groups of points in the image space, we first construct a graph whose node and edge represent a group of points and their neighborhood respectively. We then perform an agglomerative hierarchical clustering on this graph to systematically merge nodes belonging to the same plane until the plane fitting mean squared error exceeds a threshold. Finally we refine the extracted planes using pixel-wise region growing. Our experiments demonstrate that the proposed algorithm can reliably detect all major planes in the scene at a frame rate of more than 35Hz for 640×480 point clouds, which to the best of our knowledge is much faster than state-of-the-art algorithms.
Chen Feng 0002, Yuichi Taguchi, Vineet R. Kamat
ICRA2
2014 Rainbow Flash Camera: Depth Edge Extraction Using Complementary Colors
Yuichi Taguchi
Int. J. Comput. Vis.1
2013 Joint Geodesic Upsampling of Depth Images
abstract
We propose an algorithm utilizing geodesic distances to upsample a low resolution depth image using a registered high resolution color image. Specifically, it computes depth for each pixel in the high resolution image using geodesic paths to the pixels whose depths are known from the low resolution one. Though this is closely related to the all-pair-shortest-path problem which has O(n2log n) complexity, we develop a novel approximation algorithm whose complexity grows linearly with the image size and achieve realtime performance. We compare our algorithm with the state of the art on the benchmark dataset and show that our approach provides more accurate depth upsampling with fewer artifacts. In addition, we show that the proposed algorithm is well suited for upsampling depth images using binary edge maps, an important sensor fusion application.
Ming-Yu Liu 0001, Oncel Tuzel, Yuichi Taguchi
CVPR3
2013 Manhattan Junction Catalogue for Spatial Reasoning of Indoor Scenes
abstract
Junctions are strong cues for understanding the geometry of a scene. In this paper, we consider the problem of detecting junctions and using them for recovering the spatial layout of an indoor scene. Junction detection has always been challenging due to missing and spurious lines. We work in a constrained Manhattan world setting where the junctions are formed by only line segments along the three principal orthogonal directions. Junctions can be classified into several categories based on the number and orientations of the incident line segments. We provide a simple and efficient voting scheme to detect and classify these junctions in real images. Indoor scenes are typically modeled as cuboids and we formulate the problem of the cuboid layout estimation as an inference problem in a conditional random field. Our formulation allows the incorporation of junction features and the training is done using structured prediction techniques. We outperform other single view geometry estimation methods on standard datasets.
Srikumar Ramalingam, Jaishanker K. Pillai, Yuichi Taguchi
CVPR4
2013 Point-plane SLAM for hand-held 3D sensors
abstract
We present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. We show that it is possible to register 3D data in two different coordinate systems using any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points). Our algorithm uses the minimal set of primitives in a RANSAC framework to robustly compute correspondences and estimate the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D data, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of indoor spaces using a hand-held Kinect sensor.
Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002
ICRA1
2013 A Theory of Minimal 3D Point to 3D Plane Registration and Its Generalization
Srikumar Ramalingam, Yuichi Taguchi
Int. J. Comput. Vis.2
2012 A theory of multi-layer flat refractive geometry
abstract
Flat refractive geometry corresponds to a perspective camera looking through single/multiple parallel flat refractive mediums. We show that the underlying geometry of rays corresponds to an axial camera. This realization, while missing from previous works, leads us to develop a general theory of calibrating such systems using 2D-3D correspondences. The pose of 3D points is assumed to be unknown and is also recovered. Calibration can be done even using a single image of a plane. We show that the unknown orientation of the refracting layers corresponds to the underlying axis, and can be obtained independently of the number of layers, their distances from the camera and their refractive indices. Interestingly, the axis estimation can be mapped to the classical essential matrix computation and 5-point algorithm [15] can be used. After computing the axis, the thicknesses of layers can be obtained linearly when refractive indices are known, and we derive analytical solutions when they are unknown. We also derive the analytical forward projection (AFP) equations to compute the projection of a 3D point via multiple flat refractions, which allows non-linear refinement by minimizing the reprojection error. For two refractions, AFP is either 4th or 12th degree equation depending on the refractive indices. We analyze ambiguities due to small field of view, stability under noise, and show how a two layer system can be well approximated as a single layer system. Real experiments using a water tank validate our theory.
Amit K. Agrawal, Srikumar Ramalingam, Yuichi Taguchi, Visesh Chari
CVPR3
2012 Rainbow Flash Camera: Depth Edge Extraction Using Complementary Colors
Yuichi Taguchi
ECCV (6)1
2012 Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns
Yuichi Taguchi, Amit K. Agrawal, Oncel Tuzel
ECCV (5)1
2012 Variable focus video: Reconstructing depth and video for dynamic scenes
abstract
Traditional depth from defocus (DFD) algorithms assume that the camera and the scene are static during acquisition time. In this paper, we examine the effects of camera and scene motion on DFD algorithms. We show that, given accurate estimates of optical flow (OF), one can robustly warp the focal stack (FS) images to obtain a virtual static FS and apply traditional DFD algorithms on the static FS. Acquiring accurate OF in the presence of varying focal blur is a challenging task. We show how defocus blur variations cause inherent biases in the estimates of optical flow. We then show how to robustly handle these biases and compute accurate OF estimates in the presence of varying focal blur. This leads to an architecture and an algorithm that converts a traditional 30 fps video camera into a co-located 30 fps image and a range sensor. Further, the ability to extract image and range information allows us to render images with artistic depth-of field effects, both extending and reducing the depth of field of the captured images. We demonstrate experimental results on challenging scenes captured using a camera prototype.
Nitesh Shroff, Ashok Veeraraghavan, Yuichi Taguchi, Oncel Tuzel, Amit K. Agrawal, Rama Chellappa
ICCP3
2012 Convex bricks: A new primitive for visual hull modeling and reconstruction
abstract
Industrial automation tasks typically require a 3D model of the object for robotic manipulation. The ability to reconstruct the 3D model using a sample object is useful when CAD models are not available. For textureless objects, visual hull of the object obtained using silhouette-based reconstruction can avoid expensive 3D scanners for 3D modeling. We propose convex brick (CB), a new 3D primitive for modeling and reconstructing a visual hull from silhouettes. CB's are powerful in modeling arbitrary non-convex 3D shapes. Using CB, we describe an algorithm to generate a polyhedral visual hull from polygonal silhouettes; the visual hull is reconstructed as a combination of 3D convex bricks. Our approach uses well-studied geometric operations such as 2D convex decomposition and intersection of 3D convex cones using linear programming. The shape of CB can adapt to the given silhouettes, thereby significantly reducing the number of primitives required for a volumetric representation. Our framework allows easy control of reconstruction parameters such as accuracy and the number of required primitives. We present an extensive analysis of our algorithm and show visual hull reconstruction on challenging real datasets consisting of highly non-convex shapes. We also show real results on pose estimation of an industrial part in a bin-picking system using the reconstructed visual hull.
Visesh Chari, Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam
ICRA3
2012 Voting-based pose estimation for robotic assembly using a 3D sensor
abstract
We propose a voting-based pose estimation algorithm applicable to 3D sensors, which are fast replacing their 2D counterparts in many robotics, computer vision, and gaming applications. It was recently shown that a pair of oriented 3D points, which are points on the object surface with normals, in a voting framework enables fast and robust pose estimation. Although oriented surface points are discriminative for objects with sufficient curvature changes, they are not compact and discriminative enough for many industrial and real-world objects that are mostly planar. As edges play the key role in 2D registration, depth discontinuities are crucial in 3D. In this paper, we investigate and develop a family of pose estimation algorithms that better exploit this boundary information. In addition to oriented surface points, we use two other primitives: boundary points with directions and boundary line segments. Our experiments show that these carefully chosen primitives encode more information compactly and thereby provide higher accuracy for a wide class of industrial parts and enable faster computation. We demonstrate a practical robotic bin-picking system using the proposed algorithm and a 3D sensor.
Changhyun Choi, Yuichi Taguchi, Oncel Tuzel, Ming-Yu Liu 0001, Srikumar Ramalingam
ICRA2
2012 SLAM using both points and planes for hand-held 3D sensors
abstract
We present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. Our algorithm uses any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points) in a RANSAC framework to efficiently compute the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D scenes, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of office spaces using a hand-held Kinect sensor.
Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002
ISMAR1
2011 Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors
abstract
Catadioptric cameras are widely used to increase the field of view using mirrors. Central catadioptric systems having an effective single viewpoint are easy to model and use, but severely constraint the camera positioning with respect to the mirror. On the other hand, non-central catadioptric systems allow greater flexibility in camera placement, but are often approximated using central or linear models due to the lack of an exact model. We bridge this gap and describe an exact projection model for non-central catadioptric systems. We derive an analytical `forward projection' equation for the projection of a 3D point reflected by a quadric mirror on the imaging plane of a perspective camera, with no restrictions on the camera placement, and show that it is an 8thdegree equation in a single unknown. While previous non-central catadioptric cameras primarily use an axial configuration where the camera is placed on the axis of a rotationally symmetric mirror, we allow off-axis (any) camera placement. Using this analytical model, a non-central catadioptric camera can be used for sparse as well as dense 3D reconstruction similar to perspective cameras, using well-known algorithms such as bundle adjustment and plane sweeping. Our paper is the first to show such results for off-axis placement of camera with multiple quadric mirrors. Simulation and real results using parabolic mirrors and an off-axis perspective camera are demonstrated.
Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam
CVPR2
2011 Finding a needle in a specular haystack
abstract
Progress in machine vision algorithms has led to widespread adoption of these techniques to automate several industrial assembly tasks. Nevertheless, shiny or specular objects which are common in industrial environments still present a great challenge for vision systems. In this paper, we take a step towards this problem under the context of vision-aided robotic assembly. We show that when the illumination source moves, the specular highlights remain in a region whose radius is inversely proportional to the surface curvature. This allows us to extract regions of the object that have high surface curvature. These points of high curvature can be used as features for specular objects. Further, an inexpensive multi-flash camera (MFC) design can be used to reliably extract these features. We show that one can use multiple views of the object using the MFC in order to triangulate and obtain the 3D location and pose of the shiny objects. Finally, we show a system consisting of a robot arm with an MFC that can perform automated detection and pose estimation of shiny screws within a cluttered bin, achieving position and orientation errors less than 0.5 mm and 0.8° respectively.
Nitesh Shroff, Yuichi Taguchi, Oncel Tuzel, Ashok Veeraraghavan, Srikumar Ramalingam, Haruhisa Okuda
ICRA2
2011 Entropy-based motion selection for touch-based registration using Rao-Blackwellized particle filtering
abstract
To achieve versatile locomotion in complex amphibious environments, a robot should be capable of performing different gaits. In this paper we present such a versatile amphibious robot based on a novel eccentric paddle mechanism (ePaddle). We first illustrate the concept of the ePaddle with five major possible gaits and conceptual gait sequences. We then summarize five types of configurations from these gaits. Based on these configurations, two motion behaviors are found and modeled by using kinematic equations for the future gait planning tasks. To verify the proposed ideas, we develop an ePaddle prototype module. Several simulations on these gaits are performed to verify the conceptual locomotion gait and the developed kinematic models. Experiments on five possible configurations demonstrate the valid of the ePaddle concept and the prototype design.
Yuichi Taguchi, Tim K. Marks, John R. Hershey
IROS1
2010 Axial light field for curved mirrors: Reflect your perspective, widen your view
abstract
Mirrors have been used to enable wide field-of-view (FOV) catadioptric imaging. The mapping between the incoming and reflected light rays depends non-linearly on the mirror shape and has been well-studied using caustics. We analyze this mapping using two-plane light field parameterization, which provides valuable insight into the geometric structure of reflected rays. Using this analysis, we study the problem of generating a single-viewpoint virtual perspective image for catadioptric systems, which is unachievable for several common configurations. Instead of minimizing distortions appearing in a single image, we propose to capture all the rays required to generate a virtual perspective by capturing a light field. We consider rotationally symmetric mirrors and show that a traditional planar light field results in significant aliasing artifacts. We propose axial light field, captured by moving the camera along the mirror rotation axis, for efficient sampling and to remove aliasing artifacts. This allows us to computationally generate wide FOV virtual perspectives using a wider class of mirrors than before, without using scene priors or depth estimation. We analyze the relationship between the axial light field parameters and the FOV/resolution of the resulting virtual perspective. Real results using a spherical mirror demonstrate generating 140° FOV virtual perspective using multiple 30° FOV images.
Yuichi Taguchi, Amit K. Agrawal, Srikumar Ramalingam, Ashok Veeraraghavan
CVPR1
2010 Analytical Forward Projection for Axial Non-central Dioptric and Catadioptric Cameras
Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam
ECCV (3)2
2010 P2Pi: A Minimal Solution for Registration of 3D Points to 3D Planes
Srikumar Ramalingam, Yuichi Taguchi, Tim K. Marks, Oncel Tuzel
ECCV (5)2
2010 Rao-Blackwellized particle filtering for probing-based 6-DOF localization in robotic assembly
abstract
This paper presents a probing-based method for probabilistic localization in automated robotic assembly. We consider peg-in-hole problems in which a needle-like peg has a single point of contact with the object that contains the hole, and in which the initial uncertainty in the relative pose (3D position and 3D angle) between the peg and the object is much greater than the required accuracy (assembly clearance). We solve this 6 degree-of-freedom (6-DOF) localization problem using a Rao-Blackwellized particle filter, in which the probability distribution over the peg's pose is factorized into two components: The distribution over position (3-DOF) is represented by particles, while the distribution over angle (3-DOF) is approximated as a Gaussian distribution for each particle, updated using an extended Kalman filter. This factorization reduces the number of particles required for localization by orders of magnitude, enabling real-time online 6-DOF pose estimation. Each measurement is simply the contact position obtained by randomly repositioning the peg and moving towards the object until there is contact. To compute the likelihood of each measurement, we use as a map a mesh model of the object that is based on the CAD model but also explicitly models the uncertainty in the map. The mesh uncertainty model makes our system robust to cases in which the actual measurement is different from the expected one. We demonstrate the advantages of our approach over previous methods using simulations as well as physical experiments with a robotic arm and a metal peg and object.
Yuichi Taguchi, Tim K. Marks, Haruhisa Okuda
ICRA1
2010 Axial-cones: modeling spherical catadioptric cameras for wide-angle light field rendering
abstract
Catadioptric imaging systems are commonly used for wide-angle imaging, but lead to multi-perspective images which do not allow algorithms designed for perspective cameras to be used. Efficient use of such systems requires accurate geometric ray modeling as well as fast algorithms. We present accurate geometric modeling of the multi-perspective photo captured with a spherical catadioptric imaging system usingaxial-cone cameras:multiple perspective cameras lying on an axis each with a different viewpoint and a different cone of rays. This modeling avoids geometric approximations and allows several algorithms developed for perspective cameras to be applied to multi-perspective catadioptric cameras. We demonstrate axial-cone modeling in the context of rendering wide-angle light fields, captured using a spherical mirror array. We present several applications such as spherical distortion correction, digital refocusing for artistic depth of field effects in wide-angle scenes, and wide-angle dense depth estimation. Our GPU implementation using axial-cone modeling achieves up to three orders of magnitude speed up over ray tracing for these applications.
Yuichi Taguchi, Amit K. Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam, Ramesh Raskar
ACM Trans. Graph.1
2009 TransCAIP: A Live 3D TV System Using a Camera Array and an Integral Photography Display with Interactive Control of Viewing Parameters
abstract
The system described in this paper provides a real-time 3D visual experience by using an array of 64 video cameras and an integral photography display with 60 viewing directions. The live 3D scene in front of the camera array is reproduced by the full-color, full-parallax autostereoscopic display with interactive control of viewing parameters. The main technical challenge is fast and flexible conversion of the data from the 64 multicamera images to the integral photography format. Based on image-based rendering techniques, our conversion method first renders 60 novel images corresponding to the viewing directions of the display, and then arranges the rendered pixels to produce an integral photography image. For real-time processing on a single PC, all the conversion processes are implemented on a GPU with GPGPU techniques. The conversion method also allows a user to interactively control viewing parameters of the displayed image for reproducing the dynamic 3D scene with desirable parameters. This control is performed as a software process, without reconfiguring the hardware system, by changing the rendering parameters such as the convergence point of the rendering cameras and the interval between the viewpoints of the rendering cameras.
Yuichi Taguchi, Takafumi Koike, Keita Takahashi 0001, Takeshi Naemura
IEEE Trans. Vis. Comput. Graph.1
2008 Stereo reconstruction with mixed pixels using adaptive over-segmentation
abstract
We present an over-segmentation based, dense stereo algorithm that jointly estimates segmentation and depth. For mixed pixels on segment boundaries, the algorithm computes foreground opacity (alpha), as well as color and depth for the foreground and background. We model the scene as a collection of fronto-parallel planar segments in a reference view, and use a generative model for image formation that handles mixed pixels at segment boundaries. Our method iteratively updates the segmentation based on color, depth and shape constraints using MAP estimation. Given a segmentation, the depth estimates are updated using belief propagation. We show that our method is competitive with the state-of-the-art based on the new Middlebury stereo evaluation, and that it overcomes limitations of traditional segmentation based methods while properly handling mixed pixels. Z-keying results show the advantages of combining opacity and depth estimation.
Yuichi Taguchi, Bennett Wilburn, C. Lawrence Zitnick
CVPR1
2007 Rendering-Oriented Decoding for Distributed Multi-View Coding System
abstract
This paper discusses a system in which multi-view images are captured and encoded in a distributed fashion and a viewer synthesizes a novel view from this data. We developed an efficient method for such system that combines decoding and rendering process to directly synthesize the novel image without reconstructing all the input images. Our method jointly performs disparity compensation in decoding process and geometry estimation in rendering process, because they are essentially equivalent if the camera parameters for the input images are known. It achieves low-complexity for both encoder and decoder in distributed multi-view coding system. Experimental results show superior coding performance of our method compared to a conventional intra-coding method especially at low bit rate.
Yuichi Taguchi, Takeshi Naemura
ICIP (1)1
2006 View-Dependent Coding of Light Fields Based on Free-Viewpoint Image Synthesis
abstract
This paper proposes a view-dependent light field coding scheme using some image-based rendering techniques prior to coding. The proposed coder first synthesizes an image at a given viewpoint, which is called a representative viewpoint, and then predicts all input images by using the synthesized image as a reference. It can produce a view-dependent scalable bitstream. This means that the quality of synthesized views around the representative viewpoint is kept high even at extremely low bit rates, and the quality of views away from there is improved according to the increase of the bit rate. Our experimental results show that this coding scheme also achieves good coding efficiency for both multi-camera images and integral photography, which are common light field representations.
Yuichi Taguchi, Takeshi Naemura
ICIP1
1999 Japanese large-vocabulary continuous-speech recognition using a newspaper corpus and broadcast news
Katsutoshi Ohtsuki, Tatsuo Matsuoka, Takeshi Mori, Kotaro Yoshida, Yuichi Taguchi, Sadaoki Furui, Katsuhiko Shirai
Speech Commun.5
1997 Toward automatic transcription of Japanese broadcast news
abstract
In this paper, we report on the automatic recognition of Japanese broadcast-news speech. We have been working on largevocabulary continuous speech recognition (LVCSR) for Japanese newspaper speech transcription and have achieved good performance. We have recently applied our LVCSR system to transcribing Japanese broadcast-news speech. We extended the vocabulary from 7k words to 20k words and trained the language models using newspaper texts and broadcast-news manuscripts. These two language models were applied to our evaluation speech sets. The language model trained using broadcast-news manuscripts achieved better results for broadcast-news speech than the language model trained using newspaper texts, which achieved better results for newspaper speech. We achieved a word error rate of 19.7% for anchor-speaker’s speech by using a bigram language model and a trigram language model both trained using broadcast-news manuscripts.
Tatsuo Matsuoka, Yuichi Taguchi, Katsutoshi Ohtsuki, Sadaoki Furui, Katsuhiko Shirai
EUROSPEECH2