EDBT 2026 Demo / reviewers in the wild / expert
Pedro Miraldo
dblp:71/10771
· DBLP profile ↗
36ranked-venue papers
11as first author
13since 2021 · last 2025
0000-0002-8551-2448ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 8 since 2021Systems, architecture and hardware · 12 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SurfR: Surface Reconstruction with Multi-Scale AttentionabstractWe propose a fast and accurate surface reconstruction algorithm for unorganized point clouds using an implicit representation. Recent learning methods are either singleobject representations with small neural models that allow for high surface details but require per-object training or generalized representations that require larger models and generalize to newer shapes but lack details, and inference is slow. We propose a new implicit representation for general 3D shapes that is faster than all the baselines at their optimum resolution, with only a marginal loss in performance compared to the state-of-the-art. We achieve the best accuracy-speed trade-off using three key contributions. Many implicit methods extract features from the point cloud to classify whether a query point is inside or outside the object. First, to speed up the reconstruction, we show that this feature extraction does not need to use the query point at an early stage (lazy query). Second, we use a parallel multi-scale grid representation to develop robust features for different noise levels and input resolutions. Finally, we show that attention across scales can provide improved reconstruction results. The code will be made available. Siddhant Ranade, Gonçalo Dias Pais, Ross Tyler Whitaker, Jacinto C. Nascimento, Pedro Miraldo, Srikumar Ramalingam |
3DV | 5 |
| 2025 | SAC-GNC: Sample Consensus for Adaptive Graduated Non-Convexity
Valter Piedade, Chitturi Sidhartha, Joseé Gaspar, Venu Madhav Govindu, Pedro Miraldo |
ICCV | 5 |
| 2025 | RAPTR: Radar-based 3D Pose Estimation using TransformerabstractRadar-based indoor 3D human pose estimation typically relied on fine-grained 3D keypoint labels, which are costly to obtain especially in complex indoor settings involving clutter, occlusions, or multiple people. In this paper, we propose \textbf{RAPTR} (RAdar Pose esTimation using tRansformer) under weak supervision, using only 3D BBox and 2D keypoint labels which are considerably easier and more scalable to collect.
Our RAPTR is characterized by a two-stage pose decoder architecture with a pseudo-3D deformable attention to enhance (pose/joint) queries with multi-view radar features: a pose decoder estimates initial 3D poses with a 3D template loss designed to utilize the 3D BBox labels and mitigate depth ambiguities; and a joint decoder refines the initial poses with 2D keypoint labels and a 3D gravity loss.
Evaluated on two indoor radar datasets, RAPTR outperforms existing methods, reducing joint position error by $34.3$\% on HIBER and $76.9$\% on MMVR. Our implementation is available at \url{https://github.com/merlresearch/radar-pose-transformer}. Sorachi Kato, Ryoma Yataka, Pu Wang 0004, Pedro Miraldo, Takuya Fujihashi, Petros Boufounos |
NeurIPS | 4 |
| 2024 | Oriented-grid Encoder for 3D Implicit RepresentationsabstractEncoding 3D points is one of the primary steps in learning-based implicit scene representation. Using features that gather information from neighbors with multi-resolution grids has proven to be the best geometric encoder for this task. However, prior techniques do not exploit some characteristics of most objects or scenes, such as surface normals and local smoothness. This paper is the first to exploit those 3D characteristics in 3D geometric encoders explicitly. In contrast to prior work on using multiple levels of details, regular cube grids, and trilinear interpolation, we propose 3D-oriented grids with a novel cylindrical volumetric interpolation for modeling local planar invariance. In addition, we explicitly include a local feature aggregation for feature regularization and smoothing of the cylindrical interpolation features. We evaluate our approach on ABC, Thingi10k, ShapeNet, and Matterport3D, for object and scene representation. Compared to the use of regular grids, our geometric encoder is shown to converge in fewer steps and obtain sharper 3D surfaces. When compared to the prior techniques, our method gets state-of-the-art results. The code is available at https://github.com/merlresearch/oriented-implicit-representation. Arihant Gaur, Gonçalo Dias Pais, Pedro Miraldo |
3DV | 3 |
| 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal SamplingabstractExtensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit their ubiquity: ( i) a significant reduction in reconstruction quality when the computing budget is limited, and (ii) a lack of semantic understanding of the underlying scenes. To address these issues, we introduce Gear-NeRF, which leverages semantic information from powerful image segmentation models. Our approach presents a principled way for learning a spatio-temporal (4D) semantic embedding, based on which we introduce the concept of gears to allow for stratified modeling of dynamic regions of the scene based on the extent of their motion. Such differentiation allows us to adjust the spatio- temporal sampling resolution for each region in proportion to its motion scale, achieving more photo-realistic dynamic novel view synthesis. At the same time, almost for free, our approach enables free-viewpoint tracking of objects of interest - a functionality not yet achieved by existing NeRF-based methods. Empirical studies validate the effectiveness of our method, where we achieve state-of-the-art rendering and tracking performance on multiple challenging datasets. The project page is available at: https://merl.com/research/highlights/gear-nerf Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang, Pedro Miraldo, Suhas Lohit, Moitreya Chatterjee |
CVPR | 4 |
| 2024 | A Probability-Guided Sampler for Neural Implicit Surface Rendering
Gonçalo Dias Pais, Valter Piedade, Moitreya Chatterjee, Marcus Greiff, Pedro Miraldo |
ECCV (37) | 5 |
| 2023 | Bayesian Sensor Fusion for Joint Vehicle Localization and Road Mapping Using Onboard SensorsabstractWe propose a method for joint estimation of a host vehicle state and a map of the road based on global navigation satellite system (GNSS) and camera measurements. We model the road using a spline representation described by a parameter vector having a Gaussian prior representing the uncertainty of the prior map. Both GNSS and camera measurements, such as lane-mark measurements, have noise characteristics that vary in time. To adapt to the changing noise levels and hence improve positioning performance, we combine the sensor information in an interacting multiple-model (IMM) setting to choose the best combination of the estimators with the vehicle state and the parameter vector of the map as the state vector. In a simulation study, we compare vehicle models with varying complexity, and on a real road segment we show that the proposed method can accurately adjust to changing noise conditions and correct for errors in the prior map. Karl Berntorp, Marcus Greiff, Stefano Di Cairano, Pedro Miraldo |
FUSION | 4 |
| 2023 | Robust Frame-to-Frame Camera Rotation Estimation in Crowded ScenesabstractWe present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable speed in this setting. Because the setting is not addressed well by other datasets, we provide a new dataset and benchmark, with high-accuracy, rigorously verified ground truth, on 17 video sequences. Methods developed for wide baseline stereo (e.g., 5-point methods) perform poorly on monocular video. On the other hand, methods used in autonomous driving (e.g., SLAM) leverage specific sensor setups, specific motion models, or local optimization strategies (lagging batch processing) and do not generalize well to handheld video. Finally, for dynamic scenes, commonly used robustification techniques like RANSAC require large numbers of iterations, and become prohibitively slow. We introduce a novel generalization of the Hough transform on SO(3) to efficiently and robustly find the camera rotation most compatible with optical flow. Among comparably fast methods, ours reduces error by almost 50% over the next best, and is more accurate than any method, irrespective of speed. This represents a strong new performance point for crowded scenes, an important setting for computer vision. The code and the dataset are available at https://fabiendelattre.com/robustrotation-estimation. Fabien Delattre, David Dirnfeld, Phat Nguyen, Stephen Scarano, Michael J. Jones 0001, Pedro Miraldo, Erik G. Learned-Miller |
ICCV | 6 |
| 2023 | BANSAC: A dynamic BAyesian Network for adaptive SAmple ConsensusabstractRANSAC-based algorithms are the standard techniques for robust estimation in computer vision. These algorithms are iterative and computationally expensive; they alternate between random sampling of data, computing hypotheses, and running inlier counting. Many authors tried different approaches to improve efficiency. One of the major improvements is having a guided sampling, letting the RANSAC cycle stop sooner. This paper presents a new adaptive sampling process for RANSAC. Previous methods either assume no prior information about the inlier/outlier classification of data points or use some previously computed scores in the sampling. In this paper, we derive a dynamic Bayesian network that updates individual data points’ inlier scores while iterating RANSAC. At each iteration, we apply weighted sampling using the updated scores. Our method works with or without prior data point scorings. In addition, we use the updated inlier/outlier scoring for deriving a new stopping criterion for the RANSAC loop. We test our method in multiple real-world datasets for several applications and obtain state-of-the-art results. Our method outperforms the baselines in accuracy while needing less computational time. The code is available at https://github.com/merlresearch/bansac. Valter Piedade, Pedro Miraldo |
ICCV | 2 |
| 2023 | Fast and Accurate 3D Registration from Line Intersection Constraints
André Mateus 0001, Siddhant Ranade, Srikumar Ramalingam, Pedro Miraldo |
Int. J. Comput. Vis. | 4 |
| 2022 | A Unified Model for Line Projections in Catadioptric Cameras with Rotationally Symmetric MirrorsabstractLines are among the most used computer vision features, in applications such as camera calibration to object detection. Catadioptric cameras with rotationally symmetric mirrors are omnidirectional imaging devices, capturing up to a 360 degrees field of view. These are used in many applications ranging from robotics to panoramic vision. Although known for some specific configurations, the modeling of line projection was never fully solved for general central and non-central catadioptric cameras. We start by taking some general point reflection assumptions and derive a line reflection constraint. This constraint is then used to define a line projection into the image. Next, we compare our model with previous methods, showing that our general approach outputs the same polynomial degrees as previous configuration-specific systems. We run several experiments using synthetic and real-world data, validating our line projection model. Lastly, we show an application of our methods to an absolute camera pose problem. Pedro Miraldo, José Pedro Iglesias |
CVPR | 1 |
| 2022 | An observer cascade for velocity and multiple line estimationabstractPrevious incremental estimation methods consider estimating a single line, requiring as many observers as the number of lines to be mapped. This leads to the need for having at least 4N state variables, with N being the number of lines. This paper presents the first approach for multi-line incremental estimation. Since lines are common in structured environments, we aim to exploit that structure to reduce the state space. The modeling of structured environments proposed in this paper reduces the state space to 3N + 3 and is also less susceptible to singular configurations. An assumption the previous methods make is that the camera velocity is available at all times. However, the velocity is usually retrieved from odometry, which is noisy. With this in mind, we propose coupling the camera with an Inertial Measurement Unit (IMU) and an observer cascade. A first observer retrieves the scale of the linear velocity and a second observer for the lines mapping. The stability of the entire system is analyzed. The cascade is shown to be asymptotically stable and shown to converge in experiments with simulated data. André Mateus 0001, Pedro U. Lima, Pedro Miraldo |
ICRA | 3 |
| 2022 | On Incremental Structure from Motion Using LinesabstractHumans tend to build environments with structure, which consists of mainly planar surfaces. From the intersection of planar surfaces arise straight lines. Lines have more degrees of freedom than points. Thus, line-based structure-from-motion (SfM) provides more information about the environment. In this article, we present solutions for SfM using lines, namely, incremental SfM. These approaches consist of designing state observers for a camera’s dynamical visual system looking at a 3-D line. We start by presenting a model that uses spherical coordinates for representing the line’s moment vector. We show that this parameterization has singularities, and, therefore, we introduce a more suitable model that considers the line’s moment and shortest viewing ray. Concerning the observers, we present two different methodologies. The first uses a memory-less state-of-the-art framework for dynamic visual systems. Since the previous states of the robotic agent are accessible—while performing the 3-D mapping of the environment—the second approach aims at exploiting the use of memory to improve the estimation accuracy and convergence speed. The two models and the two observers are evaluated in simulation and real data, where mobile and manipulator robots are used. André Mateus 0001, Omar Tahri, A. Pedro Aguiar, Pedro U. Lima, Pedro Miraldo |
IEEE Trans. Robotics | 5 |
| 2020 | Mapping of Sparse 3D Data Using Alternating Projection
Siddhant Ranade, Xin Yu 0003, Shantnu Kakkar, Pedro Miraldo, Srikumar Ramalingam |
ACCV (1) | 4 |
| 2020 | Minimal Solvers for 3D Scan Alignment With Pairs of Intersecting LinesabstractWe explore the possibility of using line intersection constraints for 3D scan registration. Typical 3D registration algorithms exploit point and plane correspondences, while line intersection constraints have not been used in the context of 3D scan registration before. Constraints from a match of pairs of intersecting lines in two 3D scans can be seen as two 3D line intersections, a plane correspondence, and a point correspondence. In this paper, we present minimal solvers that combine these different type of constraints: 1) three line intersections and one point match; 2) one line intersection and two point matches; 3) three line intersections and one plane match; 4) one line intersection and two plane matches; and 5) one line intersection, one point match, and one plane match. To use all the available solvers, we present a hybrid RANSAC loop. We propose a non-linear refinement technique using all the inliers obtained from the RANSAC. Vast experiments with simulated data and two real-data data-sets show that the use of these features and the combined solvers improve the accuracy. The code is available. André Mateus 0001, Srikumar Ramalingam, Pedro Miraldo |
CVPR | 3 |
| 2020 | 3DRegNet: A Deep Neural Network for 3D Point RegistrationabstractWe present 3DRegNet, a novel deep learning architecture for the registration of 3D scans. Given a set of 3D point correspondences, we build a deep neural network to address the following two challenges: (i) classification of the point correspondences into inliers/outliers, and (ii) regression of the motion parameters that align the scans into a common reference frame. With regard to regression, we present two alternative approaches: (i) a Deep Neural Network (DNN) registration and (ii) a Procrustes approach using SVD to estimate the transformation. Our correspondence-based approach achieves a higher speedup compared to competing baselines. We further propose the use of a refinement network, which consists of a smaller 3DRegNet as a refinement to improve the accuracy of the registration. Extensive experiments on two challenging datasets demonstrate that we outperform other methods and achieve state-of-the-art results. Gonçalo Dias Pais, Srikumar Ramalingam, Venu Madhav Govindu, Jacinto C. Nascimento, Rama Chellappa, Pedro Miraldo |
CVPR | 6 |
| 2020 | Active Depth Estimation: Stability Analysis and its ApplicationsabstractRecovering the 3D structure of the surrounding environment is an essential task in any vision-controlled Structure-from-Motion (SfM) scheme. This paper focuses on the theoretical properties of the SfM, known as the incremental active depth estimation. The term incremental stands for estimating the 3D structure of the scene over a chronological sequence of image frames. Active means that the camera actuation is such that it improves estimation performance. Starting from a known depth estimation filter, this paper presents the stability analysis of the filter in terms of the control inputs of the camera. By analyzing the convergence of the estimator using the Lyapunov theory, we relax the constraints on the projection of the 3D point in the image plane when compared to previous results. Nonetheless, our method is capable of dealing with the cameras' limited field-of-view constraints. The main results are validated through experiments with simulated data. Rômulo T. Rodrigues, Pedro Miraldo, Dimos V. Dimarogonas, A. Pedro Aguiar |
ICRA | 2 |
| 2020 | Fast Model Predictive Image-Based Visual Servoing for QuadrotorsabstractThis paper studies the problem of Image-Based Visual Servo Control (IBVS) for quadrotors. Although the control of quadrotors has been extensively studied in the last decades, combining the IBVS module with the quadrotor's dynamics is still hard, mainly due to the under-actuation issues related to the quadrotor control as opposed to the 6 DoF control outputs generated by the IBVS modules. We propose an alternative formulation to solve this problem, by particularly using linear Model Predictive Control (MPC), that allows us to relax the UAVs under-actuation issues. Stability guarantees of the proposed scheme are presented. The proposed model is validated with synthetic data and tested in a real UAV's setup. Pedro Roque, Elisa Bin, Pedro Miraldo, Dimos V. Dimarogonas |
IROS | 3 |
| 2019 | Minimal Solvers for Mini-Loop Closures in 3D Multi-Scan Alignmentabstract3D scan registration is a classical, yet a highly useful problem in the context of 3D sensors such as Kinect and Velodyne. While there are several existing methods, the techniques are usually incremental where adjacent scans are registered first to obtain the initial poses, followed by motion averaging and bundle-adjustment refinement. In this paper, we take a different approach and develop minimal solvers for jointly computing the initial poses of cameras in small loops such as 3-, 4-, and 5-cycles. Note that the classical registration of 2 scans can be done using a minimum of 3 point matches to compute 6 degrees of relative motion. On the other hand, to jointly compute the 3D registrations in n-cycles, we take 2 point matches between the first n-1 consecutive pairs (i.e., Scan 1 & Scan 2, ... , and Scan n-1 & Scan n) and 1 or 2 point matches between Scan 1 and Scan n. Overall, we use 5, 7, and 10 point matches for 3-, 4-, and 5-cycles, and recover 12, 18, and 24 degrees of transformation variables, respectively. Using simulations and real-data we show that the 3D registration using mini n-cycles are computationally efficient, and can provide alternate and better initial poses compared to standard pairwise methods. Pedro Miraldo, Surojit Saha, Srikumar Ramalingam |
CVPR | 1 |
| 2019 | POSEAMM: A Unified Framework for Solving Pose Problems using an Alternating Minimization Method
João Campos, João R. Cardoso, Pedro Miraldo |
ICRA | 3 |
| 2019 | OmniDRL: Robust Pedestrian Detection using Deep Reinforcement Learning on Omnidirectional CamerasabstractPedestrian detection is one of the most explored topics in computer vision and robotics. The use of deep learning methods allowed the development of new and highly competitive algorithms. Deep Reinforcement Learning has proved to be within the state-of-the-art in terms of both detection in perspective cameras and robotics applications. However, for detection in omnidirectional cameras, the literature is still scarce, mostly because of their high levels of distortion. This paper presents a novel and efficient technique for robust pedestrian detection in omnidirectional images. The proposed method uses deep Reinforcement Learning that takes advantage of the distortion in the image. By considering the 3D bounding boxes and their distorted projections into the image, our method is able to provide the pedestrian’s position in the world, in contrast to the image positions provided by most state-of-the-art methods for perspective cameras. Our method avoids the need of pre-processing steps to remove the distortion, which is computationally expensive. Beyond the novel solution, our method compares favorably with the state-of-the-art methodologies that do not consider the underlying distortion for the detection task. Gonçalo Dias Pais, Tiago J. Dias, Jacinto C. Nascimento, Pedro Miraldo |
ICRA | 4 |
| 2019 | A Framework for Depth Estimation and Relative Localization of Ground Robots using Computer VisionabstractThe 3D depth estimation and relative pose estimation problem within a decentralized architecture is a challenging problem that arises in missions that require coordination among multiple vision-controlled robots. The depth estimation problem aims at recovering the 3D information of the environment. The relative localization problem consists of estimating the relative pose between two robots, by sensing each other's pose or sharing information about the perceived environment. Most solutions for these problems use a set of discrete data without taking into account the chronological order of the events. This paper builds on recent results on continuous estimation to propose a framework that estimates the depth and relative pose between two non-holonomic vehicles. The basic idea consists in estimating the depth of the points by explicitly considering the dynamics of the camera mounted on a ground robot, and feeding the estimates of 3D points observed by both cameras in a filter that computes the relative pose between the robots. We evaluate the convergence for a set of simulated scenarios and show experimental results validating the proposed framework. Rômulo T. Rodrigues, Pedro Miraldo, Dimos V. Dimarogonas, A. Pedro Aguiar |
IROS | 2 |
| 2018 | Analytical Modeling of Vanishing Points and Curves in Catadioptric CamerasabstractVanishing points and vanishing lines are classical geometrical concepts in perspective cameras that have a lineage dating back to 3 centuries. A vanishing point is a point on the image plane where parallel lines in 3D space appear to converge, whereas a vanishing line passes through 2 or more vanishing points. While such concepts are simple and intuitive in perspective cameras, their counterparts in catadioptric cameras (obtained using mirrors and lenses) are more involved. For example, lines in the 3D space map to higher degree curves in catadioptric cameras. The projection of a set of 3D parallel lines converges on a single point in perspective images, whereas they converge to more than one point in catadioptric cameras. To the best of our knowledge, we are not aware of any systematic development of analytical models for vanishing points and vanishing curves in different types of catadioptric cameras. In this paper, we derive parametric equations for vanishing points and vanishing curves using the calibration parameters, mirror shape coefficients, and direction vectors of parallel lines in 3D space. We show compelling experimental results on vanishing point estimation and absolute pose estimation for a wide range of catadioptric cameras in both simulations and real experiments. Pedro Miraldo, Francisco Girbal Eiras, Srikumar Ramalingam |
CVPR | 1 |
| 2018 | A Minimal Closed-Form Solution for Multi-perspective Pose Estimation using Points and Lines
Pedro Miraldo, Tiago J. Dias, Srikumar Ramalingam |
ECCV (16) | 1 |
| 2018 | Active Structure-from-Motion for 3D Straight LinesBehaviors* This work was partially supported by the Portuguese FCT grants PD/Bd/135015/2017 (through the NETSys Doctoral Program) & SFRH/BPD/111495/2015, and ISRILARSyS Strategic Funding by the FCT project PEst-OE/EEI/LA0009/2013abstractA reliable estimation of 3D parameters is a must for several applications like planning and control, in which is included Image-Based Visual Servoing. This control scheme depends directly on 3D parameters, e.g. depth of points, and/or depth and direction of 3D straight lines. Recently, a framework for Active Structure-from-Motion was proposed, addressing the former feature type. However, straight lines were not addressed. These are 1D objects, which allow for more robust detection, and tracking. In this work, the problem of Active Structure-from-Motion for 3D straight lines is addressed. An explicit representation of these features is presented, and a change of variables is proposed. The latter allows the dynamics of the line to respect the conditions for observability of the framework. A control law is used with the purpose of keeping the control effort reasonable, while achieving a desired convergence rate. The approach is validated first in simulation for a single line, and second using a real robot setup. The latter set of experiments are conducted first for a single line, and then for three lines. André Mateus 0001, Omar Tahri, Pedro Miraldo |
IROS | 3 |
| 2016 | Plücker correction problem: Analysis and improvements in efficiencyabstractA given six dimensional vector represents a 3D straight line in Plücker coordinates if its coordinates satisfy the Klein quadric constraint. In many problems aiming to find the Plücker coordinates of lines, noise in the data and other type of errors contribute for obtaining 6D vectors that do not correspond to lines, because of that constraint. A common procedure to overcome this drawback is to find the Plücker coordinates of the lines that are closest to those vectors. This is known as the Plücker correction problem. In this article we propose a simple, closed-form, and global solution for this problem. When compared with the state-of-the-art method, one can conclude that our algorithm is easier and requires much less operations than previous techniques (it does not require Singular Value Decomposition techniques). João R. Cardoso, Pedro Miraldo, Helder Araújo |
ICPR | 2 |
| 2016 | Towards an omnidirectional catadioptric RGB-D cameraabstractIn this paper we address the 3D reconstruction of points, on a non-central catadioptric system, composed by a mirror, a projector, and a perspective camera. The goal of the paper is to propose a framework to build an omnidirectional depth camera, towards an omnidirectional RGB-D camera system. The main contributions are: an efficient technique to project 3D points from the world to an image of a general non-central catadioptric camera; the definition of the template pattern (for both the projector and camera's images); and the matching between the projection of these features to the world and its respective images. The 3D depth is directly recovered using the template matching approach. In conclusion, we apply some filtering techniques to improve the results. To evaluate the proposed framework, we test the method using synthetic data, under different levels and types of noises, proving that the framework is robust to noise and, thus, can be put into practice. José Pedro Iglesias, Pedro Miraldo, Rodrigo M. M. Ventura |
IROS | 2 |
| 2016 | Efficient object search for mobile robots in dynamic environments: Semantic map as an input for the decision makerabstractIn this work we study the efficient search of objects in domestic environments, using probabilistic logic to represent uncertainty about object location and partially observable Markov decision processes (POMDP) for the decision-making process regarding the movements to be carried out by the robot to improve its belief about the object locations. We propose the use of a semantic map that stores information about the knowledge in the system and updates it, by an inference process, with sensor information received from the object recognition module. However, semantic maps are not capable of actively search for more information in the environment. For that reason a decision-making module, based on a POMDP framework, is integrated in the system. Several experiments were made in a realistic apartment test bed using every day objects and a mobile robot, showing that this hybrid solution makes the search process more efficient. Tiago Veiga, Pedro Miraldo, Rodrigo M. M. Ventura, Pedro U. Lima |
IROS | 2 |
| 2015 | Augmented reality on robot navigation using non-central catadioptric camerasabstractIn this paper we present a framework for the application of augmented reality to a mobile robot, using non-central camera systems. Considering a virtual object in the world with known local 3D coordinates, the goal is to project this object into the image of a non-central catadioptric imaging device. We propose a solution to this problem which allows us to project textured objects to the image in real-time (up to 20 fps): projection of 3D segments to the image; occlusions; and illumination. In addition, since we are considering that the imaging device is on a mobile robot, one needs to take into account the real-time localization of the robot. To the best of our knowledge this is the first time that this problem is addressed (all state-of-the-art methods are derived for central camera systems). To evaluate the proposed framework we test the solution using a mobile robot and a non-central catadioptric camera (using a spherical mirror). Tiago J. Dias, Pedro Miraldo, Nuno Gonçalves 0001, Pedro U. Lima |
IROS | 2 |
| 2015 | Generalized essential matrix: Properties of the singular value decomposition
Pedro Miraldo, Helder Araújo |
Image Vis. Comput. | 1 |
| 2015 | Direct Solution to the Minimal Generalized PoseabstractPose estimation is a relevant problem for imaging systems whose applications range from augmented reality to robotics. In this paper we propose a novel solution for the minimal pose problem, within the framework of generalized camera models and using a planar homography. Within this framework and considering only the geometric elements of the generalized camera models, an imaging system can be modeled by a set of mappings associating image pixels to 3-D straight lines. This mapping is defined in a 3-D world coordinate system. Pose estimation performs the computation of the rigid transformation between the original 3-D world coordinate system and the one in which the camera was calibrated. Using synthetic data, we compare the proposed minimal-based method with the state-of-the-art methods in terms of numerical errors, number of solutions and processing time. From the experiments, we conclude that the proposed method performs better, especially because there is a smaller variation in numerical errors, while results are similar in terms of number of solutions and computation time. To further evaluate the proposed approach we tested our method with real data. One of the relevant contributions of this paper is theoretical. When compared to the state-of-the-art approaches, we propose a completely new parametrization of the problem that can be solved in four simple steps. In addition, our approach does not require any predefined transformation of the dataset, which yields a simpler solution for the problem. Pedro Miraldo, Helder Araújo |
IEEE Trans. Cybern. | 1 |
| 2015 | Pose Estimation for General Cameras Using LinesabstractIn this paper, we address the problem of pose estimation under the framework of generalized camera models. We propose a solution based on the knowledge of the coordinates of 3-D straight lines (expressed in the world coordinate frame) and their corresponding image pixels. Previous approaches used the knowledge of the coordinates of 3-D points (zero dimensional elements) and their corresponding images (zero dimensional elements). In this paper, pixels belonging to the image of 3-D lines are used. There is no need to establish correspondences between pixels and 3-D points. Correspondences are established between 3-D lines and their images. There is no need to identify individual pixels. The use of correspondences between pixels (that belong to the images of the 3-D lines) and 3-D lines facilitates the correspondence problem when compared to the use of world and image points. This is one of the contributions of this paper. The approach is both evaluated and validated using synthetic data and also real images. Pedro Miraldo, Helder Araújo, Nuno Gonçalves 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | A simple and robust solution to the minimal general pose estimationabstractIn this article we address the problem of minimal pose under the framework of the generalized camera models. Previous approaches were based on geometric properties, such as the preservation of distance between points. In this paper we propose a novel formulation of the problem using an algebraic-based approach. We represent the pose by a 3×3 matrix. Using both the algebraic relationship between three incident 3D points and straight lines and the underlying constraints of the pose matrix, pose can be computed. In terms of experimental results, the main contribution of the proposed method is the robustness to critical configurations. In addition, a full comparison and analysis between state-of-the-art methods is made (so far, there is no published comparison between state-of-the-art methods published). Pedro Miraldo, Helder Araújo |
ICRA | 1 |
| 2014 | Planar pose estimation for general cameras using known 3D linesabstractIn this article, we address the pose estimation for planar motion under the framework of generalized camera models. We assume the knowledge of the coordinates of 3D straight lines in the world coordinate system. Pose is estimated using the images of the 3D lines. This approach does not require the determination of correspondences between pixels and 3D world points. Instead, and for each pixel, it is only required that we determine to which 3D line it is associated with. Instead of identifying individual pixels, it is only necessary to establish correspondences between the pixels that belong to the images of the 3D lines, and the 3D lines. Moreover and using the assumption that the motion is planar, this paper presents a novel method for the computation of the pose using general imaging devices and assuming the knowledge of the coordinates of 3D straight lines. The approach is evaluated and validated using both synthetic data and real images. The experiments are performed using a mobile robot equipped with a non-central camera. Pedro Miraldo, Helder Araújo |
IROS | 1 |
| 2013 | Calibration of Smooth Camera ModelsabstractGeneric imaging models can be used to represent any camera. Current generic models are discrete and define a mapping between each pixel in the image and a straight line in 3D space. This paper presents a modification of the generic camera model that allows the simplification of the calibration procedure. The only requirement is that the coordinates of the 3D projecting lines are related by functions that vary smoothly across space. Such a model is obtained by modifying the general imaging model using radial basis functions (RBFs) to interpolate image coordinates and 3D lines, thereby allowing both an increase in resolution (due to their continuous nature) and a more compact representation. Using this variation of the general imaging model, we also develop a calibration procedure. This procedure only requires that a 3D point be matched to each pixel. In addition, not all the pixels need to be calibrated. As a result, the complexity of the procedure is significantly decreased. Normalization is applied to the coordinates of both image and 3D points, which increases the accuracy of the calibration. Results with both synthetic and real datasets show that the model and calibration procedure are easily applicable and provide accurate calibration results. Pedro Miraldo, Helder Araújo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Point-based calibration using a parametric representation of the general imaging modelabstractGeneric imaging models can be used to represent any camera. These models are specially suited for non-central cameras for which closed-form models do not exist. Current models are discrete and define a mapping between each pixel in the image and a straight line in 3D space. Due to difficulties in the calibration procedure and model complexity these methods have not been used in practice. The focus of our work was to relax these drawbacks. In this paper we modify the general imaging model using radial basis functions to interpolate image coordinates and 3D lines allowing both an increase in resolution (due to their continuous nature) and a more compact representation. Using this new variation of the general imaging model we also develop a new linear calibration procedure. In this process it is only required to match one 3D point to each image pixel. Also it is not required the calibration of every image pixel. As a result the complexity of the procedure is significantly decreased. Pedro Miraldo, Helder Araújo, Joao Queiro |
ICCV | 1 |