VLDB 2026 Research / reviewers in the wild / expert
Roberto Manduchi
dblp:m/RobertoManduchi
· DBLP profile ↗
78ranked-venue papers
18as first author
12since 2021 · last 2026
0000-0003-2640-302XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 28 · 6 first-authorHuman-computer interaction and ubiquitous computing · 14 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gaze-Based Automatic Scrolling for Readers Using Screen Magnification
Glenn Grant-Richards, Roberto Manduchi, Suzana Chung |
ETRA | 2 |
| 2026 | PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation ModelabstractIndoor localization in GPS-denied environments is crucial for applications like emergency response and assistive navigation. Vision-based methods such as PALMS enable infrastructure-free localization using only a floor plan and a stationary scan, but are limited by the short range of smartphone LiDAR and ambiguity in indoor layouts. We propose PALMS+, a modular, image-based system that addresses these challenges by reconstructing scale-aligned 3D point clouds from posed RGB images using a foundation monocular depth estimation model (Depth Pro), followed by geometric layout matching via convolution with the floor plan. PALMS+ outputs a posterior over the location and orientation, usable for direct or sequential localization. Evaluated on the Structured3D and a custom campus dataset consisting of 80 observations across four large campus buildings, PALMS+ outperforms PALMS and F3Loc in stationary localization accuracy—without requiring any training. Furthermore, when integrated with a particle filter for sequential localization on 33 real-world trajectories, PALMS+ achieved lower localization errors compared to other methods, demonstrating robustness for camera-free tracking and its potential for infrastructure-free applications. Code and data are available at https://github.com/Head-inthe-Cloud/PALMS-Plane-based-Accessible-Indoor-Localization-Using-Mobile-Smartphones. Yunqian Cheng, Benjamin Princen, Roberto Manduchi |
WACV | 3 |
| 2024 | Reading with Screen Magnification: Eye Movement Analysis Using Compensated Gaze TracksabstractEye movements while reading with screen magnification (which requires manual scrolling to center the magnified portion of the screen within the viewport) pose interpretation challenges. Standard representations in terms of alternating fixations and saccades don't apply to this case. This is because, during scrolling, eyes often track a moving text element, generating a movement akin to smooth pursuit. We propose a new representation that uses information from the mouse (which the reader uses to move the center of magnification) to undo the effect of magnification and scrolling. After this "compensation" operation, gaze tracks can again be described as alternating fixations and saccades. We present an analysis of gaze tracks obtained by applying this transformation on an existing dataset, recorded from low vision readers using two modalities of screen magnification. This analysis highlights similarities and differences in terms of dynamic properties of compensated gaze tracks vis-à-vis gaze during regular reading. Seongsil Heo, Roberto Manduchi, Susana T. L. Chung |
ETRA | 2 |
| 2024 | A Functional Usability Analysis of Appearance-Based Gaze Tracking for AccessibilityabstractAppearance-based gaze tracking algorithms, which compute gaze direction from user face images, are an attractive alternative to infrared-based external devices. Their accuracy has greatly benefited by using powerful machine-learning techniques. The performance of appearance-based algorithms is normally evaluated on standard benchmarks typically involving users fixating at points on the screen. However, these metrics do not easily translate into functional usability characteristics. In this work, we evaluate a state-of-the-art algorithm, FAZE, in a number of tasks of interest to the human-computer interaction community. Specifically, we study how gaze measured by FAZE could be used for dwell-based selection and reading progression (line identification and progression along a line) - key functionalities for users facing motor and visual impairments. We compared the gaze data quality from 7 participants using FAZE against that from an infrared tracker (Tobii Pro Spark). Our analysis highlights the usability of appearance-based gaze tracking for such applications. Youn-Soo Park, Roberto Manduchi |
ETRA | 2 |
| 2024 | Step Length Estimation for Blind Walkers
Fatemeh Elyasi, Roberto Manduchi |
ICCHP (1) | 2 |
| 2024 | PALMS: Plane-based Accessible Indoor Localization Using Mobile SmartphonesabstractIn this paper, we present PALMS, an innovative indoor global localization and relocalization system for mobile smartphones that utilizes publicly available floor plans. Unlike most vision-based methods that require constant visual input, our system adopts a dynamic form of localization that considers a single instantaneous observation and odometry data. The core contribution of this work is the introduction of a particle filter initialization method that leverages the Certainly Empty Space (CES) constraint along with principal orientation matching. This approach creates a spatial probability distribution of the device's location, significantly improving localization accuracy and reducing particle filter convergence time. Our experimental evaluations demonstrate that PALMS outperforms traditional methods with uniformly initialized particle filters, providing a more efficient and accessible approach to indoor wayfinding. By eliminating the need for prior environmental fingerprinting, PALMS provides a scalable and practical approach to indoor navigation. Yunqian Cheng, Roberto Manduchi |
IPIN | 2 |
| 2024 | Robust Indoor Pedestrian Backtracking Using Magnetic Signatures and Inertial DataabstractNavigating unfamiliar environments can be challenging for visually impaired individuals due to difficulties in recognizing distant landmarks or visual cues. This work focuses on a particular form of wayfinding, specifically backtracking a previously taken path, which can be useful for blind pedestrians. We propose a hands-free indoor navigation solution using a smartphone without relying on pre-existing maps or external infrastructure. Our hybrid matching method integrates machine learning to enhance positioning accuracy, addressing real-life challenges such as odometry errors or deviations from the correct path. Testing with datasets from visually impaired individuals demonstrates the potential of our approach in providing reliable backtracking assistance. Chia Hsuan Tsai, Roberto Manduchi |
IPIN | 2 |
| 2023 | Experiments with RouteNav, A Wayfinding App for Blind Travelers in a Transit HubabstractRouteNav is an iOS app designed to support wayfinding for blind travelers in an indoor/outdoor transit hub. It doesn't rely on external infrastructure (such as BLE beacons); instead, localization is obtained by fusing spatial information from inertial dead reckoning and GPS (when available) via particle filtering. Routes are expressed as sequences of "tiles", where each tile may contain relevant points of interest. Redundant modalities are used to guide users to switching goalposts within tiles. In this paper, we describe the different components of RouteNav, and report on a user study with seven blind participants, who traversed three challenging routes in a transit hub while receiving input from the app. Peng Ren 0003, Jonathan Lam, Roberto Manduchi, Fatemeh Mirzaei |
ASSETS | 3 |
| 2023 | Screen Magnification for Readers with Low Vision: A Study on Usability and PerformanceabstractWe present a study with 20 participants with low vision who operated two types of screen magnification (lens and full) on a laptop computer to read two types of document (text and web page). Our purposes were to comparatively assess the two magnification modalities, and to obtain some insight into how people with low vision use the mouse to control the center of magnification. These observations may inform the design of systems for the automatic control of the center of magnification. Our results show that there were no significant differences in reading performances or in subjective preferences between the two magnification modes. However, when using the lens mode, our participants adopted more consistent and uniform mouse motion patterns, while longer and more frequent pauses and shorter overall path lengths were measured using the full mode. Analysis of the distribution of gaze points (as measured by a gaze tracker) using the full mode shows that, when reading a text document, most participants preferred to move the area of interest to a specific region of the screen. Meini Tang, Roberto Manduchi, Susana T. L. Chung, Raquel Prado |
ASSETS | 2 |
| 2023 | Step Length Is a More Reliable Measurement Than Walking Speed for Pedestrian Dead-Reckoning*abstractPedestrian dead reckoning (PDR) relies on the estimation of the length of each step taken by the walker in a path from inertial data (e.g. as recorded by a smartphone). Existing algorithms either estimate step lengths directly, or predict walking speed, which can then be integrated over a step period to obtain step length. We present an analysis, using a common architecture formed by an LSTM followed by four fully connected layers, of the quality of reconstruction when predicting step length vs. walking speed. Our experiments, conducted on a data set collected by twelve participants, strongly suggest that step length can be predicted more reliably than average walking speed over each step. Fatemeh Elyasi, Roberto Manduchi |
IPIN | 2 |
| 2022 | Tracker/Camera Calibration for Accurate Automatic Gaze Annotation of Images and VideosabstractModern appearance-based gaze tracking algorithms require vast amounts of training data, with images of a viewer annotated with "ground truth" gaze direction. The standard approach to obtain gaze annotations is to ask subjects to fixate at specific known locations, then use a head model to determine the location of "origin of gaze". We propose using an IR gaze tracker to generate gaze annotations in natural settings that do not require the fixation of target points. This requires prior geometric calibration of the IR gaze tracker with the camera, such that the data produced by the IR tracker can be expressed in the camera's reference frame. This contribution introduces a simple tracker/camera calibration procedure based on the PnP algorithm and demonstrates its use to obtain a full characterization of gaze direction that can be used for ground truth annotation. Swati Jindal, Harsimran Kaur, Roberto Manduchi |
ETRA | 3 |
| 2021 | Subject Guided Eye Image Synthesis with Application to Gaze RedirectionabstractWe propose a method for synthesizing eye images from segmentation masks with a desired style. The style encompasses attributes such as skin color, texture, iris color, and personal identity. Our approach generates an eye image that is consistent with a given segmentation mask and has the attributes of the input style image. We apply our method to data augmentation as well as to gaze redirection. The previous techniques of synthesizing real eye images from synthetic eye images for data augmentation lacked control over the generated attributes. We demonstrate the effectiveness of the proposed method in synthesizing realistic eye images with given characteristics corresponding to the synthetic labels for data augmentation, which is further useful for various tasks such as gaze estimation, eye image segmentation, pupil detection, etc. We also show how our approach can be applied to gaze redirection using only synthetic gaze labels, improving the previous state of the art results. The main contributions of our paper are i) a novel approach for Style-Based eye image generation from segmentation mask; ii) the use of this approach for gaze-redirection without the need for gaze annotated real eye images. Harsimran Kaur, Roberto Manduchi |
WACV | 2 |
| 2020 | A Multi-scale Embossed Map Authoring Tool for Indoor Environments
Viet Trinh, Roberto Manduchi |
ICCHP (1) | 2 |
| 2020 | EyeGAN: Gaze-Preserving, Mask-Mediated Eye Image SynthesisabstractAutomatic synthesis of realistic eye images with prescribed gaze direction is important for multiple application domains. We introduce EyeGAN, an algorithm to generate eye images in the style of a desired target domain, that inherit annotations available in images from a source domain. EyeGAN takes in input ternary masks, which are used as domain-independent proxies for gaze direction. We evaluate EyeGAN against competing eye image synthesis algorithms by measuring a specific gaze consistency index. In addition, we present results from multiple experiments (involving eye region segmentation, pupil localization, and gaze direction estimation) showing that the use of EyeGANgenerated images with inherited annotations for network training leads to superior performances compared to other domain transfer algorithms. Harsimran Kaur, Roberto Manduchi |
WACV | 2 |
| 2019 | Scene text access: a comparison of mobile OCR modalities for blind usersabstractWe present a study with seven blind participants using three different mobile OCR apps to find text posted in various indoor environments. The first app considered was Microsoft SeeingAI in its Short Text mode, which reads any text in sight with a minimalistic interface. The second app was Spot+OCR, a custom application that separates the task of text detection from OCR proper. Upon detection of text in the image, Spot+OCR generates a short vibration; as soon as the user stabilizes the phone, a high-resolution snapshot is taken and OCR-processed. The third app, Guided OCR, was designed to guide the user in taking several pictures in a 360° span at the maximum resolution available by the camera, with minimum overlap between pictures. Quantitative results (in terms of true positive ratios and traversal speed) were recorded. Along with the qualitative observation and outcomes from an exit survey, these results allow us to identify and assess the different strategies used by our participants, as well as the challenges of operating these systems without sight. Leo Neat, Peng Ren 0003, Siyang Qin, Roberto Manduchi |
IUI | 4 |
| 2018 | Multi-planar Monocular Reconstruction of Manhattan Indoor ScenesabstractWe present a novel algorithm for geometry and camera pose reconstruction from image sequences that is specialized for indoor Manhattan scenes. Unlike general-purpose SfM/SLAM, our system represents geometric primitives in terms of canonically oriented planes. The algorithm starts by computing multi-planar segmentation and motion estimation from image pairs using constrained homographies. It then proceeds to recover the relative scale at each frame and to determine chains of match clusters, where each cluster is associated with a plane in the scene. Motion and scene geometry (expressed in terms of planar models) are then optimized using a novel formulation of Bundle Adjustment. Compared with other state-of-the-art SfM/SLAM algorithms, our technique is shown to produce superior and realistic surface reconstruction for a monocular indoor scene. Seongdo Kim 0001, Roberto Manduchi, Siyang Qin |
3DV | 2 |
| 2018 | Automatic Semantic Content Removal by Learning to Neglect
Siyang Qin, Jiahui Wei, Roberto Manduchi |
BMVC | 3 |
| 2018 | Easy Return: An App for Indoor Backtracking AssistanceabstractWe present a system that, implemented as an iPhone app controllable from an Apple Watch, can help a blind person backtrack a route taken in a building. This system requires no maps of the building or environment modifications. While traversing a path from a starting location to a destination, the system builds and records a path representation in terms of a sequence of turns and of step counts between turns. If the user wants to backtrack the same path, the system can provide assistance by tracking the user's location in the recorded path, and producing directional information in speech form about the next turns and step counts to follow. The system was tested with six blind participants in a controlled indoor experiment. German H. Flores, Roberto Manduchi |
CHI | 2 |
| 2018 | Public Transit Accessibility: Blind Passengers Speak Out
Fatemeh Mirzaei, Roberto Manduchi, Sri Hastuti Kurniawan |
ICCHP (2) | 2 |
| 2018 | Robust and Accurate Text Stroke SegmentationabstractWe propose a new technique for the accurate segmentation of text strokes from an image. The algorithm takes in a cropped image containing a word. It first performs a coarse segmentation using a Fully Convolutional Network (FCN). While not accurate, this initial segmentation can usually identify most of the text stroke content even in difficult situations, with uneven lighting and non-uniform background. The segmentation is then refined using a fully connected Conditional Random Field (CRF) with a novel kernel definition that includes stroke width information. In order to train the network, we created a new synthetic data set with 100K text images. Tested against standard benchmarks with pixellevel annotation (ICDAR 2003, ICDAR 2011, and SVT) our algorithm outperforms the state of the art by a noticeable margin. Siyang Qin, Peng Ren 0003, Seongdo Kim 0001, Roberto Manduchi |
WACV | 4 |
| 2017 | Cascaded Segmentation-Detection Networks for Word-Level Text SpottingabstractWe introduce an algorithm for word-level text spotting that is able to accurately and reliably determine the bounding regions of individual words of text "in the wild". Our system is formed by the cascade of two convolutional neural networks. The first network is fully convolutional and is in charge of detecting areas containing text. This results in a very reliable but possibly inaccurate segmentation of the input image. The second network (inspired by the popular YOLO architecture) analyzes each segment produced in the first stage, and predicts oriented rectangular regions containing individual words. No post-processing (e.g. text line grouping) is necessary. With execution time of 450 ms for a 1000 × 560 image on a Titan X GPU, our system achieves good performance on the ICDAR 2013, 2015 benchmarks [2], [1]. Siyang Qin, Roberto Manduchi |
ICDAR | 2 |
| 2017 | Automatic skin and hair masking using fully convolutional networksabstractSelfies have become commonplace. More and more people take pictures of themselves, and enjoy enhancing these pictures using a variety of image processing techniques. One specific functionality of interest is automatic skin and hair segmentation, as this allows for processing one's skin and hair separately. Traditional approaches require user input in the form of fully specified trimaps, or at least of “scribbles” indicating foreground and background areas, with high-quality masks then generated via matting. Manual input, however, can be difficult or tedious, especially on a smartphone's small screen. In this paper, we propose the use of fully convolutional networks (FCN) and fully-connected CRF to perform pixel-level semantic segmentation into skin, hair and background. The trimap thus generated is given as input to a standard matting algorithm, resulting in accurate skin and hair alpha masks. Our method achieves state-of-the-art performance on the LFW Parts dataset [1]. The effectiveness of our method is also demonstrated with a specific application case. Siyang Qin, Seongdo Kim 0001, Roberto Manduchi |
ICME | 3 |
| 2017 | Velocity and shape from tightly-coupled LiDAR and cameraabstractIn this paper, we propose a multi-object tracking and reconstruction approach through measurement-level fusion of LiDAR and camera. The proposed method, regardless of object class, estimates 3D motion and structure for all rigid obstacles. Using an intermediate surface representation, measurements from both sensors are processed within a joint framework. We combine optical flow, surface reconstruction, and point-to-surface terms in a tightly-coupled non-linear energy function, which is minimized using Iterative Reweighted Least Squares (IRLS). We demonstrate the performance of our model on different datasets (KITTI with Velodyne HDL-64E and our collected data with 4-layer ScaLa Ibeo), and show an improvement in velocity error and crispness over state-of-the-art trackers. Mohammad Hossein Daraei, Anh Vu, Roberto Manduchi |
Intelligent Vehicles Symposium | 3 |
| 2017 | Multi-planar Fitting in an Indoor ManhattanWorldabstractWe present an algorithm that finds planar structures in a Manhattan world from two pictures taken from different viewpoints with unknown baseline. The Manhattan world assumption constrains the homographies induced by the visible planes on the image pair, thus enabling robust reconstruction. We extend the T-linkage algorithm for multistructure discovery to account for constrained homographies, and introduce algorithms for sample point selection and orientation-preserving cluster merging. Results are presented on three indoor data set, showing the benefit of the proposed constraints and algorithms. Seongdo Kim 0001, Roberto Manduchi |
WACV | 2 |
| 2017 | Indoor Manhattan spatial layout recovery from monocular videos via line matching
Chelhwon Kim, Roberto Manduchi |
Comput. Vis. Image Underst. | 2 |
| 2016 | WeAllWalk: An Annotated Data Set of Inertial Sensor Time Series from Blind WalkersabstractWe introduce WeAllWalk, a data set of inertial sensor time series collected from blind walkers using a long cane or a guide dog. Blind participants walked through fairly long and complex indoor routes that included obstacles to be avoided and doors to be opened. Inertial data was recorded by two iPhone 6s carried by our participants in their pockets and carefully annotated. Ground truth heel strike times were measured by two small inertial sensor units clipped to the participants' shoes. We also show comparative examples of application of step counting and turn detection algorithms to selected data from WeAllWalk. German H. Flores, Roberto Manduchi |
ASSETS | 2 |
| 2016 | Experiments with a Public Transit Assistant for Blind Passengers
German H. Flores, Roberto Manduchi |
ICCHP (2) | 2 |
| 2016 | A fast and robust text spotterabstractWe introduce an algorithm for text detection and localization ("spotting") that is computationally efficient and produces state-of-the-art results. Our system uses multi-channel MSERs to detect a large number of promising regions, then subsamples these regions using a clustering approach. Representatives of region clusters are binarized and then passed on to a deep network. A final line grouping stage forms word-level segments. On the ICDAR 2011 and 2015 benchmarks, our algorithm obtains an F-score of 82% and 83%, respectively, at a computational cost of 1.2 seconds per frame. We also introduce a version that is three times as fast, with only a slight reduction in performance. Siyang Qin, Roberto Manduchi |
WACV | 2 |
| 2015 | Zebra Crossing Spotter: Automatic Population of Spatial Databases for Increased Safety of Blind Travelersabstractin urban settings. Knowing the location of crosswalks is critical for a blind person planning a trip that includes street crossing. By augmenting existing spatial databases (such as Google Maps or OpenStreetMap) with this information, a blind traveler may make more informed routing decisions, resulting in greater safety during independent travel. Our algorithm first searches for zebra crosswalks in satellite images; all candidates thus found are validated against spatially registered Google Street View images. This cascaded approach enables fast and reliable discovery and localization of zebra crosswalks in large image datasets. While fully automatic, our algorithm could also be complemented by a final crowdsourcing validation stage for increased accuracy. Dragan Ahmetovic, Roberto Manduchi, James M. Coughlan, Sergio Mascetti |
ASSETS | 2 |
| 2015 | Towards Mobile OCR: How to Take a Good Picture of a Document Without SightabstractThe advent of mobile OCR (optical character recognition) applications on regular smartphones holds great promise for enabling blind people to access printed information. Unfortunately, these systems suffer from a problem: in order for OCR output to be meaningful, a well-framed image of the document needs to be taken, something that is difficult to do without sight. This contribution presents an experimental investigation of how blind people position and orient a camera phone while acquiring document images. We developed experimental software to investigate if verbal guidance aids in the acquisition of OCR-readable images without sight. We report on our participant's feedback and performance before and after assistance from our software. Michael Patrick Cutter, Roberto Manduchi |
DocEng | 2 |
| 2014 | Probabilistic Phase Unwrapping for Single-Frequency Time-of-Flight Range CamerasabstractThis paper proposes a solution to the 2-D phase unwrapping problem, inherent to time-of-flight range sensing technology due to the cyclic nature of phase. Our method uses a single frequency capture period to improve frame rate and decrease the presence of motion artifacts encountered in multiple frequency solutions. We present a probabilistic framework that considers intensity image in addition to the phase image. The phase unwrapping problem is cast in terms of global optimization of a carefully chosen objective function. Comparative experimental results confirm the effectiveness of the proposed approach. Ryan Crabb, Roberto Manduchi |
3DV | 2 |
| 2014 | Planar Structures from Line Correspondences in a Manhattan World
Chelhwon Kim, Roberto Manduchi |
ACCV (1) | 2 |
| 2014 | The last meter: blind visual guidance to a targetabstractSmartphone apps can use object recognition software to provide information to blind or low vision users about objects in the visual environment. A crucial challenge for these users is aiming the camera properly to take a well-framed picture of the desired target object. We investigate the effects of two fundamental constraints of object recognition - frame rate and camera field of view - on a blind person's ability to use an object recognition smartphone app. The app was used by 18 blind participants to find visual targets beyond arm's reach and approach them to within 30 cm. While we expected that a faster frame rate or wider camera field of view should always improve search performance, our experimental results show that in many cases increasing the field of view does not help, and may even hurt, performance. These results have important implications for the design of object recognition systems for blind users. Roberto Manduchi, James M. Coughlan |
CHI | 1 |
| 2014 | Transit Information Access for Persons with Visual or Cognitive Impairments
German H. Flores, Benjamin Cizdziel, Roberto Manduchi, Katia Obraczka, Julie Do, Tyler Esser, Sri Hastuti Kurniawan |
ICCHP (1) | 3 |
| 2014 | A novel approach for color barcode decoding using smart phonesabstractThe use of colors increases the information storage capacities in barcodes. Increasing the number of colors to encode information makes the decoding a challenging task due to the dependency of the surface color on the illuminant spectrum, viewing parameters, printing device and material, color fading in addition to other nuisance parameters. A popular solution is the use of a color palette of reference colors printed with barcode. The decoding becomes more challenging if a mobile phone as a decoding device is used due to the capture of images from different distances as well as angles. In addition, the barcode images are often blurry because of incorrect focus or camera shake. We present an iterative decoding algorithm that decodes the colors of all barcode patches across the barcode by minimizing the overall observation error. Our method is able to decode colors in presence of blur using a small number of colors yet ensuring high information density. Homayoun Bagherinia, Roberto Manduchi |
ICIP | 2 |
| 2013 | Real Time Camera Phone Guidance for Compliant Document Image Acquisition without SightabstractHere we present an evaluation of an ideal document acquisition guidance system.Guidance is provided to help someone take a picture of a document capable of Optical Character Recognition (OCR).Our method infers the pose of the camera by detecting a pattern of fiduciary markers on a printed page.The guidance system offers a corrective trajectory based on the current pose, by optimizing the requirements for complete OCR.We evaluate the effectiveness of our software by measuring the quality of the image captured when we vary the experimental setting.After completing a user study with eight participants, we found that our guidance system is effective at helping the user position the phone in such a way that a compliant image is captured.This is based on an evaluation of a one way analysis of variance comparing the percentage of successful trials in each experimental setting.Negative Helmert Contrast is applied in order to tolerate only one ordering of experimental settings: no guidance (control), confirmation, and full guidance with confirmation. Michael Patrick Cutter, Roberto Manduchi |
ICDAR | 2 |
| 2012 | Mobile Vision as Assistive Technology for the Blind: An Experimental Study
Roberto Manduchi |
ICCHP (2) | 1 |
| 2012 | Metering for Exposure StacksabstractAbstract When creating a High‐Dynamic‐Range (HDR) image from a sequence of differently exposed Low‐Dynamic‐Range (LDR) images, the set of LDR images is usually generated by sampling the space of exposure times with a geometric progression and without explicitly accounting for the distribution of irradiance values of the scene. We argue that this choice can produce sub‐optimal results both in terms of the number of acquired pictures and the quality of the resulting HDR image. This paper presents a method to estimate the full irradiance histogram of a scene, and a strategy to select the set of exposures that need to be acquired. Our selection usually requires a smaller or equal set of LDRs, yet produces higher quality HDR images. Orazio Gallo, Marius Tico, Roberto Manduchi, Natasha Gelfand, Kari Pulli |
Comput. Graph. Forum | 3 |
| 2011 | Fast image motion segmentation for surveillance applications
Xiaoye Lu, Roberto Manduchi |
Image Vis. Comput. | 2 |
| 2011 | Reading 1D Barcodes with Mobile Phones Using Deformable TemplatesabstractCamera cellphones have become ubiquitous, thus opening a plethora of opportunities for mobile vision applications. For instance, they can enable users to access reviews or price comparisons for a product from a picture of its barcode while still in the store. Barcode reading needs to be robust to challenging conditions such as blur, noise, low resolution, or low-quality camera lenses, all of which are extremely common. Surprisingly, even state-of-the-art barcode reading algorithms fail when some of these factors come into play. One reason resides in the early commitment strategy that virtually all existing algorithms adopt: The image is first binarized and then only the binary data are processed. We propose a new approach to barcode decoding that bypasses binarization. Our technique relies on deformable templates and exploits all of the gray-level information of each pixel. Due to our parameterization of these templates, we can efficiently perform maximum likelihood estimation independently on each digit and enforce spatial coherence in a subsequent step. We show by way of experiments on challenging UPC-A barcode images from five different databases that our approach outperforms competing algorithms. Implemented on a Nokia N95 phone, our algorithm can localize and decode a barcode on a VGA image (640 × 480, JPEG compressed) in an average time of 400-500 ms. Orazio Gallo, Roberto Manduchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | CC-RANSAC: Fitting planes in the presence of multiple surfaces in range data
Orazio Gallo, Roberto Manduchi, Abbas Rafii |
Pattern Recognit. Lett. | 2 |
| 2010 | A general education course on universal access, disability, technology and societyabstractThis paper reports on a General Education course called Universal Access: Disability, Technology and Society that enables students from all majors to learn more about disability and the issues that surround it, as well as how Assistive Technology facilitates effective participation of those with disabilities in society. Guest lectures, meant to give the students different perspectives on disability, are integral part of the course. Guest lecturers include experts in disability studies, professionals working with people with disabilities, and persons with disability. To gain practical knowledge, the students carry out group projects or volunteering activities that involves people with disabilities. Since its first introduction in 2006, the course had always filled to capacity. A survey with 75 students conducted in Winter 2010 revealed that students felt that their knowledge about universal access and disabilities had improved significantly, and that they had become aware of accessibility in everyday life. Sri Hastuti Kurniawan, Sonia M. Arteaga, Roberto Manduchi |
ASSETS | 3 |
| 2010 | Blind guidance using mobile computer vision: a usability studyabstractWe present a study focusing on the usability of a wayfinding and localization system for persons with visual impairment. This system uses special color markers, placed at key locations in the environment, that can be detected by a regular camera phone. Three blind participants tested the system in various indoor locations and under different system settings. Quantitative performance results are reported. Roberto Manduchi, Sri Hastuti Kurniawan, Homayoun Bagherinia |
ASSETS | 1 |
| 2010 | An Ultra-Low-Power Contrast-Based Integrated Camera Node and its Application as a People CounterabstractWe describe the implementation in a self-standing system of a novel contrast-based binary CMOS imaging sensor.This sensor is characterized by very low power consumption and wide dynamic range, which makes it attractive for wireless camera network applications. In our implementation,the sensor is interfaced with a Flash-based FPGA processor,which handles data readout and image processing.This self-standing camera node is configured as a system for counting persons walking through a corridor. Simple features are extracted from each image in a video stream at 30 fps. A classifier is designed based on the temporal evolution of these features, which is modeled as a Markov chain. The video stream is then segmented into intervals corresponding to individual persons crossing through the field of view. Experimental results are shown in cross-validated tests over real sequences acquired by the camera. Leonardo Gasparini, Roberto Manduchi, Massimo Gottardi |
AVSS | 2 |
| 2010 | One-Shot Optimal Exposure Control
David Ilstrup, Roberto Manduchi |
ECCV (1) | 2 |
| 2009 | Reading challenging barcodes with camerasabstractCurrent camera-based barcode readers do not work well when the image has low resolution, is out of focus, or is motion-blurred. One main reason is that virtually all existing algorithms perform some sort of binarization, either by gray scale thresholding or by finding the bar edges. We propose a new approach to barcode reading that never needs to binarize the image. Instead, we use deformable barcode digit models in a maximum likelihood setting. We show that the particular nature of these models enables efficient integration over the space of deformations. Global optimization over all digits is then performed using dynamic programming. Experiments with challenging UPC-A barcode images show substantial improvement over other state-of-the-art algorithms. Orazio Gallo, Roberto Manduchi |
WACV | 2 |
| 2008 | Cellphone Accessible Information Via Bluetooth Beaconing for the Visually Impaired
S. Bohonos, A. Malik, C. Thai, Roberto Manduchi |
ICCHP | 5 |
| 2008 | Portable and Mobile Systems in Assistive Technology
Roberto Manduchi, James M. Coughlan |
ICCHP | 1 |
| 2008 | Search Strategies of Visually Impaired Persons Using a Camera Phone Wayfinding System
Roberto Manduchi, James M. Coughlan, Volodymyr Ivanchenko |
ICCHP | 1 |
| 2007 | Accessible spaces: navigating through a marked environment with a camera phoneabstractWe demonstrate a system designed to assist a visually impaired individual while moving in an unfamiliar environment. Small and economical color markers are placed in key locations, possibly in the vicinity of other signs (bar codes or text). The user can detect these markers by means of a cell phone equipped with a camera. Our demonstration highlights a number of novel features, including: improved acoustic interfaces; estimation of the distance to the marker, which is communicated to the user via text-to-speech (TTS); increased robustness via rotation invariance, which makes the system easier to use for users with reduced dexterity. Kee-Yip Chan, Roberto Manduchi, James M. Coughlan |
ASSETS | 2 |
| 2007 | On the Bayes fusion of visual features
Xiaojin Shi, Roberto Manduchi |
Image Vis. Comput. | 2 |
| 2006 | Computer Vision-Based Terrain Sensors for Blind Wheelchair Users
James M. Coughlan, Roberto Manduchi, Huiying Shen |
ICCHP | 2 |
| 2006 | Learning Outdoor Color ClassificationabstractWe present an algorithm for color classification with explicit illuminant estimation and compensation. A Gaussian classifier is trained with color samples from just one training image. Then, using a simple diagonal illumination model, the illuminants in a new scene that contains some of the surface classes seen in the training image are estimated in a maximum likelihood framework using the Expectation Maximization algorithm. We also show how to impose priors on the illuminants, effectively computing a maximum a posteriori estimation. Experimental results are provided to demonstrate the performance of our classification algorithm in the case of outdoor images. Roberto Manduchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Rotational Invariant Operators Based on Steerable Filter BanksabstractWe introduce a technique for designing rotation invariant operators based on steerable filter banks. Steerable filters are widely used in computer vision as local descriptors for texture analysis. Rotation invariance has been shown to improve texture-based classification in certain contexts. Our approach to invariance is based on solving the partial differential equation associated with the formulation of invariance in a Lie group framework. Xiaojin Shi, Alex L. Ribeiro-Castro, Roberto Manduchi, Richard Montgomery 0002 |
IEEE Signal Process. Lett. | 3 |
| 2005 | Hybrid Joint-Separable Multibody TrackingabstractStatistical models for tracking different moving bodies must be able to reason about occlusions in order to be effective. Representing the joint statistics across different bodies is computationally hard, since the size of the representation grows exponentially with the number of bodies being tracked. Separable tracking, with one tracker per body, cannot deal with occlusions effectively. We propose a new model, dubbed Hybrid Joint-Separable (HJS), that uses a representation size that grows linearly with the number of bodies, and a computational complexity that grows quadratically. This model can reason explicitly about occlusions. We describe a particle filter implementation of this model, and present promising experimental results. Oswald Lanz, Roberto Manduchi |
CVPR (1) | 2 |
| 2005 | Dynamic Environment Exploration Using a Virtual White CaneabstractThe virtual white cane is a range sensing device based on active triangulation, that can measure distances at a rate of 15 measurements/second. A blind person can use this device for sensing the environment, pointing it as if it was a flashlight. Beside measuring distances, this device can detect surface discontinuities, such as the foot of a wall, a step, or a drop-off. This is obtained by analyzing the range data collected as the user swings the device around, tracking planar patches and finding discontinuities. In this paper we briefly describe the range sensing device, and present an online surface tracking algorithm, based on a Jump-Markov model. We show experimental results proving the robustness of the tracking system in real-world conditions. Dan Yuan, Roberto Manduchi |
CVPR (1) | 2 |
| 2005 | Detection and Localization of Curbs and Stairways Using Stereo VisionabstractWe present algorithms to detect and precisely localize curbs and stairways for autonomous navigation. These algorithms combine brightness information (in the form of edgels) with 3-D data from a commercial stereo system. The overall system (including stereo computation) runs at about 4 Hz on a 1 GHz laptop. We show experimental results and discuss advantages and shortcomings of our approach. Xiaoye Lu, Roberto Manduchi |
ICRA | 2 |
| 2004 | Wide Baseline Feature Matching Using the Cross-Epipolar Ordering Constraint
Xiaoye Lu, Roberto Manduchi |
CVPR (1) | 2 |
| 2004 | Invariant Operators, Small Samples, and the Bias-Variance Dilemma
Xiaojin Shi, Roberto Manduchi |
CVPR (2) | 2 |
| 2004 | Learning Outdoor Color Classification from Just One Training Image
Roberto Manduchi |
ECCV (4) | 1 |
| 2003 | Visual curb localization for autonomous navigationabstractPercept-referenced commanding is an attractive paradigm for autonomous navigation over long distances. Rather than relying on precise prior environment maps and self-localization, this approach uses high-level primitives that refer to environmental features. The assumption is that the sensing and processing system onboard the robot should be able to extract such features reliably. In this context, we present an algorithm for the visual-based localization of curbs in an urban scenario. This represents a basic sensing capability that would enable behaviors such as "follow this road keeping at a certain distance from its edge". Our approach to curb localization relies on both photometry and range information (as obtained by stereopsis). Candidate image points are first selected using a combination of cues, and then input to a weighted Hough transform, which determines the line (or lines) that are most likely to belong to the curb's edge. Our experimental results show that, as long as the range data has acceptable quality, our system performs very robustly in real-world situations. Riccardo Turchetto, Roberto Manduchi |
IROS | 2 |
| 2003 | Obstacle Detection in Foliage with Ladar and Radar
Larry H. Matthies, Chuck Bergh, Andres Castano, Jose Macedo, Roberto Manduchi |
ISRR | 5 |
| 2002 | Autonomous terrain characterisation and modelling for dynamic control of unmanned vehiclesabstractWe discuss techniques to predict the dynamic vehicle response to various natural obstacles. This method can then be used to adjust the vehicle dynamics to optimize performance (e.g. speed) while ensuring that the vehicle is not damaged. This capability opens up a new area of obstacle negotiation for UGVs, where the vehicle moves over certain obstacles, rather than avoiding them, thereby resulting in more effective achievement of objectives. Robust obstacle negotiation and vehicle dynamics prediction requires several key technologies that are discussed in this paper. We detect and segment (label) obstacles using a novel 3D obstacle algorithm. The material of each labelled obstacle (rock, vegetation, etc) is then determined using a texture or color classification scheme. Terrain load-bearing surface models are then constructed using vertical springs to model the compressibility and traversability of each obstacle in front of the vehicle. The terrain model is then combined with the vehicle suspension model to yield an estimate of the maximum safe velocity, and predict the vehicle dynamics as the vehicle follows a path. This end-to-end obstacle negotiation system is envisioned to be useful in optimized path planning and vehicle navigation in terrain conditions cluttered with vegetation, bushes, rocks, etc. Results on natural terrain with various natural materials are presented. Ashit Talukder, Roberto Manduchi, Rebecca Castaño, Ken Owens, Larry H. Matthies, Andres Castano, Robert W. Hogg |
IROS | 2 |
| 2000 | Mixture Models and the Segmentation of Multimodal TexturesabstractA problem with using mixture-of-Gaussian models for unsupervised texture segmentation is that a "multimodal" texture (such as can often be encountered in natural images) cannot be well represented by a single Gaussian cluster. We propose a divide-and-conquer method that groups together Gaussian clusters (estimated via Expectation Maximization) into homogeneous texture classes. This method allows to successfully segment even rather complex textures, as demonstrated by experimental tests on natural images. Roberto Manduchi |
CVPR | 1 |
| 2000 | Semantic Progressive Transmission for Deep Space CommunicationsabstractSummary form only given. We present an integrated system for the intelligent progressive transmission of data for deep space communications. This work is motivated by the realization that much more information can be collected by imaging and remote sensing equipment than can be transmitted through downlink channels. The goal of this work was to extend the idea of progressive transmission by incorporating semantic value to the "importance" attributes of encoded data. In particular, we define a simple measure of science return based on information theoretic considerations as well as on the estimated scientific value of the transmitted data. The task of determining the relative scientific importance of segments of data is carried out by an onboard science processing module, designed according to guidelines provided by the remote user (the science community). This module pre-processes the image and provides input to the progressive encoder in the form of a suitable "classification map". By combining the semantic characterization produced by the science processing module with the content-blind data organization criteria of the traditional progressive encoder, we obtain a new measure of the importance of each segment of data. Algorithms that best utilize the available resources can thus be designed and their performances in terms of science return assessed. Roberto Manduchi, Samuel Dolinar, Fabrizio Pollara, Adina Matache |
Data Compression Conference | 1 |
| 2000 | A Cluster Grouping Technique for Texture SegmentationabstractWe propose an algorithm for texture segmentation based on a divide-and-conquer strategy of statistical modeling. Selected sets of Gaussian clusters, estimated via expectation maximization on the texture features, are grouped together to form composite texture classes. Our cluster grouping technique exploits the inherent local spatial correlation among posterior distributions of clusters belonging to the same texture class. Despite its simplicity, this algorithm can model even very complex distributions, typical of natural outdoor images. Roberto Manduchi |
ICPR | 1 |
| 1999 | Bayesian Fusion of Color and Texture SegmentationsabstractIn many applications one would like to use information from both color and texture features in order to segment an image. We propose a novel technique to combine "soft" segmentations computed for two or more features independently. Our algorithm merges models according to a maximum descriptiveness criterion, and allows us to choose any number of classes for the final grouping. This technique also allows us to improve the quality of supervised classification based on one feature (e.g. color) by merging information from unsupervised segmentation based on another feature (e.g., texture). Roberto Manduchi |
ICCV | 1 |
| 1999 | Independent Component Analysis of TexturesabstractA common method for texture representation is to use the marginal probability densities over the outputs of a set of multi-orientation, multi-scale filters as a description of the texture. We propose a technique, based on independent component analysis, for choosing the set of filters that yield the most informative marginals, meaning that the product over the marginals most closely approximates the joint probability density function of the filter outputs. The algorithm is implemented using a steerable filter space. Experiments involving both texture classification and synthesis show that compared to principal components analysis, ICA provides superior performance for modeling of natural and synthetic textures. Roberto Manduchi, Javier Portilla |
ICCV | 1 |
| 1998 | Bilateral Filtering for Gray and Color ImagesabstractBilateral filtering smooths images while preserving edges, by means of a nonlinear combination of nearby image values. The method is noniterative, local, and simple. It combines gray levels or colors based on both their geometric closeness and their photometric similarity, and prefers near values to distant values in both domain and range. In contrast with filters that operate on the three bands of a color image separately, a bilateral filter can enforce the perceptual metric underlying the CIE-Lab color space, and smooth colors and preserve edges in a way that is tuned to human perception. Also, in contrast with standard filtering, bilateral filtering produces no phantom colors along edges in color images, and reduces phantom colors where they appear in the original image. Carlo Tomasi, Roberto Manduchi |
ICCV | 2 |
| 1998 | Stereo Matching as a Nearest-Neighbor ProblemabstractWe propose a representation of images, called intrinsic curves, that transforms stereo matching from a search problem into a nearest-neighbor problem. Intrinsic curves are the paths that a set of local image descriptors trace as an image scanline is traversed from left to right. Intrinsic curves are ideally invariant with respect to disparity. Stereo correspondence then becomes a trivial lookup problem in the ideal case. We also show how to use intrinsic curves to match real images in the presence of noise, brightness bias, contrast fluctuations, moderate geometric distortion, image ambiguity, and occlusions. In this case, matching becomes a nearest-neighbor problem, even for very large disparity values. Carlo Tomasi, Roberto Manduchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Stereo Without Search
Carlo Tomasi, Roberto Manduchi |
ECCV (1) | 2 |
| 1996 | Some properties of generalized factorable 2-D FIR filtersabstractA paper by Chen and Vaidyanathan (1993) describes a method for designing multidimensional (MD) filters with spectral support in the shape of a parallelepiped, starting from M 1-D prototypes. Such filters require a number of elementary operations per input sample (OPS) that grows linearly with the size of each 1-D prototype. As a matter of fact, they belong to the general class of generalized factorable (GF) filters, which we define here. Beside the computational efficiency (in both the design and the implementation), we show in this work that GF filters share some interesting properties (relative to symmetry and frequency response constraints). Roberto Manduchi |
ICASSP | 1 |
| 1995 | Pyramidal implementation of deformable kernelsabstractIn computer vision and increasingly, in rendering and image processing, it is useful to filter images with continuous rotated and scaled families of filters. For practical implementations, one can think of using a discrete family of filters, and then to interpolate from their outputs to produce the desired filtered version of the image. We propose a multirate implementation of deformable kernels, capable to further reduce the computational weight. The "basis" filters are applied to the different levels of a pyramidal decomposition. The new system is not shift-invariant-it suffers from "aliasing". We introduce a new quadratic error criterion which keeps into account the inherent system aliasing. By using hypermatrix and Kronecker algebra, we are able to cast the global optimization task into a multilinear problem. An iterative procedure ("pseudo-SVD") is used to minimize the overall quadratic approximation error. Roberto Manduchi, Pietro Perona |
ICIP | 1 |
| 1995 | 2-D IFIR Structures Using Generalized Factorable FiltersabstractThe extension of the idea of Interpolated FIR filters to the two-dimensional case is presented. Such systems allow for lower computational weight, in terms of number of elementary operations per input sample. We have considered 2-D IFIR filters with parallelogram-shaped spectral support. To design the filters in the two stages, we have used a technique recently developed by Chen and Vaidyanathan. The resulting filters belong to the class of Generalized Factorable filters, for which an efficient implementation exists. An interesting problem peculiar to the multidimensional case is the choice of the sublattice which represents the definition support of the first-stage filter. We present a strategy for choosing (given the spectral support of the desired frequency response) the optimal sublattice, and to design the second-stage (interpolator) filter in order to achieve low overall computational complexity. Roberto Manduchi |
ISCAS | 1 |
| 1995 | Spectral characteristics and motion-compensated restoration of composite framesabstractThe practice of superimposing the fields of a frame is applied in various fields, for example, thermographic and biomedical imaging. The pictures obtained in this way, which are termed composite frames, are severely degraded if the scene's objects are not perfectly still. The restoration of composite frames affected by motion-induced blurring requires the ability to estimate the field displacement from composite frames. The frequency domain analysis of composite frames proposed in the paper suggests a displacement estimation technique of a phase-correlation type that can be applied to composite frames. Roberto Manduchi, Guido M. Cortelazzo |
IEEE Trans. Image Process. | 1 |
| 1993 | Accuracy analysis for correlation-based image registration algorithms
Roberto Manduchi, Gian Antonio Mian |
ISCAS | 1 |
| 1993 | On the determination of all the sublattices of preassigned index and its application to multidimensional subsamplingabstractSubsampling offers an effective and simple way to implement a data compression technique. Its use in the video context is rather attractive, as its application to current video broadcasting schemes, such as MUSE and HD-MAC, clearly shows. A sensible exploitation of subsampling's potential requires the systematic evaluation of the image degradation effects due to the specific subsampling lattice. Such a possibility is equivalent to the determination of all the sublattices of a given index within a lattice, a problem that can be solved, as the work shows, on the basis of the Hermite normal-form theorem. The practical use of the result in the subsampling context is exemplified for a data compression case.> Guido M. Cortelazzo, Roberto Manduchi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1993 | Multistage sampling structure conversion of video signalsabstractThe work extends multistage implementation of sampling structure conversion to the multidimensional (M-D) case. The issues arising in this task are usefully addressed on the basis of lattice theory. Numerical data supporting the advantages of multistage sampling conversion are presented, and the case of format conversion from the 4/3 to the 16/9 aspect ratio is examined as a study case. The main indication of the present work is that multistage implementation, in the case of systems for sampling structure conversion of video signals, may improve the system characteristics and visual rendition.> Roberto Manduchi, Guido M. Cortelazzo, Gian Antonio Mian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |