EDBT 2026 Demo / reviewers in the wild / expert
James M. Coughlan
dblp:43/6505
· DBLP profile ↗
56ranked-venue papers
15as first author
1since 2021 · last 2024
0000-0003-2775-4083ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 11 first-authorApplied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-authorHuman-computer interaction and ubiquitous computing · 11 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Probabilistic and Bayesian machine learning · 42% 3D vision · 34% Segmentation and scene understanding · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
Accessibility and assistive technology · 77% Interaction techniques and input · 23% | |
| Computer graphics and multimedia
7 papers |
Image and video processing · 53% Geometric modeling and processing · 17% Rendering · 16% | |
| Theoretical computer science
3 papers |
Computational complexity · 44% Information theory · 30% Algorithms and data structures · 26% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Accessibility and assistive technology › visual impairment
blind and low vision users |
0.2 | 1 | 2014 | The last meter: blind visual guidance to a target · CHI 2014 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.1 | 4 | 2000 | Fundamental Limits of Bayesian Inference: Order Parameters and Phase Transitions for Road Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2000 The Manhattan World Assumption: Regularities in Scene Statistics which Enable Bayesian Inference · NIPS 2000 Order Parameters for Minimax Entropy Distributions: When Does High Level Knowledge Help? · CVPR 2000 |
Image and video processing
edge detection |
0.1 | 2 | 2003 | Statistical Edge Detection: Learning and Evaluating Edge Cues · IEEE Trans. Pattern Anal. Mach. Intell. 2003 Fundamental Bounds on Edge Detection: An Information Theoretic Evaluation of Different Edge Cues · CVPR 1999 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.1 | 2 | 2001 | The g Factor: Relating Distributions on Features to Distributions on Images · NIPS 2001 A Phase Space Approach to Minimax Entropy Learning and the Minutemax Approximations · NIPS 1998 |
Computer vision › 3D vision › 3d scene understanding › scene geometry
manhattan world assumption |
0.1 | 2 | 2000 | The Manhattan World Assumption: Regularities in Scene Statistics which Enable Bayesian Inference · NIPS 2000 Manhattan World: Compass Direction from a Single Image by Bayesian Inference · ICCV 1999 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.1 | 2 | 2000 | The Manhattan World Assumption: Regularities in Scene Statistics which Enable Bayesian Inference · NIPS 2000 Manhattan World: Compass Direction from a Single Image by Bayesian Inference · ICCV 1999 |
Computer vision › Image recognition and object detection
visual search |
0.1 | 2 | 2000 | Order Parameters for Minimax Entropy Distributions: When Does High Level Knowledge Help? · CVPR 2000 High-Level and Generic Models for Visual Search: When Does High Level Knowledge Help? · CVPR 1999 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network |
0.0 | 1 | 2003 | A Bayesian Network Framework for Relational Shape Matching · ICCV 2003 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
bethe free energy |
0.0 | 1 | 2003 | A Bayesian Network Framework for Relational Shape Matching · ICCV 2003 |
Computer vision › 3D vision
shape matching |
0.0 | 1 | 2003 | A Bayesian Network Framework for Relational Shape Matching · ICCV 2003 |
Computer vision › 3D vision
shape perception |
0.0 | 1 | 2003 | The Generic Viewpoint Assumption and Planar Bias · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.0 | 1 | 2003 | The Generic Viewpoint Assumption and Planar Bias · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation |
0.0 | 1 | 2002 | Finding Deformable Shapes Using Loopy Belief Propagation · ECCV (3) 2002 |
Geometric modeling and processing › shape matching
deformable shape matching |
0.0 | 1 | 2002 | Finding Deformable Shapes Using Loopy Belief Propagation · ECCV (3) 2002 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.0 | 1 | 2001 | The g Factor: Relating Distributions on Features to Distributions on Images · NIPS 2001 |
Computational photography and imaging › shape and reflectance estimation
shape from shading |
0.0 | 1 | 2001 | The KGBR Viewpoint-Lighting Ambiguity and its Resolution by Generic Constraints · ICCV 2001 |
Computer vision › 3D vision
image registration |
0.0 | 2 | 1997 | Spline-Based Image Registration · Int. J. Comput. Vis. 1997 Hierarchical spline-based image registration · CVPR 1994 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.0 | 1 | 2000 | Fundamental Limits of Bayesian Inference: Order Parameters and Phase Transitions for Road Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2000 |
Computational complexity
phase transition |
0.0 | 1 | 2000 | Fundamental Limits of Bayesian Inference: Order Parameters and Phase Transitions for Road Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2000 |
Computer vision › 3D vision
camera pose estimation |
0.0 | 1 | 1999 | Manhattan World: Compass Direction from a Single Image by Bayesian Inference · ICCV 1999 |
Computer vision › 3D vision › pose estimation
orientation estimation |
0.0 | 1 | 1999 | Manhattan World: Compass Direction from a Single Image by Bayesian Inference · ICCV 1999 |
Information theory
statistical inference |
0.0 | 1 | 1999 | Fundamental Bounds on Edge Detection: An Information Theoretic Evaluation of Different Edge Cues · CVPR 1999 |
Rendering › appearance modeling › reflectance and appearance modeling
reflectance and illumination modeling |
0.0 | 2 | 2003 | The Generic Viewpoint Assumption and Planar Bias · IEEE Trans. Pattern Anal. Mach. Intell. 2003 The KGBR Viewpoint-Lighting Ambiguity and its Resolution by Generic Constraints · ICCV 2001 |
Image and video processing › image segmentation
contour detection |
0.0 | 1 | 1998 | Convergence Rates of Algorithms for Visual Search: Detecting Visual Contours · NIPS 1998 |
Image and video processing › image matching
deformable template matching |
0.0 | 1 | 1998 | Efficient Optimization of a Deformable Template Using Dynamic Programming · CVPR 1998 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
prior modeling |
0.0 | 1 | 2003 | The Generic Viewpoint Assumption and Planar Bias · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Medical and health informatics › medical imaging
medical image analysis |
0.0 | 1 | 2003 | A Bayesian Network Framework for Relational Shape Matching · ICCV 2003 |
Rendering › reflectance modeling
lambertian reflectance |
0.0 | 1 | 2003 | The Generic Viewpoint Assumption and Planar Bias · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Computer vision › 3D vision › correspondence estimation
optical flow and stereo matching |
0.0 | 1 | 1994 | Hierarchical spline-based image registration · CVPR 1994 |
Computer vision › 3D vision
structure from motion |
0.0 | 1 | 1994 | Hierarchical spline-based image registration · CVPR 1994 |
Methods — techniques the papers use, named apart from their topics
user study · 0.2object recognition · 0.2orthographic projection · 0.1order parameters · 0.1bethe free energy · 0.1bayesian prior analysis · 0.1bayesian network · 0.1affine warp · 0.1bayesian inference · 0.1minimax entropy learning · 0.0saliency model · 0.0information maximization · 0.0likelihood ratio test · 0.0chernoff information · 0.0ROC analysis · 0.0loopy belief propagation · 0.0lambertian reflectance · 0.0generic viewpoint and lighting constraints · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Accessible Point-and-Tap Interaction for Acquiring Detailed Information About Tactile Graphics and 3D Models
Andrea Narcisi, Huiying Shen, Dragan Ahmetovic, Sergio Mascetti, James M. Coughlan |
ICCHP (1) | 5 |
| 2020 | An Audio-Based 3D Spatial Guidance AR System for Blind Users
James M. Coughlan, Brandon Biggs, Marc-Aurèle Rivière, Huiying Shen |
ICCHP (1) | 1 |
| 2020 | An Indoor Navigation App Using Computer Vision and Sign Recognition
Giovanni Fusco 0003, Seyed Ali Cheraghi, Leo Neat, James M. Coughlan |
ICCHP (1) | 4 |
| 2018 | Indoor Localization Using Computer Vision and Visual-Inertial OdometryabstractIndoor wayfinding is a major challenge for people with visual impairments, who are often unable to see visual cues such as informational signs, land-marks and structural features that people with normal vision rely on for wayfinding. We describe a novel indoor localization approach to facilitate wayfinding that uses a smartphone to combine computer vision and a dead reckoning technique known as visual-inertial odometry (VIO). The approach uses sign recognition to estimate the user's location on the map whenever a known sign is recognized, and VIO to track the user's movements when no sign is visible. The ad-vantages of our approach are (a) that it runs on a standard smartphone and re-quires no new physical infrastructure, just a digital 2D map of the indoor environment that includes the locations of signs in it; and (b) it allows the user to walk freely without having to actively search for signs with the smartphone (which is challenging for people with severe visual impairments). We report a formative study with four blind users demonstrating the feasibility of the approach and suggesting areas for future improvement. Giovanni Fusco 0003, James M. Coughlan |
ICCHP (2) | 2 |
| 2017 | Evaluating Author and User Experience for an Audio-Haptic System for Annotation of Physical ModelsabstractWe describe three usability studies involving a prototype system for creation and haptic exploration of labeled locations on 3D objects. The system uses a computer, webcam, and fiducial markers to associate a physical 3D object in the camera's view with a predefined digital map of labeled locations ("hotspots"), and to do real-time finger tracking, allowing a blind or visually impaired user to explore the object and hear individual labels spoken as each hotspot is touched. This paper describes: (a) a formative study with blind users exploring pre-annotated objects to assess system usability and accuracy; (b) a focus group of blind participants who used the system and, through structured and unstructured discussion, provided feedback on its practicality, possible applications, and real-world potential; and (c) a formative study in which a sighted adult used the system to add labels to on-screen images of objects, demonstrating the practicality of remote annotation of 3D models. These studies and related literature suggest potential for future iterations of the system to benefit blind and visually impaired users in educational, professional, and recreational contexts. James M. Coughlan, Joshua A. Miele |
ASSETS | 1 |
| 2017 | JustPoint: Identifying Colors with a Natural User InterfaceabstractPeople with severe visual impairments usually have no way of identifying the colors of objects in their environment. While existing smartphone apps can recognize colors and speak them aloud, they require the user to center the object of interest in the camera's field of view, which is challenging for many users. We developed a smartphone app to address this problem that reads aloud the color of the object pointed to by the user's fingertip, without confusion from background colors. We evaluated the app with nine people who are blind, demonstrating the app's effectiveness and suggesting directions for improvements in the future. Sergio Mascetti, Andrea Gerino, Cristian Bernareggi, Silvia D'Acquisto, Mattia Ducci, James M. Coughlan |
ASSETS | 6 |
| 2016 | Towards a Sign-Based Indoor Navigation System for People with Visual ImpairmentsabstractNavigation is a challenging task for many travelers with visual impairments. While a variety of GPS-enabled tools can provide wayfinding assistance in outdoor settings, GPS provides no useful localization information indoors. A variety of indoor navigation tools are being developed, but most of them require potentially costly physical infrastructure to be installed and maintained, or else the creation of detailed visual models of the environment. We report development of a new smartphone-based navigation aid, which combines inertial sensing, computer vision and floor plan information to estimate the user's location with no additional physical infrastructure and requiring only the locations of signs relative to the floor plan. A formative study was conducted with three blind volunteer participants demonstrating the feasibility of the approach and highlighting the areas needing improvement. Alejandro Rituerto, Giovanni Fusco 0003, James M. Coughlan |
ASSETS | 3 |
| 2015 | Zebra Crossing Spotter: Automatic Population of Spatial Databases for Increased Safety of Blind Travelersabstractin urban settings. Knowing the location of crosswalks is critical for a blind person planning a trip that includes street crossing. By augmenting existing spatial databases (such as Google Maps or OpenStreetMap) with this information, a blind traveler may make more informed routing decisions, resulting in greater safety during independent travel. Our algorithm first searches for zebra crosswalks in satellite images; all candidates thus found are validated against spatially registered Google Street View images. This cascaded approach enables fast and reliable discovery and localization of zebra crosswalks in large image datasets. While fully automatic, our algorithm could also be complemented by a final crowdsourcing validation stage for increased accuracy. Dragan Ahmetovic, Roberto Manduchi, James M. Coughlan, Sergio Mascetti |
ASSETS | 3 |
| 2015 | Appliance Displays: Accessibility Challenges and Proposed SolutionsabstractPeople who are blind or visually impaired face difficulties using a growing array of everyday appliances because they are equipped with inaccessible electronic displays. We report developments on our "Display Reader" smartphone app, which uses computer vision to help a user acquire a usable image of a display and have the contents read aloud, to address this problem. Drawing on feedback from past and new studies with visually impaired volunteer participants, as well as from blind accessibility experts, we have improved and simplified our user interface and have also added the ability to read seven-segment digit displays. Our system works fully automatically and in real time, and we compare it with general-purpose assistive apps such as Be My Eyes, which recruit remote sighted assistants (RSAs) to answer questions about video captured by the user. Our discussions and preliminary experiment highlight the advantages and disadvantages of fully automatic approaches compared with RSAs, and suggest possible hybrid approaches to investigate in the future. Giovanni Fusco 0003, Ender Tekin, Nicholas A. Giudice, James M. Coughlan |
ASSETS | 4 |
| 2014 | Using computer vision to access appliance displaysabstractPeople who are blind or visually impaired face difficulties accessing a growing array of everyday appliances, needed to perform a variety of daily activities, because they are equipped with electronic displays. We are developing a "Display Reader" smartphone app, which uses computer vision to help a user acquire a usable image of a display, to address this problem. The current prototype analyzes video from the smartphone's camera, providing real-time feedback to guide the user until a satisfactory image is acquired, based on automatic estimates of image blur and glare. Formative studies were conducted with several blind and visually impaired participants, whose feedback is guiding the development of the user interface. The prototype software has been released as a Free and Open Source (FOSS) project. Giovanni Fusco 0003, Ender Tekin, Richard E. Ladner, James M. Coughlan |
ASSETS | 4 |
| 2014 | The last meter: blind visual guidance to a targetabstractSmartphone apps can use object recognition software to provide information to blind or low vision users about objects in the visual environment. A crucial challenge for these users is aiming the camera properly to take a well-framed picture of the desired target object. We investigate the effects of two fundamental constraints of object recognition - frame rate and camera field of view - on a blind person's ability to use an object recognition smartphone app. The app was used by 18 blind participants to find visual targets beyond arm's reach and approach them to within 30 cm. While we expected that a faster frame rate or wider camera field of view should always improve search performance, our experimental results show that in many cases increasing the field of view does not help, and may even hurt, performance. These results have important implications for the design of object recognition systems for blind users. Roberto Manduchi, James M. Coughlan |
CHI | 2 |
| 2014 | Determining a Blind Pedestrian's Location and Orientation at Traffic Intersections
Giovanni Fusco 0003, Huiying Shen, Vidya N. Murali, James M. Coughlan |
ICCHP (1) | 4 |
| 2014 | An Investigation into Incorporating Visual Information in Audio Processing
Ender Tekin, James M. Coughlan, Helen J. Simon |
ICCHP (1) | 2 |
| 2013 | CamIO: a 3D computer vision system enabling audio/haptic interaction with physical objects by blind usersabstractCamIO (short for "Camera Input-Output") is a novel camera system designed to make physical objects (such as documents, maps, devices and 3D models) fully accessible to blind and visually impaired persons, by providing real-time audio feedback in response to the location on an object that the user is pointing to. The project will have wide ranging impact on access to graphics, tactile literacy, STEM education, independent travel and wayfinding, access to devices, and other applications to increase the independent functioning of blind, low vision and deaf-blind individuals. We describe our preliminary results with a prototype CamIO system consisting of the Microsoft Kinect camera connected to a laptop computer. An experiment with a blind user demonstrates the feasibility of the system, which issues Text-to-Speech (TTS) annotations whenever the user's fingers approach any pre-defined "hotspot" regions on the object. Huiying Shen, Owen Edwards, Joshua A. Miele, James M. Coughlan |
ASSETS | 4 |
| 2012 | The Crosswatch Traffic Intersection Analyzer: A Roadmap for the Future
James M. Coughlan, Huiying Shen |
ICCHP (2) | 1 |
| 2012 | Towards a Real-Time System for Finding and Reading Signs for Visually Impaired Users
Huiying Shen, James M. Coughlan |
ICCHP (2) | 2 |
| 2011 | Localizing blurry and low-resolution text in natural imagesabstractThere is a growing body of work addressing the problem of localizing printed text regions occurring in natural scenes, all of it focused on images in which the text to be localized is resolved clearly enough to be read by OCR. This paper introduces an alternative approach to text localization based on the fact that it is often useful to localize text that is identifiable as text but too blurry or small to be read, for two reasons. First, an image can be decimated and processed at a coarser resolution than usual, resulting in faster localization before OCR is performed (at full resolution, if needed). Second, in real-time applications such as a cell phone app to find and read text, text may initially be acquired from a lower-resolution video image in which it appears too small to be read; once the text's presence and location have been established, a higher-resolution image can be taken in order to resolve the text clearly enough to read it.We demonstrate proof of concept of this approach by describing a novel algorithm for binarizing the image and extracting candidate text features, called "blobs," and grouping and classifying the blobs into text and non-text categories. Experimental results are shown on a variety of images in which the text is resolved too poorly to be clearly read, but is still identifiable by our algorithm as text. Pannag R. Sanketi, Huiying Shen, James M. Coughlan |
WACV | 3 |
| 2011 | Real-time detection and reading of LED/LCD displays for visually impaired personsabstractModern household appliances, such as microwave ovens and DVD players, increasingly require users to read an LED or LCD display to operate them, posing a severe obstacle for persons with blindness or visual impairment. While OCR-enabled devices are emerging to address the related problem of reading text in printed documents, they are not designed to tackle the challenge of finding and reading characters in appliance displays. Any system for reading these characters must address the challenge of first locating the characters among substantial amounts of background clutter; moreover, poor contrast and the abundance of specular highlights on the display surface - which degrade the image in an unpredictable way as the camera is moved - motivate the need for a system that processes images at a few frames per second, rather than forcing the user to take several photos, each of which can take seconds to acquire and process, until one is readable.We describe a novel system that acquires video, detects and reads LED/LCD characters in real time, reading them aloud to the user with synthesized speech. The system has been implemented on both a desktop and a cell phone. Experimental results are reported on videos of display images, demonstrating the feasibility of the system. Ender Tekin, James M. Coughlan, Huiying Shen |
WACV | 2 |
| 2010 | Anti-blur feedback for visually impaired users of smartphone camerasabstractA wide range of smartphone applications are emerging that employ image processing and computer vision algorithms to interpret the contents of images acquired by the phone's built-in camera, including applications that read product barcodes and recognize a variety of documents and other objects. However, almost all of these applications are designed for normally sighted users; a major barrier for visually impaired users (who might benefit greatly from such applications) is the difficulty of taking good-quality images. To overcome this barrier, this paper focuses on reducing the incidence of motion blur, caused by camera shake and other movements, which is a common cause of poor-quality, unusable images. We propose a simple technique for detecting camera shake, using the smartphone's built-in accelerometer (i.e. tilt sensor) to alert the user in real-time to any shake, providing feedback that enables him/her to hold the camera more steadily. A preliminary experiment with a blind iPhone user demonstrates the feasibility of the approach. Pannag R. Sanketi, James M. Coughlan |
ASSETS | 2 |
| 2010 | Real-Time Walk Light Detection with a Mobile Phone
Volodymyr Ivanchenko, James M. Coughlan, Huiying Shen |
ICCHP (2) | 2 |
| 2010 | A Mobile Phone Application Enabling Visually Impaired Users to Find and Read Product Barcodes
Ender Tekin, James M. Coughlan |
ICCHP (2) | 2 |
| 2009 | Elevation-based MRF stereo implemented in real-time on a GPUabstractWe describe a novel framework for calculating dense, accurate elevation maps from stereo, in which the height of each point in the scene is estimated relative to the ground plane. The key to our framework's ability to estimate elevation accurately is an MRF formulation of stereo that directly represents elevation at each pixel instead of the usual disparity. By enforcing smoothness of elevation rather than disparity (using pairwise interactions in the MRF), the usual fronto-parallel bias is transformed into a horizontal (parallel to the ground) bias - a bias that is more appropriate for scenes characterized by a dominant ground plane viewed from an angle. This horizontal bias amounts to a more informative prior for such scenes, which results in more accurate surface reconstruction, with sub-pixel accuracy. We apply this framework to the problem of finding small obstacles, such as curbs and other small deviations from the ground plane, a few meters in front of a vehicle (such as a wheelchair or robot) that are missed by standard real-time correlation stereo algorithms. We demonstrate a real-time implementation of our framework on a GPU (we have made the code publicly available), which processes a 640 × 480 stereo image pair in 160 ms using either our elevation model or a standard disparity-based model (with 32 elevation or disparity levels), and describe experimental results. Volodymyr Ivanchenko, Huiying Shen, James M. Coughlan |
WACV | 3 |
| 2009 | An algorithm enabling blind users to find and read barcodesabstractMost camera-based systems for finding and reading barcodes are designed to be used by sighted users (e.g. the Red Laser iPhone app), and assume the user carefully centers the barcode in the image before the barcode is read. Blind individuals could benefit greatly from such systems to identify packaged goods (such as canned goods in a supermarket), but unfortunately in their current form these systems are completely inaccessible because of their reliance on visual feedback from the user.To remedy this problem, we propose a computer vision algorithm that processes several frames of video per second to detect barcodes from a distance of several inches; the algorithm issues directional information with audio feedback (e.g. "left," "right") and thereby guides a blind user holding a webcam or other portable camera to locate and home in on a barcode. Once the barcode is detected at sufficiently close range, a barcode reading algorithm previously developed by the authors scans and reads aloud the barcode and the corresponding product information. We demonstrate encouraging experimental results of our proposed system implemented on a desktop computer with a webcam held by a blindfolded user; ultimately the system will be ported to a camera phone for use by visually impaired users. Ender Tekin, James M. Coughlan |
WACV | 2 |
| 2009 | Figure-ground segmentation using factor graphs
Huiying Shen, James M. Coughlan, Volodymyr Ivanchenko |
Image Vis. Comput. | 2 |
| 2008 | Computer vision-based clear path guidance for blind wheelchair usersabstractWe describe a system for guiding blind and visually impaired wheelchair users along a clear path that uses computer vision to sense the presence of obstacles or other terrain features and warn the user accordingly. Since multiple terrain features can be distributed anywhere on the ground, and their locations relative to a moving wheelchair are continually changing, it is challenging to communicate this wealth of spatial information in a way that is rapidly comprehensible to the user. The main contribution of our system is the development of a novel user interface that allows the user to interrogate the environment by sweeping a standard (unmodified) white cane back and forth: the system continuously tracks the cane location and sounds an alert if a terrain feature is detected in the direction the cane is pointing. Experiments are described demonstrating the feasibility of the approach. Volodymyr Ivanchenko, James M. Coughlan, William Gerrey, Huiying Shen |
ASSETS | 2 |
| 2008 | Crosswatch: A Camera Phone System for Orienting Visually Impaired Pedestrians at Traffic Intersections
Volodymyr Ivanchenko, James M. Coughlan, Huiying Shen |
ICCHP | 2 |
| 2008 | Portable and Mobile Systems in Assistive Technology
Roberto Manduchi, James M. Coughlan |
ICCHP | 2 |
| 2008 | Search Strategies of Visually Impaired Persons Using a Camera Phone Wayfinding System
Roberto Manduchi, James M. Coughlan, Volodymyr Ivanchenko |
ICCHP | 2 |
| 2007 | Accessible spaces: navigating through a marked environment with a camera phoneabstractWe demonstrate a system designed to assist a visually impaired individual while moving in an unfamiliar environment. Small and economical color markers are placed in key locations, possibly in the vicinity of other signs (bar codes or text). The user can detect these markers by means of a cell phone equipped with a camera. Our demonstration highlights a number of novel features, including: improved acoustic interfaces; estimation of the distance to the marker, which is communicated to the user via text-to-speech (TTS); increased robustness via rotation invariance, which makes the system easier to use for users with reduced dexterity. Kee-Yip Chan, Roberto Manduchi, James M. Coughlan |
ASSETS | 3 |
| 2007 | Dynamic quantization for belief propagation in sparse spaces
James M. Coughlan, Huiying Shen |
Comput. Vis. Image Underst. | 1 |
| 2006 | Computer Vision-Based Terrain Sensors for Blind Wheelchair Users
James M. Coughlan, Roberto Manduchi, Huiying Shen |
ICCHP | 1 |
| 2004 | An Information Maximization Model of Eye MovementsabstractWe propose a sequential information maximization model as a general strategy for programming eye movements. The model reconstructs high-resolution visual information from a sequence of fixations, taking into account the fall-off in resolution from the fovea to the periphery. From this framework we get a simple rule for predicting fixation sequences: after each fixation, fixate next at the location that minimizes uncertainty (maximizes information) about the stimulus. By comparing our model performance to human eye movement data and to predictions from a saliency and random model, we demonstrate that our model is best at predicting fixation locations. Modeling additional biological constraints will improve the prediction of fixation sequences. Our results suggest that information maximization is a useful principle for programming eye movements. Laura Walker Renninger, James M. Coughlan, Preeti Verghese, Jitendra Malik |
NIPS | 2 |
| 2003 | A Bayesian Network Framework for Relational Shape MatchingabstractA Bayesian network formulation for relational shape matching is presented. The main advantage of the relational shape matching approach is the obviation of the nonrigid spatial mappings used by recent nonrigid matching approaches. The basic variables that need to be estimated in the relational shape matching objective function are the global rotation and scale and the local displacements and correspondences. The new Bethe free energy approach is used to estimate the pairwise correspondences between links of the template graphs and the data. The resulting framework is useful in both registration and recognition contexts. Results are shown on hand-drawn templates and on 2D transverse T1-weighted MR images. Anand Rangarajan 0001, James M. Coughlan, Alan L. Yuille |
ICCV | 2 |
| 2003 | Algorithms from statistical physics for generative models of images
James M. Coughlan, Alan L. Yuille |
Image Vis. Comput. | 1 |
| 2003 | A statistical approach to multi-scale edge detection
Scott Konishi, Alan L. Yuille, James M. Coughlan |
Image Vis. Comput. | 3 |
| 2003 | Manhattan World: Orientation and Outlier Detection by Bayesian InferenceabstractThis letter argues that many visual scenes are based on a "Manhattan" three-dimensional grid that imposes regularities on the image statistics. We construct a Bayesian model that implements this assumption and estimates the viewer orientation relative to the Manhattan grid. For many images, these estimates are good approximations to the viewer orientation (as estimated manually by the authors). These estimates also make it easy to detect outlier structures that are unaligned to the grid. To determine the applicability of the Manhattan world model, we implement a null hypothesis model that assumes that the image statistics are independent of any three-dimensional scene structure. We then use the log-likelihood ratio test to determine whether an image satisfies the Manhattan world assumption. Our results show that if an image is estimated to be Manhattan, then the Bayesian model's estimates of viewer direction are almost always accurate (according to our manual estimates), and vice versa. James M. Coughlan, Alan L. Yuille |
Neural Comput. | 1 |
| 2003 | Statistical Edge Detection: Learning and Evaluating Edge CuesabstractWe formulate edge detection as statistical inference. This statistical edge detection is data driven, unlike standard methods for edge detection which are model based. For any set of edge detection filters (implementing local edge cues), we use presegmented images to learn the probability distributions of filter responses conditioned on whether they are evaluated on or off an edge. Edge detection is formulated as a discrimination task specified by a likelihood ratio test on the filter responses. This approach emphasizes the necessity of modeling the image background (the off-edges). We represent the conditional probability distributions nonparametrically and illustrate them on two different data sets of 100 (Sowerby) and 50 (South Florida) images. Multiple edges cues, including chrominance and multiple-scale, are combined by using their joint distributions. Hence, this cue combination is optimal in the statistical sense. We evaluate the effectiveness of different visual cues using the Chernoff information and Receiver Operator Characteristic (ROC) curves. This shows that our approach gives quantitatively better results than the Canny edge detector when the image background contains significant clutter. In addition, it enables us to determine the effectiveness of different edge cues and gives quantitative measures for the advantages of multilevel processing, for the use of chrominance, and for the relative effectiveness of different detectors. Furthermore, we show that we can learn these conditional distributions on one data set and adapt them to the other with only slight degradation of performance without knowing the ground truth on the second data set. This shows that our results are not purely domain specific. We apply the same approach to the spatial grouping of edge cues and obtain analogies to nonmaximal suppression and hysteresis. Scott Konishi, Alan L. Yuille, James M. Coughlan, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | The Generic Viewpoint Assumption and Planar BiasabstractWe show that generic viewpoint and lighting assumptions resolve standard visual ambiguities by biasing toward planar surfaces. Our model uses orthographic projection with a two-dimensional affine warp and Lambertian reflectance functions, including cast and attached shadows. We use uniform priors on nuisance variables such as viewpoint direction and the light source. Limitations of using uniform priors on nuisance variables are discussed. Alan L. Yuille, James M. Coughlan, Scott Konishi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Finding Deformable Shapes Using Loopy Belief Propagation
James M. Coughlan, Sabino J. Ferreira |
ECCV (3) | 1 |
| 2002 | Bayesian A* Tree Search with Expected O(N) Node Expansions: Applications to Road TrackingabstractMany perception, reasoning, and learning problems can be expressed as Bayesian inference. We point out that formulating a problem as Bayesian inference implies specifying a probability distribution on the ensemble of problem instances. This ensemble can be used for analyzing the expected complexity of algorithms and also the algorithm-independent limits of inference. We illustrate this problem by analyzing the complexity of tree search. In particular, we study the problem of road detection, as formulated by Geman and Jedynak (1996). We prove that the expected convergence is linear in the size of the road (the depth of the tree) even though the worst-case performance is exponential. We also put a bound on the constant of the convergence and place a bound on the error rates. James M. Coughlan, Alan L. Yuille |
Neural Comput. | 1 |
| 2001 | The KGBR Viewpoint-Lighting Ambiguity and its Resolution by Generic ConstraintsabstractWe describe a novel viewpoint-lighting ambiguity which we call the KGBR. This ambiguity assumes orthographic projecting or an affine camera, and uses Lambertian reflectance functions including case/attached shadows and multiple light sources. A KGBR transform alters the geometry (by a three-dimensional affine transformation) and albedo properties of objects. If two objects are related by a KGBR transform then for any viewpoint and lighting of the first object there exists a corresponding viewpoint and lighting of the second object so that the images are identical up to an affine transformation. The Generalized Bas Relief (GBR) ambiguity is obtained as a special case of the KGBR. We describe generic viewpoint and lighting assumptions and show that either, or both, resolve this ambiguity by biasing towards objects with planar geometry. Alan L. Yuille, James M. Coughlan, Scott Konishi |
ICCV | 2 |
| 2001 | The g Factor: Relating Distributions on Features to Distributions on ImagesabstractWe describe the g-factor, which relates probability distributions on image features to distributions on the images themselves. The g-factor depends only on our choice of features and lattice quanti(cid:173) zation and is independent of the training image data. We illustrate the importance of the g-factor by analyzing how the parameters of Markov Random Field (i.e. Gibbs or log-linear) probability models of images are learned from data by maximum likelihood estimation. In particular, we study homogeneous MRF models which learn im(cid:173) age distributions in terms of clique potentials corresponding to fea(cid:173) ture histogram statistics (d. Minimax Entropy Learning (MEL) by Zhu, Wu and Mumford 1997 [11]) . We first use our analysis of the g-factor to determine when the clique potentials decouple for different features . Second, we show that clique potentials can be computed analytically by approximating the g-factor. Third, we demonstrate a connection between this approximation and the Generalized Iterative Scaling algorithm (GIS), due to Darroch and Ratcliff 1972 [2], for calculating potentials. This connection en(cid:173) ables us to use GIS to improve our multinomial approximation, using Bethe-Kikuchi[8] approximations to simplify the GIS proce(cid:173) dure. We support our analysis by computer simulations. James M. Coughlan, Alan L. Yuille |
NIPS | 1 |
| 2001 | Order Parameters for Detecting Target Curves in Images: When Does High Level Knowledge Help?
Alan L. Yuille, James M. Coughlan, Ying Nian Wu, Song-Chun Zhu |
Int. J. Comput. Vis. | 2 |
| 2000 | Order Parameters for Minimax Entropy Distributions: When Does High Level Knowledge Help?abstractMany problems in vision can be formulated as Bayesian inference. It is important to determine the accuracy of these inferences and how they depend on the problem domain. In recent work, Coughlan and Yuille showed that, for a restricted class of problems, the performance of Bayesian inference could be summarized by an order parameter K which depends on the probability distributions which characterize the problem domain. In this paper we generalize the theory of order parameters so that it applies to domains for which the probability models can be obtained by Minimax Entropy learning theory. By analyzing order parameters it is possible to determine whether a target can be detected using a general purpose "generic" model or whether a more specific "high-level" model is needed. At critical values of the order parameters the problem becomes unsolvable without the addition of extra prior knowledge. Alan L. Yuille, James M. Coughlan, Song-Chun Zhu, Ying Nian Wu |
CVPR | 2 |
| 2000 | The Manhattan World Assumption: Regularities in Scene Statistics which Enable Bayesian InferenceabstractPreliminary work by the authors made use of the so-called "Man(cid:173) hattan world" assumption about the scene statistics of city and indoor scenes. This assumption stated that such scenes were built on a cartesian grid which led to regularities in the image edge gra(cid:173) dient statistics. In this paper we explore the general applicability of this assumption and show that, surprisingly, it holds in a large variety of less structured environments including rural scenes. This enables us, from a single image, to determine the orientation of the viewer relative to the scene structure and also to detect target ob(cid:173) jects which are not aligned with the grid. These inferences are performed using a Bayesian model with probability distributions (e.g. on the image gradient statistics) learnt from real data. James M. Coughlan, Alan L. Yuille |
NIPS | 1 |
| 2000 | Efficient Deformable Template Detection and Localization without User Initialization
James M. Coughlan, Alan L. Yuille, Camper English, Daniel Snow |
Comput. Vis. Image Underst. | 1 |
| 2000 | Fundamental Limits of Bayesian Inference: Order Parameters and Phase Transitions for Road TrackingabstractThere is a growing interest in formulating vision problems in terms of Bayesian inference and, in particular, the maximum a posteriori (MAP) estimator. In this paper, we consider the special case of detecting roads from aerial images and demonstrate that analysis of this ensemble enables us to determine fundamental bounds on the performance of the MAP estimate. We demonstrate that there is a phase transition at a critical value of the order parameter; below this phase transition, it is impossible to detect the road by any algorithm. We derive closely related order parameters which determine the time and memory complexity of search and the accuracy of the solution using the n* search strategy. Our approach can be applied to other vision problems, and we briefly summarize the results when the model uses the "wrong prior". We comment on how our work relates to studies of the complexity of visual search and the critical behaviour in the computational cost of solving NP-complete problems. Alan L. Yuille, James M. Coughlan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | An A* perspective on deterministic optimization for deformable templates
Alan L. Yuille, James M. Coughlan |
Pattern Recognit. | 2 |
| 1999 | Fundamental Bounds on Edge Detection: An Information Theoretic Evaluation of Different Edge CuesabstractWe treat the problem of edge detection as one of statistical inference. Local edge cues, implemented by filters, provide information about the likely positions of edges which can be used as input to higher-level models. Different edge cues can be evaluated by the statistical effectiveness of their corresponding filters evaluated on a dataset of 100 presegmented images. We use information theoretic measures to determine the effectiveness of a variety of different edge detectors working at multiple scales on black and white and color images. Our results give quantitative measures for the advantages of multi-level processing, for the use of chromaticity in addition to greyscale, and for the relative effectiveness of different detectors. Scott Konishi, Alan L. Yuille, James M. Coughlan, Song-Chun Zhu |
CVPR | 3 |
| 1999 | High-Level and Generic Models for Visual Search: When Does High Level Knowledge Help?abstractWe analyze the problem of detecting a road target in background clutter and investigate the amount of prior (i.e. target specific) knowledge needed to perform this search task. The problem is formulated in terms of Bayesian inference and we define a Bayesian ensemble of problem instances. This formulation implies that the performance measures of different models depend on order parameters which characterize the problem. This demonstrates that if there is little clutter then only weak knowledge about the target is required in order to detect the target. However at a critical value of the order parameters there is a phase transition and it becomes effectively impossible to detect the target unless high-level target specific knowledge is used. These phase transitions determine different regimes within which different search strategies will be effective. These results have implications for bottom-up and top-down theories of vision. Alan L. Yuille, James M. Coughlan |
CVPR | 2 |
| 1999 | Manhattan World: Compass Direction from a Single Image by Bayesian InferenceabstractWhen designing computer vision systems for the blind and visually impaired it is important to determine the orientation of the user relative to the scene. We observe that most indoor and outdoor (city) scenes are designed on a Manhattan three-dimensional grid. This Manhattan grid structure puts strong constraints on the intensity gradients in the image. We demonstrate an algorithm for detecting the orientation of the user in such scenes based on Bayesian inference using statistics which we have learnt in this domain. Our algorithm requires a single input image and does not involve pre-processing stages such as edge detection and Hough grouping. We demonstrate strong experimental results on a range of indoor and outdoor images. We also show that estimating the grid structure makes it significantly easier to detect target objects which are not aligned with the grid. James M. Coughlan, Alan L. Yuille |
ICCV | 1 |
| 1998 | Efficient Optimization of a Deformable Template Using Dynamic ProgrammingabstractA novel deformable template is presented which detects the boundary of an open hand in a grayscale image. A dynamic programming algorithm enhanced by pruning techniques finds the hand contour in the image in as little as 19 seconds without initialization by the user. The template is translation- and rotation-invariant and accommodates shape deformation, significant occlusion and background clutter, and the presence of multiple hands. James M. Coughlan, Alan L. Yuille, Camper English, Daniel Snow |
CVPR | 1 |
| 1998 | A Phase Space Approach to Minimax Entropy Learning and the Minutemax Approximations
James M. Coughlan, Alan L. Yuille |
NIPS | 1 |
| 1998 | Convergence Rates of Algorithms for Visual Search: Detecting Visual Contours
Alan L. Yuille, James M. Coughlan |
NIPS | 2 |
| 1997 | Spline-Based Image Registration
Richard Szeliski, James M. Coughlan |
Int. J. Comput. Vis. | 2 |
| 1994 | Hierarchical spline-based image registrationabstractThe problem of image registration subsumes a number of topics in multiframe image analysis, including the computation of optic flow (general pixel-based motion), stereo correspondence, structure from motion, and feature tracking. We present a new registration algorithm based on a spline representation of the displacement field which can be specialized to solve all of the above mentioned problems. In particular, we show how to compute local flow, global (parametric) flow, rigid flow resulting from camera egomotion, and multiframe versions of the above problems. Using a spline-based description of the flow removes the need for overlapping correlation windows, and produces an explicit measure of the correlation between adjacent flow estimates. We demonstrate our algorithm on multiframe image registration and the recovery of 3D projective scene geometry. We also provide results on a number of standard motion sequences.> Richard Szeliski, James M. Coughlan |
CVPR | 2 |