EDBT 2026 Demo / reviewers in the wild / expert
Charles V. Stewart
dblp:43/471
· DBLP profile ↗
72ranked-venue papers
16as first author
16since 2021 · last 2026
0000-0001-6532-6675ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 13 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 11 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Combining Animal Re-Identification Models to Address Small DatasetsabstractAbstract Recent advancements in the automatic re-identification of animal individuals from images have opened up new possibilities for studying wildlife through camera traps and citizen science projects. Existing methods leverage distinct and permanent visual body markings, such as fur patterns or scars, and typically employ one of two approaches: local features or end-to-end learning. The end-to-end learning-based methods outperform local feature-based methods given a sufficient amount of good-quality training data, but the challenge of gathering such datasets for wildlife animals means that local feature-based methods remain a more practical approach for many species. In this study, we aim to achieve two goals: (1) to obtain a better understanding of the impact of training-set size on animal re-identification, and (2) to explore ways to combine various methods to leverage the advantages of their approaches for re-identification. In the work, we conduct comprehensive experiments across six different methods and six animal species with various training set sizes. Furthermore, we propose a simple yet effective combination strategy and show that a properly selected method combinations outperform the individual methods with both small and large training sets up to 30%. Additionally, the proposed combination strategy offers a generalizable framework to improve accuracy across species and address the challenges posed by small datasets, which are common in ecological research. This work lays the foundation for more robust and accessible tools to support wildlife conservation, population monitoring, and behavioral studies. Aleksandr Algasov, Ekaterina A. Nepovinnykh, Fedor Zolotarev, Tuomas Eerola, Heikki Kälviäinen, Charles V. Stewart, Lasha Otarashvili, Jason Holmberg |
Int. J. Comput. Vis. | 6 |
| 2025 | Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained AnalysisabstractWe present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pretrained ViTs, such as DINO, have demonstrated remarkable capabilities in extracting localized, discriminative features. However, saliency maps like Grad-CAM often fail to identify these traits, producing blurred, coarse heatmaps that highlight entire objects instead. We propose a novel approach, Prompt Class Attention Map (Prompt-CAM), to address this limitation. Prompt-CAM learns class-specific prompts for a pre-trained ViT and uses the corresponding outputs for classification. To correctly classify an image, the true-class prompt must attend to unique image patches not present in other classes’ images (i.e., traits). As a result, the true class’s multi-head attention maps reveal traits and their locations. Implementation-wise, Prompt-CAM is almost a "free lunch," requiring only a modification to the prediction head of Visual Prompt Tuning (VPT). This makes Prompt-CAM easy to train and apply, in stark contrast to other interpretable methods that require designing specific models and training processes. Extensive empirical studies on a dozen datasets from various domains (e.g., birds, fishes, insects, fungi, flowers, food, and cars) validate the superior interpretation capability of Prompt-CAM. The source code and demo are available at https://github.com/Imageomics/Prompt_CAM. Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu, Kazi Sajeed Mehrab, Elizabeth G. Campolongo, Daniel I. Rubenstein, Charles V. Stewart, Anuj Karpatne, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao |
CVPR | 9 |
| 2025 | Re-identification of patterned animals by multi-image feature aggregation and geometric similarityabstractAbstract Image‐based re‐identification of animal individuals allows gathering of information such as population size and migration patterns of the animals over time. This, together with large image volumes collected using camera traps and crowdsourcing, opens novel possibilities to study animal populations. For many species, the re‐identification can be done by analysing the permanent fur, feather, or skin patterns that are unique to each individual. In this paper, the authors study pattern feature aggregation based re‐identification and consider two ways of improving accuracy: (1) aggregating pattern image features over multiple images and (2) combining the pattern appearance similarity obtained by feature aggregation and geometric pattern similarity. Aggregation over multiple database images of the same individual allows to obtain more comprehensive and robust descriptors while reducing the computation time. On the other hand, combining the two similarity measures allows to efficiently utilise both the local and global pattern features, providing a general re‐identification approach that can be applied to a wide variety of different pattern types. In the experimental part of the work, the authors demonstrate that the proposed method achieves promising re‐identification accuracies for Saimaa ringed seals and whale sharks without species‐specific training or fine‐tuning. Ekaterina A. Nepovinnykh, Veikka Immonen, Tuomas Eerola, Charles V. Stewart, Heikki Kälviäinen |
IET Comput. Vis. | 4 |
| 2025 | Adapting the Re-ID Challenge for Static SensorsabstractABSTRACT The Grévy's zebra, an endangered species native to Kenya and southern Ethiopia, has been the target of sustained conservation efforts in recent years. Accurately monitoring Grévy's zebra populations is essential for ecologists to evaluate ongoing conservation initiatives. Recently, in both 2016 and 2018, a full census of the Grévy's zebra population was enabled by the Great Grévy's Rally (GGR), a citizen science event that combines teams of volunteers to capture data with computer vision algorithms that help experts estimate the number of individuals in the population. A complementary, scalable, cost‐effective and long‐term Grévy's population monitoring approach involves deploying a network of camera traps, which we have done at the Mpala Research Centre in Laikipia County, Kenya. In both scenarios, a substantial majority of the images of zebras are not usable for individual identification due to ‘in‐the‐wild’ imaging conditions—occlusions from vegetation or other animals, oblique views, low image quality and animals that appear in the far background and are thus too small to identify. Camera trap images, without an intelligent human photographer to select the framing and focus on the animals of interest, are of even poorer quality, with high rates of occlusion and high spatiotemporal similarity within image bursts. We employ an image filtering pipeline incorporating animal detection, species identification, viewpoint estimation, quality evaluation and temporal subsampling to compensate for these factors and obtain individual crops from camera trap and GGR images of suitable quality for re‐ID. We then employ the local clusterings and their alternatives (LCA) algorithm, a hybrid computer vision and graph clustering method for animal re‐ID, on the resulting high‐quality crops. Our method processed images taken during GGR‐16 and GGR‐18 in Meru County, Kenya, into 4142 highly comparable annotations, requiring only 120 contrastive same‐vs‐different‐individual decisions from a human reviewer to produce a population estimate of 349 individuals (within 4.6 of the ground truth count in Meru County). Our method also efficiently processed 8.9M unlabelled camera trap images from 70 camera traps at Mpala over 2 years into 685 encounters of 173 unique individuals, requiring only 331 contrastive decisions from a human reviewer. Avirath Sundaresan, Jason Parham, Jonathan P. Crall, Rosemary Warungu, Timothy Muthami, Jackson Miliko, Margaret Mwangi, Jason Holmberg, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Charles V. Stewart, Sara Beery |
IET Comput. Vis. | 11 |
| 2025 | BaboonLand Dataset: Tracking Primates in the Wild and Automating Behaviour Recognition from Drone Videos
Isla Duporge, Maksim Kholiavchenko, Roi Harel, Scott Wolf, Daniel I. Rubenstein, Margaret Crofoot, Tanya Y. Berger-Wolf, Stephen J. Lee, Julie Barreau, Jenna Kline, Michelle Ramirez, Charles V. Stewart |
Int. J. Comput. Vis. | 12 |
| 2025 | Correction: BaboonLand Dataset: Tracking Primates in the Wild and Automating Behaviour Recognition from Drone Videos
Isla Duporge, Maksim Kholiavchenko, Roi Harel, Scott Wolf, Daniel I. Rubenstein, Margaret Crofoot, Tanya Y. Berger-Wolf, Stephen J. Lee, Julie Barreau, Jenna Kline, Michelle Ramirez, Charles V. Stewart |
Int. J. Comput. Vis. | 12 |
| 2025 | Deep dive into KABR: a dataset for understanding ungulate behavior from in-situ drone video
Maksim Kholiavchenko, Jenna Kline, Maksim Kukushkin, Otto Brookes, Samuel Stevens 0001, Isla Duporge, Alec Sheets, Reshma Ramesh Babu, Namrata Banerji, Elizabeth G. Campolongo, Matthew J. Thompson, Nina Van Tiel, Jackson Miliko, Eduardo Bessa, Majid Mirmehdi, Thomas Schmid 0003, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Tilo Burghardt, Charles V. Stewart |
Multim. Tools Appl. | 20 |
| 2024 | Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
Vardaan Pahuja, Weidi Luo, Yu Gu 0016, Cheng-Hao Tu 0001, Hong-You Chen, Tanya Y. Berger-Wolf, Charles V. Stewart, Song Gao 0001, Wei-Lun Chao, Yu Su 0001 |
CIKM | 7 |
| 2024 | BioCLIP: A Vision Foundation Model for the Tree of LifeabstractImages of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an ex-plosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general or-ganismal biology questions on images is of timely need. To approach this, we curate and release Tree Of Life-10m, the largest and most diverse ML-ready dataset of biology images. We then develop Bioclip, a foundation model for the tree of life, leveraging the unique properties of bi-ology captured by Treeoflife-10m, namely the abun-dance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on di-verse fine-grained biology classification tasks and find that BloCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BloCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability.11imageomics.github.io/bioclip has models, data and code. Samuel Stevens 0001, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Carlyn, Wasila M. Dahdul, Charles V. Stewart, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001 |
CVPR | 9 |
| 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species EvolutionabstractAbstract A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion ) Mridul Khurana, Arka Daw, M. Maruf, Josef C. Uyeda, Wasila M. Dahdul, Caleb Charpentier, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Anuj Karpatne |
ECCV (89) | 13 |
| 2024 | A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisabstractWe present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn ''class-specific'' queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via ''multi-head'' cross-attention, INTR could identify different ''attributes'' of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR. Dipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang, David Carlyn, Samuel Stevens 0001, Kaiya Provost, Anuj Karpatne, Bryan Carstens, Daniel I. Rubenstein, Charles V. Stewart, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao |
ICLR | 11 |
| 2024 | Fine-Tuning is Fine, if CalibratedabstractFine-tuning is arguably the most straightforward way to tailor a pre-trained model (e.g., a foundation model) to downstream applications, but it also comes with the risk of losing valuable knowledge the model had learned in pre-training. For example, fine-tuning a pre-trained classifier capable of recognizing a large number of classes to master a subset of classes at hand is shown to drastically degrade the model's accuracy in the other classes it had previously learned. As such, it is hard to further use the fine-tuned model when it encounters classes beyond the fine-tuning data. In this paper, we systematically dissect the issue, aiming to answer the fundamental question, "What has been damaged in the fine-tuned model?" To our surprise, we find that the fine-tuned model neither forgets the relationship among the other classes nor degrades the features to recognize these classes. Instead, the fine-tuned model often produces more discriminative features for these other classes, even if they were missing during fine-tuning! What really hurts the accuracy is the discrepant logit scales between the fine-tuning classes and the other classes, implying that a simple post-processing calibration would bring back the pre-trained model's capability and at the same time unveil the feature improvement over all classes. We conduct an extensive empirical study to demonstrate the robustness of our findings and provide preliminary explanations underlying them, suggesting new directions for future theoretical analysis. Zheda Mai, Arpita Chowdhury, Ping Zhang 0016, Cheng-Hao Tu 0001, Hong-You Chen, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 9 |
| 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological ImagesabstractImages are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images. M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James P. Balhoff, Yasin Bakis, Bahadir Altintas, Matthew J. Thompson, Elizabeth G. Campolongo, Josef C. Uyeda, Hilmar Lapp, Henry L. Bart Jr., Paula M. Mabee, Yu Su 0001, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Wasila M. Dahdul, Anuj Karpatne |
NeurIPS | 19 |
| 2024 | Species-Agnostic Patterned Animal Re-identification by Aggregating Deep Local FeaturesabstractAbstract Access to large image volumes through camera traps and crowdsourcing provides novel possibilities for animal monitoring and conservation. It calls for automatic methods for analysis, in particular, when re-identifying individual animals from the images. Most existing re-identification methods rely on either hand-crafted local features or end-to-end learning of fur pattern similarity. The former does not need labeled training data, while the latter, although very data-hungry typically outperforms the former when enough training data is available. We propose a novel re-identification pipeline that combines the strengths of both approaches by utilizing modern learnable local features and feature aggregation. This creates representative pattern feature embeddings that provide high re-identification accuracy while allowing us to apply the method to small datasets by using pre-trained feature descriptors. We report a comprehensive comparison of different modern local features and demonstrate the advantages of the proposed pipeline on two very different species. Ekaterina A. Nepovinnykh, Ilja Chelak, Tuomas Eerola, Veikka Immonen, Heikki Kälviäinen, Maksim Kholiavchenko, Charles V. Stewart |
Int. J. Comput. Vis. | 7 |
| 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural NetworksabstractDiscovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne |
KDD | 15 |
| 2023 | Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target DataabstractWe propose a learning problem involving adapting a pre-trained source model to the target domain for classifying all classes that appeared in the source data, using target data that covers only a partial label space. This problem is practical, as it is unrealistic for the target end-users to collect data for all classes prior to adaptation. However, it has received limited attention in the literature. To shed light on this issue, we construct benchmark datasets and conduct extensive experiments to uncover the inherent challenges. We found a dilemma --- on the one hand, adapting to the new target domain is important to claim better performance; on the other hand, we observe that preserving the classification accuracy of classes missing in the target adaptation data is highly challenging, let alone improving them. To tackle this, we identify two key directions: 1) disentangling domain gradients from classification gradients, and 2) preserving class relationships. We present several effective solutions that maintain the accuracy of the missing classes and enhance the overall performance, establishing solid baselines for holistic transfer of pre-trained models with partial target data. Cheng-Hao Tu 0001, Hong-You Chen, Zheda Mai, Jike Zhong, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 8 |
| 2020 | Extracting identifying contours for African elephants and humpback whales using a learned appearance modelabstractThis paper addresses the problem of identifying individual animals in images based on extracting and matching contours, focusing in particular on the trailing edges of humpback whale flukes and the outline of the ears of African savanna elephants. A coarse-grained FCNN is learned to isolate the contour in an image, and a fine-grained FCNN is learned to provide more precise boundary information. The latter is trained by generating synthetic boundaries from coarse, easily-extracted training data, avoiding tedious manual effort. An A* algorithm extracts the final contour, which is converted to set of digital curvature descriptors and matched against a database of descriptors using local-naive Bayes nearest neighbors. We show that using the learned fine-grained FCNN produces more accurate contours than using image gradients for fine localization, especially for elephant ears where the boundaries are primarily texture. Matching using contours extracted using the fine-grained FCNN improves top-1 accuracy from 80% to 85% for flukes and 78% to 84% for ears. Hendrik J. Weideman, Charles V. Stewart, Jason R. Parham, Jason Holmberg, Kiirsten Flynn, John Calambokidis, D. Barry Paul, Anka Bedetti, Michelle Henley, Jerenimo Lepirei, Frank G. Pope |
WACV | 2 |
| 2018 | An Animal Detection Pipeline for IdentificationabstractThis paper proposes a 5-component detection pipeline for use in a computer vision-based animal recognition system. The end result of our proposed pipeline is a collection of novel annotations of interest (AoI) with species and view-point labels. These AoIs, for example, could be fed as the focused input data into an appearance-based animal identification system. The goal of our method is to increase the reliability and automation of animal censusing studies and to provide better ecological information to conservationists. Our method is able to achieve a localization mAP of 81.67%, a species and viewpoint annotation classification accuracy of 94.28% and 87.11%, respectively, and an AoI accuracy of 72.75% across 6 animal species of interest. We also introduce the Wildlife Image and Localization Dataset (WILD), which contains 5,784 images and 12,007 labeled annotations across 28 classification species and a variety of challenging, real-world detection scenarios. Jason R. Parham, Charles V. Stewart, Jonathan P. Crall, Daniel I. Rubenstein, Jason Holmberg, Tanya Y. Berger-Wolf |
WACV | 2 |
| 2013 | HotSpotter - Patterned species instance recognitionabstractWe present HotSpotter, a fast, accurate algorithm for identifying individual animals against a labeled database. It is not species specific and has been applied to Grevy's and plains zebras, giraffes, leopards, and lionfish. We describe two approaches, both based on extracting and matching keypoints or “hotspots”. The first tests each new query image sequentially against each database image, generating a score for each database image in isolation, and ranking the results. The second, building on recent techniques for instance recognition, matches the query image against the database using a fast nearest neighbor search. It uses a competitive scoring mechanism derived from the Local Naive Bayes Nearest Neighbor algorithm recently proposed for category recognition. We demonstrate results on databases of more than 1000 images, producing more accurate matches than published methods and matching each query image in just a few seconds. Jonathan P. Crall, Charles V. Stewart, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Siva R. Sundaresan |
WACV | 2 |
| 2013 | Automatic scallop detection in benthic environmentsabstractAs a multi-billion dollar industry, scallop fisheries world-wide rely on maintaining healthy off-shore populations. Recent developments in the collection of optical images from extended areas of the ocean floor has opened the possibility of assessing scallop populations from imagery. The shear volume of data - upwards of 20,000 images per hour - implies that automatic image analysis is necessary. This paper presents a computer vision software system to identify and count scallops. For each image, the system generates initial candidate regions of potential scallops, extracts image features in the candidate regions, and then applies one of several different trained Adaboost classifiers to determine the strength of each region as a scallop. In making the final classification decision, the strength of the scallop classifier output is compared to the output of other classifiers trained to detect sand dollars, clams and other “distractors”. Matthew Dawkins, Charles V. Stewart, Scott Gallager, Amber York |
WACV | 2 |
| 2012 | Physical Scale Keypoints: Matching and Registration for Combined Intensity/Range Images
Eric R. Smith, Richard J. Radke, Charles V. Stewart |
Int. J. Comput. Vis. | 3 |
| 2010 | Location registration and recognition (LRR) for serial analysis of nodules in lung CT scans
Michal Sofka, Charles V. Stewart |
Medical Image Anal. | 2 |
| 2008 | Location Registration and Recognition (LRR) for Longitudinal Evaluation of Corresponding Regions in CT Volumes
Michal Sofka, Charles V. Stewart |
MICCAI (2) | 2 |
| 2008 | Registration of combined range-intensity scans: Initialization through verification
Eric R. Smith, Bradford J. King, Charles V. Stewart, Richard J. Radke |
Comput. Vis. Image Underst. | 3 |
| 2008 | Automated Retinal Image Analysis Over the InternetabstractRetinal clinicians and researchers make extensive use of images, and the current emphasis is on digital imaging of the retinal fundus. The goal of this paper is to introduce a system, known as retinal image vessel extraction and registration system, which provides the community of retinal clinicians, researchers, and study directors an integrated suite of advanced digital retinal image analysis tools over the Internet. The capabilities include vasculature tracing and morphometry, joint (simultaneous) montaging of multiple retinal fields, cross-modality registration (color/red-free fundus photographs and fluorescein angiograms), and generation of flicker animations for visualization of changes from longitudinal image sequences. Each capability has been carefully validated in our previous research work. The integrated Internet-based system can enable significant advances in retina-related clinical diagnosis, visualization of the complete fundus at full resolution from multiple low-angle views, analysis of longitudinal changes, research on the retinal vasculature, and objective, quantitative computer-assisted scoring of clinical trials imagery. It could pave the way for future screening services from optometry facilities. Chia-Ling Tsai, Benjamin Madore, Matthew J. Leotta, Michal Sofka, Gehua Yang, Anna Majerovics, Howard L. Tanenbaum, Charles V. Stewart, Badrinath Roysam |
IEEE Trans. Inf. Technol. Biomed. | 8 |
| 2007 | Keypoint Descriptors for Matching Across Multiple Image Modalities and Non-linear Intensity VariationsabstractIn this paper, we investigate the effect of substantial inter-image intensity changes and changes in modality on the performance of keypoint detection, description, and matching algorithms in the context of image registration. In doing so, we modify widely-used keypoint descriptors such as SIFT and shape contexts, attempting to capture the insight that some structural information is indeed preserved between images despite dramatic appearance changes. These extensions include (a) pairing opposite-direction gradients in the formation of orientation histograms and (b) focusing on edge structures only. We also compare the stability of MSER, Laplacian-of-Gaussian, and Harris corner keypoint location detection and the impact of detection errors on matching results. Our experiments on multimodal image pairs and on image pairs with significant intensity differences show that indexing based on our modified descriptors produces more correct matches on difficult pairs than current techniques at the cost of a small decrease in performance on easier pairs. This extends the applicability of image registration algorithms such as the Dual-Bootstrap which rely on correctly matching only a small number of keypoints. Avi Kelman, Michal Sofka, Charles V. Stewart |
CVPR | 3 |
| 2007 | Simultaneous Covariance Driven Correspondence (CDC) and Transformation Estimation in the Expectation Maximization FrameworkabstractThis paper proposes a new registration algorithm, Co-variance Driven Correspondences (CDC), that depends fundamentally on the estimation of uncertainty in point correspondences. This uncertainty is derived from the covariance matrices of the individual point locations and from the covariance matrix of the estimated transformation parameters. Based on this uncertainty, CDC uses a robust objective function and an EM-like algorithm to simultaneously estimate the transformation parameters, their covariance matrix, and the likely correspondences. Unlike the Robust Point Matching (RPM) algorithm, CDC requires neither an annealing schedule nor an explicit outlier process. Experiments on synthetic and real images using a polynomial transformation models in 2D and in 3D show that CDC has a broader domain of convergence than the well-known Iterative Closest Point (ICP) algorithm and is more robust to missing or extraneous structures in the data than RPM. Michal Sofka, Gehua Yang, Charles V. Stewart |
CVPR | 3 |
| 2007 | Registration of Challenging Image Pairs: Initialization, Estimation, and DecisionabstractOur goal is an automated 2D-image-pair registration algorithm capable of aligning images taken of a wide variety of natural and man-made scenes as well as many medical images. The algorithm should handle low overlap, substantial orientation and scale differences, large illumination variations, and physical changes in the scene. An important component of this is the ability to automatically reject pairs that have no overlap or have too many differences to be aligned well. We propose a complete algorithm, including techniques for initialization, for estimating transformation parameters, and for automatically deciding if an estimate is correct. Keypoints extracted and matched between images are used to generate initial similarity transform estimates, each accurate over a small region. These initial estimates are rank-ordered and tested individually in succession. Each estimate is refined using the Dual-Bootstrap ICP algorithm, driven by matching of multiscale features. A three-part decision criteria, combining measurements of alignment accuracy, stability in the estimate, and consistency in the constraints, determines whether the refined transformation estimate is accepted as correct. Experimental results on a data set of 22 challenging image pairs show that the algorithm effectively aligns 19 of the 22 pairs and rejects 99.8% of the misalignments that occur when all possible pairs are tried. The algorithm substantially out-performs algorithms based on keypoint matching alone. Gehua Yang, Charles V. Stewart, Michal Sofka, Chia-Ling Tsai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Erratum to "Retinal Vessel Centerline Extraction Using Multiscale Matched Filters, Confidence and Edge Measures"abstractIn order to detect vessels at a variety of widths, we apply the matched filter at multiple scales (i.e., compute for multiple values and then combine the responses across scales). Unfortunately, the output amplitudes of spatial operators such as derivatives or matched filters generally decrease with increasing scale. To compensate for this effect, Lindeberg introduced gamma-normalized derivatives. We use this notion to define a gamma-normalized matched filter, Michal Sofka, Charles V. Stewart |
IEEE Trans. Medical Imaging | 2 |
| 2006 | Automatic robust image registration system: Initialization, estimation, and decisionabstractOur goal is a highly-reliable, fully-automated image registration technique that takes two images and correctly aligns them or decides that they can not be aligned. The technique should handle image pairs having low overlap, variations in scale, large illumination differences (e.g. day and night), substantial scene changes, and different modalities. Our approach is a combination of algorithms for initialization, estimation and refinement, and decision-making. It starts by extracting and matching keypoints. Rank-ordered matches are tested individually in succession. Each is used to generate a similarity transformation estimate in a small region of each image surrounding the matched keypoints. A generalization of the recently developed Dual-Bootstrap algorithm is then applied to generate an image-wide transformation estimate through a combination of matching and reestimation, model selection, and region growing, all driven by a new multiscale feature extraction technique. After convergence of the Dual-Bootstrap, the transformation is accepted if it passes a correctness test that combines measures of accuracy, stability and non-randomness; otherwise the process starts over with the next keypoint match. Experimental results on a suite of challenging image pairs shows the effectivenss of the complete system. Gehua Yang, Charles V. Stewart, Michal Sofka, Chia-Ling Tsai |
ICVS | 2 |
| 2006 | A Correspondence-Based Software Toolkit for Image RegistrationabstractThis paper presents a correspondence-based toolkit for image registration. Written in C++, the toolkit complements the capabilities of the insight toolkit (ITK). Major components include features, feature sets, match generators, error scale estimators, robust transformation estimators, and convergence testers, all combined and controlled by several different registration engines. Correspondence-based algorithms which can be implemented using the toolkit extend from ICP to hybrids of intensity-based and feature-based registration. The toolkit is being used both as an education tool and the foundation for developing new algorithms. Chia-Ling Tsai, Charles V. Stewart, A. G. Amitha Perera, Ying-Lin Lee, Gehua Yang, Michal Sofka |
SMC | 2 |
| 2006 | Retinal Vessel Centerline Extraction Using Multiscale Matched Filters, Confidence and Edge MeasuresabstractMotivated by the goals of improving detection of low-contrast and narrow vessels and eliminating false detections at nonvascular structures, a new technique is presented for extracting vessels in retinal images. The core of the technique is a new likelihood ratio test that combines matched-filter responses, confidence measures and vessel boundary measures. Matched filter responses are derived in scale-space to extract vessels of widely varying widths. A vessel confidence measure is defined as a projection of a vector formed from a normalized pixel neighborhood onto a normalized ideal vessel profile. Vessel boundary measures and associated confidences are computed at potential vessel boundaries. Combined, these responses form a six-dimensional measurement vector at each pixel. A training technique is used to develop a mapping of this vector to a likelihood ratio that measures the "vesselness" at each pixel. Results comparing this vesselness measure to matched filters alone and to measures based on the Hessian of intensities show substantial improvements, both qualitatively and quantitatively. The Hessian can be used in place of the matched filter to obtain similar but less-substantial improvements or to steer the matched filter by preselecting kernel orientations. Finally, the new vesselness likelihood ratio is embedded into a vessel tracing framework, resulting in an efficient and effective vessel centerline extraction algorithm. Michal Sofka, Charles V. Stewart |
IEEE Trans. Medical Imaging | 2 |
| 2005 | Teaching medical image analysis with the Insight Toolkit
Damion Shelton, George D. Stetten, Stephen R. Aylward, Luis Ibáñez, Aaron Cois, Charles V. Stewart |
Medical Image Anal. | 6 |
| 2004 | Multiple Kernel Tracking with SSD
Gregory D. Hager, Maneesh Dewan, Charles V. Stewart |
CVPR (1) | 3 |
| 2004 | Covariance-Driven Mosaic Formation from Sparsely-Overlapping Image Sets with Application to Retinal Image Mosaicing
Gehua Yang, Charles V. Stewart |
CVPR (1) | 2 |
| 2004 | An Uncertainty-Driven Hybrid of Intensity-Based and Feature-Based Registration with Application to Retinal and Lung CT Images
Charles V. Stewart, Ying-Lin Lee, Chia-Ling Tsai |
MICCAI (1) | 1 |
| 2004 | Model-Based Method for Improving the Accuracy and Repeatability of Estimating Vascular Bifurcations and Crossovers From Retinal Fundus ImagesabstractA model-based algorithm, termed exclusion region and position refinement (ERPR), is presented for improving the accuracy and repeatability of estimating the locations where vascular structures branch and cross over, in the context of human retinal images. The goal is two fold. First, accurate morphometry of branching and crossover points (landmarks) in neuronal/vascular structure is important to several areas of biology and medicine. Second, these points are valuable as landmarks for image registration, so improved accuracy and repeatability in estimating their locations and signatures leads to more reliable image registration for applications such as change detection and mosaicing. The ERPR algorithm is shown to reduce the median location error from 2.04 pixels down to 1.1 pixels, while improving the median spread (a measure of repeatability) from 2.09 pixels down to 1.05 pixels. Errors in estimating vessel orientations were similarly reduced from 7.2 degrees down to 3.8 degrees. Chia-Ling Tsai, Charles V. Stewart, Howard L. Tanenbaum, Badrinath Roysam |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2004 | Efficient Migration of Complex Off-Line Computer Vision Software to Real-Time System Implementation on Generic Computer HardwareabstractThis paper addresses the problem of migrating large and complex computer vision code bases that have been developed off-line, into efficient real-time implementations avoiding the need for rewriting the software, and the associated costs. Creative linking strategies based on Linux loadable kernel modules are presented to create a simultaneous realization of real-time and off-line frame rate computer vision systems from a single code base. In this approach, systemic predictability is achieved by inserting time-critical components of a user-level executable directly into the kernel as a virtual device driver. This effectively emulates a single process space model that is nonpreemptable, nonpageable, and that has direct access to a powerful set of system-level services. This overall approach is shown to provide the basis for building a predictable frame-rate vision system using commercial off-the-shelf hardware and a standard uniprocessor Linux operating system. Experiments on a frame-rate vision system designed for computer-assisted laser retinal surgery show that this method reduces the variance of observed per-frame central processing unit cycle counts by two orders of magnitude. The conclusion is that when predictable application algorithms are used, it is possible to efficiently migrate to a predictable frame-rate computer vision system. James Alexander Tyrrell, Justin M. LaPre, Christopher D. Carothers, Badrinath Roysam, Charles V. Stewart |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2003 | Disease-Oriented Evaluation of Dual-Bootstrap Retinal Image Registration
Chia-Ling Tsai, Anna Majerovics, Charles V. Stewart, Badrinath Roysam |
MICCAI (1) | 3 |
| 2003 | Frame-Rate Spatial Referencing Based on Invariant Indexing and Alignment with Application to Online Retinal Image RegistrationabstractThis paper describes an algorithm to continually and accurately estimate the absolute location of a diagnostic or surgical tool (such as a laser) pointed at the human retina, from a series of image frames. We treat the problem as a registration problem using diagnostic images to build a spatial map of the retina and then registering each online image against this map. Since the image location where the laser strikes the retina is easily found, this registration determines the position of the laser in the global coordinate system defined by the spatial map. For each online image, the algorithm computes similarity invariants, locally valid despite the curved nature of the retina, from constellations of vascular landmarks. These are detected using a high-speed algorithm that iteratively traces the blood vessel structure. Invariant indexing establishes initial correspondences between landmarks from the online image and landmarks stored in the spatial map. Robust alignment and verification steps extend the similarity transformation computed from these initial correspondences to a global, high-order transformation. In initial experimentation, the method has achieved 100 percent success on 1024 /spl times/ 1024 retina images. With a version of the tracing algorithm optimized for speed on 512 /spl times/ 512 images, the computation time is only 51 milliseconds per image on a 900MHz PentiumIII processor and a 97 percent success rate is achieved. The median registration error in either case is about 1 pixel. Hong Shen 0003, Charles V. Stewart, Badrinath Roysam, Howard L. Tanenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | The Dual Bootstrap Iterative Closest Point Algorithm with Application to Retinal Image RegistrationabstractMotivated by the problem of retinal image registration, this paper introduces and analyzes a new registration algorithm called Dual-Bootstrap Iterative Closest Point (Dual-Bootstrap ICP). The approach is to start from one or more initial, low-order estimates that are only accurate in small image regions, called bootstrap regions. In each bootstrap region, the algorithm iteratively: 1) refines the transformation estimate using constraints only from within the bootstrap region; 2) expands the bootstrap region; and 3) tests to see if a higher order transformation model can be used, stopping when the region expands to cover the overlap between images. Steps 1): and 3), the bootstrap steps, are governed by the covariance matrix of the estimated transformation. Estimation refinement [Step 2)] uses a novel robust version of the ICP algorithm. In registering retinal image pairs, Dual-Bootstrap ICP is initialized by automatically matching individual vascular landmarks, and it aligns images based on detected blood vessel centerlines. The resulting quadratic transformations are accurate to less than a pixel. On tests involving approximately 6000 image pairs, it successfully registered 99.5% of the pairs containing at least one common landmark, and 100% of the pairs containing at least one common landmark and at least 35% image overlap. Charles V. Stewart, Chia-Ling Tsai, Badrinath Roysam |
IEEE Trans. Medical Imaging | 1 |
| 2002 | A Feature-Based, Robust, Hierarchical Algorithm for Registering Pairs of Images of the Curved Human RetinaabstractThis paper describes a robust hierarchical algorithm for fully-automatic registration of a pair of images of the curved human retina photographed by a fundus microscope. Accurate registration is essential for mosaic synthesis, change detection, and design of computer-aided instrumentation. Central to the algorithm is a 12-parameter interimage transformation derived by modeling the retina as a rigid quadratic surface with unknown parameters. The parameters are estimated by matching vascular landmarks by recursively tracing the blood vessel structure. The parameter estimation technique, which could be generalized to other applications, is a hierarchy of models and methods, making the algorithm robust to unmatchable image features and mismatches between features caused by large interframe motions. Experiments involving 3,000 image pairs from 16 different healthy eyes were performed. Final registration errors less than a pixel are routinely achieved. The speed, accuracy, and ability to handle small overlaps compare favorably with retinal image registration techniques published in the literature. Ali Can, Charles V. Stewart, Badrinath Roysam, Howard L. Tanenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A Feature-Based Technique for Joint Linear Estimation of High-Order Image-to-Mosaic Transformations: Mosaicing the Curved Human RetinaabstractAn algorithm for constructing image mosaics from multiple, uncalibrated, weak-perspective views of the human retina is presented and analyzed. It builds on an algorithm for registering pairs of retinal images using a noninvertible, 12-parameter, quadratic image transformation model and hierarchical, robust estimation. The major innovation presented is a linear, feature-based, noniterative method for jointly estimating consistent transformations of all images onto the mosaic "anchor image." Constraints for this estimation are derived from pairwise registration both directly with the anchor image and indirectly between pairs of nonanchor images. An incremental, graph-based technique constructs the set of registered image pairs used in the solution. The estimation technique allows images that do not overlap the anchor frame to be successfully mosaiced, a valuable capability for mosaicing images of the retinal periphery. Experimental analysis on data sets from 16 eyes shows the average overall median transformation error in final mosaic to be 0.76 pixels. The technique is simpler, more accurate, and offers broader coverage than previously published methods. Ali Can, Charles V. Stewart, Badrinath Roysam, Howard L. Tanenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Frame-Rate Spatial Referencing Based on Invariant Indexing and Alignment with Application to Laser Retinal SurgeryabstractThe paper describes a fast feature-based algorithm for accurately estimating the absolute location of a surgical tool (a laser), pointed at the curved human retina, from a series of image frames. The method is capable of a 91% success rate with just a third of the feature extraction computation, taking 37 ms overall per image on a 900 MHz Pentium III. The success rate approaches 100% when the feature extraction is allowed to run to completion. The median error is 0.92 pixels with 512/spl times/512 8-bit image frames. Making a significant break from prior incremental tracking-based efforts, we propose a framework that involves extensive offline precomputation to build a "spatial map" and high-speed online "spatial referencing" to rapidly register each surgical image with this spatial map. The spatial referencing technique is designed around the idea of quasi-invariant indexing. Similarity invariants, locally valid despite the curved nature of the retina, are computed from constellations of vascular landmarks. These are detected using a high-speed algorithm that recursively traces the blood vessel structure. Invariant indexing establishes initial correspondences between landmarks from the online image and landmarks stored in the spatial map. Alignment and verification steps gradually extend the similarity transformation computed from these initial correspondences to a global, high-order transformation. The spatial map is precomputed to contain mosaics, distance maps and an invariant database, all designed to make these spatial referencing computations extremely fast. Hong Shen 0003, Charles V. Stewart, Badrinath Roysam, Howard L. Tanenbaum |
CVPR (1) | 2 |
| 2001 | Optimal scheduling of tracing computations for real-time vascular landmark extraction from retinal fundus imagesabstractRecently, this group published fast algorithms for automatic tracing (vectorization) of the vasculature in live retinal angiograms, and for the extraction of visual landmarks formed by vascular bifurcations and crossings. These landmarks are used for feature-based image matching for controlling a computer-assisted laser retinal surgery instrument currently under development. This paper describes methods to schedule the vascular tracing computations to maximize the rate of growth of quality of the partial tracing results within a frame cycle. There are two main advantages. First, progressive image matching from partially extracted landmark sets can be faster, and provide an earlier indication of matching failure. Second, the likelihood of successful image matching is greatly improved since the extracted landmarks are of the highest quality for the given computational budget. The scheduling method is based on quantitative measures for the computational work and the quality of landmarks. A coarse grid-based analysis of the image is used to generate seed points for the tracing computations, along with estimates of local edge strengths, orientations, and vessel thickness. These estimates are used to define criteria for real-time preemptive scheduling of the tracing computations. It is shown that the optimal schedule can only be achieved in perfect hindsight, and is thus unrealizable. This leads to scheduling heuristics that approximate the behavior of the optimal algorithm. One such approximation produced approximately 400% improvement in the quality of the partial results at a defined milestone, as compared to random scheduling. The resulting algorithm can be readily implemented on conventional and multiple-processor systems, and is being applied to computer-assisted laser retinal surgery. Hong Shen 0003, Badrinath Roysam, Charles V. Stewart, James N. Turner, Howard L. Tanenbaum |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2000 | A Feature-Based Technique for Joint, Linear Estimation of High-Order Image-to-Mosaic Transformations: Application to Mosaicing the Curved Human RetinaabstractMethods are presented for increasing the coverage and accuracy of image mosaics constructed from multiple, uncalibrated, weak-perspective views of the human retina. Extending our previous algorithm for registering pairs of images using a non-invertible, 12-parameter, quadratic image transformation model and a hierarchical, robust estimation technique, two important innovations are presented. The first is a linear, non-iterative method for jointly estimating the transformations of all images onto the mosaic. This employs constraints derived from pairwise matching between the non-mosaic image frames. It allows the transformations to be estimated for images that do not overlap the mosaic anchor frame, and results in mutually consistent transformations for all images. This means the mosaics can cover a much broader area of the retinal surface, even though the transformation model is not closed under composition. This capability is particularly valuable for mosaicing the retinal periphery in the context of diseases such as AIDS/CMV. The second innovation is a method to improve the accuracy of the pairwise matches as well as the joint estimation by refining the feature locations and by adding new features based on the transformation estimates themselves. For matching image frames of size 1024/spl times/1024, this cuts the registration error from the range of 1 to 3 pixels to about 0.55 pixels. The overall transformation error in final mosaic construction is 0.80 pixels based on experiments over a large set of eyes. Ali Can, Charles V. Stewart, Badrinath Roysam, Howard L. Tanenbaum |
CVPR | 2 |
| 2000 | Maintaining Valid Topology with Active Contours: Theory and ApplicationabstractWe develop and prove correct an algorithm that enables active contours to correctly represent regions that undergo topology changes as the contours evolve. Using the incremental motion typical of active contours, we introduce the concept of motion regions to determine the new topology. When the topology changes (e.g. by contours intersecting), the motion regions are used to delete and and reconnect the contours to accurately describe the new region. Contour intersections can also occur without topology changes. These are also appropriately handled. The algorithm to perform this task is proved correct in a general framework that makes few assumptions about the contour representation. We describe how this algorithm is applied to a polygonal representation of the contours, and argue that it does not significantly affect execution time. Finally, this polygonal implementation is used in surface extraction and phytoplankton classification. A. G. Amitha Perera, Chia-Ling Tsai, Robin Y. Flatland, Charles V. Stewart |
CVPR | 4 |
| 2000 | Extending range queries and nearest neighbors
Robin Y. Flatland, Charles V. Stewart |
Comput. Geom. | 2 |
| 2000 | Model Selection Techniques and Merging Rules for Range Data Segmentation Algorithms
Kishore Bubna, Charles V. Stewart |
Comput. Vis. Image Underst. | 2 |
| 2000 | Robust Computer Vision: An Interdisciplinary Challenge
Peter Meer, Charles V. Stewart, David E. Tyler |
Comput. Vis. Image Underst. | 2 |
| 2000 | Real time tracking of borescope tip pose
Ken Martin 0001, Charles V. Stewart |
Image Vis. Comput. | 2 |
| 1999 | Estimating Model Parameters and Boundaries By Minimizing a Joint, Robust Objective FunctionabstractMany problems in computer vision require estimation of both model parameters and boundaries, which limits the usefulness of standard estimation techniques from statistics. Example problems include surface reconstruction from range data, estimation of parametric motion models, fitting circular or elliptic arcs to edgel data, and many others. This paper introduces a new estimation technique, called the "Domain Bounding M-Estimator", which is a generalization of ordinary M-estimators combining error measures on model parameters and boundaries in a joint, robust objective function. Minimization of the objective function given a rough initialization yields simultaneous estimates of parameters and boundaries. The DBM-Estimator has been applied to estimating line segments, surfaces, and the symmetry transformation between two edgel chains. It is unaffected by outliers and prevents boundary estimates from crossing even small magnitude discontinuities. Charles V. Stewart, Kishore Bubna, A. G. Amitha Perera |
CVPR | 1 |
| 1999 | Robust Hierarchical Algorithm for Constructing a Mosaic from Images of the Curved Human RetinaabstractThis paper describes computer vision algorithms to assist in retinal laser surgery, which is widely used to treat leading blindness causing conditions but only has a 50% success rate, mostly due to a lack of spatial mapping and reckoning capabilities in current instruments. The novel technique described here automatically constructs a composite (mosaic) image of the retina from a sequence of incomplete views. This mosaic will be useful to ophthalmologists for both diagnosis and surgery. The new technique goes beyond published methods in both the medical and computer vision literatures because it is fully automated, models the patient-dependent curvature of the retina, handles large interframe motions, and does not require calibration. At the heart of the technique is a 12-parameter image transformation model derived by modeling the retina as a quadratic surface and assuming a weak perspective camera, and rigid motion. Estimating the parameters of this transformation model requires robustness to unmatchable image features and mismatches between features caused by large interframe motions. The described estimation technique is a hierarchy of models and methods: the initial match set is pruned based on a 0th order transformation estimated using a similarity-weighted histogram; a 1st order affine transformation is estimated using the reduced match set and least-median of squares; and the final, 2nd order 12-parameter transformation is estimated using an M-estimator initialized from the 1st order results. Initial experimental results show the method to be robust and accurate in accounting for the unknown retinal curvature in a fully automatic manner while preserving image details. Charles V. Stewart, Badrinath Roysam, Ali Can |
CVPR | 1 |
| 1998 | Model Selection and Surface Merging in Reconstruction AlgorithmsabstractThe problem of model selection is relevant to many areas of computer vision. Model selection criteria have been used in the vision literature and many more have been proposed in statistics, but the relative strengths of these criteria have not been analyzed in vision. More importantly, suitable extensions to these criteria must be made to solve problems unique to computer vision. Using the problem of surface reconstruction as our context, we analyze existing criteria using simulations and sensor data, introduce new criteria from statistics, develop novel criteria capable of handling unknown error distributions and outliers, and extend model selection criteria to apply to the surface merging problem. The new surface merging rules improve upon previous results, and work well even at small step heights (h=3/spl sigma/) and crease discontinuities. Our results show that a Bayesian criteria and its bootstrapped variant perform the best, although for time-sensitive applications, a variant of the Akaike criterion may be a better choice. Unfortunately, none of the criteria work reliably for small region sizes, implying that model selection and surface merging should be avoided unless the region size is sufficiently large. Kishore Bubna, Charles V. Stewart |
ICCV | 2 |
| 1998 | Recognition of Plane Projective SymmetryabstractA novel approach to grouping symmetrical planar curves under a projective transform is described. Symmetric curves are important as a generic model for object recognition where an object class is defined by the set of symmetries that any object in the class obeys. In this paper, a new algorithm is presented for grouping curves based on their correspondence under a plane projectivity. The correspondence between curves is established from an initial correspondence between two pairs of distinguished lines, such as lines tangent to inflection points. This initial correspondence leads to a reduced dimensional form for the projective mapping between the curves and a natural method for establishing correspondence between all points on the curves. A saliency measure is introduced which permits grouping results to be ordered in terms of the degree of symmetry supported by each curve pair. This saliency measure provides a basis for recognition in the case of approximate symmetry. Rupert W. Curwen, Charles V. Stewart, Joseph L. Mundy |
ICCV | 2 |
| 1997 | Prediction Intervals for Surface Growing Range SegmentationabstractThe surface growing framework presented by P. Besl and R. Jain (1988) has served as the basis for many range segmentation techniques. It has been augmented with alternative fitting techniques, model selection criteria, and solid modelling components. All of these approaches, however require global thresholds and large isolated seed regions. Range scenes typically do not satisfy the global threshold assumption since it requires data noise characteristics to be constant throughout the scene. Furthermore, as scene complexity increases, the number of surfaces, discontinuities, and outliers increase, hindering the identification of large seed regions. We present statistical criteria based on multivariate regression to replace the traditional decision criteria used in surface growing. We use local estimates and their uncertainties to construct criteria which capture the uncertainty in extrapolating estimated fits. We restrict surface expansion to very localized extrapolations, increasing the sensitivity to discontinuities and allowing regions to refine their estimates and uncertainties. Our approach uses a small number of parameters which are either statistical thresholds or cardinality measures, i.e. we do not use thresholds defined by specific range distances or orientation angles. James V. Miller, Charles V. Stewart |
CVPR | 2 |
| 1997 | Bias in robust estimation caused by discontinuities and multiple structuresabstractWhen fitting models to data containing multiple structures, such as when fitting surface patches to data taken from a neighborhood that includes a range discontinuity, robust estimators must tolerate both gross outliers and pseudo outliers. Pseudo outliers are outliers to the structure of interest, but inliers to a different structure. They differ from gross outliers because of their coherence. Such data occurs frequently in computer vision problems, including motion estimation, model fitting, and range data analysis. The focus in this paper is the problem of fitting surfaces near discontinuities in range data. To characterize the performance of least median of the squares, least trimmed squares, M-estimators, Hough transforms, RANSAC, and MINPRAN on this type of data, the "pseudo outlier bias" metric is developed using techniques from the robust statistics literature, and it is used to study the error in robust fits caused by distributions modeling various types of discontinuities. The results show each robust estimator to be biased at small, but substantial, discontinuities. They also show the circumstances under which different estimators are most effective. Most importantly, the results imply present estimators should be used with care, and new estimators should be developed. Charles V. Stewart |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | MUSE: Robust Surface Fitting using Unbiased Scale EstimatesabstractDespite many successful applications of robust statistics, they have yet to be completely adapted to many computer vision problems. Range reconstruction, particularly in unstructured environments, requires a robust estimator that not only tolerates a large outlier percentage but also tolerates several discontinuities, extracting multiple surfaces in an image region. Observing that random outliers and/or points from across discontinuities increase a hypothesized fit's scale estimate (standard deviation of the noise), our new operator; called MUSE (Minimum Unbiased Scale Estimator), evaluates a hypothesized fit over potential inlier sets via an objective function of unbiased scale estimates. MUSE extracts the single best fit from the data by minimizing its objective function over a set of hypothesized fits and can sequentially extract multiple surfaces from an image region. We show MUSE to be effective on synthetic data modelling small scale discontinuities and in preliminary experiments on complicated range data. James V. Miller, Charles V. Stewart |
CVPR | 2 |
| 1996 | Real time tracking of borescope tip poseabstractThe authors present a technique for the real-time tracking of borescope tip pose. While borescopes are used on a regular basis to inspect machinery for wear or damage, knowing the exact location of a borescope is difficult due to its flexibility. They present a technique for incremental borescope pose determination consisting of off-line feature extraction and on-line pose determination. The feature extraction precomputes from a CAD model of the object the features visible in a selected set of views. The on-line pose determination starts from a current pose estimate, determines the visible model features, projects them into a two-dimensional image coordinate system, matches each to the current borescope video image (without explicitly extracting features from this image), and uses the differences between the predicted and matched feature positions in a gradient descent technique to iteratively refine the pose estimate. The approach supports the mixed use of both matched feature positions and errors along the gradient within the pose determination. The on-line system is designed to execute at video frame rates, providing a continual indication of borescope tip pose. Ken Martin 0001, Charles V. Stewart, Rich Hammond |
WACV | 2 |
| 1996 | Geometric constraints and stereo disparity computation
Charles V. Stewart, Robin Y. Flatland, Kishore Bubna |
Int. J. Comput. Vis. | 1 |
| 1995 | Expected Performance of Robust Estimators Near DiscontinuitiesabstractIn extracting a polynomial surface patch near an intensity or range discontinuity, a robust estimator must tolerate not only the truly random bad data ("random outliers"), but also the coherently structured points ("pseudo outliers") that belong to a different surface. To characterize the performance of least median of squares, M estimators, Hough transforms, RANSAC, and MINPRAN on data containing both random and pseudo outliers, we develop two analytical measures, "pseudo outlier bias" and "pseudo outlier breakdown". Using these measures, we find that each robust estimator has surprisingly poor performance, even under the best possible circumstances, implying that present estimators should be used with care and new estimators should be developed.> Charles V. Stewart |
ICCV | 1 |
| 1995 | MINPRAN: A New Robust Estimator for Computer VisionabstractMINPRAN is a new robust estimator capable of finding good fits in data sets containing more than 50% outliers. Unlike other techniques that handle large outlier percentages, MINPRAN does not rely on a known error bound for the good data. Instead, it assumes the bad data are randomly distributed within the dynamic range of the sensor. Based on this, MINPRAN uses random sampling to search for the fit and the inliers to the fit that are least likely to have occurred randomly. It runs in time O(N/sup 2/+SN log N), where S is the number of random samples and N is the number of data points. We demonstrate analytically that MINPRAN distinguished good fits to random data and MINPRAN finds accurate fits and nearly the correct number of inliers, regardless of the percentage of true inliers. We confirm MINPRAN's properties experimentally on synthetic data and show it compares favorably to least median of squares. Finally, we apply MINPRAN to fitting planar surface patches and eliminating outliers in range data taken from complicated scenes.> Charles V. Stewart |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | A new robust operator for computer vision: theoretical analysisabstractMINPRAN, a new robust operator, finds good fits in data sets where more than 50% of the points are outliers. Unlike other techniques that handle large outlier percentages, MINPRAN does not rely on a known error bound for the good data. Instead it assumes that the bad data are randomly (uniformly) distributed within the dynamic range of the sensor. Based on this, MINPRAN uses random sampling to search for the fit and the number of inliers to the fit that are least likely to have occurred randomly. It runs in time O(N/sup 2/+SNlogN), where S is the number of random samples and N is the number of data points. We demonstrate analytically and experimentally that MINPRAN distinguishes good fits from fits to random data, and that MINPRAN finds accurate fits and nearly the correct number of inliers, regardless of the percentage of true inliers.> Charles V. Stewart |
CVPR | 1 |
| 1994 | A new robust operator for computer vision: application to range dataabstractThe basic MINPRAN (MINimize the Probability of RANdomness) technique, introduced by C.V. Stewart (1994), is extended to handle range data taken from complex scenes. Such data often includes: (1) a large numbers of outliers, (2) points from multiple surfaces interspersed over large image regions, and (3) extended regions containing only bad data. The initial version of MINPRAN handles cases (1) and (3). For (2), given an image region containing data from more than one surface, the basic technique tends to favor a single fit that "bridges" two surfaces. We analyze the extent of this problem and introduce two modifications to solve it. The new version of the algorithm, called MINPRAN2, produces extremely good results on difficult range data.> Charles V. Stewart |
CVPR | 1 |
| 1992 | Robust focus rangingabstractDepth maps obtained from focus ranging can have numerous errors and distortions due to edge bleeding, feature shifts, image noise, and field curvature. An improved algorithm that examines an initial high depth-of-field image of the scene to identify regions susceptible to edge bleeding and image noise is given. Focus evaluation windows are adapted to local image content and optimize the tradeoff between spatial resolution and noise sensitivity. An elliptical paraboloid field curvature model is used to reduce range distortion in peripheral image areas. Spatio-temporal tracking compensates for image feature shifts. The result is a sparse but reliable depth map.> Hari N. Nair, Charles V. Stewart |
CVPR | 2 |
| 1992 | On the derivation of geometric constraints in stereoabstractProbability density functions (PDFs) are derived for many of the geometric measurements upon which stereo matching techniques are based, including orientation differences between matching line segments or curves, the gradient of disparity, the directional derivative of disparity, and disparity differences between matches. The PDFs resulting from the transformations are used to critically examine many existing stereo techniques. Several techniques based on these PDFs are proposed.> Charles V. Stewart |
CVPR | 1 |
| 1991 | An analysis of the probability of disparity changes in stereo matching and a new algorithm based on the analysisabstractThe main contributions of this research are: (1) the derivation of the probability density function (pdf) of disparity changes in stereo matching based on the pdf of depth changes in the world and on the parameters of the stereo image formation process, (2) the definition of a match support equation based on the derived pdf, and (3) the incorporation of the support equation into a relaxation matching algorithm. The derived pdf and support equation are applicable to many existing stereo algorithms.> Charles V. Stewart |
CVPR | 1 |
| 1991 | Reducing the search time of a steady state genetic algorithm using the immigration operatorabstractAn examination is made of the fundamental trade-off between exploration and exploitation in a genetic algorithm (GA). An immigration operator is introduced that infuses random members into successive GA populations. It is theorized that immigration maintains much of the exploitation of the GA while increasing exploration. To test this theory, a set of functions that often require the GA to perform an excessive number of evaluations to find the global optimum of the function is designed. For These functions, it is shown experimentally that a GA enhanced with immigration (1) reduces the number of trials that require an excessive number of evaluations and (2) decreases the average number of evaluations needed to find the optimum function.> Michael C. Moed, Charles V. Stewart, Robert B. Kelley |
ICTAI | 2 |
| 1990 | Experimental analysis of a number of stereo matching components using LMAabstractA stereo matching algorithm is described that combines new techniques with features of existing algorithms. The final version of the algorithm, called LMA (local matching algorithm), includes (1) a modified disparity gradient support computation that gathers support for each candidate match by examining other candidate matches in an adaptively sized window surrounding that match, (2) a new form of multiresolution that derives support for each match by searching adjacent resolution levels for candidate matches that have similar positions, orientations, and disparities, and (3) a new final consistency check called the area rule that tests each high-confidence match to ensure consistency with at least a minimum percentage of its neighboring high-confidence matches. The algorithm is built incrementally, and the effectiveness of each proposed component of the algorithm is demonstrated experimentally.> Charles V. Stewart, Janice K. MacCrone |
ICPR (1) | 1 |
| 1988 | Local constraint integration in a connectionist model of stereo visionabstractThe authors' approach to stereo vision involves: (1) designing a general-purpose low-level matching algorithm: (2) testing it on a wide range of stereo pairs; and (3) adding mechanisms to the low-level algorithm that solve any remaining problems without interfering with its successful behavior. They begin by building the general support algorithm (GSA), a low-level matching algorithm that integrates the influence of a number of a locally defined constraints cooperatively and in parallel using only positive constraint influences (except for uniqueness). The constraints include uniqueness, coarse-to-fine and fine-to-coarse multiresolution, detailed match, figural continuity, and the disparity gradient. The GSA is implemented in a connectionist network. When tested on a wide-range of natural and synthetic images it produces a high percentage (97%) of correct matching decisions. The errors arise in situations that can not be identified using locally-defined constraints. These include partially occluded periodic regions, occlusions, and significant structural differences between the images.> Charles V. Stewart, Charles R. Dyer |
CVPR | 1 |
| 1988 | The Trinocular General Support Algorithm: A Three-camera Stereo Algorithm For Overcoming Binocular Matching ErrorsabstractThe combined use of binocular and new trinocular matching constraints in the Trinocular General Support Algorithm's (TGSA) parallel relaxation computation is shown to overcome many of the problems in binocular sterm matching. These problems include: (I) ambiguity in matching in periodic regions, especially when such a region is partially-occluded, (2) erroneous matches near occluded regions, and (3) missing and erronwus matches due to significant structural variations between the images. The TGSA employs cameras positioned at the vertices of an isosceles right triangle. Matching takes place between the horizontally-aligned pair of images and the vertically-aligned pair of images. Along with a variety of binocular constraints, new trinocular constraints, called trinocular uniqueness and the trinocular disparity gradient, are used to relate vertical and horizontal matches. When combined using a connectionist network relaxation algorithm, these constraints help to overcome the binocular matching problems listed above. For example, the trinocular disparity gradient provides enough information to directly resolve ambiguity in periodic regions in many cases. The TGSA has been tested on a number of image triples to demonstrate its advantages over previous binocular and trinocular stereo matching algorithms. Charles V. Stewart, Charles R. Dyer |
ICCV | 1 |
| 1988 | Scheduling Algorithms for PIPE (Pipelined Image-Processing Engine)
Charles V. Stewart, Charles R. Dyer |
J. Parallel Distributed Comput. | 1 |