EDBT 2026 Demo / reviewers in the wild / expert
Linda G. Shapiro
dblp:s/LindaGShapiro · also Linda Shapiro 0001
· DBLP profile ↗
155ranked-venue papers
20as first author
17since 2021 · last 2026
0000-0002-9495-0968ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 105 · 13 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bayesian Unsupervised Disentanglement of Anatomy and Geometry for Deep Groupwise Image RegistrationabstractThis article presents a general Bayesian learning framework for multi-modal groupwise image registration. The method builds on probabilistic modelling of the image generative process, where the underlying common anatomy and geometric variations of the observed images are explicitly disentangled as latent variables. Therefore, groupwise image registration is achieved via hierarchical Bayesian inference. We propose a novel hierarchical variational auto-encoding architecture to realise the inference procedure of the latent variables, where the registration parameters can be explicitly estimated in a mathematically interpretable fashion. Remarkably, this new paradigm learns groupwise image registration in an unsupervised closed-loop self-reconstruction process, sparing the burden of designing complex image-based similarity measures. The computationally efficient disentangled network architecture is also inherently scalable and flexible, allowing for groupwise registration on large-scale image groups with variable sizes. Furthermore, the inferred structural representations from multi-modal images via disentanglement learning are capable of capturing the latent anatomy of the observations with visual semantics. Extensive experiments were conducted to validate the proposed framework, including four different datasets from cardiac, brain, and abdominal medical images. The results have demonstrated the superiority of our method over conventional similarity-based approaches in terms of accuracy, efficiency, scalability, and interpretability. Xinzhe Luo, Xin Wang 0113, Linda G. Shapiro, Chun Yuan 0001, Jianfeng Feng, Xiahai Zhuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Unified and Semantically Grounded Domain Adaptation for Medical Image SegmentationabstractMost prior unsupervised domain adaptation approaches for medical image segmentation are narrowly tailored to either the source-accessible setting, where adaptation is guided by source-target alignment, or the source-free setting, which typically resorts to implicit adaptation mechanisms such as pseudo-labeling and network distillation. This substantial divergence in methodological designs between the two settings reveals an inherent flaw: the lack of an explicit, structured construction of anatomical knowledge that naturally generalizes across domains and settings. To bridge this longstanding divide, we introduce a unified, semantically grounded framework that supports both source-accessible and source-free adaptation. Fundamentally distinct from all prior works, our framework's adaptability emerges naturally as a direct consequence of the model architecture, without relying on explicit cross-domain alignment strategies. Specifically, our model learns a domain-agnostic probabilistic manifold as a global space of anatomical regularities, mirroring how humans establish visual understanding. Thus, the structural content in each image can be interpreted as a canonical anatomy retrieved from the manifold and a spatial transformation capturing individual-specific geometry. This disentangled, interpretable formulation enables semantically meaningful prediction with intrinsic adaptability. Extensive experiments on challenging cardiac and abdominal datasets show that our framework achieves state-of-the-art results in both settings, with source-free performance closely approaching its source-accessible counterpart, a level of consistency rarely observed in prior works. Beyond quantitative improvement, we demonstrate strong interpretability of the proposed framework via manifold traversal for smooth shape manipulation. The results provide a principled foundation for anatomically informed, interpretable, and unified solutions for domain adaptation in medical imaging. The code is available at https://github.com/wxdrizzle/remind. Xin Wang 0113, Jiamin Xia, Niranjan Balu, Mahmud Mossa-Basha, Linda G. Shapiro, Chun Yuan 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Perception Tokens Enhance Visual Reasoning in Multimodal Language ModelsabstractMultimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning about 2D object instances benefits from object detection. Yet, MLMs can not produce intermediate depth or boxes to reason over. Fine-tuning MLMs on relevant data doesn’t generalize well and outsourcing computation to specialized vision tools is too compute-intensive and memory-inefficient. To address this, we introduce Perception Tokens, intrinsic image representations designed to assist reasoning tasks where language is insufficient. Perception tokens act as auxiliary reasoning tokens, akin to chain-of-thought prompts in language models. For example, in a depth-related task, an MLM augmented with perception tokens can reason by generating a depth map as tokens, enabling it to solve the problem effectively. We propose Aurora, a training method that augments MLMs with perception tokens for improved reasoning over visual inputs. AURORA leverages a VQVAE to transform intermediate image representations, such as depth maps into a tokenized format and bounding box tokens, which are then used in a multi-task training framework. AURORA achieves notable improvements across counting benchmarks: +10.8% on BLINK, +11.3% on CVBench, and +8.3% on SEED-Bench, outperforming fine-tuning approaches in generalization across datasets. It also improves on relative depth: over +6% on BLINK. With perception tokens, Aurora expands the scope of MLMs beyond language-based reasoning, paving the way for more effective visual reasoning capabilities. Code and data will be released at the project page. Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G. Shapiro, Ranjay Krishna |
CVPR | 6 |
| 2025 | BADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan ReconstructionabstractReconstructing precise camera poses and floor plan layouts from wide-baseline RGB panoramas is a difficult and unsolved problem. We introduce BADGR, a novel diffusion model that jointly performs reconstruction and bundle adjustment (BA) to refine poses and layouts from a coarse state, using 1D floor boundary predictions from dozens of sparsely captured images. Unlike guided diffusion models, BADGR is conditioned on dense per-column outputs from a single-step Levenberg Marquardt (LM) optimizer and is trained to predict camera and wall positions, while minimizing reprojection errors for view consistency. The objective of layout generation from denoising diffusion process complements BA optimization by providing additional learned layout-structural constraints on top of the co-visible features across images. These constraints help BADGR make plausible guesses about spatial relationships, which constrain the pose graph, such as wall adjacency and collinearity, while also learning to mitigate errors from dense boundary observations using global context. BADGR trains exclusively on 2D floor plans, simplifying data acquisition, enabling robust augmentation, and supporting a variety of input densities. Our experiments validate our method, which significantly outperforms the state-of-the-art pose and floor plan layout reconstruction with different input densities. Visit project website at: https://badgr-diffusion.github.io. Yuguang Li, Ivaylo Boyadzhiev, Zixuan Liu 0001, Linda G. Shapiro, Alex Colburn |
CVPR | 4 |
| 2025 | PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to HistopathologyabstractDiagnosing diseases through histopathology whole slide images (WSIs) is fundamental in modern pathology but is challenged by the gigapixel scale and complexity of WSIs. Trained histopathologists overcome this challenge by navigating the WSI, looking for relevant patches, taking notes, and compiling them to produce a final holistic diagnostic. Traditional AI approaches, such as multiple instance learning and transformer-based models, fail short of such a holistic, iterative, multi-scale diagnostic procedure, limiting their adoption in the real-world. We introduce PathFinder, a multi-modal, multi-agent framework that emulates the decision-making process of expert pathologists. PathFinder integrates four AI agents, the Triage Agent, Navigation Agent, Description Agent, and Diagnosis Agent, that collaboratively navigate WSIs, gather evidence, and provide comprehensive diagnoses with natural language explanations. The Triage Agent classifies the WSI as benign or risky; if risky, the Navigation and Description Agents iteratively focus on significant regions, generating importance maps and descriptive insights of sampled patches. Finally, the Diagnosis Agent synthesizes the findings to determine the patient's diagnostic classification. Our Experiments show that PathFinder outperforms state-of-the-art methods in skin melanoma diagnosis by 8% while offering inherent explainability through natural language descriptions of diagnostically relevant patches. Qualitative analysis by pathologists shows that the Description Agent's outputs are of high quality and comparable to GPT-4o. PathFinder is also the first AI-based system to surpass the average performance of pathologists in this challenging melanoma classification task by 9%, setting a new record for efficient, accurate, and interpretable AI-assisted diagnostics in pathology. Data, code and models available at https://pathfinder-dx.github.io/ Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom Oluchi Ikezogwo, Beibin Li, Tejoram Vivekanandan, Joann G. Elmore, Ranjay Krishna, Linda G. Shapiro |
ICCV | 9 |
| 2025 | GrInAdapt: Source-Free Multi-target Domain Adaptation for Retinal Vessel Segmentation
Zixuan Liu 0001, Aaron Honjaya, Yuekai Xu, Hefu Pan, Xin Wang 0113, Linda G. Shapiro, Sheng Wang 0012, Ruikang K. Wang |
MICCAI (5) | 7 |
| 2024 | Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology VideosabstractDiagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multimodal models. Training multi-model models for histopathology requires instruction tuning datasets, which currently contain information for individual image patches, without a spatial grounding of the concepts within each patch and without a wider view of the WSI. To bridge this gap, we introduce QUILT-INSTRUCT, a large-scale dataset of107, 131 histopathology-specific instruction question/answer pairs, grounded within diagnostically relevant image patches that make up the WSI. Our dataset is collected by leveraging educational histopathology videos from YouTube, which provides spatial localization of narrations by automatically extracting the narrators' cursor positions. QUILT-INSTRUCT supports contextual reasoning by extracting diagnosis and supporting facts from the entire WSI. Using QUILT-INSTRUCT, we train QUILT-LLAVA, which can reason beyond the given single image patch, enabling diagnostic reasoning across patches. To evaluate QUILT-LLAVA, we propose a compre-hensive evaluation dataset created from 985 images and 1283 human-generated question-answers. We also thor-oughly evaluate QUILT-LLAVA using public histopathology datasets, where QUILT-LLAVA significantly outperforms SOTA by over 10% on relative GPT-4 score and 4% and 9% on open and closed set VQA11Our code, data, and model is publicly accessible at quilt-llava.github.io.. Mehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna, Linda G. Shapiro |
CVPR | 5 |
| 2024 | Semantics-Aware Attention Guidance for Diagnosing Whole Slide Images
Kechun Liu, Joann G. Elmore, Linda G. Shapiro |
MICCAI (5) | 4 |
| 2023 | Quilt-1M: One Million Image-Text Pairs for HistopathologyabstractRecent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online. However, the scarcity of analogous data in the medical field, specifically in histopathology, has slowed comparable progress. To enable similar representation learning for histopathology, we turn to YouTube, an untapped resource of videos, offering $1,087$ hours of valuable educational histopathology videos from expert clinicians.From YouTube, we curate QUILT: a large-scale vision-language dataset consisting of $802, 144$ image and text pairs.QUILT was automatically curated using a mixture of models, including large language models, handcrafted algorithms, human knowledge databases, and automatic speech recognition.In comparison, the most comprehensive datasets curated for histopathology amass only around $200$K samples.We combine QUILT with datasets from other sources, including Twitter, research papers, and the internet in general, to create an even larger dataset: QUILT-1M, with $1$M paired image-text samples, marking it as the largest vision-language histopathology dataset to date. We demonstrate the value of QUILT-1M by fine-tuning a pre-trained CLIP model. Our model outperforms state-of-the-art models on both zero-shot and linear probing tasks for classifying new histopathology images across $13$ diverse patch-level datasets of $8$ different sub-pathologies and cross-modal retrieval tasks. Wisdom Oluchi Ikezogwo, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Stefan Chan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, Linda G. Shapiro |
NeurIPS | 8 |
| 2023 | VSGD-Net: Virtual Staining Guided Melanocyte Detection on Histopathological ImagesabstractDetection of melanocytes serves as a critical prerequisite in assessing melanocytic growth patterns when diagnosing melanoma and its precursor lesions on skin biopsy specimens. However, this detection is challenging due to the visual similarity of melanocytes to other cells in routine Hematoxylin and Eosin (H&E) stained images, leading to the failure of current nuclei detection methods. Stains such as Sox10 can mark melanocytes, but they require an additional step and expense and thus are not regularly used in clinical practice. To address these limitations, we introduce VSGD-Net, a novel detection network that learns melanocyte identification through virtual staining from H&E to Sox10. The method takes only routine H&E images during inference, resulting in a promising approach to support pathologists in the diagnosis of melanoma. To the best of our knowledge, this is the first study that investigates the detection problem using image synthesis features between two distinct pathology stainings. Extensive experimental results show that our proposed model outperforms state-of-the-art nuclei detection methods for melanocyte detection. The source code and pre-trained model are available at: https://github.com/kechunl/VSGD-Net. Kechun Liu, Beibin Li, Caitlin J. May, Oliver Chang, Stevan Knezevich, Lisa M. Reisch, Joann G. Elmore, Linda G. Shapiro |
WACV | 9 |
| 2022 | Calibration Error Prediction: Ensuring High-Quality Mobile Eye-TrackingabstractGaze calibration is common in traditional infrared oculographic eye tracking. However, it is not well studied in visible-light mobile/remote eye tracking. We developed a lightweight real-time gaze error estimator and analyzed calibration errors from two perspectives: facial feature-based and Monte Carlo-based. Both methods correlated with gaze estimation errors, but the Monte Carlo method associated more strongly. Facial feature associations with gaze error were interpretable, relating movements of the face to the visibility of the eye. We highlight the degradation of gaze estimation quality in a sample of children with autism spectrum disorder (as compared to typical adults), and note that calibration methods may improve Euclidean error by 10%. Beibin Li, James C. Snider, Quan Wang 0003, Sachin Mehta, Claire E. Foster, Erin Barney, Linda G. Shapiro, Pamela Ventola, Frédérick Shic |
ETRA | 7 |
| 2022 | Brain-Aware Replacements for Supervised Contrastive Learning in Detection of Alzheimer's Disease
Mehmet Saygin Seyfioglu, Zixuan Liu 0001, Pranav Kamath, Sadjyot Gangolli, Sheng Wang 0012, Thomas J. Grabowski, Linda G. Shapiro |
MICCAI (1) | 7 |
| 2022 | End-to-End diagnosis of breast biopsy images with transformers
Sachin Mehta, Ximing Lu, Donald L. Weaver, Hannaneh Hajishirzi, Joann G. Elmore, Linda G. Shapiro |
Medical Image Anal. | 7 |
| 2021 | Semi-Supervised Synthesis of High-Resolution Editable Textures for 3D HumansabstractWe introduce a novel approach to generate diverse high fidelity texture maps for 3D human meshes in a semi-supervised setup. Given a segmentation mask defining the layout of the semantic regions in the texture map, our network generates high-resolution textures with a variety of styles, that are then used for rendering purposes. To accomplish this task, we propose a Region-adaptive Adversarial Variational AutoEncoder (ReAVAE) that learns the probability distribution of the style of each region individually so that the style of the generated texture can be controlled by sampling from the region-specific distributions. In addition, we introduce a data generation technique to augment our training set with data lifted from single-view RGB inputs. Our training strategy allows the mixing of reference image styles with arbitrary styles for different regions, a property which can be valuable for virtual try-on AR/VR applications. Experimental results show that our method synthesizes better texture maps compared to prior work while enabling independent layout and style controllability. Bindita Chaudhuri, Nikolaos Sarafianos, Linda G. Shapiro, Tony Tung |
CVPR | 3 |
| 2021 | Contextual Emotion Learning ChallengeabstractEmotion recognition via vision has been deeply associated with facial expressions, and the inference of emotions has, more often than not, been based on the same. However, context, both environmental and social, plays an imperative role in emotion recognition but has not been incorporated widely so far. The meaning of emotion might entirely switch when shifted from one setting to another if only facial expressions are taken into account. Moreover, there exists no study in the Indian context about the same. To cater to this issue, we generate and introduce the Indian Contextual Emotion Recognition (ICER) dataset based on the multi-ethnic Indian context. This paper summarises the Contextual Emotion Learning Challenge (CELC 2021) organized in conjunction with the 16th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2021. We outline the tasks posed in the challenge, the novel dataset, along with its challenges and the evaluation method. Lastly, we conclude by discussing the possible future directions. Jainendra Shukla, Puneet Gupta 0002, Aniket Bera, Arka Sarkar, Prakhar Goel, Shubhangi Butta, Anup Kumar Gupta 0001, Snehil Sanyal, Debanga Raj Neog, Manas Kamal Bhuyan, Kalyani Marathe, Linda G. Shapiro, Alex Colbrn, Varchita Lalwani |
FG | 12 |
| 2021 | Learning Oculomotor Behaviors from ScanpathabstractIdentifying oculomotor behaviors relevant for eye-tracking applications is a critical but often challenging task. Aiming to automatically learn and extract knowledge from existing eye-tracking data, we develop a novel method that creates rich representations of oculomotor scanpaths to facilitate the learning of downstream tasks. The proposed stimulus-agnostic Oculomotor Behavior Framework (OBF) model learns human oculomotor behaviors from unsupervised and semi-supervised tasks, including reconstruction, predictive coding, fixation identification, and contrastive learning tasks. The resultant pre-trained OBF model can be used in a variety of applications. Our pre-trained model outperforms baseline approaches and traditional scanpath methods in autism spectrum disorder and viewed-stimulus classification tasks. Ablation experiments further show our proposed method could achieve even better results with larger model sizes and more diverse eye-tracking training datasets, supporting the model’s potential for future eye-tracking applications. Open source code: http://github.com/BeibinLi/OBF. Beibin Li, Nicholas Nuechterlein, Erin Barney, Claire E. Foster, Minah Kim, Monique Mahony, Adham Atyabi, Quan Wang 0003, Pamela Ventola, Linda G. Shapiro, Frédérick Shic |
ICMI | 11 |
| 2021 | Deep Feature Representations for Variable-Sized Regions of Interest in Breast HistopathologyabstractOBJECTIVE: Modeling variable-sized regions of interest (ROIs) in whole slide images using deep convolutional networks is a challenging task, as these networks typically require fixed-sized inputs that should contain sufficient structural and contextual information for classification. We propose a deep feature extraction framework that builds an ROI-level feature representation via weighted aggregation of the representations of variable numbers of fixed-sized patches sampled from nuclei-dense regions in breast histopathology images. METHODS: First, the initial patch-level feature representations are extracted from both fully-connected layer activations and pixel-level convolutional layer activations of a deep network, and the weights are obtained from the class predictions of the same network trained on patch samples. Then, the final patch-level feature representations are computed by concatenation of weighted instances of the extracted feature activations. Finally, the ROI-level representation is obtained by fusion of the patch-level representations by average pooling. RESULTS: Experiments using a well-characterized data set of 240 slides containing 437 ROIs marked by experienced pathologists with variable sizes and shapes result in an accuracy score of 72.65% in classifying ROIs into four diagnostic categories that cover the whole histologic spectrum. CONCLUSION: The results show that the proposed feature representations are superior to existing approaches and provide accuracies that are higher than the average accuracy of another set of pathologists. SIGNIFICANCE: The proposed generic representation that can be extracted from any type of deep convolutional architecture combines the patch appearance information captured by the network activations and the diagnostic relevance predicted by the class-specific scoring of patches for effective modeling of variable-sized ROIs. Caner Mercan, Bulut Aygünes, Selim Aksoy, Ezgi Mercan, Linda G. Shapiro, Donald L. Weaver, Joann G. Elmore |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Personalized Face Modeling for Improved Face Reconstruction and Motion Retargeting
Bindita Chaudhuri, Noranart Vesdapunt, Linda G. Shapiro, Baoyuan Wang |
ECCV (5) | 3 |
| 2020 | Selection of Eye-Tracking Stimuli for Prediction by Sparsely Grouped Input Variables for Neural Networks: towards Biomarker Refinement for AutismabstractEye tracking has become a powerful tool in the study of autism spectrum disorder (ASD). Current, large-scale efforts aim to identify specific eye-tracking stimuli to be used as biomarkers for ASD, with the intention of informing the diagnostic process, monitoring therapeutic response, predicting outcomes, or identifying subgroups with the spectrum. However, there are hundreds of candidate experimental paradigms, each of which contains dozens or even hundreds of individual stimuli. Each stimuli is associated with an array of potential derived outcome variables, thus the number of variables to consider can be enormous. Standard variable selection techniques are not applicable to this problem, because selection must be done at the level of stimuli and not individual variables. In other words, this is a grouped variable selection problem. In this work, we apply lasso, group lasso, and a new technique, Sparsely Grouped Input Variables for Neural Network (SGIN), to select experimental stimuli for group discrimination and regression with clinical variables. Using a dataset obtained from children with and without ASD who were administered a battery containing 109 different stimuli presentations involving 9647 features, we are able to retain strong group separation even with only 11 out of the 109 stimuli. This work sets the stage for concerted techniques designed around engines to iteratively refine and define next-generation biomarkers using eye tracking for psychiatric conditions. http://github.com/beibinli/SGIN Beibin Li, Erin Barney, Caitlin Hudac, Nicholas Nuechterlein, Pamela Ventola, Linda G. Shapiro, Frédérick Shic |
ETRA | 6 |
| 2020 | Classifying Breast Histopathology Images with a Ductal Instance-Oriented PipelineabstractIn this study, we propose the Ductal Instance-Oriented Pipeline (DIOP) that contains a duct-level instance segmentation model, a tissue-level semantic segmentation model, and three-levels of features for diagnostic classification. Based on recent advancements in instance segmentation and the Mask RCNN model, our duct-level segmenter tries to identify each ductal individual inside a microscopic image; then, it extracts tissue-level information from the identified ductal instances. Leveraging three levels of information obtained from these ductal instances and also the histopathology image, the proposed DIOP outperforms previous approaches (both feature-based and CNN-based) in all diagnostic tasks; for the four-way classification task, the DIOP achieves comparable performance to general pathologists in this unique dataset. The proposed DIOP only takes a few seconds to run in the inference time, which could be used interactively on most modern computers. More clinical explorations are needed to study the robustness and generalizability of this system in the future. Beibin Li, Ezgi Mercan, Sachin Mehta, Stevan Knezevich, Corey W. Arnold, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
ICPR | 8 |
| 2020 | Leveraging Unlabeled Data for Glioma Molecular Subtype and Survival PredictionabstractIn this paper, we address two long-standing radio-genomic challenges in glioma subtype and survival prediction: (1) how to leverage large amounts of unlabeled magnetic resonance (MR) imaging data and (2) how to unite MR data and genomic data. We propose a novel application of multi-task learning (MTL) that leverages unlabeled MR data by jointly learning an auxiliary tumor segmentation task with glioma subtype prediction and that can learn from patients with and without genomic data. We analyze multi-parametric MR data from 542 patients in the combined training, validation, and testing sets of the 2018 Multimodal Brain Tumor Segmentation Challenge and somatic copy number alteration (SCNA) data from 1090 patients in The Cancer Genome Atlas' (TCGA) lower-grade glioma and glioblastoma projects. Our MTL model significantly outperforms comparable classification models trained only on labeled MR data for both IDH1/2 mutation and 1p/19q co-deletion subtype prediction tasks. We also show that embeddings produced by our MTL models improve survival predictions beyond MR or SCNA on their own. Our code is available at https://github.com/nknuecht/glioma_mtl. Nicholas Nuechterlein, Beibin Li, Mehmet Saygin Seyfioglu, Sachin Mehta, Patrick J. Cimino, Linda G. Shapiro |
ICPR | 6 |
| 2019 | ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural NetworkabstractWe introduce a light-weight, power efficient, and general purpose convolutional neural network, ESPNetv2, for modeling visual and sequential data. Our network uses group point-wise and depth-wise dilated separable convolutions to learn representations from a large effective receptive field with fewer FLOPs and parameters. The performance of our network is evaluated on four different tasks: (1) object classification, (2) semantic segmentation, (3) object detection, and (4) language modeling. Experiments on these tasks, including image classification on the ImageNet and language modeling on the PenTree bank dataset, demonstrate the superior performance of our method over the state-of-the-art methods. Our network outperforms ESPNet by 4-5% and has 2-4x fewer FLOPs on the PASCAL VOC and the Cityscapes dataset. Compared to YOLOv2 on the MS-COCO object detection, ESPNetv2 delivers 4.4% higher accuracy with 6x fewer FLOPs. Our experiments show that ESPNetv2 is much more power efficient than existing state-of-the-art efficient methods including ShuffleNets and MobileNets. Our code is open-source and available at https://github.com/sacmehta/ESPNetv2. Sachin Mehta, Mohammad Rastegari, Linda G. Shapiro, Hannaneh Hajishirzi |
CVPR | 3 |
| 2019 | A Facial Affect Analysis System for Autism Spectrum DisorderabstractIn this paper, we introduce an end-to-end machine learning-based system for classifying autism spectrum disorder (ASD) using facial attributes such as expressions, action units, arousal, and valence. Our system classifies ASD using representations of different facial attributes from convolutional neural networks, which are trained on images in the wild. Our experimental results show that different facial attributes used in our system are statistically significant and improve sensitivity, specificity, and F1 score of ASD classification by a large margin. In particular, the addition of different facial attributes improves the performance of ASD classification by about 7% which achieves a F1 score of 76%. Beibin Li, Sachin Mehta, Deepali Aneja, Claire E. Foster, Pamela Ventola, Frédérick Shic, Linda G. Shapiro |
ICIP | 7 |
| 2018 | ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation
Sachin Mehta, Mohammad Rastegari, Anat Caspi, Linda G. Shapiro, Hannaneh Hajishirzi |
ECCV (10) | 4 |
| 2018 | Efficient and Accurate Mitosis Detection - A Lightweight RCNN Approach
Yuguang Li, Ezgi Mercan, Stevan Knezevich, Joann G. Elmore, Linda G. Shapiro |
ICPRAM | 5 |
| 2018 | Automated Diagnosis of Breast Cancer and Pre-invasive Lesions on Digital Whole Slide Images
Ezgi Mercan, Sachin Mehta, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
ICPRAM | 6 |
| 2018 | Y-Net: Joint Segmentation and Classification for Diagnosis of Breast Biopsy Images
Sachin Mehta, Ezgi Mercan, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
MICCAI (2) | 6 |
| 2018 | Learning to Generate 3D Stylized Character Expressions from HumansabstractWe present ExprGen, a system to automatically generate 3D stylized character expressions from humans in a perceptually valid and geometrically consistent manner. Our multi-stage deep learning system utilizes the latent variables of human and character expression recognition convolutional neural networks to control a 3D animated character rig. This end-to-end system takes images of human faces and generates the character rig parameters that best match the human's facial expression. ExprGen generalizes to multiple characters, and allows expression transfer between characters in a semi-supervised manner. Qualitative and quantitative evaluation of our method based on Mechanical Turk tests show the high perceptual accuracy of our expression transfer results. Deepali Aneja, Bindita Chaudhuri, Alex Colburn, Gary Faigin, Linda G. Shapiro, Barbara Mones |
WACV | 5 |
| 2018 | Learning to Segment Breast Biopsy Whole Slide ImagesabstractWe trained and applied an encoder-decoder model to semantically segment breast biopsy images into biologically meaningful tissue labels. Since conventional encoderdecoder networks cannot be applied directly on large biopsy images and the different sized structures in biopsies present novel challenges, we propose four modifications: (1) an input-aware encoding block to compensate for information loss, (2) a new dense connection pattern between encoder and decoder, (3) dense and sparse decoders to combine multi-level features, (4) a multi-resolution network that fuses the results of encoder-decoders run on different resolutions. Our model outperforms a feature-based approach and conventional encoder-decoders from the literature. We use semantic segmentations produced with our model in an automated diagnosis task and obtain higher accuracies than a baseline approach that employs an SVM for featurebased segmentation, both using the same segmentationbased diagnostic features. Sachin Mehta, Ezgi Mercan, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
WACV | 6 |
| 2018 | Detection and classification of cancer in whole slide breast histopathology images using deep convolutional networks
Baris Gecer, Selim Aksoy, Ezgi Mercan, Linda G. Shapiro, Donald L. Weaver, Joann G. Elmore |
Pattern Recognit. | 4 |
| 2018 | Multi-Instance Multi-Label Learning for Multi-Class Classification of Whole Slide Breast Histopathology ImagesabstractDigital pathology has entered a new era with the availability of whole slide scanners that create the high-resolution images of full biopsy slides. Consequently, the uncertainty regarding the correspondence between the image areas and the diagnostic labels assigned by pathologists at the slide level, and the need for identifying regions that belong to multiple classes with different clinical significances have emerged as two new challenges. However, generalizability of the state-of-the-art algorithms, whose accuracies were reported on carefully selected regions of interest (ROIs) for the binary benign versus cancer classification, to these multi-class learning and localization problems is currently unknown. This paper presents our potential solutions to these challenges by exploiting the viewing records of pathologists and their slide-level annotations in weakly supervised learning scenarios. First, we extract candidate ROIs from the logs of pathologists' image screenings based on different behaviors, such as zooming, panning, and fixation. Then, we model each slide with a bag of instances represented by the candidate ROIs and a set of class labels extracted from the pathology forms. Finally, we use four different multi-instance multi-label learning algorithms for both slide-level and ROI-level predictions of diagnostic categories in whole slide breast histopathology images. Slide-level evaluation using 5-class and 14-class settings showed average precision values up to 81% and 69%, respectively, under different weakly labeled learning scenarios. ROI-level predictions showed that the classifier could successfully perform multi-class localization and classification within whole slide images that were selected to include the full range of challenging diagnostic categories. Caner Mercan, Selim Aksoy, Ezgi Mercan, Linda G. Shapiro, Donald L. Weaver, Joann G. Elmore |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Video to fully automatic 3D hair modelabstractImagine taking a selfie video with your mobile phone and getting as output a 3D model of your head (face and 3D hair strands) that can be later used in VR, AR, and any other domain. State of the art hair reconstruction methods allow either a single photo (thus compromising 3D quality) or multiple views, but they require manual user interaction (manual hair segmentation and capture of fixed camera views that span full 360°). In this paper, we describe a system that can completely automatically create a reconstruction from any video (even a selfie video), and we don't require specific views, since taking your -90°, 90°, and full back views is not feasible in a selfie capture. In the core of our system, in addition to the automatization components, hair strands are estimated and deformed in 3D (rather than 2D as in state of the art) thus enabling superior results. We provide qualitative, quantitative, and Mechanical Turk human studies that support the proposed system, and show results on a diverse variety of videos (8 different celebrity videos, 9 selfie mobile videos, spanning age, gender, hair length, type, and styling). Shu Liang, Xiufeng Huang, Xianyu Meng, Kunyao Chen, Linda G. Shapiro, Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 5 |
| 2017 | Closing the Loop for Edge Detection and Object ProposalsabstractEdge grouping and object perception are unified procedures in perceptual organization. However the computer vision literature classifies them as independent tasks. In this paper, we argue that edge detection and object proposals should benefit one another. To achieve this, we go beyond bounding boxes and extract closed contours that represent potential objects within. A novel objectness metric is proposed to score and rank the proposal boxes by considering the sizes and edge intensities of the closed contours. To improve the edge detector given the top-down object proposals, we group local closed contours and construct global object hierarchies and segmentations. The edge detector is retrained and enhanced using these hierarchical segmentations as additional feature channels. In the experiments we show that by closing the loop for edge detection and object proposals, we observe improvements for both tasks. Unifying edges and object proposals is valid and useful. Yao Lu 0028, Linda G. Shapiro |
AAAI | 2 |
| 2016 | Modeling Stylized Character Expressions via Deep Learning
Deepali Aneja, Alex Colburn, Gary Faigin, Linda G. Shapiro, Barbara Mones |
ACCV (2) | 4 |
| 2016 | Coherent Parametric Contours for Interactive Video Object SegmentationabstractInteractive video segmentation systems aim at producing sub-pixel-level object boundaries for visual effect applications. Recent approaches mainly focus on using sparse user input (i.e. scribbles) for efficient segmentation, however, the quality of the final object boundaries is not satisfactory for the following reasons: (1) the boundary on each frame is often not accurate, (2) boundaries across adjacent frames wiggle around inconsistently, causing temporal flickering, and (3) there is a lack of direct user control for fine tuning. We propose Coherent Parametric Contours, a novel video segmentation propagation framework that addresses all the above issues. Our approach directly models the object boundary using a set of parametric curves, providing direct user controls for manual adjustment. A spatiotemporal optimization algorithm is employed to produce object boundaries that are spatially accurate and temporally stable. We show that existing evaluation datasets are limited and demonstrate a new set to cover the common cases in professional rotoscoping. A new metric for evaluating temporal consistency is proposed. Results show that our approach generates higher quality, more coherent segmentation results than previous methods. Yao Lu 0028, Linda G. Shapiro, Jue Wang 0001 |
CVPR | 3 |
| 2016 | Head Reconstruction from Internet Photos
Shu Liang, Linda G. Shapiro, Ira Kemelmacher-Shlizerman |
ECCV (2) | 2 |
| 2015 | Automated Detection of 3D Landmarks for the Elimination of Non-Biological Variation in Geometric Morphometric AnalysesabstractLandmark-based morphometric analyses are used by anthropologists, developmental and evolutionary biologists to understand shape and size differences (eg. in the cranioskeleton) between groups of specimens. The standard, labor intensive approach is for researchers to manually place landmarks on 3D image datasets. As landmark recognition is subject to inaccuracies of human perception, digitization of landmark coordinates is typically repeated (often by more than one person) and the mean coordinates are used. In an attempt to improve efficiency and reproducibility between researchers, we have developed an algorithm to locate landmarks on CT mouse hemi-mandible data. The method is evaluated on 3D meshes of 28-day old mice, and results compared to landmarks manually identified by experts. Quantitative shape comparison between two inbred mouse strains demonstrate that data obtained using our algorithm also has enhanced statistical power when compared to data obtained by manual landmarking. Deepali Aneja, Siddharth R. Vora, Esra D. Camci, Linda G. Shapiro, Timothy C. Cox |
CBMS | 4 |
| 2015 | Generating Notifications for Missing Actions: Don't Forget to Turn the Lights Off!abstractWe all have experienced forgetting habitual actions among our daily activities. For example, we probably have forgotten to turn the lights off before leaving a room or turn the stove off after cooking. In this paper, we propose a solution to the problem of issuing notifications on actions that may be missed. This involves learning about interdependencies between actions and being able to predict an ongoing action while segmenting the input video stream. In order to show a proof of concept, we collected a new egocentric dataset, in which people wear a camera while making lattes. We show promising results on the extremely challenging task of issuing correct and timely reminders. We also show that our model reliably segments the actions, while predicting the ongoing one when only a few frames from the beginning of the action are observed. The overall prediction accuracy is 46.2% when only 10 frames of an action are seen (2/3 of a sec). Moreover, the overall recognition and segmentation accuracy is shown to be 72.7% when the whole activity sequence is observed. Finally, the online prediction and segmentation accuracy is 68.3% when the prediction is made at every time step. Bilge Soran, Ali Farhadi, Linda G. Shapiro |
ICCV | 3 |
| 2014 | 3D Face Hallucination from a Single Depth FrameabstractWe present an algorithm that takes a single frame of a person's face from a depth camera, e.g., Kinect, and produces a high-resolution 3D mesh of the input face. We leverage a dataset of 3D face meshes of 1204 distinct individuals ranging from age 3 to 40, captured in a neutral expression. We divide the input depth frame into semantically significant regions (eyes, nose, mouth, cheeks) and search the database for the best matching shape per region. We further combine the input depth frame with the matched database shapes into a single mesh that results in a highresolution shape of the input person. Our system is fully automatic and uses only depth data for matching, making it invariant to imaging conditions. We evaluate our results using ground truth shapes, as well as compare to state-of-the-art shape estimation methods. We demonstrate the robustness of our local matching approach with high-quality reconstruction of faces that fall outside of the dataset span, e.g., faces older than 40 years old, facial expressions, and different ethnicities. Shu Liang, Ira Kemelmacher-Shlizerman, Linda G. Shapiro |
3DV | 3 |
| 2014 | Action Recognition in the Presence of One Egocentric and Multiple Static Cameras
Bilge Soran, Ali Farhadi, Linda G. Shapiro |
ACCV (5) | 3 |
| 2014 | Classifying Craniosynostosis with a 3D Projection-Based Feature Extraction SystemabstractCraniosynostosis, a disorder in which one or more fibrous joints of the skull fuse prematurely, causes skull deformity and is associated with increased intracranial pressure and developmental delays. Although clinicians can easily diagnose craniosynostosis and can classify its type, being able to quantify the condition is an important problem in craniofacial research. While several papers have attempted this quantification through statistical models, the methods have not been intuitive to biomedical researchers and clinicians who want to use them. The goal of this work was to develop a general platform upon which new quantification measures could be developed and tested. The features reported in this paper were developed as basic shape measures, both single-valued and vector-valued, that are extracted from a single plane projection of the 3D skull. This technique allows us to process images that would otherwise be eliminated in previous systems due to poor resolution, noise or imperfections on their CT scans. We test our new features on classification tasks and also compare their performance to previous research. In spite of its simplicity, the classification accuracy of our new features is significantly higher than previous results on head CT scan data from the same research studies. Irma Lam, Michael L. Cunningham, Matthew L. Speltz, Linda G. Shapiro |
CBMS | 4 |
| 2014 | Localization of Diagnostically Relevant Regions of Interest in Whole Slide ImagesabstractWhole slide imaging technology enables pathologists to screen biopsy images and make a diagnosis in a digital form. This creates an opportunity to understand the screening patterns of expert pathologists and extract the patterns that lead to accurate and efficient diagnoses. For this purpose, we are taking the first step to interpret the recorded actions of world-class expert pathologists on a set of digitized breast biopsy images. We propose an algorithm to extract regions of interest from the logs of image screenings using zoom levels, time and the magnitude of panning motion. Using diagnostically relevant regions marked by experts, we use the visual bag-of-words model with texture and color features to describe these regions and train probabilistic classifiers to predict similar regions of interest in new whole slide images. The proposed algorithm gives promising results for detecting diagnostically relevant regions. We hope this attempt to predict the regions that attract pathologists' attention will provide the first step in a more comprehensive study to understand the diagnostic patterns in histopathology. Ezgi Mercan, Selim Aksoy, Linda G. Shapiro, Donald L. Weaver, Tad T. Brunyé, Joann G. Elmore |
ICPR | 3 |
| 2014 | Learning to Rank the Severity of Unrepaired Cleft Lip Nasal Deformity on 3D Mesh DataabstractCleft lip is a birth defect that results in deformity of the upper lip and nose. Its severity is widely variable and the results of treatment are influenced by the initial deformity. Objective assessment of severity would help to guide prognosis and treatment. However, most assessments are subjective. The purpose of this study is to develop and test quantitative computer-based methods of measuring cleft lip severity. In this paper, a grid-patch based measurement of symmetry is introduced, with which a computer program learns to rank the severity of cleft lip on 3D meshes of human infant faces. Three computer-based methods to define the midfacial reference plane were compared to two manual methods. Four different symmetry features were calculated based upon these reference planes, and evaluated. The result shows that the rankings predicted by the proposed features were highly correlated with the ranking orders provided by experts that were used as the ground truth. Jia Wu 0003, Raymond Tse, Linda G. Shapiro |
ICPR | 3 |
| 2013 | Supervised Semantic Gradient Extraction Using Linear-Time OptimizationabstractThis paper proposes a new supervised semantic edge and gradient extraction approach, which allows the user to roughly scribble over the desired region to extract semantically-dominant and coherent edges in it. Our approach first extracts low-level edge lets (small edge clusters) from the input image as primitives and build a graph upon them, by jointly considering both the geometric and appearance compatibility of edge lets. Given the characteristics of the graph, it cannot be effectively optimized by commonly-used energy minimization tools such as graph cuts. We thus propose an efficient linear algorithm for precise graph optimization, by taking advantage of the special structure of the graph. %Optimal parameter settings of the model are learnt from a dataset. Objective evaluations show that the proposed method significantly outperforms previous semantic edge detection algorithms. Finally, we demonstrate the effectiveness of the system in various image editing tasks. Shulin Yang, Jue Wang 0001, Linda G. Shapiro |
CVPR | 3 |
| 2012 | 3D shape isometric correspondence by spectral assignment
Linda G. Shapiro |
ICPR | 2 |
| 2012 | Tremor detection using motion filtering and SVM
Bilge Soran, Jenq-Neng Hwang, Su-In Lee, Linda G. Shapiro |
ICPR | 4 |
| 2012 | Unsupervised Template Learning for Fine-Grained Object RecognitionabstractFine-grained recognition refers to a subordinate level of recognition, such are recognizing different species of birds, animals or plants. It differs from recognition of basic categories, such as humans, tables, and computers, in that there are global similarities in shape or structure shared within a category, and the differences are in the details of the object parts. We suggest that the key to identifying the fine-grained differences lies in finding the right alignment of image regions that contain the same object parts. We propose a template model for the purpose, which captures common shape patterns of object parts, as well as the co-occurence relation of the shape patterns. Once the image regions are aligned, extracted features are used for classification. Learning of the template model is efficient, and the recognition results we achieve significantly outperform the state-of-the-art algorithms. Shulin Yang, Liefeng Bo, Jue Wang 0001, Linda G. Shapiro |
NIPS | 4 |
| 2011 | Groupwise pose normalization for craniofacial applicationsabstractA general framework is proposed for solving groupwise pose normalization problems and is analyzed in detail under different feature spaces. The analysis shows that using principal component analysis for pose normalization is a special case of using the proposed framework under a special feature space. The experimental results on two craniofacial datasets show the proposed method achieved promising results for solving groupwise pose normalization problems for craniofacial applications. Jiun-Hung Chen, Linda G. Shapiro |
WACV | 2 |
| 2011 | Stacked spatial-pyramid kernel: An object-class recognition method to combine scores from random treesabstractThe combination of local features, complementary feature types, and relative position information has been successfully applied to many object-class recognition tasks. Stacking is a common classification approach that combines the results from multiple classifiers, having the added benefit of allowing each classifier to handle a different feature space. However, the standard stacking method by its own nature discards any spatial information contained in the features, because only the combination of raw classification scores are input to the final classifier. The object-class recognition method proposed in this paper combines different feature types in a new stacking framework that efficiently quantizes input data and boosts classification accuracy, while allowing the use of spatial information. This classification method is applied to the task of automated insect-species identification for biomonitoring purposes. The test data set for this work contains 4722 images with 29 insect species, belonging to the three most common orders used to measure stream water quality, several of which are closely related and very difficult to distinguish. The specimens are in different 3D positions, different orientations, and different developmental and degradation stages with wide intra-class variation. On this very challenging data set, our new algorithm outperforms other classifiers, showing the benefits of using spatial information in the stacking framework with multiple dissimilar feature types. Natalia Larios, Junyuan Lin, Mengzi Zhang, David A. Lytle, Andrew Moldenke, Linda G. Shapiro, Thomas G. Dietterich |
WACV | 6 |
| 2011 | An ontology-based comparative anatomy information system
Ravensara S. Travillian, Kremena Diatchka, Tejinder K. Judge, Katarzyna Wilamowska, Linda G. Shapiro |
Artif. Intell. Medicine | 5 |
| 2010 | 3D Point Correspondence by Minimum Description Length in Feature Space
Jiun-Hung Chen, Ke Colin Zheng, Linda G. Shapiro |
ECCV (3) | 3 |
| 2010 | Robust interactive image segmentation with automatic boundary refinementabstractWe propose an effective image segmentation approach with a novel automatic boundary refinement procedure that requires little user interaction and makes the object cutout process more robust and convenient. It achieves these goals by the following three steps. First, merge over-segmented regions according to the maximal similarity rule using a few marking strokes as input. Second, detect possible erroneous low-contrast object boundaries by analyzing image content. Third, automatically refine those boundary regions using both local and global information. Experimental results are good even on very complex images. Dingding Liu, Yingen Xiong, Linda G. Shapiro, Kari Pulli |
ICIP | 3 |
| 2010 | The Use of Genetic Programming for Learning 3D Craniofacial Shape QuantificationsabstractCraniofacial disorders commonly result in various head shape dysmorphologies. The goal of this work is to quantify the various 3D shape variations that manifest in the different facial abnormalities in individuals with a craniofacial disorder called 22q11.2 Deletion Syndrome. Genetic programming (GP) is used to learn the different 3D shape quantifications. Experimental results show that the GP method achieves a higher classification rate than those of human experts and existing computer algorithms [1], [2]. Indriyati Atmosukarto, Linda G. Shapiro, Carrie Heike |
ICPR | 2 |
| 2010 | Haar Random Forest Features and SVM Spatial Matching Kernel for Stonefly Species IdentificationabstractThis paper proposes an image classification method based on extracting image features using Haar random forests and combining them with a spatial matching kernel SVM. The method works by combining multiple efficient, yet powerful, learning algorithms at every stage of the recognition process. On the task of identifying aquatic stonefly larvae, the method has state-of-the-art or better performance, but with much higher efficiency. Natalia Larios, Bilge Soran, Linda G. Shapiro, Gonzalo Martínez-Muñoz, Junyuan Lin, Thomas G. Dietterich |
ICPR | 3 |
| 2010 | 3D object classification using salient point patterns with application to craniofacial research
Indriyati Atmosukarto, Katarzyna Wilamowska, Carrie Heike, Linda G. Shapiro |
Pattern Recognit. | 4 |
| 2009 | Dictionary-free categorization of very similar objects via stacked evidence treesabstractCurrent work in object categorization discriminates among objects that typically possess gross differences which are readily apparent. However, many applications require making much finer distinctions. We address an insect categorization problem that is so challenging that even trained human experts cannot readily categorize images of insects considered in this paper. The state of the art that uses visual dictionaries, when applied to this problem, yields mediocre results (16.1% error). Three possible explanations for this are (a) the dictionaries are unsupervised, (b) the dictionaries lose the detailed information contained in each keypoint, and (c) these methods rely on hand-engineered decisions about dictionary size. This paper presents a novel, dictionary-free methodology. A random forest of trees is first trained to predict the class of an image based on individual keypoint descriptors. A unique aspect of these trees is that they do not make decisions but instead merely record evidence-i.e., the number of descriptors from training examples of each category that reached each leaf of the tree. We provide a mathematical model showing that voting evidence is better than voting decisions. To categorize a new image, descriptors for all detected keypoints are “dropped” through the trees, and the evidence at each leaf is summed to obtain an overall evidence vector. This is then sent to a second-level classifier to make the categorization decision. We achieve excellent performance (6.4% error) on the 9-class STONEFLY9 data set. Also, our method achieves an average AUC of 0.921 on the PASCAL06 VOC, which places it fifth out of 21 methods reported in the literature and demonstrates that the method also works well for generic object categorization. Gonzalo Martínez-Muñoz, Natalia Larios, Eric N. Mortensen, Wei Zhang 0014, Asako Yamamuro, Robert Paasch, Nadia Payet, David A. Lytle, Linda G. Shapiro, Sinisa Todorovic, Andrew Moldenke, Thomas G. Dietterich |
CVPR | 9 |
| 2008 | Medical image segmentation via min s-t cuts with sides constraintsabstractGraph cut algorithms (i.e., min s-t cuts) [3][10][15] are useful in many computer vision applications. In this paper we develop a formulation that allows the addition of side constraints to the min s-t cuts algorithm in order to improve its performance. We apply this formulation to foreground/background segmentation and provide empirical evidence to support its usefulness. From our experiments on medical image segmentation, the graph cut with constraints achieve significantly better performance than that without any constraint. Although the constrained min s-t cut problem is generally NP-hard, our approximation algorithm that uses linear programming relaxation and a simple rounding technique as a heuristic produces good results in a few seconds with our unoptimized code. Jiun-Hung Chen, Linda G. Shapiro |
ICPR | 2 |
| 2008 | Automated insect identification through concatenated histograms of local appearance features: feature vector generation and region detection for deformable objects
Natalia Larios, Hongli Deng, Wei Zhang 0014, Matt Sarpola, Jenny Yuen, Robert Paasch, Andrew Moldenke, David A. Lytle, Salvador Ruiz-Correa, Eric N. Mortensen, Linda G. Shapiro, Thomas G. Dietterich |
Mach. Vis. Appl. | 11 |
| 2007 | A Model Browser for Biosimulation
Gary Yngve, James F. Brinkley, Daniel L. Cook, Linda G. Shapiro |
AMIA | 4 |
| 2007 | Principal Curvature-Based Region Detector for Object RecognitionabstractThis paper presents a new structure-based interest region detector called principal curvature-based regions (PCBR) which we use for object class recognition. The PCBR interest operator detects stable watershed regions within the multi-scale principal curvature image. To detect robust watershed regions, we "clean" a principal curvature image by combining a grayscale morphological close with our new "eigenvectorflow" hysteresis threshold. Robustness across scales is achieved by selecting the maximally stable regions across consecutive scales. PCBR typically detects distinctive patterns distributed evenly on the objects and it shows significant robustness to local intensity perturbations and intra-class variations. We evaluate PCBR both qualitatively (through visual inspection) and quantitatively (by measuring repeatability and classification accuracy in real-world object-class recognition problems). Experiments on different benchmark datasets show that PCBR is comparable or superior to state-of-art detectors for both feature matching and object recognition. Moreover, we demonstrate the application of PCBR to symmetry detection. Hongli Deng, Wei Zhang 0014, Eric N. Mortensen, Thomas G. Dietterich, Linda G. Shapiro |
CVPR | 5 |
| 2007 | Automated Insect Identification through Concatenated Histograms of Local Appearance FeaturesabstractThis paper describes a fully automated stone fly-larvae classification system using a local features approach. It compares the three region detectors employed by the system: the Hessian-affine detector, the Kadir entropy detector and a new detector we have developed called the principal curvature based region detector (PCBR). It introduces a concatenated feature histogram (CFH) methodology that uses histograms of local region descriptors as feature vectors for classification and compares the results using this methodology to that of Opelt [Opelt, A, et.al., 2006.] on three stonefly identification tasks. Our results indicate that the PCBR detector outperforms the other two detectors on the most difficult discrimination task and that the use of all three detectors outperforms any other configuration. The CFH methodology also outperforms the Opelt methodology in these tasks Natalia Larios, Hongli Deng, Wei Zhang 0014, Matt Sarpola, Jenny Yuen, Robert Paasch, Andrew Moldenke, David A. Lytle, Ruiz Correa, Eric N. Mortensen, Linda G. Shapiro, Thomas G. Dietterich |
WACV | 11 |
| 2006 | A Graphical User Interface for a Comparative Anatomy Information System: Design, Implementation and Usage Scenarios
Ravensara S. Travillian, Kremena Diatchka, Tejinder K. Judge, Katarzyna Wilamowska, Linda G. Shapiro |
AMIA | 5 |
| 2006 | Automatic Segmentation of Neck CT ImagesabstractIn this era of cross-sectional imaging, it is useful to think of the neck in terms of adjacent anatomical spaces separated by fascial layers extended from the skull base to the thoracic inlet. Although these layers are not usually visible in CT or MR images, their locations can be inferred by knowing their relationships to various anatomical structures that are visible in cross-sectional images. Identifying these anatomical structures in the neck region can facilitate applications such as automated diagnosis by finding abnormality, or computer assisted radiation therapy planning by inferring cervical lymph node regions from these anatomical structures. We are proposing a system that can automatically segment the neck region from CT images by identifying the cranial and caudal bounding slices and then locating and labeling various anatomical structures in the region. Chia-Chi Teng, Linda G. Shapiro, Ira J. Kalet |
CBMS | 2 |
| 2006 | Symbolic Signatures for Deformable ShapesabstractRecognizing classes of objects from their shape is an unsolved problem in machine vision that entails the ability of a computer system to represent and generalize complex geometrical information on the basis of a finite amount of prior data. A practical approach to this problem is particularly difficult to implement, not only because the shape variability of relevant object classes is generally large, but also because standard sensing devices used to capture the real world only provide a partial view of a scene, so there is partial information pertaining to the objects of interest. In this work, we develop an algorithmic framework for recognizing classes of deformable shapes from range data. The basic idea of our component-based approach is to generalize existing surface representations that have proven effective in recognizing specific 3D objects to the problem of object classes using our newly introduced symbolic-signature representation that is robust to deformations, as opposed to a numeric representation that is often tied to a specific shape. Based on this approach, we present a system that is capable of recognizing and classifying a variety of object shape classes from range data. We demonstrate our system in a series of large-scale experiments that were motivated by specific applications in scene analysis and medical diagnosis. Salvador Ruiz-Correa, Linda G. Shapiro, Marina Meila, Gabriel Berson, Michael L. Cunningham, Raymond W. Sze |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | A Symbolic Shape-based Retrieval of Skull Images
H. Jill Lin, Salvador Ruiz-Correa, Linda G. Shapiro, Michael L. Cunningham, Raymond W. Sze |
AMIA | 3 |
| 2005 | Of Mice and Men: Design of a Comparative Anatomy Information System
Ravensara S. Travillian, John H. Gennari, Linda G. Shapiro |
AMIA | 3 |
| 2005 | Classifying Craniosynostosis Deformations by Skull Shape ImagingabstractCraniosynostosis is a serious and common disease of children, caused by premature fusion of the sutures of the skull. The resulting abnormal skull growth can lead to severe deformity, increased intra-cranial pressure, vision, hearing and breathing problems. In this work we develop an algorithmic framework to accurately classify deformations caused by sagittal craniosynostosis. The basic idea is to combine our novel cranial image shape descriptors and off-the-shelf classification technologies to encode morphological variations that characterize the synostotic skull. We demonstrate the efficacy of our approach in a series of large-scale classification experiments that compare the performance of our proposed image descriptors to those of traditional clinical indices and Fourier-based measurements. Salvador Ruiz-Correa, Raymond W. Sze, H. Jill Lin, Linda G. Shapiro, Matthew L. Speltz, Michael L. Cunningham |
CBMS | 4 |
| 2005 | A SIFT Descriptor with Global ContextabstractMatching points between multiple images of a scene is a vital component of many computer vision tasks. Point matching involves creating a succinct and discriminative descriptor for each point. While current descriptors such as SIFT can find matches between features with unique local neighborhoods, these descriptors typically fail to consider global context to resolve ambiguities that can occur locally when an image has multiple similar regions. This paper presents a feature descriptor that augments SIFT with a global context vector that adds curvilinear shape information from a much larger neighborhood, thus reducing mismatches when multiple local descriptors are similar. It also provides a more robust method for handling 2D nonrigid transformations since points are more effectively matched individually at a global scale rather than constraining multiple matched points to be mapped via a planar homography. We have tested our technique on various images and compare matching accuracy between the SIFT descriptor with global context to that without. Eric N. Mortensen, Hongli Deng, Linda G. Shapiro |
CVPR (1) | 3 |
| 2005 | A Generative/Discriminative Learning Algorithm for Image ClassificationabstractWe have developed a two-phase generative/discriminative learning procedure for the recognition of classes of objects and concepts in outdoor scenes. Our method uses both multiple types of object features and context within the image. The generative phase normalizes the description length of images, which can have an arbitrary number of extracted features of each type. In the discriminative phase, a classifier learns which images, as represented by this fixed-length description, contain the target object. We have tested the approach by comparing it to several other approaches in the literature and by experimenting with several different data sets and combinations of features. Our results, using color, texture, and structure features, show a significant improvement over previously published results in image retrieval. Using salient region features, we are competitive with recent results in object recognition. Linda G. Shapiro, Jeff A. Bilmes |
ICCV | 2 |
| 2005 | Application of Information Technology: Processes and Problems in the Formative Evaluation of an Interface to the Foundational Model of Anatomy Knowledge BaseabstractThe Digital Anatomist Foundational Model of Anatomy (FMA) is a large semantic network of more than 100,000 terms that refer to the anatomical entities, which together with 1.6 million structural relationships symbolically represent the physical organization of the human body. Evaluation of such a large knowledge base by domain experts is challenging because of the sheer size of the resource and the need to evaluate not just classes but also relationships. To meet this challenge, the authors have developed a relation-centric query interface, called Emily, that is able to query the entire range of classes and relationships in the FMA, yet is simple to use by a domain expert. Formative evaluation of this interface considered the ability of Emily to formulate queries based on standard anatomy examination questions, as well as the processing speed of the query engine. Results show that Emily is able to express 90% of the examination questions submitted to it and that processing time is generally 1 second or less, but can be much longer for complex queries. These results suggest that Emily will be a very useful tool, not only for evaluating the FMA, but also for querying and evaluating other large semantic networks. Linda G. Shapiro, Emily Chung, Landon Fridman Detwiler, José L. V. Mejino Jr., Augusto V. Agoncillo, James F. Brinkley, Cornelius Rosse |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | An Approach to the Anatomical Correlation of Species through the Foundational Model of Anatomy
Ravensara S. Travillian, Cornelius Rosse, Linda G. Shapiro |
AMIA | 3 |
| 2003 | A New Paradigm for Recognizing 3-D Object Shapes from Range DataabstractMost of the work on 3D object recognition from range data has used an alignment-verification approach in which a specific 3D object is matched to an exact instance of the same object in a scene. This approach has been successfully used in industrial machine vision, but it is not capable of dealing with the complexities of recognizing classes of similar objects. This paper undertakes this task by proposing and testing a component-based methodology encompassing three main ingredients: 1) a new way of learning and extracting shape-class components from surface shape information; 2) a new shape representation called a symbolic surface signature that summarizes the geometric relationships among components; and 3) an abstract representation of shape classes formed by a hierarchy of classifiers that learn object-class parts and their spatial relationships from examples. Salvador Ruiz-Correa, Linda G. Shapiro, Marina Meila |
ICCV | 2 |
| 2003 | Discriminating Deformable Shape ClassesabstractWe present and empirically test a novel approach for categorizing 3-D free form ob- ject shapes represented by range data . In contrast to traditional surface-signature based systems that use alignment to match specific objects, we adapted the newly introduced symbolic-signature representation to classify deformable shapes [10]. Our approach con- structs an abstract description of shape classes using an ensemble of classifiers that learn object class parts and their corresponding geometrical relationships from a set of numeric and symbolic descriptors. We used our classification engine in a series of large scale dis- crimination experiments on two well-defined classes that share many common distinctive features. The experimental results suggest that our method outperforms traditional numeric signature-based methodologies. 1 Salvador Ruiz-Correa, Linda G. Shapiro, Marina Meila, Gabriel Berson |
NIPS | 2 |
| 2003 | A hierarchical multiple classifier learning algorithm
Yu-Yu Chou, Linda G. Shapiro |
Pattern Anal. Appl. | 2 |
| 2003 | Estimating Piecewise-Smooth Optical Flow with Global Matching and Graduated OptimizationabstractThis paper presents a new method for estimating piecewise-smooth optical flow. We propose a global optimization formulation with three-frame matching and local variation and develop an efficient technique to minimize the resultant global energy. This technique takes advantage of local gradient, global gradient, and global matching methods and alleviates their limitations. Experiments on various synthetic and real data show that this method achieves highly competitive accuracy. Robert M. Haralick, Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | A rule-based model for local and regional tumor spread
Ira J. Kalet, Mark Whipple, Silvia Pessah, Jerry Barker, Mary M. Austin-Seymour, Linda G. Shapiro |
AMIA | 6 |
| 2002 | Head and neck lymph node region delineation with 3-D CT image registration
Chia-Chi Teng, Mary M. Austin-Seymour, Jerry Barker, Ira J. Kalet, Linda G. Shapiro, Mark Whipple |
AMIA | 5 |
| 2002 | Estimating optical flow using a global matching formulation and graduated optimizationabstractIn this paper we consider the problem of optimal optical flow estimation assuming brightness conservation and piecewise smoothness. We propose a formulation based on three-frame matching and global optimization allowing local variation. It is superior to popular gradient-based models and more justifiable than existing global methods. We also develop an efficient technique to minimize the resultant global energy. It takes advantage of local gradient, global gradient and global matching methods and overcomes their limitations. Experiments on various synthetic and real data and comparison with state-of-the-art techniques show that the method achieves discontinuity preserving capability and sub-pixel accuracy. Linda G. Shapiro, Robert M. Haralick |
ICIP (2) | 2 |
| 2001 | A New Signature-Based Method for Efficient 3-D Object RecognitionabstractThe paper considers the problem of shape-based recognition and pose estimation of 3D free-form objects in scenes that contain occlusion and clutter. Our approach is based on a novel set of discriminating descriptors called spherical spin images, which encode the shape information conveyed by classes of distributions of surface points constructed with respect to reference points on the surface of an object. The key to this approach is the relationship that exists between the l/sub 2/ metric, which compares n-dimensional signatures in Euclidean space, and the metric of the compact space on which the class representatives (spherical spin images) are defined. The connection allows us to efficiently utilize the linear correlation coefficient to discriminate scene points which have spherical spin images that are similar to the spherical spin images of points on the object being sought. The paper also addresses the problem of compressed spherical-spin-image representation by means of a random projection of the original descriptors that reduces the dimensionality without a significant loss of recognition/localization performance. Finally, the efficacy of the proposed representation is validated in a comparative study of the two algorithms presented that use uncompressed and compressed spherical spin images versus two previous spin image algorithms reported previously (A.E. Johnson and M. Hebert, 1999). The results of 2012 experiments suggest that the performance of our proposed algorithms is significantly better with respect to accuracy and speed than the performance of the other algorithms tested. Salvador Ruiz-Correa, Linda G. Shapiro, Marina Meila |
CVPR (1) | 2 |
| 2000 | A Hierarchical Multiple Classifier Learning AlgorithmabstractWe describe a hierarchical multiple classifier learning algorithm that was developed to provide a tool for classifier construction in medical applications. The mechanism uses data clustering and sub-class labeling to reduce the overhead of training and enhance the classification accuracy. The results of using this method are compared to a hand-built classifier and to other multiple classifier algorithms. Yu-Yu Chou, Linda G. Shapiro |
ICPR | 2 |
| 2000 | A Symbolic Representation for 3-D Object Feature DetectionabstractWe define a spatial symbolic model that can be used to describe classes of 3D objects (anatomical and man-made) and a method for finding correspondences between the features of the symbolic models and point sets of 3D mesh data. An abstract symbolic model is used to describe spatial object classes in terms of parts, boundaries, and spatial associations. A working model is a mechanism to link the symbolic model to geometric information found in a sensed instance of the class, represented by a 3D mesh data set. Matching is performed in a three-step procedure that first finds working sets of points in the mesh, then fits constructed features to these sets, and finally selects a subset of these constructed features that best correspond to the features of the working model. Pamela J. Neal, Linda G. Shapiro |
ICPR | 2 |
| 2000 | Surface Reconstruction and Display from Range and Color Data
Kari Pulli, Linda G. Shapiro |
Graph. Model. | 2 |
| 2000 | 3D Object Recognition and Pose with Relational IndexingabstractThis paper addresses the problem of recognizing 3D objects from 2D intensity images. It describes the object recognition system named RIO (relational indexing of objects), which contains a number of new techniques. RIO begins with an edge image obtained from a pair of intensity images taken with a single camera and two different lightings. From the edge image, a set of new high-level features and relationships are extracted, and a technique called relational indexing is used to efficiently recall 2D view-class object models that have similar relational descriptions from a potentially large database of models. Once a model has been hypothesized, pairs of 2D–3D corresponding features, including point pairs, line–segment pairs, and ellipse–circle pairs, are used in a new linear pose estimation framework to produce a hypothesized transformation from a 3D mesh model of the object to the image. The transformation is either accepted or rejected by a verification procedure that projects the 3D model wireframe to the image and computes a Hausdorff-like distance measure between the projected model and the edge image. The resultant object recognition system is able to recognize 3D objects having planar, cylindrical, and threaded surfaces in complex, multiobject scenes. Mauro S. Costa, Linda G. Shapiro |
Comput. Vis. Image Underst. | 2 |
| 1999 | A Flexible Image Database System for Content-Based Retrieval
Andrew Berman, Linda G. Shapiro |
Comput. Vis. Image Underst. | 2 |
| 1999 | An Integrated Linear Technique for Pose Estimation from Different Geometric FeaturesabstractExisting linear solutions for the pose estimation (or exterior orientation) problem suffer from a lack of robustness and accuracy partially due to the fact that the majority of the methods utilize only one type of geometric entity and their frameworks do not allow simultaneous use of different types of features. Furthermore, the orthonormality constraints are weakly enforced or not enforced at all. We have developed a new analytic linear least-squares framework for determining pose from multiple types of geometric features. The technique utilizes correspondences between points, between lines and between ellipse–circle pairs. The redundancy provided by different geometric features improves the robustness and accuracy of the least-squares solution. A novel way of approximately imposing orthonormality constraints on the sought rotation matrix within the linear framework is presented. Results from experimental evaluation of the new technique using both synthetic data and real images reveal its improved robustness and accuracy over existing direct methods. Mauro S. Costa, Robert M. Haralick, Linda G. Shapiro |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 1999 | 3D object identification with color and curvature signatures
Adnan A. Y. Mustafa, Linda G. Shapiro, Mark A. Ganter |
Pattern Recognit. | 2 |
| 1998 | The digital anatomist structural abstraction: a scheme for the spatial description of anatomical entities
Pamela J. Neal, Linda G. Shapiro, Cornelius Rosse |
AMIA | 2 |
| 1998 | The digital anatomist foundational model: principles for defining and structuring its concept domain
Cornelius Rosse, Linda G. Shapiro, James F. Brinkley |
AMIA | 2 |
| 1998 | A flexible image database system for content-based retrievalabstractThere is a growing need for the ability to query image databases based on similarity of image content rather than strict keyword search. As distance computations can be expensive, there is a need for indexing systems and algorithms that can eliminate candidate images without performing distance calculations. As user needs may change from session to session, there is also a need for run-time creation of distance measures. In this paper we present FIDS (flexible image database system). FIDS allows the user to query the database based on complex combinations of dozens of pre-defined distance measures. Using an indexing scheme and algorithms based on the triangle inequality, FIDS can often return matches to the query image without directly comparing the query image to more than a small percentage of the database. Andrew Berman, Linda G. Shapiro |
ICPR | 2 |
| 1998 | Probabilistic relational indexingabstractWe describe a new pattern matching methodology called probabilistic relational indexing that extends the work of Costa and Shapiro (1995, 1996) to handle uncertainty in pattern recognition. The new technique uses relational models, but avoids the complexity of full graph matching while incorporating probabilistic information that decreases the sensitivity to noise and errors in the data. The probabilistic relational indexing algorithm is compared to two popular decision tree classifiers and with the original discrete relational indexing algorithm. Yu-Yu Chou, Linda G. Shapiro |
ICPR | 2 |
| 1998 | Acquisition and visualization of colored 3D objectsabstractThis paper presents a complete system for scanning the geometry and surface color of a 3D object and for displaying realistic images of the object from arbitrary viewpoints. A stereo system with active light produces several views of dense range and color data. The data is registered and a surface that approximates the data is constructed. The surface estimate can be fairly coarse, as the appearance of fine detail is recreated by view-dependent texturing of the surface using color images. Kari Pulli, Habib Abi-Rached, Tom Duchamp, Linda G. Shapiro, Werner Stuetzle |
ICPR | 4 |
| 1997 | Object Identification with Surface Signatures
Adnan A. Y. Mustafa, Linda G. Shapiro, Mark A. Ganter |
CAIP | 2 |
| 1997 | PERFORM: A Fast Object Recognition Method Using Intersection of Projection Error RegionsabstractThis paper describes an object recognition methodology called PERFORM that finds matches by establishing correspondences between model and image features using this formulation. PERFORM evaluates correspondences by intersecting error regions in the image space. The algorithm is analyzed with respect to theoretical complexity as well as actual running times. When a single solution to the matching problem is sought, the time complexity of the sequential matching algorithm for 2D-2D matching using point features is of the order O(l/sup 3/ N/sup 2/), where N is the number of model features and l is the number of image features. When line features are used, the sequential complexity is of the order O(l/sup 2/ N/sup 2/). When a single solution is sought, PERFORM runs faster than the fastest known algorithm to solve the bounded-error matching problem. The PERFORM method is shown to be easily realizable on both SIMD and MIMD architectures. Bharath R. Modayur, Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | A Parallel Algorithm for Graph Matching and Its MasPar ImplementationabstractSearch of discrete spaces is important in combinatorial optimization. Such problems arise in artificial intelligence, computer vision, operations research, and other areas. For realistic problems, the search spaces to be processed are usually huge, necessitating long computation times, pruning heuristics, or massively parallel processing. We present an algorithm that reduces the computation time for graph matching by employing both branch-and-bound pruning of the search tree and massively-parallel search of the as-yet-unpruned portions of the space. Most research on parallel search has assumed that a multiple-instruction-stream/multiple-data-stream (MIMD) parallel computer is available. Since massively parallel stream (SIMD) computers are much less expensive than MIMD systems with equal numbers of processors, the question arises as to whether SIMD systems can efficiently handle state-space search problems. We demonstrate that the answer is yes, and in particular, that graph matching has a natural and efficient implementation on SIMD machines. Luigi Cinque, Steven L. Tanimoto, Linda G. Shapiro, Dean Yasuda |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 1996 | 3D matching using statistically significant groupingsabstractVision programming is defined as the task of constructing explicit object models to be used in object recognition. These object models specify the features to be used in recognizing the object as well as the exact order in which they have to be used. In this article, we describe a vision programming approach to matching 3D models to 2D images. Our system considers feature clusters instead of individual features and dynamically orders unmatched feature clusters based on the existing state of the match. The dynamic feature cluster ordering is achieved through the use of a new dynamic cost function. The automatic vision programming framework is general enough to be used by any feature-based recognition system, and in this article, it is shown to lead to dramatic improvements in the performance of a correspondence-based object recognition system. Bharath R. Modayur, Linda G. Shapiro |
ICPR | 2 |
| 1996 | 3D object recognition from color intensity imagesabstractWe describe a color-based identification system for 3D objects. Given a set of models with known attributes, and a scene containing one or more of these objects, the system identifies which objects are apparent in the scene. An important aspect of the system is that it integrates the use of curvature and spectral (color) attributes. Furthermore, the system employs sets of intensify images obtained through an inexpensive acquisition setup consisting of a single CCD camera and a set of three color filters. Surface signatures extracted from the scene, through color photometric stereo, are the main features employed for matching and identification. The system has been tested with excellent results on 95 observed-surfaces and 77 objects, appearing in 7 synthetic scenes and 14 real scenes. Adnan A. Y. Mustafa, Linda G. Shapiro, Mark A. Ganter |
ICPR | 2 |
| 1996 | Triplet-based object recognition using synthetic and real probability modelsabstractWe describe a model-based object recognition system that uses a probabilistic model for recognizing and locating objects. For each major view class of each 3D object, a probability model consisting of triplets of visible features, their parametrization, and their frequency of detection is constructed from a set of synthetic training images. These synthetic probability models are used to recognize and locate the 3D object from real 2D camera images. The features captured from the real images are then used to create a new, more accurate probability model. Kari Pulli, Linda G. Shapiro |
ICPR | 2 |
| 1996 | An improved algorithm for relational distance graph matching
Luigi Cinque, Dean Yasuda, Linda G. Shapiro, Steven L. Tanimoto |
Pattern Recognit. | 3 |
| 1995 | Optimal Sensor and Light Source Positioning for Machine Vision
Seungku Yi, Robert M. Haralick, Linda G. Shapiro |
Comput. Vis. Image Underst. | 3 |
| 1995 | Proteus: A reconfigurable computational network for computer vision
Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos |
Mach. Vis. Appl. | 6 |
| 1995 | Knowledge-based organ identification from CT images
Masaharu Kobashi, Linda G. Shapiro |
Pattern Recognit. | 2 |
| 1994 | Object representation for object recognitionabstractThis paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.> Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman |
CVPR | 9 |
| 1994 | Fast parallel object recognitionabstractThe problem of model-based object recognition is one of identifying occurrences of objects known a priori in an image. Not all the existing algorithms lend themselves well to parallel implementations. In this paper, we describe a new formulation of the recognition problem that is amenable to a naturally parallel solution. The method that we describe solves the bounded error recognition problem accurately by incorporating an explicit noise model. The time complexity of the sequential matching algorithm using point features is of the order O(I/sup 2/NI), where N is the number of model features and I is the number of image features. The corresponding parallel algorithm using O(I/sup 2/) processors has O(NI) complexity. When line features are used, the sequential complexity is of the order O(I/sup 2/N) and the parallel algorithm, utilizing O(I) processors has O(NI) complexity. Results are presented for a sequential version running on a Sun as well as a parallel version running on a 1024-processor MasPar MP-1. Bharath R. Modayur, Linda G. Shapiro |
ICPR (3) | 2 |
| 1994 | Three-dimensional shape from color photometric stereo
Per H. Christensen, Linda G. Shapiro |
Int. J. Comput. Vis. | 2 |
| 1994 | Error propagation in machine vision
Seungku Yi, Robert M. Haralick, Linda G. Shapiro |
Mach. Vis. Appl. | 3 |
| 1993 | Determining the shape of multi-colored dichromatic surface using color photometric stereoabstractThe utilization of color information to increase the accuracy of 3-D shapes derived from photometric stereo is described. Multicolored objects under white illumination are considered. Using color images rather than gray-scale makes it simple to separate the specular and diffuse reflection components. The diverse properties of diffusive and specular reflection yield supplementing constraints on the surface orientation, making the method less sensitive to noise.> Per H. Christensen, Linda G. Shapiro |
CVPR | 2 |
| 1993 | MUSER: A prototype musical score recognition system using mathematical morphology
Bharath R. Modayur, Visvanathan Ramesh, Robert M. Haralick, Linda G. Shapiro |
Mach. Vis. Appl. | 4 |
| 1993 | Representative patterns for model-based matching
Jorja G. Henikoff, Linda G. Shapiro |
Pattern Recognit. | 2 |
| 1992 | Visual inspection of machined partsabstractA CAD-model-based machine vision system for dimensional inspection of machine parts is described, with emphasis on the theory behind the system. The original contributions of this work are: (1) the use of precise definitions of geometric tolerances suitable for use in image processing, (2) the development of measurement algorithms corresponding directly to these definitions, (3) the derivation of the uncertainties in the measurement tasks, and (4) the use of this uncertainty information in the decision-making process. Initial experimental results have verified the uncertainty derivations statistically and proved that the error probabilities obtained by propagating uncertainties are lower than those obtainable without uncertainty propagation.> Bharath R. Modayur, Linda G. Shapiro, Robert M. Haralick |
CVPR | 2 |
| 1992 | Proteus: a reconfigurable computational network for computer visionabstractThe Proteus architecture is a highly parallel MIMD, multiple instruction, multiple-data machine, optimized for large granularity tasks such as machine vision and image processing. The system can achieve 20 Giga-flops (80 Giga-flops peak). It accepts data via multiple serial links at a rate of up to 640 megabytes/second. The system employs a hierarchical reconfigurable interconnection network with the highest level being a circuit switched Enhanced Hypercube serial interconnection network for internal data transfers. The system is designed to use 256 to 1024 RISC processors. The processors use one megabyte external Read/Write Allocating Caches for reduced multiprocessor contention. The system detects, locates, and replaces faulty subsystems using redundant hardware to facilitate fault tolerance.> Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos |
ICPR (4) | 6 |
| 1992 | Knowledge-based matching for 3D radiotherapy planningabstractDevelops a system that performs segmentation and recognition of major organs in CT images of human abdomen for 3D organ reconstruction used in radiotherapy. A knowledge-based system has been developed that features the use of constraint-based dynamic thresholding, negative-shape constraints to rapidly rule out infeasible segmentations, and progressive landmarking that takes advantage of the different degrees of certainty of unsuccessful identification of each organ. The results of the experiments indicate that the knowledge-based approach is promising.> Masaharu Kobashi, Linda G. Shapiro |
ICPR (1) | 2 |
| 1992 | Automated inspection of machine partsabstractDescribes a CAD-model-based machine vision system for dimensional inspection of machined parts, with emphasis on the theory behind the system. The original contributions of the work are: the use of precise definitions of geometric tolerances suitable for use in image processing, the development of measurement algorithms corresponding directly to these definitions; the derivation of the uncertainties in the measurement tasks; and the use of this uncertainty information in the decision-making process. Experimental results have verified the uncertainty derivations statistically and proved that the error probabilities obtained by propagating uncertainties are lower than those obtainable without uncertainty propagation.> Bharath R. Modayur, Linda G. Shapiro |
ICPR (1) | 2 |
| 1992 | Object Recognition Using Prediction And Probabilistic MatchabstractPREMIO is a CAD-based object recognition and localization system that uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the predictions to guide the search and reduce the search space. The purpose of this paper is to describe the matching algorithm and illustrative results. I Introduction Most feature-based matching schemes assume that all the features that are potentially visible in a view of an object will appear with equal probability. The resultant matching algorithms have to allow for “errors” without really understanding what the errors mean. PREMIO [2] is an object recognition/localization system that attempts to model some of the physical processes that can cause these “errors”. It uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. From these predictions, PREMIO calculates probabilities for each feature of being detected as a whole, being missed entirely, or breaking into pieces and conditional probabilities of the detection of one feature given the detection or nondetection of other features. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the probabilities to guide the search and prune the tree. The matching algorithm represents a large theoretical effort that is actually independent of the PREMIO system. The algorithm has been implemented as a C program and teated on data specifically generated to fit the abstract paradigm for the probabilistic search. The purpose of this paper is to describe the theory, the algorithm, and illustrative results. Octavia I. Camps, Linda G. Shapiro, Robert M. Haralick |
IROS | 2 |
| 1992 | A Cad-based System For Automated Inspection Of Machined PartsabstractAlthough many special purpose inspection systems have been developed, general purpose systems utalizing CAD models of the parts are still in the research stage. While it is easy to define ad hoc algorithms for inspection, it is much more diflcult to justify the algorithms with solid theory. In this paper we describe a CAD-model-based machine vision system for dimensional inspection of machined parts, with emphasis on the theory behind the system, The original contributions of our work are: 1) the use of precise definitions of geometric tolerances suitable for use in image processing, 2) the development of measurement algorithms corresponding directly to these definitions, 3) the derivation of the uncertainties in the measurement tasks, and 4) the use of this uncertainty information in the decision-making process. Our experimental results have verified the uncertainty derivations statistically, proved that the error probabilities obtained by propagating uncertainties are lower than those obtaanable without uncertainty propagation, and demonstrated that the inspection system responds in a predictable manner when applied to deformed objects. Bharath R. Modayur, Linda G. Shapiro |
IROS | 2 |
| 1991 | Glossary of computer vision terms
Robert M. Haralick, Linda G. Shapiro |
Pattern Recognit. | 2 |
| 1990 | Interesting patterns for model-based machine visionabstractThe author's work builds on D.G. Lowe's (1987) theory of perceptual groupings. Minimal processing is applied to an image to extract edges. The edges are then represented as well-defined two-dimensional patterns that the authors call interesting patterns. No attempt is made to infer three-dimensional structure from the patterns, and they are matched against two-dimensional models which are projections of characteristic views of three-dimensional objects. The patterns are built up from modular building blocks called triples.> Jorja G. Henikoff, Linda G. Shapiro |
ICCV | 2 |
| 1990 | Toward the automatic generation of mathematical morphology procedures using predicate logicabstractA discussion is presented of the design of a system that can input a vision task specification and use its knowledge of the operations of mathematical morphology to automatically construct a procedure that can execute the task. To do this, the authors develop a predicate calculus representation to describe the essence of the states of all the images that are created during the execution of the morphological procedure and the states of the relationships among them. The authors translate the English descriptions of morphological procedures into predicate logic. In so doing they gain an understanding of the goal of each procedure and the exact conditions under which a procedure achieves its goal. With this knowledge of the operations of mathematical morphology represented in predicate logic, a search procedure can be used to automatically produce vision procedures.> Hyonam Joo, Robert M. Haralick, Linda G. Shapiro |
ICCV | 3 |
| 1990 | Optimal affine-invariant point matchingabstractApplication of the affine-invariant point-matching scheme proposed by R. Hummel and H. Wolfson (1988) to the problem of recognizing and determining the pose of sheet metal parts is discussed. Attention is given to errors that can occur with this method due to quantization, stability, symmetry, and noise problems. These errors make the original affine-invariant matching technique unsuitable for use on the factory floor. An explicit noise model, which the Hummel and Wolfson technique lacks, is used. An optimal approach which overcomes these problems is then derived. The performance of the proposed algorithm under the influence of several distorting parameters is evaluated.> Mauro S. Costa, Robert M. Haralick, Linda G. Shapiro |
ICPR (1) | 3 |
| 1990 | Automatic sensor and light source positioning for machine visionabstractThe authors discuss an optimization approach to automatic sensor and light source positioning for a machine vision task where geometric measurement and/or object verification is important. The goal of the vision task is assumed to be specified in terms of edge visibility. There are two types of edge visibility: geometric edge visibility tells how much of the given edge is not occluded, and photometric visibility tells how much of the given edge has enough contrast to be detected in the image. A heuristic optimality criterion for the optimal sensor and light source position is defined in terms of these two edge visibilities. A preliminary experiment has been conducted to demonstrate the feasibility of the optimization approach. The result shows that the optimization problem formulated can be solved by mathematical programming techniques.> Seungku Yi, Robert M. Haralick, Linda G. Shapiro |
ICPR (1) | 3 |
| 1990 | Accumulator-based inexact matching using relational summaries
Linda G. Shapiro, Haiyuan Lu |
Mach. Vis. Appl. | 1 |
| 1989 | Shape from shading using the facet model
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro |
Pattern Recognit. | 3 |
| 1989 | Matching topographic structures in stereo vision
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro |
Pattern Recognit. Lett. | 3 |
| 1988 | The use of relational pyramid representation for view classes in a CAD-to-vision systemabstractA CAD-to-vision system with applications in robot guidance, docking and tracking, and inspection tasks is being developed. The goal of the system is to start with CAD models from a geometric modeling system, convert to models that are appropriate for object recognition in vision, and use these vision models for online vision tasks. The vision models consist of two parts: a three-dimensional component, and a set of two-dimensional structures called view classes, each of which relates back to the three-dimensional component. A view class represents a range of views of the object that have similar features in similar relationships. The definition of a view class, its representation as a relational pyramid, and an algorithm for rapidly selecting the view class that best matches an unknown view of the object are discussed.> Linda G. Shapiro, Haiyuan Lu |
ICPR | 1 |
| 1988 | Digital halftoning : Robert Ulichney
Linda G. Shapiro |
Comput. Vis. Graph. Image Process. | 1 |
| 1988 | Processor arrays: Architecture and applications : Terry Fountain
Linda G. Shapiro |
Comput. Vis. Graph. Image Process. | 1 |
| 1987 | Insight: a Dataflow Language for Programming Vision Algorithms in a Reconfigurable Computational NetworkabstractMachine vision systems used in industrial applications must execute their algorithms in real time to perform such tasks as inspecting a wire bond or guiding a robot to install a part on a car body moving along a conveyer. The real time speed is achieved by employing simple-minded algorithms and by designing parallel architectures and parallel algorithms for some tasks. The majority of the work on parallel architectures has been limited to architectures that support image processing, but not mid- or high-level vision In order for more complex vision algorithms to execute in real time, a more flexible architecture is needed. Our conceptual approach to the problem is a reconfigurable computational network. Each configuration of the network implements an algorithm or class of algorithms A high-level language expresses the algorithms in a relational form that can be easily translated to the specification for a configuration. The language must be able to encode low-, mid-, and high-level vision algorithms and to efficiently handle not only pixel data, but also higher level structures. In this paper we describe a dataflow language called INSIGHT, which we have designed to meet these needs, and give several examples of parallel machine vision algorithms expressed in the language. Linda G. Shapiro, Robert M. Haralick, Michael Goulish |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1987 | Ordered structural shape matching with primitive extraction by mathematical morphology
Linda G. Shapiro, Robert S. MacDonald, Stanley R. Sternberg |
Pattern Recognit. | 1 |
| 1987 | Morphologic edge detectionabstractEdge operators based on gray-scale morphologic operations are introduced. These operators can be efficiently implemented in near real time machine vision systems which have special hardware support for gray-scale morphologic operations. The simplest morphologic edge detectors are the dilation residue and erosion residue operators. The underlying motivation for these and some of their combinations are discussed and justified. Finally, the blur-minimum morphologic edge operator is defined. Its inherent noise sensitivity is less than the dilation or the erosion residue operators. Some experimental results are provided to show the validity of these morphologic operators. When compared with the enhancement/thresholding edge detectors and the cubic facet second derivative zero-crossing edge operator, the results show that all the edge operators have similar performance when the noise is small. However, as the noise increases, the second derivative zero-crossing edge operator and the blur-minimum morphologic edge operator have much better performance than the rest of the operators. The advantage of the blur-minimum edge operator is that it is less computationally complex than the facet edge operator. James Shih-Jong Lee, Robert M. Haralick, Linda G. Shapiro |
IEEE J. Robotics Autom. | 3 |
| 1986 | Shape from perspective: A rule-based approach
Prasanna G. Mulgaonkar, Linda G. Shapiro, Robert M. Haralick |
Comput. Vis. Graph. Image Process. | 2 |
| 1986 | Special issue on current issues and trends in computer vision
Linda G. Shapiro, Avinash C. Kak |
Comput. Vis. Graph. Image Process. | 1 |
| 1985 | Computer Architecture for Solving Consistent Labelling ProblemsabstractConsistent labelling problems are a family of NP-complete constraint satisfaction problems such as school timetabling, for which a conventional computer may be too slow. There are a variety of techniques for reducing the elapsed time to find one or all solutions to a consistent labelling problem. In this paper we discuss and illustrate solutions consisting of special hardware to accomplish the required constraint propagation and an asynchronous network of intercommunicating computers to accomplish the tree search in parallel. Julian R. Ullmann, Robert M. Haralick, Linda G. Shapiro |
Comput. J. | 3 |
| 1985 | Image segmentation techniques
Robert M. Haralick, Linda G. Shapiro |
Comput. Vis. Graph. Image Process. | 2 |
| 1985 | A Metric for Comparing Relational DescriptionsabstractRelational models are frequently used in high-level computer vision. Finding a correspondence between a relational model and an image description is an important operation in the analysis of scenes. In this paper the process of finding the correspondence is formalized by defining a general relational distance measure that computes a numeric distance between any two relational descriptions-a model and an image description, two models, or two image descriptions. The distance measure is proved to be a metric, and is illustrated with examples of distance between object models. A variant measure used in our past studies is shown not to be a metric. Linda G. Shapiro, Robert M. Haralick |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1985 | Shape estimation from topographic primal sketch
Ting-Chuen Pong, Linda G. Shapiro, Robert M. Haralick |
Pattern Recognit. | 2 |
| 1984 | A hierarchical relational model for automated inspection tasksabstractAn F-15 bulkhead is to be inspected by a computer system employing television cameras for vision and robot arms with tactile sensors for precise measurements. The system requires a suitable model of the object to be inspected. The model must contain very precise, accurate information for low-level vision and measurement processes and at the same time be useful to high-level vision and planning processes. This requirement makes most existing three-dimensional models useless. In this paper we define a hierarchical, relational model that we have developed to be used by the robot inspection system. Linda G. Shapiro, Robert M. Haralick |
ICRA | 1 |
| 1984 | Experiments in segmentation using a facet model region grower
Ting-Chuen Pong, Linda G. Shapiro, Layne T. Watson, Robert M. Haralick |
Comput. Vis. Graph. Image Process. | 2 |
| 1984 | Matching 'sticks, plates and blobs' objects using geometric and relational constraints
Prasanna G. Mulgaonkar, Linda G. Shapiro, Robert M. Haralick |
Image Vis. Comput. | 2 |
| 1984 | Matching wire frame objects from their two dimensional perspective projections
Robert M. Haralick, Yu Hong Chu, Layne T. Watson, Linda G. Shapiro |
Pattern Recognit. | 4 |
| 1984 | Matching three-dimensional objects using a relational paradigm
Linda G. Shapiro, John D. Moriarty, Robert M. Haralick, Prasanna G. Mulgaonkar |
Pattern Recognit. | 1 |
| 1983 | A new connected components algorithm for virtual memory computers
Ronald Lumia, Linda G. Shapiro, Oscar A. Zuniga |
Comput. Vis. Graph. Image Process. | 2 |
| 1983 | Texture analysis of aerial photographs
Ronald Lumia, Robert M. Haralick, Oscar A. Zuniga, Linda G. Shapiro, Ting-Chuen Pong, Far-Peing Wang |
Pattern Recognit. | 4 |
| 1982 | Organization of Relational Models for Scene AnalysisabstractRelational models are commonly used in scene analysis systems. Most such systems are experimental and deal with only a small number of models. Unknown objects to be analyzed are usually sequentially compared to each model. In this paper, we present some ideas for organizing a large database of relational models. We define a simple relational distance measure, prove it is a metric, and using this measure, describe two organizational/access methods: clustering and binary search trees. We illustrate these methods with a set of randomly generated graphs. Linda G. Shapiro, Robert M. Haralick |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1982 | Identification of Space Curves from Two-Dimensional Perspective ViewsabstractThis paper describes a new method to be used for matching three-dimensional objects with curved surfaces to two-dimensional perspective views. The method requires for each three-dimensional object a stored model consisting of a closed space curve representing some characteristic connected curved edges of the object. The input is a two-dimensional perspective projection of one of the stored models represented by an ordered sequence of points. The input is converted to a spline representation which is sampled at equal intervals to derive a curvature function. The Fourier transform of the curvature function is used to represent the shape. The actual matching is reduced to a minimization problem which is handled by the Levenberg-Marquardt algorithm [3]. Layne T. Watson, Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1982 | The nearest neighbor problem in an abstract metric space
Charles Feustel, Linda G. Shapiro |
Pattern Recognit. Lett. | 2 |
| 1982 | Design and Architectural Implications of a Spatial Information SystemabstractImage analysis, at the higher levels, works with extracted regions and line segments and their properties, not with the original raster data. Thus, a spatial information system must be able to store points, lines, and areas as well as their properties and interrelationships. In a previous paper (Shapiro and Haralick [17]), we proposed for this purpose an entity-oriented relational database system. In this paper, we describe our first experimental spatial information system which employs these concepts to store and retrieve watershed data for a portion of the state of Virginia. We describe the logical and physical design of the system and discuss the architectural implications. Prashant D. Vaidya, Linda G. Shapiro, Robert M. Haralick, Gary J. Minden |
IEEE Trans. Computers | 2 |
| 1981 | Structural Descriptions and Inexact MatchingabstractIn this paper we formally define the structural description of an object and the concepts of exact and inexact matching of two structural descriptions. We discuss the problems associated with a brute-force backtracking tree search for inexact matching and develop several different algorithms to make the tree search more efficient. We develop the formula for the expected number of nodes in the tree for backtracking alone and with a forward checking algorithm. Finally, we present experimental results showing that forward checking is the most efficient of the algorithms tested. Linda G. Shapiro, Robert M. Haralick |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1980 | Sticks, Plates, and Blobs: A Three-Dimensional Object Representation for Scene Analysis
Linda G. Shapiro, John D. Moriarty, Prasanna G. Mulgaonkar, Robert M. Haralick |
AAAI | 1 |
| 1980 | The Consistent Labeling Problem: Part IIabstractIn this second part of a two-part paper, we explore the power and complexity of the g=fKP and g=vKP class of look-ahead operators which can be used to speed up the tree search in the consistent labeling problem. For a specified K and P we show that the fixedpoint power of g=fKP and g=vKP is the same, that g=fKP+1 is at least as powerful as g=fKP, and that g=vK+1p is at least as powerful at g=fKP. Finally, we define a minimal compatibility relation and show how the standard tree search procedure for finding all the consistent labelings is quicker for a minimal relation. This leads to the concept of grading the complexity of compatibility relations according to how much look-ahead work it requires to reduce them to minimal relations and suggests that the reason look-ahead operators, such as Waltz filtering, work so well is that the compatibility relations used in practice are not very complex and are reducible to minimal or near minimal relations by a g=fKP or g=vKP look-ahead operator with small value for parameter P. Robert M. Haralick, Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1980 | A Structural Model of ShapeabstractShape description and recognition is an important and interesting problem in scene analysis. Our approach to shape description is a formal model of a shape consisting of a set of primitives, their properties, and their interrelationships. The primitives are the simple parts and intrusions of the shape which can be derived through the graph-theoretic clustering procedure described in [31]. The interrelationships are two ternary relations on the primitives: the intrusion relation which relates two simple parts that join to the intrusion they surround and the protrusion relation which relates two intrusions to the protrusion between them. Using this model, a shape matching procedure that uses a tree search with look-ahead to find mappings from a prototype shape to a candidate shape has been developed. An experimental Snobol4 implementation has been used to test the program on hand-printed character data with favorable results. Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1979 | The Consistent Labeling Problem: Part IabstractIn this first part of a two-part paper we introduce a general consistent labeling problem based on a unit constraint relation T containing N-tuples of units which constrain one another, and a compatibility relation R containing N-tuples of unit-label pairs specifying which N-tuples of units are compatible with which N-tuples of labels. We show that Latin square puzzles, finding N-ary relations, graph or auto-mata homomorphisms, graph colorings, as well as determining satisfiability of propositional logic statements and solving scene and edge labeling problems, are all special cases of the general consistent labeling problem. We then discuss the various approaches that researchers have used to speed up the tree search required to find consistent labelings. Each of these approaches uses a particular look-ahead operator to help eliminate backtracking in the tree search. Finally, we define the ¿KP two-parameter class of look-ahead operators which includes, as special cases, the operators other researchers have used. Robert M. Haralick, Linda G. Shapiro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1979 | Decomposition of Two-Dimensional Shapes by Graph-Theoretic ClusteringabstractThis paper describes a technique for transforming a twodimensional shape into a binary relation whose clusters represent the intuitively pleasing simple parts of the shape. The binary relation can be defined on the set of boundary points of the shape or on the set of line segments of a piecewise linear approximation to the boundary. The relation includes all pairs of vertices (or segments) such that the line segment joining the pair lies entirely interior to the boundary of the shape. The graph-theoretic clustering method first determines dense regions, which are local regions of high compactness, and then forms clusters by merging together those dense regions having high enough overlap. Using this procedure on handdrawn colon shapes copied from an X-ray and on handprinted characters, the parts determined by the clustering often correspond well to decompositions that a human might make. Linda G. Shapiro, Robert M. Haralick |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1978 | Data structures for picture processingabstractA variety of algorithms have been invented for use in picture processing.An important aspect of the algorithms are the data structures employed.A data structure may be chosen to represent a particular structural relationship, to save space, or to allow for fast access to data.This paper surveys four major classes of data structures used in current picture processing research and gives several examples of the use of each type of structure in particular algorithms or systems.The structures surveyed are linear lists, hierarchic structures, graph structures, and recursive structures. Linda G. Shapiro |
SIGGRAPH | 1 |
| 1978 | Inexact matching of line drawings in a syntactic pattern recognition system
Linda G. Shapiro |
Pattern Recognit. | 1 |
| 1977 | ESP³: A Language for Pattern Description and a System for Pattern RecognitionabstractExtended Snobol picture pattern processor (ESP3) is a programming language and pattern recognition system which was designed for generating, recognizing, and manipulating two-dimensional line drawings. An ESP3 picture pattern describes a class of line drawings just as a Snobol pattern describes a class of strings. During pattern matching, a subject picture is searched for the occurrence of a sub-picture which matches a given picture pattern. The search of the subject picture is ordered left-to-right and top-to-bottom, and the search program uses scanner guidance information found in the picture pattern to limit the area of the subject picture to be searched. An experimental implementation of ESP3 has been developed to test the feasibility of the system. This paper describes the ESP3 picture patterns, the pattern matching operation, and the experimental implementation of ESP3. Linda G. Shapiro, Robert J. Baron |
IEEE Trans. Software Eng. | 1 |
| 1975 | ESP3: a high-level graphics languageabstractMost graphics languages are composed of a primitive set of commands which allow for the creation and manipulation of graphical objects. These commands are generally at a low level, in that each command causes one operation to be performed. Often the commands are to subroutines embedded in an algorithmic language so that the arithmetic and control features of the higher level language may be used.ESP3 (Extended SNOBOL Picture Pattern Processor) is a new high-level graphics and pattern recognition language. ESP3 was designed in an effort to provide simple, natural, and efficient manipulation of line drawings. ESP3 differs from present graphics languages in the following ways:1) It provides a high-level method for picture construction. The evaluation of a picture expression (analagous to the SNOBOL4 string-valued expression) causes the construction of a picture.2) It provides extensive referencing facilities for naming and accessing points, subpictures, and attributes of pictures.3) It provides predicates for testing attributes of and relationships among pictures and points.4) It provides a means for defining picture patterns that describe classes of line drawings in much the same way that SNOBOL4 patterns describe classes of strings. Picture pattern matching is a built-in facility.ESP3 is based on the premise that structural descriptions are an essential part of both picture construction and pattern recognition. The concept of a structural description of a picture has its origin with the linguistic-approach to pattern recognition. In the linguistic approach, formal grammars are used as a mechanism for picture description. [Kirsch (1964), Narasimhan (1964, 1966,1970), Anderson (1968), Evans (1968), Miller and Shaw (1969), Fu and Swain (1971), Shaw (1970, 1972), Chien and Ribak (1972), Thomason and Gonzalez (1975)]. Stanton (1970) described a graphics language based on linguistic pattern recognition. ESP3 incorporates and extends many ideas from the above work, and includes all of the features of SNOBOL4 to provide a high-level graphics and pattern recognition language. Some suggested applications of ESP3 are the generation of graphical output, AI programs with imaging capabilities, pattern recognition systems, and scene analysis programs.This paper will describe picture construction and pattern recognition in ESP with emphasis on picture construction. Some tests performed with an experimental version of ESP3 will also be discussed. For a more detailed description of ESP3, see Shapiro (1974). Linda G. Shapiro |
SIGGRAPH | 1 |