Robert M. Haralick

dblp:80/2840 · also Robert Martin Haralick · DBLP profile ↗
← Back
250ranked-venue papers
68as first author
1since 2021 · last 2021
0000-0002-2021-4327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 171 · 46 first-authorGraphics, computer vision, multimedia, augmented reality and games · 115 · 29 first-authorDatabases, data management, data science and information retrieval · 21 · 1 first-authorHuman-computer interaction and ubiquitous computing · 11 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-authorSystems, architecture and hardware · 8 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorTheory of computation · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
25 papers
3D vision · 48% Representation and self-supervised learning · 30% Learning theory · 15%
Computer graphics and multimedia
37 papers
Image and video processing · 74% Multimedia analysis and retrieval · 14% Geometric modeling and processing · 9%
Theoretical computer science
14 papers
Mathematical optimization · 72% Computational geometry · 8% Computational complexity · 6%
Software engineering, system software, and programming languages
2 papers
Software testing · 98% Requirements engineering and software design · 2%
Databases, data mining, and information retrieval
5 papers
Data mining · 54% Information retrieval · 38% Indexing and storage engines · 5%

Topics — the 30 heaviest of 115, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.112010
Second-Order Bilinear Discriminant Analysis · J. Mach. Learn. Res. 2010
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
discriminant analysis
0.112010
Second-Order Bilinear Discriminant Analysis · J. Mach. Learn. Res. 2010
Computer vision › 3D vision › motion estimation
optical flow
0.142003
Estimating Piecewise-Smooth Optical Flow with Global Matching and Graduated Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Local Gradient, Global Matching, Piecewise-Smooth Optical Flow · CVPR (2) 2001
Two-Stage Robust Optical Flow Estimation · CVPR 2000
Image and video processing
mathematical morphology
0.182000
Recursive binary dilation and erosion using digital line structuring elements in arbitrary orientations · IEEE Trans. Image Process. 2000
Gray-scale structuring element decomposition · IEEE Trans. Image Process. 1996
Recursive erosion, dilation, opening, and closing transforms · IEEE Trans. Image Process. 1995
Image and video processing
document image analysis
0.132001
An Optimization Methodology for Document Structure Extraction on Latin Character Documents · IEEE Trans. Pattern Anal. Mach. Intell. 2001
A Statistical, Nonparametric Methodology for Document Degradation Model Validation · IEEE Trans. Pattern Anal. Mach. Intell. 2000
An Automatic Closed-Loop Methodology for Generating Character Groundtruth for Scanned Documents · IEEE Trans. Pattern Anal. Mach. Intell. 1999
Computer vision › 3D vision
motion estimation
0.132001
Local Gradient, Global Matching, Piecewise-Smooth Optical Flow · CVPR (2) 2001
Two-Stage Robust Optical Flow Estimation · CVPR 2000
From depth and optical flow to rigid body motion · CVPR 1988
Image and video processing
image restoration
0.112005
A Bayesian framework for noise covariance estimation using the facet model · IEEE Trans. Image Process. 2005
Software testing › fault analysis
error propagation
0.112005
On the Use of Error Propagation for Statistical Validation of Computer Vision Software · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Multimedia analysis and retrieval
image retrieval
0.122000
Probabilistic vs. Geometric Similarity Measures for Image Retrieval · CVPR 2000
Graph-Theoretic Clustering for Image Grouping and Retrieval · CVPR 1999
Image and video processing
edge detection
0.062000
Assignment Problem in Edge Detection Performance Evaluation · CVPR 2000
Context dependent edge detection · CVPR 1988
Morphologic edge detection · IEEE J. Robotics Autom. 1987
Mathematical optimization
global optimization
0.012003
Estimating Piecewise-Smooth Optical Flow with Global Matching and Graduated Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Computer vision › 3D vision › motion estimation › optical flow
robust optical flow
0.012000
Two-Stage Robust Optical Flow Estimation · CVPR 2000
Image and video processing › edge detection
edge detector evaluation
0.012000
Assignment Problem in Edge Detection Performance Evaluation · CVPR 2000
Mathematical optimization › combinatorial optimization
assignment problem
0.012000
Assignment Problem in Edge Detection Performance Evaluation · CVPR 2000
Data mining
clustering
0.021999
Graph-Theoretic Clustering for Image Grouping and Retrieval · CVPR 1999
Organization of Relational Models for Scene Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 1982
Data mining › clustering
graph clustering
0.011999
Graph-Theoretic Clustering for Image Grouping and Retrieval · CVPR 1999
Multimedia analysis and retrieval
image analysis
0.011998
Breakpoint Detection Using Covariance Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Computer vision › 3D vision › low-level vision › feature detection
corner detection
0.011997
Corner Detection with Covariance Propagation · CVPR 1997
Computer vision › 3D vision › low-level vision
feature detection
0.011997
Corner Detection with Covariance Propagation · CVPR 1997
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.012005
A Bayesian framework for noise covariance estimation using the facet model · IEEE Trans. Image Process. 2005
Geometric modeling and processing › 3d reconstruction
photogrammetry
0.012005
On the Use of Error Propagation for Statistical Validation of Computer Vision Software · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Computer vision › 3D vision › multi-view geometry
triangulation
0.011996
Finding Corresponding Points Based on Bayesian Triangulation · CVPR 1996
Image and video processing › mathematical morphology
structuring element decomposition
0.011996
Gray-scale structuring element decomposition · IEEE Trans. Image Process. 1996
Computer vision › 3D vision
camera pose estimation
0.021994
Review and analysis of solutions of the three point perspective pose estimation problem · Int. J. Comput. Vis. 1994
Two view motion analysis, stereo vision and a moving camera's positioning their equivalence and a new solution procedure · ICRA 1985
Image and video processing
motion estimation
0.012003
Estimating Piecewise-Smooth Optical Flow with Global Matching and Graduated Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Natural language and speech › Information extraction and text analysis
document understanding
0.011994
Document image understanding: geometric and logical layout · CVPR 1994
Computer vision › 3D vision › camera pose estimation
perspective pose estimation
0.011994
Review and analysis of solutions of the three point perspective pose estimation problem · Int. J. Comput. Vis. 1994
Image and video processing › image restoration
image denoising
0.011994
Quantitative performance evaluation of thinning algorithms under noisy conditions · CVPR 1994
Computational geometry › geometric inference
geometric estimation
0.011994
Review and analysis of solutions of the three point perspective pose estimation problem · Int. J. Comput. Vis. 1994
Geometric modeling and processing
shape decomposition
0.021992
Morphological decomposition of restricted domains: a vector space solution · CVPR 1992
Decomposition of Two-Dimensional Shapes by Graph-Theoretic Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 1979

Methods — techniques the papers use, named apart from their topics

global matching · 0.2statistical hypothesis testing · 0.1three-frame matching · 0.1local variation · 0.1graduated optimization · 0.1facet model · 0.1generalized inverted wishart · 0.1expectation-maximization · 0.1error propagation analysis · 0.1bayesian estimation · 0.1relaxation · 0.1probabilistic partitioning · 0.1robust regression · 0.1hypothesis testing · 0.0combinatorial assignment · 0.0k-nearest neighbor · 0.0graph clustering · 0.0review · 0.0
YearPublicationVenuePosition
2021 The N-Tuple Subspace Classifier: Extensions and Survey
abstract
This article is written in recognition of W. Bledsoe, who with Browning, introduced the N-tuple subspace classifier in 1959. This 1959 article was the first article to introduce subspace classifiers and the sum rule to combine the outputs of the classifiers. A mathematical notation is given to easily express in a precise and unambiguous way everything going on in the N-tuple subspace classifier. Extensions of the N-tuple method are discussed using a generalized product expression and we relate the generalization to graphical models. We discuss the sum rule, the product rule, and the plurality voting rule for combining the scores of the subspace classifiers. We selected a representative sample of papers that the 1959 N-tuple subspace classifier inspired. Some of the papers introduced specialized improvements. Many of the papers showed the value of the N-tuple subspace classifier in all kinds of applications and compared the results of one or more varieties of the N-tuple subspace classifier with other state-of-the-art classifiers. Their experiments showed that the N-tuple subspace classifier was competitive with the state-of-the-art classifiers and often had a higher accuracy. Finally, we highlight some papers that describe experiments of the N-tuple subspace classifier executing in a quantum computer.
Robert M. Haralick, Ahmet Cem Yuksel
IEEE Trans. Syst. Man Cybern. Syst.1
2020 Partial Monotone Dependence
abstract
We present a new measure of dependence suitable for time series forecasting: Partial Monotone Correlation (PMC) that generalizes Monotone Correlation. Unlike the Monotone Correlation, the new measure of dependence uses piecewise strictly monotone transformations that increase the value of the correlation coefficient. We explore its properties, its relationship with Monotone and Maximal Correlation, and present an algorithm that calculates it based on the Simultaneous Perturbation Stochastic Approximation method. We also demonstrate how to apply Partial Monotone Correlation for time series analysis and forecasting introducing Partial Monotone Autoregressive model (PMAR) of order 1. Its performance is then evaluated against the baseline of linear and nonlinear autoregressive models (AR, LSTAR) on 150 time series produced from 3 datasets: Yellow Taxi pickups, Citi Bike pickups, and Cellular Network hits. Overall, PMAR model outperforms the baseline with the average sMAPE of about 1.7-4% lower.
Denis Khryashchev, Huy T. Vo, Robert M. Haralick
ICPR3
2019 Dependence
Robert M. Haralick
Pattern Recognit. Lett.1
2019 A method for discovering knowledge in texts
Minhua Huang 0001, Robert M. Haralick
Pattern Recognit. Lett.2
2016 Inexact MDL for linear manifold clusters
abstract
We present a regularization technique based on the minimum description length (MDL) principle for the linear manifold clustering. We suggest an inexact minimum description length method based on describing the data structure as linear manifold clusters. We examine the behavior of the proposed method and compare it performance against simulated clustering results of various dimensionality and structure. Finally, we empirically evaluate the proposed technique on a climate data.
Robert M. Haralick, Art Diky, Nancy Y. Kiang
ICPR1
2014 Quadratic Discriminant Revisited
abstract
In this study, we revisit quadratic discriminant analysis (QDA). For this purpose, we present a majorize-minimize (MM) optimization algorithm to estimate parameters for generative classifiers, of which conditional distributions are from the exponential family. Furthermore, we propose a block-coordinate descent algorithm to sequentially update parameters of QDA in each iteration of the MM algorithm, for each update, we apply a trust region method, of which each iteration has a simple closed form solution. Numerical experiments show that: when compared with conjugate gradient method, the new proposed method is faster in 9 of 10 benchmark data sets, when compared with other widely used quadratic classifiers in the literature, QDA trained with the proposed method is either the best or not statistically significantly different from the best ones in 8 of 10 benchmark data sets.
Wenbo Cao, Robert M. Haralick
ICPR2
2014 GIST: Graphical Interactive Display Tools Defining a Model for Interactive Search
abstract
Information retrieval methods represent query results in a ranked, one dimensional list without revealing connections among documents and document groups. We propose a new model of document representation and extend the notion of similarity to consider document length and word synonyms to organize documents into topically relevant groups. Matches to a user query are presented in an intuitive, interactive map that facilitates browsing information and finding the most relevant matches within the overall landscape of results. Current research is aimed at finding the best metric to group documents, and in defining the visual model. Our system, named GIST, is built using the PREFUSE toolkit for graphical display [1].
Ingrid Montealegre, Robert M. Haralick
ICPR2
2013 Multiscale 3D feature extraction and matching with an application to 3D face recognition
Hadi Fadaifard, George Wolberg, Robert M. Haralick
Graph. Model.3
2012 Developing an Algorithm for Mining Semantics in Texts
Minhua Huang 0001, Robert M. Haralick
CICLing (2)2
2010 Second-Order Bilinear Discriminant Analysis
Christoforos Christoforou, Robert M. Haralick, Paul Sajda, Lucas C. Parra
J. Mach. Learn. Res.2
2009 Affine feature extraction: A generalization of the Fukunaga-Koontz transformation
Wenbo Cao, Robert M. Haralick
Eng. Appl. Artif. Intell.2
2007 Mining Subspace Correlations
abstract
In recent applications of clustering such as gene expression microarray analysis, collaborative filtering, and Web mining, object similarity is no longer measured by physical distance, but rather by the behavior patterns objects manifest or the magnitude of correlations they induce. Current state of the art algorithms aiming at this type of clustering typically postulate specific cluster models that are able to capture only specific behavior patterns or correlations, and omit the possibility that other information carrying patterns or correlations may coexist in the data. We cast the problem of searching for pattern clusters or clusters that induce large correlations in some subset of features into the problem of searching for groups of points embedded in lines. The advantage of this approach is that is allows the clustering of different patterns or correlations simultaneously. It also allows the clustering of patterns and correlations that are overlooked by existing methods. A formal stochastic line cluster model is presented and its connection to correlation is established. Based on this model an algorithm, which uses feature selection to search for line clusters embedded in subspaces of the data is presented
Rave Harpaz, Robert M. Haralick
CIDM2
2007 Maximum Likelihood Quantization of Genomic Features Using Dynamic Programming
abstract
Dynamic programming is introduced to quantize a continuous random variable into a discrete random variable. Quantization is often useful before statistical analysis or reconstruction of large network models among multiple random variables. The quantization, through dynamic programming, finds the optimal discrete representation of the original probability density function of a random variable by maximizing the likelihood for the observed data. This algorithm is highly applicable to study genomic features such as the recombination rate across the chromosomes and the statistical properties of non-coding elements such as LINE1. In particular, the recombination rate obtained by quantization is studied for LINE1 elements that are grouped also using quantization by length. The exact and density-preserving quantization approach provides an alternative superior to the inexact and distance-based k-means clustering algorithm for discretization of a single variable.
Mingzhou Song 0001, Robert M. Haralick, Stéphane Boissinot
ICMLA2
2007 Linear manifold clustering in high dimensional spaces by stochastic search
Robert M. Haralick, Rave Harpaz
Pattern Recognit.1
2006 A hybrid search algorithm for the Whitehead Minimization problem
Alex D. Myasnikov, Robert M. Haralick
J. Symb. Comput.2
2006 Document zone content classification and its performance evaluation
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick
Pattern Recognit.3
2005 Towards a Formal Concept Analysis Approach to Exploring Communities on the World Wide Web
Jayson E. Rome, Robert M. Haralick
ICFCA2
2005 On the Use of Error Propagation for Statistical Validation of Computer Vision Software
abstract
Computer vision software is complex, involving many tens of thousands of lines of code. Coding mistakes are not uncommon. When the vision algorithms are run on controlled data which meet all the algorithm assumptions, the results are often statistically predictable. This renders it possible to statistically validate the computer vision software and its associated theoretical derivations. In this paper, we review the general theory for some relevant kinds of statistical testing and then illustrate this experimental methodology to validate our building parameter estimation software. This software estimates the 3D positions of buildings vertices based on the input data obtained from multi-image photogrammetric resection calculations and 3D geometric information relating some of the points, lines and planes of the buildings to each other.
Xufei Liu, Tapas Kanungo, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 A Bayesian framework for noise covariance estimation using the facet model
abstract
In image processing literature, thus far, researchers have assumed the perturbation in the data to be white (or uncorrelated) having a covariance matrix sigma2I, i.e., assumption of equal variance for all the data samples and that no correlation exists between the data samples. However, there have been very few attempts to estimate noise characteristics under the assumption that there is a correlation between data samples. In this work, we propose a new and a novel approach for the simultaneous Bayesian estimation of the unknown colored or correlated noise (population) covariance matrix and the hyperparameters of the covariance model using the well-known facet model. We also estimate the facet model coefficients. We use the facet model because of its simple, yet elegant, mathematical formulation. We use the generalized inverted Wishart density as the prior model for the noise covariance matrix. We place a structure on the covariance matrix using the parameters of a correlation filter. These hyperparameters are estimated by a new extension of the expectation-maximization algorithm called the generalized constrained expectation maximization algorithm that we developed.
Desikachari Nadadur, Robert M. Haralick, David E. Gustafson
IEEE Trans. Image Process.2
2004 Practical Aspects of Efficient Forward Selection in Decomposable Graphical Models
abstract
We discuss efficient forward selection in the class of decomposable graphical models. This subclass of graphical models has a number of desirable properties. The contributions of This work are twofold. First we improve an existing algorithm by addressing cases previously not considered. Second we extend the algorithm to reflect model graphs with multiple disconnected components. We further present experimental results that apply this approach to a real dataset and discuss its properties. We belief that the presented approach is applicable to a wide area of fields and problems.
Stephan Altmueller, Robert M. Haralick
ICTAI2
2004 Table structure understanding and its performance evaluation
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick
Pattern Recognit.3
2003 Estimating Piecewise-Smooth Optical Flow with Global Matching and Graduated Optimization
abstract
This paper presents a new method for estimating piecewise-smooth optical flow. We propose a global optimization formulation with three-frame matching and local variation and develop an efficient technique to minimize the resultant global energy. This technique takes advantage of local gradient, global gradient, and global matching methods and alleviates their limitations. Experiments on various synthetic and real data show that this method achieves highly competitive accuracy.
Robert M. Haralick, Linda G. Shapiro
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 A Study on the Document Zone Content Classification Problem
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick
Document Analysis Systems3
2002 Table Detection via Probability Optimization
Yalin Wang 0001, Ihsin T. Phillips, Robert M. Haralick
Document Analysis Systems3
2002 Estimating optical flow using a global matching formulation and graduated optimization
abstract
In this paper we consider the problem of optimal optical flow estimation assuming brightness conservation and piecewise smoothness. We propose a formulation based on three-frame matching and global optimization allowing local variation. It is superior to popular gradient-based models and more justifiable than existing global methods. We also develop an efficient technique to minimize the resultant global energy. It takes advantage of local gradient, global gradient and global matching methods and overcomes their limitations. Experiments on various synthetic and real data and comparison with state-of-the-art techniques show that the method achieves discontinuity preserving capability and sub-pixel accuracy.
Linda G. Shapiro, Robert M. Haralick
ICIP (2)3
2002 Special Issue on "Performance Evaluation: Theory, Practice, and Impact"
Tapas Kanungo, Henry S. Baird, Robert M. Haralick
Int. J. Document Anal. Recognit.3
2002 Efficient facet edge detection and quantitative performance evaluation
Robert M. Haralick
Pattern Recognit.2
2002 Optimal matching problem in detection and recognition performance evaluation
Robert M. Haralick
Pattern Recognit.2
2002 Integrated Surface Optimization for 3-D Freehand Echocardiography
abstract
The major obstacle of three-dimensional (3-D) echocardiography is that the ultrasound image quality is too low to reliably detect features locally. Almost all available surface-finding algorithms depend on decent quality boundaries to get satisfactory surface models. We formulate the surface model optimization problem in a Bayesian framework, such that the inference made about a surface model is based on the integration of both the low-level image evidence and the high-level prior shape knowledge through a pixel class prediction mechanism. We model the probability of pixel classes instead of making explicit decisions about them. Therefore, we avoid the unreliable edge detection or image segmentation problem and the pixel correspondence problem. An optimal surface model best explains the observed images such that the posterior probability of the surface model for the observed images is maximized. The pixel feature vector as the image evidence includes several parameters such as the smoothed grayscale value and the minimal second directional derivative. Statistically, we describe the feature vector by the pixel appearance probability model obtained by a nonparametric optimal quantization technique. Qualitatively, we display the imaging plane intersections of the optimized surface models together with those of the ground-truth surfaces reconstructed from manual delineations. Quantitatively, we measure the projection distance error between the optimized and the ground-truth surfaces. In our experiment, we use 20 studies to obtain the probability models offline. The prior shape knowledge is represented by a catalog of 86 left ventricle surface models. In another set of 25 test studies, the average epicardial and endocardial surface projection distance errors are 3.2 +/- 0.85 mm and 2.6 +/- 0.78 mm, respectively.
Mingzhou Song 0001, Robert M. Haralick, Florence H. Sheehan, Richard K. Johnson
IEEE Trans. Medical Imaging2
2001 Local Gradient, Global Matching, Piecewise-Smooth Optical Flow
abstract
In this paper we discuss a hybrid technique for piecewise-smooth optical flow estimation. We first pose optical flow estimation as a gradient-based local regression problem and solve it under a high-breakdown robust criterion. Then taking the output from the first step as the initial guess, we recast the problem in a robust matching-based global optimization framework. We have developed novel fast-converging deterministic algorithms for both optimization problems and incorporated a hierarchical scheme to handle large motions. This technique inherits the good subpixel accuracy from the local gradient approach and the insensitivity to local perturbation and derivative quality from the global matching approach, and it overcomes the limitations of both. Significant advantages over competing techniques are demonstrated on various standard synthetic and real image sequences.
Robert M. Haralick
CVPR (2)2
2001 Automatic Table Ground Truth Generation and a Background-Analysis-Based Table Structure Extraction Method
abstract
We first describe an automatic table ground truth generation system which can efficiently generate a large amount of accurate table ground truth suitable for the development of table detection algorithms. Then a novel background analysis-based, coarse-to-fine table identification algorithm and an X-Y cut table decomposition algorithm are described. We discuss an experimental protocol to evaluate the table detection algorithms. For a total of 1,125 document pages having 518 table entities and a total of 10,941 cell entities, our table detection algorithm takes line, word segmentation results as input and obtains around 90% cell correct detection rates.
Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
2001 Zone Content Classification and its Performance Evaluation
abstract
We present an improved zone content classification method and its performance evaluation. We added two new features to the feature vector from one previously published method (Sivaramakrishnan et al., 1995). We assumed different independence relationships in two zone sets. We used an optimized binary decision tree to estimate the maximum zone content class probability in one set while using the Viterbi algorithm to find the optimal solution for a zone sequence in the other set. The training, pruning and testing data set for the algorithm include 1,600 images drawn from the UWCDROM III document image database. The classifier is able to classify each given scientific and technical document zone into one of the nine classes, 2 text classes (of font size 4 - 18pt and font size 19 - 32 pt), math, table, halftone, map/drawing, ruling, logo, and others. Compared with our previous work (Wang et al., 2000), it raised the accuracy rate to 98.52% from 97.53% and reduced the mean false alarm rate to 0.53% from 1.26%.
Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
2001 Performance Evaluation of Document Structure Extraction Algorithms
Jisheng Liang, Ihsin T. Phillips, Robert M. Haralick
Comput. Vis. Image Underst.3
2001 An Optimization Methodology for Document Structure Extraction on Latin Character Documents
abstract
In this paper, we give a formal definition of a document image structure representation, and formulate document image structure extraction as a partitioning problem: finding an optimal solution partitioning the set of glyphs of an input document image into a hierarchical tree structure where entities within the hierarchy at each level have similar physical properties and compatible semantic labels. We present a unified methodology that is applicable to construction of document structures at different hierarchical levels. An iterative, relaxation-like method is used to find a partitioning solution that maximizes the probability of the extracted structure. All the probabilities used in the partitioning process are estimated from an extensive training set of various kinds of measurements among the entities within the hierarchy. The offline probabilities estimated in the training then drive all decisions in the online document structure extraction. We have implemented a text line extraction algorithm using this framework.
Jisheng Liang, Ihsin T. Phillips, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Feature normalization and likelihood-based similarity measures for image retrieval
Selim Aksoy, Robert M. Haralick
Pattern Recognit. Lett.2
2001 Error propagation for the Hough transform
Robert M. Haralick
Pattern Recognit. Lett.2
2001 A knowledge-based boundary delineation system for contrast ventriculograms
abstract
Automated left-ventricle (LV) boundary delineation from contrast ventriculograms has been studied for decades. Unfortunately, no accurate methods have ever been reported. A new knowledge based multistage method to automatically delineate the LV boundary at end diastole (ED) and end systole (ES) is discussed in this paper. It has a mean absolute boundary error or about 2 mm and an associated ejection fraction error of about 6%. The method makes extensive use of knowledge about LV shape and movement. The processing includes a multiimage pixel region classification, shape regression, and rejection classification. The method was trained and cross-validated tested on a database of 375 studies whose ED and ES boundary had been manually traced as the ground truth. The cross-validated results presented in this paper show that the accuracy is close to and slightly above the interobserver variability.
Lei Sui, Robert M. Haralick, Florence H. Sheehan
IEEE Trans. Inf. Technol. Biomed.2
2000 Automated Left Ventricle Boundary Delineation
abstract
Automated left ventricle (LV) boundary delineation from left ventriculograms has been studied for decades. Unfortunately, no methods in terms of the accuracy about volume and ejection fraction have ever been reported. A new knowledge based multi-stage method to automatically delineate the LV boundary at end diastole and end systole is discussed in this paper: It has a mean absolute boundary error of about 2 mm and an associated ejection fraction error of about 6%. The method makes extensive use of knowledge about LV shape and movement. The processing includes a multi-image pixel region classification, a shape regression and a rejection classification. The method was trained and tested on a database of 375 studies whose ED and ES boundary have been manually traced as the ground truth. The cross-validated results presented in this paper shows that the accuracy is close to and slightly above inter-observer variability.
Lei Sui, Robert M. Haralick
BIBE2
2000 Probabilistic vs. Geometric Similarity Measures for Image Retrieval
abstract
Similarity between images in image retrieval is measured by computing distances between feature vectors. This paper presents a probabilistic approach and describes two likelihood-based similarity measures for image retrieval. Popular distance measures like the Euclidean distance implicitly assign more more weighting to features with large ranges than those with small ranges. First, we discuss the effects of five feature normalization methods on retrieval performance. Then, we show that the probabilistic methods perform significantly better than geometric approaches like the nearest neighbor rule with city-block or Euclidean distances. They are also more robust to normalization effects and using better models for the features improves the retrieval results compared to making only general assumptions. Experiments on a database of approximately 10000 images show that studying the feature distributions are important and this information should be used in designing feature normalization methods and similarity measures.
Selim Aksoy, Robert M. Haralick
CVPR2
2000 Assignment Problem in Edge Detection Performance Evaluation
abstract
We propose to use the combinatorial assignment problem to model the issue of associating ground-truth and declared edge pixels in the objective empirical performance evaluation of edge detectors. The assignment problem is adapted to the maximal assignment problem to incorporate the need for tolerating certain amount of localization error for the detected ground-truth pixels. The solution to this problem yields a maximal one-to-one association between ground-truth and declared edge pixels. Performance evaluation based on this association has the attitude of making the most positive interpretation of the declared edge map. Synthetic test data is used in the experiment to allow unambiguous subjective judgement of edge detection performance. The preciseness and reasonableness of the performance evaluation from the proposed method is observed. The usefulness of this method in other performance evaluation applications is also discussed.
Robert M. Haralick
CVPR2
2000 Two-Stage Robust Optical Flow Estimation
abstract
We formulate optical flow estimation as a two-stage regression problem. Based on characteristics of these two regression models and conclusions on modern regression methods, we choose a least trimmed squares followed by a weighted least squares estimator to solve the optical flow constraint (OFC); and at places where this one-stage robust method fails due to poor derivative quality, we use a least trimmed squares estimator to make the facet model fitting robust. This two-stage robust scheme produces significantly higher accuracy than non-robust algorithms and those only using robust methods at the OFC stage. On the synthetic data, the one-stage robust method has an average error of 7.7% against 24% of Black's and 19% of the pure LS method; and the two-stage robust method further reduces the error by half near motion boundaries. Advantages are also demonstrated on real data.
Robert M. Haralick
CVPR2
2000 Ultrasound Imaging Simulation and Echocardiographic Image Synthesis
abstract
Presents a ray tracing ultrasound imaging simulation method that accounts for the effects of reflection, scattering and attenuation. Two-dimensional echocardiographic images were synthesized using this method. Nonlinear effects introduced by the signal processing units in an ultrasound imaging system were also emulated. Three-dimensional triangular facet mesh models were employed in representing the heart structures and generating the synthetic two-dimensional echocardiographic images. The synthetic images agreed with the corresponding real ultrasound images in major ultrasound effects. No other prior work in echocardiographic image synthesis is known.
Mingzhou Song 0001, Robert M. Haralick, Florence H. Sheehan
ICIP2
2000 A Weighted Distance Approach to Relevance Feedback
abstract
Content-based image retrieval systems use low-level features like color and texture for image representation. Given these representations as feature vectors, similarity between images is measured by computing distances in the feature space. Unfortunately, these low-level features cannot always capture the high-level concept of similarity in human perception. Relevance feedback tries to improve the performance by allowing iterative retrievals where the feedback information from the user is incorporated into the database search. We present a weighted distance approach where the weights are the ratios of standard deviations of the feature values both for the whole database and also among the images selected as relevant by the user. The feedback is used for both independent and incremental updating of the weights and these weights are used to iteratively refine the effects of different features in the database search. Retrieval performance is evaluated using average precision and progress that are computed on a database of approximately 10,000 images and an average performance improvement of 19% is obtained after the first iteration.
Selim Aksoy, Robert M. Haralick, Faouzi Alaya Cheikh, Moncef Gabbouj
ICPR2
2000 Algorithm Performance Contest
abstract
This contest involved the running and evaluation of computer vision and pattern recognition techniques on different data sets with known groundwidth. The contest included three areas; binary shape recognition, symbol recognition and image flow estimation. A package was made available for each area. Each package contained either real images with manual groundtruth or programs to generate data sets of ideal as well as noisy images with known groundtruth. They also contained programs to evaluate the results of an algorithm according to the given groundtruth. These evaluation criteria included the generation of confusion matrices, computation of the misdetection and false alarm rates and other performance measures suitable for the problems. The paper summarizes the data generation for each area and experimental results for a total of six participating algorithms.
Selim Aksoy, Michael L. Schauf, Mingzhou Song 0001, Yalin Wang 0001, Robert M. Haralick, Jim R. Parker, Juraj Pivovarov, Dominik Royko, Changming Sun, Gunnar Farnebäck
ICPR6
2000 A Methodology for Special Symbol Recognitions
abstract
Presents a special symbol recognition system that incorporates the result of OCR to recognize the special symbols those nor handled by the current commercial OCR systems. Given a document image and the OCR output, we first refine the character coordinates produced by the OCR. Then, the special symbols are distinguished from the normal characters. Finally, we compute the features from the special symbol sub-images and a supervised classifier is used to assign the sub-images to one of the predefined special symbol categories. The system was tested on 5516 images from the National Library of Medicine. The evaluation results are reported in the paper.
Jisheng Liang, Vikram Chalana, Ihsin T. Phillips, Robert M. Haralick
ICPR4
2000 Vehicle Ground-Truth Database for the Vertical-View Ft. Hood Imagery
abstract
Reports work on building the ground-truth databases for the vehicles in the vertical-view Ft. Hood (VVFH) image data set. We briefly describe the protocols followed in manual annotation of the images, how the ground-truth information is inferred from the annotated image data, and the major entities in the ground-truth database. The vehicle detection performance of a few algorithms is evaluated using the data set. The entire data set of images and ground-truth is available on the Internet.
Robert M. Haralick
ICPR2
2000 Two Practical Issues in Canny's Edge Detector Implementation
abstract
We address two practical issues, namely smoothing factor selection and efficient implementation of thresholding with hysteresis, in implementing Canny's edge detector. The smoothing factor of the Gaussian kernel should be chosen to maximize the discrete version of Canny's original criteria. Thresholding with hysteresis should be implemented using an efficient connected component analysis algorithm. Following these suggestions in implementing Canny's edge detector will in general result in optimal edge detection quality and very significant reduction in running time for large images.
Robert M. Haralick
ICPR2
2000 Single View Computer Vision in Polyhedral World: Geometric Inference and Performance Characterization
abstract
An algorithm for making consistent 2-D to 3-D geometric inference in a polyhedral world using one perspective line drawing is described. Hypotheses are made on the internal angles of visible faces. The normals to the face planes are then determined. Valid normals lead to the reconstruction of the 3-D polyhedral world up to a scale factor. The performance of the algorithm is verified by using covariance matrix propagation. The experimental results show satisfactory performance. The general propagation formulae for the covariance matrix of both observed and inferred quantities are also derived.
Mingzhou Song 0001, Aiwen Guo, Robert M. Haralick
ICPR3
2000 Statistical-Based Approach to Word Segmentation
abstract
This paper presents a text word extraction algorithm that takes a set of bounding boxes of glyphs and their associated text lines of a given document and partitions the glyphs into a set of text words, using only the geometric information of the input glyphs. The algorithm is probability based. An iterative, relaxation-like method is used to find the partitioning solution that maximizes the joint probability. To evaluate the performance of our test word extraction algorithm, we used a 3-fold validation method and developed a quantitative performance measure. The algorithm was evaluated on the UW-III database of some 1600 scanned document image pages. An area-overlap measure was used to find the correspondence between the detected entities and the ground-truth. For a total of 827, 433 ground truth words, the algorithm identified and segmented 800, 149 words correctly, an accuracy of 97.43%.
Yalin Wang 0001, Robert M. Haralick, Ihsin T. Phillips
ICPR2
2000 Optical Flow from a Least-Trimmed Squares Based Adaptive Approach
abstract
Optical flow estimation can be formulated as two regression stages: derivative estimation and optical flow constraints (OFC) solving. Traditional approaches use least-squares at both stages and are sensitive to assumption violations. To improve estimation accuracy, especially near motion boundaries, we use a least trimmed squares (LTS) estimator to solve the OFC, obtaining a confidence measure for each estimate; and at place with low confidence, we use another LTS estimator to make the derivative estimation robust. This adaptive two-stage robust scheme has significantly higher accuracy than non-robust algorithms and those only using robust methods at the OFC stage. Advantages are illustrated on both synthetic and real data.
Robert M. Haralick
ICPR2
2000 Consistent Partition and Labelling of Text Blocks
Jisheng Liang, Ihsin T. Phillips, Robert M. Haralick
Pattern Anal. Appl.3
2000 Greedy Algorithm for Error Correction in Automatically Produced Boundaries from Low Contrast Ventriculograms
Jasjit S. Suri, Robert M. Haralick, Florence H. Sheehan
Pattern Anal. Appl.2
2000 A Statistical, Nonparametric Methodology for Document Degradation Model Validation
abstract
Printing, photocopying, and scanning processes degrade the image quality of a document. Statistical models of these degradation processes are crucial for document image understanding research. In this paper, we present a statistical methodology that can be used to validate local degradation models. This method is based on a nonparametric, two-sample permutation test. Another standard statistical device, the power function, is then used to choose between algorithm variables such as distance functions. Since the validation and the power function procedures are independent of the model, they can be used to validate any other degradation model. A method for comparing any two models is also described. It uses p-values associated with the estimated models to select the model that is closer to the real world.
Tapas Kanungo, Robert M. Haralick, Henry S. Baird, Werner Stuetzle, David Madigan
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Recursive binary dilation and erosion using digital line structuring elements in arbitrary orientations
abstract
Performing morphological operations such as dilation and erosion of binary images, using very long line structuring elements is computationally expensive when performed brute-force following definitions. We present two-pass algorithms that run at constant time for obtaining binary dilations and erosions with all possible length line structuring elements, simultaneously. The algorithms run at constant time for any orientation of the line structuring element. Another contribution of this paper is the use of the concept of orientation error between a continuous line and its discrete counterpart. The orientation error is used in determining the minimum length of the basic digital line structuring element used in obtaining what we call dilation and erosion transforms. The transforms are then thresholded by the length of the desired structuring element to obtain the dilation and erosion results. The algorithms require only one maximum operation for erosion transform and only one minimum operation for dilation transform, and one thresholding step and one translation step per result pixel. We tested the algorithms on Sun Sparc Station 10, on a set of 240x250 salt and pepper noise images with probability of a pixel being a 1-pixel set to 0.25, for orientations of the normals of the structuring elements in the range [pi/2,3pi/2] and lengths, in pixels, in the range [5,145]. We achieved a speed up of about 50 (and for special orientations theta in {(pi/2), (3pi/4), pi, (5pi/4), (3pi/2)} a speed up of about 100) when the structuring elements had lengths of 145 pixels, over the brute-force methods in these experiments. We compared the results of our dilation algorithm with those of the algorithm discussed by Soille et al. (see IEEE Trans. Pattern Anal. Machine Intell., vol.18, p.562-67, 1996) and showed that for binary dilation (and erosion since it is just the dilation of the background with the reflected structuring element) our algorithm performed better and achieved a speed up of about four when dilation or erosion transform alone is obtained.
Desikachari Nadadur, Robert M. Haralick
IEEE Trans. Image Process.2
1999 Linear vs. Quadratic Optimization Algorithms for Bias Correction of Left Ventricle Chamber Boundaries in Low Contrast Projection Ventriculograms Produced from Xray Cardiac Catheterization Procedure
Jasjit S. Suri, Robert M. Haralick, Florence H. Sheehan
CAIP2
1999 Graph-Theoretic Clustering for Image Grouping and Retrieval
abstract
Image retrieval algorithms are generally based on the assumption that visually similar images are located close to each other in the feature space. Since the feature vectors usually exist in a very high dimensional space, a parametric characterization of their distribution is impossible, so non-parametric approaches, like the k-nearest neighbor search, are used for retrieval. This paper introduces a graph-theoretic approach for image retrieval by formulating the database search as a graph clustering problem by using a constraint that retrieved images should be consistent with each other (close in the feature space) as well as being individually similar (close) to the query image. The experiments that compare retrieval precision with and without clustering showed an average precision of 0.76 after clustering, which is an improvement by 5.56% over the average precision before clustering.
Selim Aksoy, Robert M. Haralick
CVPR2
1999 A Statistically based, Highly Accurate Text-line Segmentation Method
abstract
This paper describes a text-line identification and segmentation technique that is probability based, where all probabilities are estimated from an extensive training set of various kind of measurements of distances between the terminal and non-terminal entities with which the algorithm works. The off-line probabilities estimated in the training then drive all decisions in the on-line segmentation algorithm. On the UW-III database of some 1600 scanned document image pages, having some 105020 text lines, the algorithm identifies and segments 104773 correctly, an accuracy of 99.76%.
Jisheng Liang, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1999 An Optimal Bayesian Hough Transform for Line Detection
abstract
In this paper, we describe a statistically efficient Hough transform technique with improved performance in accuracy and robustness. The proposed technique analytically computes the uncertainty of each feature point based on image noise, the procedure used for estimating edge orientation, and the specific parametric representation scheme of a line. Using the estimated uncertainty of each feature point, a Bayesian probabilistic scheme is introduced to compute the contribution of each feature point to the accumulator. A performance evaluation of the technique reveals its superior performance, especially for noisy images.
Robert M. Haralick
ICIP (2)2
1999 Quantitative Evaluation of Edge Detectors Using the Minimum Kernel Variance Criterion
abstract
In this paper, we introduce a new criterion for analytically evaluating different edge detectors (both gradient and zero-crossing based methods) without the need of ground-truth information. The criterion is based on the observation that most edge detectors make a decision of whether a pixel is an edgel or not based on the result of convolution of the image with a kernel. The variance of the convolution output therefore directly affects the performance of an edge detector. We show how to compute the variance of a convolution. We then describe results from comparing four well-known edge detectors using the proposed criterion.
Robert M. Haralick
ICIP (2)2
1999 A Statistically Efficient Method for Ellipse Detection
abstract
In this paper, we introduce a statistically efficient method for detecting ellipses in an image. Given a set of digital arc segments, we introduce geometric criteria to select possible pairs of arc segments belonging to the same ellipse. The selected arc pairs are subsequently validated or rejected based on certain statistical criteria via hypothesis testing. The advantages of the technique include: 1) the proposed criteria are scale-invariant; and 2) they can automatically adapt to the noise characteristics of each image and do not need to be adjusted empirically. Performance evaluation of the technique with real images demonstrates its good performance.
Robert M. Haralick
ICIP (2)2
1999 An Integrated Linear Technique for Pose Estimation from Different Geometric Features
abstract
Existing linear solutions for the pose estimation (or exterior orientation) problem suffer from a lack of robustness and accuracy partially due to the fact that the majority of the methods utilize only one type of geometric entity and their frameworks do not allow simultaneous use of different types of features. Furthermore, the orthonormality constraints are weakly enforced or not enforced at all. We have developed a new analytic linear least-squares framework for determining pose from multiple types of geometric features. The technique utilizes correspondences between points, between lines and between ellipse–circle pairs. The redundancy provided by different geometric features improves the robustness and accuracy of the least-squares solution. A novel way of approximately imposing orthonormality constraints on the sought rotation matrix within the linear framework is presented. Results from experimental evaluation of the new technique using both synthetic data and real images reveal its improved robustness and accuracy over existing direct methods.
Mauro S. Costa, Robert M. Haralick, Linda G. Shapiro
Int. J. Pattern Recognit. Artif. Intell.3
1999 An Automatic Closed-Loop Methodology for Generating Character Groundtruth for Scanned Documents
abstract
Character groundtruth for real, scanned document images is crucial for evaluating the performance of OCR systems, training OCR algorithms, and validating document degradation models. Unfortunately, manual collection of accurate groundtruth for characters in a real (scanned) document image is not practical because (i) accuracy in delineating groundtruth character bounding boxes is not high enough, (ii) it is extremely laborious and time consuming, and (iii) the manual labor required for this task is prohibitively expensive. Ee describe a closed-loop methodology for collecting very accurate groundtruth for scanned documents. We first create ideal documents using a typesetting language. Next we create the groundtruth for the ideal document. The ideal document is then printed, photocopied and then scanned. A registration algorithm estimates the global geometric transformation and then performs a robust local bitmap match to register the ideal document image to the scanned document image. Finally, groundtruth associated with the ideal document image is transformed using the estimated geometric transformation to create the groundtruth for the scanned document image. This methodology is very general and can be used for creating groundtruth for documents in typeset in any language, layout, font, and style. We have demonstrated the method by generating groundtruth for English, Hindi, and FAX document images. The cost of creating groundtruth using our methodology is minimal. If character, word or zone groundtruth is available for any real document, the registration algorithm can be used to generate the corresponding groundtruth for a rescanned version of the document.
Tapas Kanungo, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Robust extraction of characters from color scene image using mathematical morphology
abstract
Current character extraction systems for scene images are not robust for most real-world applications. In contrast, the system present here achieves robust performance by using morphological segmentation. This paper describes a new morphological segmentation algorithm-differential top-hats (DTT). In, addition, a complete system for extraction of characters from color scene images is presented. The system was verified through experiments on sequences of outdoor color images with varying external conditions. A high average extraction rate of 95% is obtained.
Lixu Gu, Toyohisa Kaneko, Naoki Tanaka, Robert M. Haralick
ICPR4
1998 The Torah code controversy
abstract
We briefly explain what a Torah code is and give an example of one. We show how to calculate the probability of its occurrence in a suitably defined population of texts. We explain how it is possible to misadvertise the computed probabilities, making it seem that the probability of a Torah code is very small. We briefly discuss the controversy and suggest that new more carefully controlled experiments are needed to resolve the controversy. We detail the protocol for a new experiment by a website reference.
Robert M. Haralick
ICPR1
1998 A statistical framework for geometric tolerancing manufactured parts
abstract
Visual inspection of a part from its image is always affected by image errors. Understanding how image errors affect measurement precision is therefore critical for accurate inspection. In this paper we lay out a statistical framework that allows one to explicitly handle image errors and characterize their impact on measurement precision. A hierarchical model is also proposed to model manufacturing and measurement errors. Based on the model, a Bayesian technique is introduced to statistically infer the geometric tolerances of a manufactured part.
Robert M. Haralick
ICPR2
1998 Using centroid covariance in target recognition
abstract
An automatic target recognition algorithm for low quality imagery is reported. Compact shaped targets are represented by their 2D silhouettes. Associated with each point on the silhouette, there is a direction roughly perpendicular to the local segment of the silhouette. The location of each silhouette point is assumed to be perturbed along that direction. A statistical technique is used to estimate the variance of that perturbation for the silhouette points of the hypothesized target. This variance is then used to estimate the location covariance of the target centroid. Target detection and recognition is based on this covariance. Target scaling, aspect, and rotation are not considered. Experiments on 31 FLIR images give a correct recognition of target identity and target location for 29 of the 31 images.
Robert M. Haralick
ICPR2
1998 Model-based shape recognition using recursive mathematical morphology
abstract
This paper introduces a size invariant method to recognize two-dimensional binary shapes using the recursive erosion transform. Using recursive morphology with multiple structuring elements, the method takes constant time per pixel regardless of the scale of the shape model, and also works on noisy images without requiring noise removal. Results from experiments on 100 noisy images show the methodology is able to detect every shape model's scale and position with 13 false alarms and five misdetections out of 254 total translated and scaled models.
Michael L. Schauf, Selim Aksoy, Robert M. Haralick
ICPR3
1998 Automatic quadratic calibration for correction of pixel classification boundaries to an accuracy of 2.5 millimeters: an application in cardiac imaging
abstract
Left ventricle (LV) boundaries estimated using automated pixel-based classifiers fall short in the apical zone and are not close the ground truth boundaries as delineated by the cardiologists. These errors are not random. They have a systematic positional and orientational bias, the boundary being under-estimated in the apex zone. They have a mean apex error of about 10.5 mm. The main reasons for bias errors are the poor contrast in the apex zone, nonhomogeneous mixing of the dye with the blood inside the left ventricle and interference of the diaphragm with inferior wall. Suri et al. (1996, 1997) showed, in earlier work, techniques to remove these bias errors using calibration schemes utilizing convex information of the left ventricle that consists of the aortic valve and the apex. This convex information was linear in nature. This paper presents an improved method of two calibration schemes using the quadratic nature of convex information. They are the quadratic identical coefficient and the quadratic independent coefficient methods. Using linear calibration, the independent coefficient method yields an error of 3.47 mm without utilizing apex, while quadratic calibration with convex information has a mean error of 2.49 mm. The goal set by the cardiologists is 2.5 mm.
Jasjit S. Suri, Robert M. Haralick, Florence H. Sheehan
ICPR2
1998 An improved Hough transform technique based on error propagation
abstract
This paper describes a Bayesian scheme for incrementing the Hough Transform (NT) accumulator to improve the performance of the HT, making it more robust to noise. The proposed technique analytically computes the uncertainty of each feature point based on its gradient and location in the image. Using the estimated uncertainty of each feature point, a Bayesian probabilistic scheme is proposed to compute the contribution of each feature point to the accumulator. A performance evaluation of our technique reveals its superior performance, especially for noisy images.
Robert M. Haralick
SMC2
1998 A segmentation-free approach to text recognition with application to Arabic text
Badr Al-Badr, Robert M. Haralick
Int. J. Document Anal. Recognit.2
1998 Breakpoint Detection Using Covariance Propagation
abstract
Presents a statistical approach for detecting breakpoints from chain encoded digital arcs. An arc point is declared as a breakpoint if the estimated orientations of the two fitted lines of the two arc segments immediately to the right and left of the arc point are significantly statistically different. The major contributions of this research include developing a method for analytically estimating the covariance matrix of the fitted line parameters and proposing a perturbation model to characterize the perturbation associated with each arc point.
Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Terrain Reconstruction from Multiple Views
Georgy L. Gimel'farb, Robert M. Haralick
CAIP2
1997 Corner Detection with Covariance Propagation
abstract
This paper presents a statistical approach for detecting corners from chain encoded digital arcs. An arc point is declared as a corner if the estimated parameters of the two fitted lines of the two arc segments immediately to the right and left of the are point are statistically significantly different. The corner detection algorithm consists of two steps: corner detection and optimization. While corner detection involves statistically identifying the most likely corner points along an arc sequence, corner optimization deals with improving the locational errors of the detected corners. The major contributions of this research include developing a method for analytically estimating the covariance matrix of the fitted line parameters and developing a hypothesis test statistic to statistically test the difference between the parameters of two fitted lines. Performance evaluation study showed that the algorithm is robust and accurate for complex images. It has an average misdetection rate of 2.5% and false alarm rate of 2.2% for the complex RADIUS images. This paper discusses the theory and performance characterization of the proposed corner detector.
Robert M. Haralick
CVPR2
1997 UW-ISL Document Image Analysis Toolbox: An Experimental Environment
abstract
A document image analysis toolbox including a collection of data structures and algorithms to support a variety of applications, is described in this paper. An experimental environment is built to allow developers to develop, test and optimize their algorithms and systems. Appropriate and quantitative performance metrics for each kind of information a document analysis technique infers have been developed. The performance of each algorithm has been evaluated based on these metrics and the UW-III document image database which contains a total of 1600 English document images randomly selected from scientific and technical journals.
Jisheng Liang, Richard Rogers, Robert M. Haralick, Ihsin T. Phillips
ICDAR3
1997 Left Ventricle Longitudinal Axis Fitting and its Apex Estimation using a Robust Algorithm and its Performance: A Parametric Apex Model
abstract
For complete automatic left ventricle border detection in a cardiac frame, the apex needs to be located. As the apex zone has less contrast and is harder to identify in the gray scale left ventriculograms, we use the left ventricle's longitudinal axis to assist in apex location. To automatically find the longitudinal axis of the left ventricle in any frame, we find the longest segment from either the anterior aspect of the aortic valve or the inferior aspect of the aortic valve to the left ventricle border. We assume that the ruled surface generated by the sequence of longitudinal axes through the cardiac cycle is of sufficiently simple form, so that the perturbation error, especially the large errors, between the automatically measured axis and the physician defined axis can, in part, be filtered out by a robust procedure. To discriminate those automatically determined axes that might differ significantly from the physician defined ground truth, we use Huber's weight function in an iterative reweighted least square robust fitting. The effects of inlier and outlier noise are discussed. We demonstrate that for 90% of 1200 frames of clinical data, the automatically determined apex location is less than an arc length distance of 1% of the ventricle border length from the ground truth apex location as delineated by the cardiologist.
Jasjit S. Suri, Robert M. Haralick, Florence H. Sheehan
ICIP (3)2
1997 Random perturbation models for boundary extraction sequence
Visvanathan Ramesh, Robert M. Haralick
Mach. Vis. Appl.2
1996 Finding Corresponding Points Based on Bayesian Triangulation
abstract
In this paper, we consider the problems of finding corresponding points from multiple perspective projection images (the correspondence problem), and estimating the 3-D point from which these points have arisen (the triangulation problem). We pose the triangulation problem as that of finding the Bayesian maximum, a posteriori estimate of the 3-D point, given its projections in N images, assuming a Gaussian error model for the image point co-ordinates and the camera parameters. We solve this by an iterative steepest descent method. We then consider the correspondence problem as a, statistical hypothesis verification problem. Given a set of 2-D points, under the hypothesis that the points are in correspondence, the MAP estimate of the 3-D point is computed. Based on the MAP estimate, we derive a statistical test for verifying this hypothesis. To find sets of corresponding points when multiple points in each of N images are given, we propose a method that does the Bayesian triangulation and hypothesis verification on each N-tuple of points, selecting those that pass the hypothesis test. We characterize the performance of the Bayesian triangulation in terms of the average distance of the triangulated 3-D point from the true 3-D point, and of the point correspondence method in terms of its misdetection and false alarm rates.
Anand S. Bedekar, Robert M. Haralick
CVPR2
1996 Zone classification using texture features
abstract
We consider the problem of zone classification in document image processing. Document blocks are labelled as text or nontext using texture features derived from a feature based interaction map (FBIM), a recently introduced general tool for texture analysis. The zone classification procedure proposed is tested on the comprehensive document image database UW-I created at the University of Washington in Seattle. Different classification procedures are considered. The performance ranges from 96% to 98% using 6 FBIM texture features only.
Dmitry Chetverikov, Jisheng Liang, József Kömüves, Robert M. Haralick
ICPR4
1996 Automatic generation of character groundtruth for scanned documents: a closed-loop approach
abstract
Character groundtruth for scanned document images is crucial for evaluating OCR system performance, training OCR algorithms, and validating document degradation models. Manual collection of accurate character groundtruth in a real (scanned) document image is not possible because (i) accuracy in delineating groundtruth character bounding boxes is too low, (ii) it is very laborious and time consuming and (iii) the manual labor required is prohibitively expensive. We present a closed-loop methodology. We first create ideal documents using a typesetting language. Next we create the groundtruth for the ideal document. The document is then printed, photocopied and scanned. A registration algorithm estimates the geometric transformation that registers the ideal document image to the scanned document image. Finally, groundtruth associated with the ideal document image is transformed using the estimated geometric transform to create the groundtruth for the scanned document image. This methodology is very general and can be used for creating groundtruth for documents typeset in any language, layout, font, and style. The cost of creating groundtruth using our methodology is minimal. We use this methodology to groundtruth 33 English documents consisting of over 62000 symbols. The procedure takes approximately 5 minutes per page on a SUN Sparc 10. We also use the method for Hindi and FAX documents.
Tapas Kanungo, Robert M. Haralick
ICPR2
1996 Correction of systematic errors in automatically produced boundaries from low-contrast ventriculograms
abstract
Poor contrast in the apex zone and nonhomogeneous mixing of the dye with the blood in the left ventricle causes the left ventricle pixel-based classifiers operating on ventriculograms to yield boundaries which are not close to ground truth boundaries as delineated by the cardiologist. They have a mean boundary error of about 6.4 mm and an error of about 12.5 mm in the apex zone. These errors have a systematic positional and orientational bias, the boundary being under-estimated in the apex zone. This paper discusses two calibration methods: the identical coefficient and the independent coefficient to remove these systematic biases. From these methods, we constitute a combined algorithm which reduces the boundary error compared to either of the calibration methods. The algorithm, in a greedy way, computes which and how many vertices of the left ventricle boundary can be taken from the computed boundary of each method to best improve the performance. The corrected boundaries have a mean error of less than 3.5 mm with a standard deviation of 3.4 mm over the approximately 6/spl times/10/sup 4/ vertices in the data set of 291 studies. Our methodology reduces the mean boundary error by 2.9 millimeters over the boundary produced by the classifier. We also show the calibration algorithm performs better in the apex zone where the dye is unable to reach. For end-diastole, it reduces the error in the apex zone by 8.5 millimeters over the pixel-based classifier boundaries.
Jasjit S. Suri, Robert M. Haralick, Florence H. Sheehan
ICPR2
1996 Document layout structure extraction using bounding boxes of different entitles
abstract
The paper presents an efficient technique for document page layout structure extraction and classification by analyzing the spatial configuration of the bounding boxes of different entities on the given image. The algorithm segments an image into a list of homogeneous zones. The classification algorithm labels each zone as test, table, line-drawing, halftone, ruling, or noise. The text lines and words are extracted within text zones and neighboring text lines are merged to form text blocks. The tabular structure is further decomposed into row and column items. Finally, the document layout hierarchy is produced from these extracted entities.
Jisheng Liang, Jaekyu Ha, Robert M. Haralick, Ihsin T. Phillips
WACV3
1996 Propagating Covariance in Computer Vision
abstract
This paper describes how to propagate approximately additive random perturbations through any kind of vision algorithm step in which the appropriate random perturbation model for the estimated quantity produced by the vision step is also an additive random perturbation. We assume that the vision algorithm step can be modeled as a calculation (linear or non-linear) that produces an estimate that minimizes an implicit scaler function of the input quantity and the calculated estimate. The only assumption is that the scaler function has finite second partial derivatives and that the random perturbations are small enough so that the relationship between the scaler function evaluated at the ideal but unknown input and output quantities and the observed input quantity and perturbed output quantity can be approximated sufficiently well by a first order Taylor series expansion. The paper finally discusses the issues of verifying that the derived statistical behavior agrees with the experimentally observed statistical behavior.
Robert M. Haralick
Int. J. Pattern Recognit. Artif. Intell.1
1996 Statistical estimation for exterior orientation from line-to-line correspondences
Chung-Nan Lee, Robert M. Haralick
Image Vis. Comput.2
1996 Gray-scale structuring element decomposition
abstract
Efficient implementation of morphological operations requires the decomposition of structuring elements into the dilation of smaller structuring elements. Zhuang and Haralick (1986) presented a search algorithm to find optimal decompositions of structuring elements in binary morphology. We use the concepts of Top of a set and Umbra of a surface to extend this algorithm to find an optimal decomposition of any arbitrary gray-scale structuring element.
Octavia I. Camps, Tapas Kanungo, Robert M. Haralick
IEEE Trans. Image Process.3
1995 Texture Anisotropy, Symmetry, Regularity: Recovering Structure and Orientation from Interaction Maps
abstract
We discuss a novel method for recovering fundamental, perceptually motivated structural features of a texture pattern: anisotropy, symmetry, and regularity. The method is based on extended spatial grey-level difference statistics which describe pairwise pixel interactions and yield an interaction map used to assess the overall two-dimensional structure of interactions and extract the significant short- and long-range interactions (intersample spacings). The new approach extends, in digital images, the notion of grey-level difference to arbitrary spacing vectors (i.e. any angle at any displacement). This provides the necessary background for precise anisotropy (or directionality) and symmetry analysis. Experimental results are shown with a set of Brodatz images that range from highly regular to patterns with weak regularity or anisotropy. A few especially interesting examples of recovering hardly visible structural features are given. Finally, the approach is applied to rotation-invariant texture classification.
Dmitry Chetverikov, Robert M. Haralick
BMVC2
1995 Segmentation-free word recognition with application to Arabic
abstract
This paper describes the design and implementation of a system that recognizes machine-printed Arabic words without prior segmentation. The technique is based on describing symbols in terms of shape primitives. At recognition time, the primitives are detected on a word image using mathematical morphology operations. The system then matches the detected primitives with symbol models. This leads to a spatial arrangement of matched symbol models. The system conducts a search in the space of spatial arrangements of models and outputs the arrangement with the highest posterior probability as the recognition of the word. The advantage of using this whole word approach versus a segmentation approach is that the result of recognition is optimized with regard to the whole word. Results of preliminary experiments using a lexicon of 42,000 words show a recognition rate of 99.4% for noise-free text and 73% for scanned text.
Badr Al-Badr, Robert M. Haralick
ICDAR2
1995 Simultaneous word segmentation from document images using recursive morphological closing transform
abstract
This paper describes a word segmentation algorithm which is based on the recursive morphological closing transform. The algorithm is trainable for any given document image population and is capable of detecting words on a document image simultaneously. We describe an experimental protocol to train and evaluate our word segmentation algorithm based on a set of layout ground-truthed document images. We also discussed a method to compare two sets of word bounding boxes-one from the ground truth and the other from the output of the word segmentation algorithm, and compute the numbers of miss, false, correct splitting, merging and spurious detections. The experimental results demonstrate that under the optimal algorithm parameter settings, the correct word detection percentage is about 95% on both the training and testing image populations. If this includes the splitting and merging detections, the detection percentage is about 99.4%.
Su S. Chen, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1995 Automatic text skew estimation in document images
abstract
This paper describes an algorithm to estimate the text skew angle in a document image. The algorithm utilizes the recursive morphological transforms and yields accurate estimates of text skew angles on a large document image data set. The algorithm computes the optimal parameter settings on the fly without any human interaction. In this automatic mode, experimental results indicate that the algorithm generates estimated text skew angles within 0.5/spl deg/ of the true text skew angles with a probability of 99%. To process a 300 dpi document image, the algorithm takes 10 seconds on SUN Sparc 10 machines.
Su S. Chen, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1995 Recursive X-Y cut using bounding boxes of connected components
abstract
A top-down page segmentation technique known as the recursive X-Y cut decomposes a document image recursively into a set of rectangular blocks. This paper proposes that the recursive X-Y cut be implemented using bounding boxes of connected components of black pixels instead of using image pixels. The advantage is that great improvement can be achieved in computation. In fact, once bounding boxes of connected components are obtained, the recursive X-Y cut is completed within an order of a second on Sparc-10 workstations for letter-sized document images scanned at 900 dpi resolution.
Jaekyu Ha, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1995 Document page decomposition by the bounding-box project
abstract
This paper describes a method for extracting words, textlines and text blocks by analyzing the spatial configuration of bounding boxes of connected component on a given document image. The basic idea is that connected components of black pixels can be used as computational units in document image analysis. In this paper, the problem of extracting words, textlines and text blocks is viewed as a clustering problem in the 2-dimensional discrete domain. Our main strategy is that profiling analysis is utilized to measure horizontal or vertical gaps of (groups of) components during the process of image segmentation. For this purpose, we compute the smallest rectangular box, called the bounding box, which circumscribes a connected component. Those boxes are projected horizontally and/or vertically, and local and global projection profiles are analyzed for word, textline and text-block segmentation. In the last step of segmentation, the document decomposition hierarchy is produced from these segmented objects.
Jaekyu Ha, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1995 Power functions and their use in selecting distance functions for document degradation model validation
abstract
Two document degradation models that model the perturbations introduced during the document printing and scanning process were proposed recently. Although degradation models are very useful, it is very important that we validate these models by comparing the synthetically generated images against real images. In recent past, two different validation procedures have also been proposed to validate such document degradation models. These validation procedures are functions of sample size and various distance functions. In this paper we outline a statistical methodology to compare the various validation schemes that result by using different distance functions. This methodology is general enough to compare any two validation schemes.
Tapas Kanungo, Robert M. Haralick, Henry S. Baird
ICDAR2
1995 Zone classification in a document using the method of feature vector generation
abstract
A document can be divided into zones on the basis of its content. For example, a zone can be either text or non-text. This paper describes an algorithm to classify each given document zone into one of nine different classes. Features for each zone such as run length mean and variance, spatial mean and variance, fraction of the total number of black pixels in the zone, and the zone width ratio for each zone are extracted. Run length related features are computed along four different canonical directions. A decision tree classifier is used to assign a zone class on the basis of its feature vector. The performance on an independent test set was 97%.
Ramaswamy Sivaramakrishnan, Ihsin T. Phillips, Jaekyu Ha, Suresh Subramanium, Robert M. Haralick
ICDAR5
1995 A Bayesian method for triangulation and its application to finding corresponding points
abstract
The problems of finding corresponding points from multiple perspective projection images, and estimating the 3-D points from which these points have arisen, are addressed. The problem of finding corresponding points is formulated as a hypothesis verification problem. Given a set of 2-D points, one from each of N perspective projection images, under the hypothesis that the points are projections of the same 3-D point, the coordinates of the 3-D point are estimated. The triangulation problem-the problem of estimating the coordinates of a 3-D point, given its projections in N perspective projection images-is posed as a Bayesian estimation problem, taking into account the uncertainties in the observed image points and the camera parameters. Based on the Bayesian estimate of the triangulated point, a statistical test is derived for verifying the hypothesis that the given set of image points is in correspondence. For finding N-tuples of corresponding points from N perspective projection images, this test can be used on each N-tuple of points to verify the hypothesis that that N-tuple of points is in correspondence, selecting those N-tuples that pass the hypothesis test. Experiments are described for characterizing the distance of the 3-D point estimated by the Bayesian triangulation from the true 3-D point, and characterizing the misdetection and false alarm rates of this method of finding corresponding points.
Anand S. Bedekar, Robert M. Haralick
ICIP2
1995 Constrained monotone regression of ROC curves and histograms using splines and polynomials
abstract
Receiver operating characteristics (ROC) curves have the property that they start at (0,1) and end at (1,0) and are monotonically decreasing. Furthermore, a parametric representation for the curves is more natural, since ROCs need not be single valued functions: they can start with infinite slope. We show how to fit parametric splines and polynomials to ROC data with the end-point and monotonicity constraints. Spline and polynomial representations provide us a way of computing derivatives at various locations of the ROC curve, which are necessary in order to find the optimal operating points. Density functions are not monotonic but the cumulative density functions are. Thus in order to fit a spline to a density function, we fit a monotonic spline to the cumulative density function and then take the derivative of the fitted spline function. Just as ROCs have end-point constraints, the density functions have end-point constraints. Furthermore, derivatives of splines are spline functions and can be computed in closed form. Thus smoothing of histograms can also be treated as a constrained monotone regression problem. The algorithms were implementation in a mathematical programming language called AMPL and results on sample data sets are given.
Tapas Kanungo, David M. Gay, Robert M. Haralick
ICIP3
1995 Receiver operating characteristic curves and optimal Bayesian operating points
abstract
The receiver operating characteristic curve is a standard method for reporting the performance of a system. In this paper we show how do choose the optimal operating point when we are given a receiver operating curve, the prior probabilities, and the economic gain matrix. Unlike earlier methods, we make no assumptions regarding underlying distributions.
Tapas Kanungo, Robert M. Haralick
ICIP (3)2
1995 Model-based morphology: the opening spectrum
Robert M. Haralick, Philip L. Katz
CVGIP Graph. Model. Image Process.1
1995 Optimal Sensor and Light Source Positioning for Machine Vision
Seungku Yi, Robert M. Haralick, Linda G. Shapiro
Comput. Vis. Image Underst.2
1995 Proteus: A reconfigurable computational network for computer vision
Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos
Mach. Vis. Appl.1
1995 A pattern recognition approach to the detection of complex edges
Dov Dori, Robert M. Haralick
Pattern Recognit. Lett.2
1995 Recursive erosion, dilation, opening, and closing transforms
abstract
A new group of recursive morphological transforms on the discrete space Z(2) are discussed. The set of transforms include the recursive erosion transform (RET), the recursive dilation transform (RDT), the recursive opening transform (ROT), and the recursive closing transform (RCT), The transforms are able to compute in constant time per pixel erosions, dilations, openings, and closings with all sized structuring elements simultaneously. They offer a solution to some vision tasks that need to perform a morphological operation but where the size of the structuring element has to be determined after a morphological examination of the content of the image. The computational complexities of the transforms show that the recursive erosion and dilation transform can be done in N+2 operations per pixel, where N is the number of pixels in the base structuring element. The recursive opening and closing transform can be done in 14N operations per pixel based on experimental results.
Su S. Chen, Robert M. Haralick
IEEE Trans. Image Process.2
1995 A methodology for quantitative performance evaluation of detection algorithms
Tapas Kanungo, Mysore Y. Jaisimha, John Palmer, Robert M. Haralick
IEEE Trans. Image Process.4
1994 Document image understanding: geometric and logical layout
abstract
Document image understanding encompasses the technology required to make paper documents equivalent to other computer exchange media like floppies, tapes, and CDROMs. The physical reader of the paper document is the scanner just like the physical reader of the floppy is the floppy drive and the physical reader of the tape cartridge is the tape cartridge drive, and the physical reader of the CDROM is the CDROM drive. In the survey presented, we restrict ourselves to documents such as business letters, forms, and scientific and technical articles such as those found in archival journals and technical conferences. Understanding such documents involves estimating the rotation skew of each document page, determining the geometric page layout, labeling blocks as text or non-text, determining the read order for text blocks, recognizing the text of text blocks through an OCR system, determining the logical page layout, and formatting the data and information of the document in a suitable way for use by a word processing system or by an information retrieval system.>
Robert M. Haralick
CVPR1
1994 Quantitative performance evaluation of thinning algorithms under noisy conditions
abstract
Thinning algorithms are an important sub-component in the construction of computer vision (especially for optical character recognition (OCR)) systems. Important criteria for the choice of a thinning algorithm include the sensitivity of the algorithms to input shape complexity and to the amount of noise. In previous work, we introduced a methodology to quantitatively analyse the performance of thinning algorithms. The methodology uses an ideal world model for thinning based on the concept of Blum ribbons. In this paper we extend upon this methodology to answer these and other experimental questions of interest. We contaminate the noise-free images using a noise model that simulates the degradation introduced by the process of xerographic copying and laser printing. We then design experiments that study how each of 16 popular thinning algorithms performs relative to the Blum ribbon gold standard and relative to itself as the amount of noise varies. We design statistical data analysis procedures for various performance comparisons. We present the results obtained from these comparisons and a discussion of their implications in this paper.>
Mysore Y. Jaisimha, Robert M. Haralick, Dov Dori
CVPR2
1994 Automatic selection of tuning parameters for feature extraction sequences
abstract
Computer vision algorithms are composed of different sub-algorithms often applied in sequence. Previous work on performance characterization illustrated how random perturbation models can be setup at various stages of an algorithm sequence for the input and output data. In this paper we address the issue of how one could utilize these random perturbation models in order to automate the selection of free parameters used in an algorithm sequence. We consider an operation sequence that involves edge finding, linking, corner finding and matching. Appropriate prior distributions for the parameters that describe the graytone/geometric characteristics of the image features are specified and validated by using an annotation process. The annotation process involves the manual specification (outlining) of the geometry and spatial extent of the image features. Statistics are gathered for parameters describing features of interest and non-interest (clutter features). The appropriate prior distributions are used to derive the theoretical expressions for feature detector performance over a given image population. These performance measures are then optimized to determine the tuning parameters for the feature detector(s).>
Visvanathan Ramesh, Robert M. Haralick, Desika C. Nadadur, Ken Thornton
CVPR2
1994 On vector quantization for fast facet edge detection
abstract
Presents an approach for performing edge detection which builds on prior work in fast facet edge detection using tree-structured vector quantization (TSVQ). The authors first extend the approach by using larger image vectors to reduce computational complexity by performing edge detection on multiple pixels at once. They then reduce the computational complexity of the edge detector without sacrificing performance by pruning the TSVQ with an edge detection-based criterion. They present results of edge detector performance on a sequence of images obtained from a mobile robot.>
Mysore Y. Jaisimha, Jill R. Goldschneider, Alexander E. Mohr, Eve A. Riskin, Robert M. Haralick
ICASSP (5)5
1994 An Automatic Algorithm for Text Skew Estimation in Document Images using Recursive Morphological Transforms
abstract
The text skew estimation algorithm utilizes recursive morphological transforms. With hand tuned parameters the algorithm produces estimated text skew angles which are within 0.1/spl deg/ of the true text skew angles 99% of the time. We also developed methodology to allow the algorithm to determine the optimal algorithm parameter settings on the fly without any human interaction. Under this automatic mode, our experimental results indicate that the algorithm generates estimated text skew angles which are within 0.5/spl deg/ of the true text skew angles 99% of the time. To process a 3300/spl times/2550 document image, the algorithm takes about 10 seconds on SUN Sparc 10 machines if discounting the document image file reading time.>
Su S. Chen, Robert M. Haralick
ICIP (1)2
1994 A Bayesian Corner Detector
abstract
A corner is modelled as the intersection of two lines. A corner point is that point on an input digital arc whose a posteriori probability of being a corner is the maximum among all the points on the arc. The performance of the corner detector is characterized by its false alarm rate, misdetection rate, and the corner location error all as a function of the noise variance, the included corner angle, and the arc length. Theoretical expressions for the quantities compare well with experimental results.>
Robert M. Haralick, Visvanathan Ramesh, Anand S. Bedekar, Ihsin T. Phillips
ICIP (2)2
1994 Propagating covariance in computer vision
abstract
This paper describes how to propagate approximately additive random perturbations through any kind of vision algorithm step in which the appropriate random perturbation model for the estimated quantity produced by the vision step is also an additive random perturbation. The author assumes that the vision algorithm step can be modeled as a calculation (linear or nonlinear) that produces an estimate that minimizes an implicit scaler function of the input quantity and the calculated estimate. The only assumption is that the scaler functions have finite second partial derivatives and that the random perturbations are small enough so that the relationship between the scaler function evaluated at the ideal but unknown input and output quantities and the observed input quantity and perturbed output quantity can be approximated sufficiently well by a first order Taylor series expansion. The paper finally discusses the issues of verifying that the derived statistical behavior agrees with the experimentally observed statistical behavior.
Robert M. Haralick
ICPR (1)1
1994 Finite random sets and morphology
abstract
In order to be able to optimally design morphological shape extraction algorithms operating on binary digital images, a probability theory is needed for finite random sets and probability relations that show how the probability changes as a finite random set is propagated through a morphological operation. In this paper, we develop such a theory for finite random sets. We then demonstrate how to apply this theory for calculating the probability that a set S perturbed by min or max noise N and dilated or eroded by a structuring element K is a subset, superset, or hits a given set R. In some cases we obtain exact results and in some cases we obtain bounds for the desired probability.
Robert M. Haralick, Su S. Chen, Xinhua Zhuang
ICPR (2)1
1994 Corner detection using the MAP technique
abstract
This paper describes a corner detection method that obtains maximum a posteriori estimates for the corner location in a given sequence of points. The authors model an ideal corner as the intersection of two ideal line segments. The perturbations on the sample points in a given line segment are assumed to be i.i.d Gaussian random variables of zero mean and variance /spl sigma//sup 2/. Further, the perturbations on the points are assumed to be orthogonal to the ideal line. The paper discusses the theory of the corner detector and an algorithm that extends the basic theory to handle multilinear segment arcs. Experiments were conducted according to a specific protocol and performance curves showing the location error versus the noise variance, the included corner angle, and the arc length, are provided. Performance characterization of the corner detector is also performed by plotting the false alarm rate and the misdetect rate versus the context window length and included corner angle. It is shown that the experimental results match the theoretical error propagation.
Robert M. Haralick, Visvanathan Ramesh
ICPR (1)2
1994 Review and analysis of solutions of the three point perspective pose estimation problem
Robert M. Haralick, Chung-Nan Lee, Karsten Ottenberg, Michael Nölle
Int. J. Comput. Vis.1
1994 Error propagation in machine vision
Seungku Yi, Robert M. Haralick, Linda G. Shapiro
Mach. Vis. Appl.2
1993 Bayesian Corner Detection
abstract
Corners play important roles in high level image understanding. They are the main features in many 2D or 3D image models associated with image understanding algorithms. The Bayesian corner detection method inputs a sequence of row-column pairs along an arc and outputs the corner positions and the corner included angles that maximize the a posteriori probability. Experiments on artificially generated sequences permit the measurement of errors of the estimated corner positions and included angles versus different noise perturbations, angles and line lengths respectively. 1
Robert M. Haralick
BMVC2
1993 Performance Characterization in Computer Vision
Robert M. Haralick
CAIP1
1993 Estimation of the position and orientation of a planar surface using multiple beams
abstract
A sensory system consisting of a camera and several laser beams is described. It is designed for estimating the parameters of a planar surface with respect to the camera. Estimation is possible when the beam positions and directions are known, as well as the image of the beam spots on a planar surface. The system is readily applicable in an industrial environment for automation, if it is attached to a robot arm and is connected to a computer.>
Jaekyu Ha, Robert M. Haralick
CVPR2
1993 Estimation of curvature from sampled noisy data
abstract
Estimation of curvature from noisy sampled data is a fundamental problem in digital arc segmentation. The facet approach to curvature estimation involves least square fitting the observed data points to a parametric cubic polynomial and the calculation of the curvature analytically from the fitted parametric coefficients. Due to the fitting, there exists systematic error or bias between curvature which is calculated analytically from the parameterization of a circle and one which is calculated analytically based on the coefficients of the fitted cubic polynomial, even when the data are sampled from a noiseless circle. It is shown how to compensate this bias by estimating it with the coefficients of the fitted cubic polynomial, which gives more accurate curvature value. Small perturbations are introduced to the sampled data from a noiseless circle, and the authors analytically trace how the perturbation propagates through coefficients of the fitted polynomials and results in perturbation error of the curvature.>
Chang-Kyu Lee, Robert M. Haralick, Koichiro Deguchi
CVPR2
1993 A quantitative methodology for analyzing the performance of detection algorithms
abstract
The authors present a methodology for designing experiments to characterize detection algorithms. The usual method is to vary parameters of the input images or parameters of the algorithms and then construct operating curves that relate the probability of misdetection and false alarm for each parameter setting. Such an analysis does not integrate the performance of the numerous operating curves. A methodology is outlined for summarizing many operating curves into a few performance curves. This methodology is adapted from the human psychophysics literature and is general to any detection algorithm. The central concept is to measure the effect of variables in terms of the equivalent effect of a critical signal variable. The methodology is demonstrated by comparing the performance of two line detection algorithms.>
Tapas Kanungo, Mysore Y. Jaisimha, John Palmer, Robert M. Haralick
ICCV4
1993 Radiographs as medical documents: automating standard measurements
abstract
A program has been initiated to automate certain radiographic measurements. The edge detection step is discussed. An edge detection model is presented for standard radiographs, which will be the basis for the measurements step. In particular, the edges of interest are modeled as a linear combination of basis functions. The problem of edge detection is then presented as a supervised pattern recognition problem in which the parameters of the basis functions are learned during the training phase, and the recognition phase uses these learned parameters to locate pixels that belong to edges. Since edges may appear in any direction, the mathematical tools to determine the gradient direction of the edge are developed.>
Dov Dori, Adam I. Harris, Gustavo Gambach, Robert M. Haralick
ICDAR4
1993 A methodology for the characterization of the performance of thinning algorithms
abstract
The authors measure the performance of thinning algorithms in the ideal world of noise-free Blum ribbons. Differences in the performance of thinning algorithms emerge even when they are applied on noise-free ribbon images. The authors define an error criterion function based on the Hausdorf distance that measures the deviation between the ideal and actual output of the thinning algorithms. They illustrate the process of performance evaluation by the application of ten state of the art thinning algorithms to the same (large) set of input images. They present results that show the mean value of the normalized error for each algorithm over each population of images. How performance of an algorithm varies over different populations of images is examined.>
Mysore Y. Jaisimha, Robert M. Haralick, Dov Dori
ICDAR2
1993 Global and local document degradation models
abstract
Two sources of document degradation are modeled: i) perspective distortion that occurs while photocopying or scanning thick, bound documents, and ii) degradation due to perturbations in the optical scanning and digitization process: speckle, blurr, jitter, thresholding. Perspective distortion is modeled by studying the underlying perspective geometry of the optical system of photocopiers and scanners. An illumination model is described to account for the nonlinear intensity change occuring across a page in a perspective-distorted document. The optical distortion process is modeled morphologically. First, a distance transform on the foreground is performed, followed by a random inversion of binary pixels where the probability of flip is a function of the distance of the pixel to the boundary of the foreground. Correlating the flipped pixels is modeled by a morphological closing operation.>
Tapas Kanungo, Robert M. Haralick, Ihsin T. Phillips
ICDAR2
1993 CD-ROM document database standard
abstract
The design of a comprehensive standard document database for machine-printed documents is presented. The effort to produce a series of carefully ground-truthed document databases to be issued on CD-ROMs is described in detail. The databases can be utilized by the OCR and document understanding community as a common platform to develop, test, and evaluate their algorithms.>
Ihsin T. Phillips, Su S. Chen, Robert M. Haralick
ICDAR3
1993 The implementation methodology for a CD-ROM English document database
abstract
Producing a database of scanned document images for development or evaluation of OCR and document image understanding algorithms is neither easy nor inexpensive. The authors first briefly describe the makeup of a database of scanned document images of scientific and technical documents written in English which are being produced in a CD-ROM format. Then, the authors concentrate on the implementation methodology used to prepare the database. The methodology gives the protocols for each step of the database preparation, and the error model used for the estimation of the ground-truth errors that may exist in the database is discussed.>
Ihsin T. Phillips, Jaekyu Ha, Robert M. Haralick, Dov Dori
ICDAR3
1993 Wigner Distribution for 2D Motion Estimation from Noisy Images
Vinay G. Vaidya, Robert M. Haralick
J. Vis. Commun. Image Represent.2
1993 MUSER: A prototype musical score recognition system using mathematical morphology
Bharath R. Modayur, Visvanathan Ramesh, Robert M. Haralick, Linda G. Shapiro
Mach. Vis. Appl.3
1992 Performance Characterization in Computer Vision
Robert M. Haralick
BMVC1
1992 Predicting expected gray level statistics of opened signals
abstract
The opening of a model signal with a convex, zero-height structuring element is studied empirically. Experiments are performed in which the input signal model parameters and the opening length are varied over an acceptable range and the corresponding grey level distributions in the opened signal are fit to Pearson distributions. Regressions are then used to relate the Pearson distribution parameters to the input parameters, resulting in equations that may be used to predict the effect of an opening. Characterization experiments show that the maximum absolute errors between actual and predicted cumulative distributions using these regression equations have a mean of 0.036 and a standard deviation of 0.011 (for a range of zero to one); the worst-case maximum absolute error encountered between the cumulative distributions is 0.066.>
Wendy Swan Costa, Robert M. Haralick
CVPR2
1992 Recursive opening transform
abstract
The opening transformation on N-dimensional discrete space Z/sup N/ is discussed. The transform efficiently computes the binary opening (closing) with any size structuring element. It also provides a quick way to calculate the pattern spectrum of an image. The pattern spectrum is found to be nothing more than a histogram of the opening transform. An efficient two-pass recursive opening transform algorithm is developed and implemented. The correctness of the algorithm is proved, and some experimental results are given. The results show that the execution time of the algorithm is a linear function of n, where n is the product of the number of points in the structuring element. When the input binary image size is 256*256 and 50% of the image is covered by the binary-one pixels, it takes approximately 250 ms to do an arbitrary sized line opening and approximately 500 ms to do an arbitrary size box opening on the Sun/Sparc II workstation (with C compiler optimization flag on).>
Robert M. Haralick, Su S. Chen, Tapas Kanungo
CVPR1
1992 Morphological decomposition of restricted domains: a vector space solution
abstract
Restricted domains, which are a restricted class of 2-D shapes, are defined. It is proved that any restricted domain can be decomposed as n-fold dilations of thirteen basis structuring elements and hence can be represented in a thirteen-dimensional space. This thirteen-dimensional space is spanned by the thirteen basis structuring elements comprising of lines, triangles, and a rhombus. It is shown that there is a linear transformation from this thirteen-dimensional space to an eight-dimensional space wherein a restricted domain is represented in terms of its side lengths. Furthermore, the decomposition in general is not unique, and all the decompositions can be constructed by finding the homogeneous solutions of the transformation and adding it to a particular solution. An algorithm for finding all possible decompositions is provided.>
Tapas Kanungo, Robert M. Haralick
CVPR2
1992 Visual inspection of machined parts
abstract
A CAD-model-based machine vision system for dimensional inspection of machine parts is described, with emphasis on the theory behind the system. The original contributions of this work are: (1) the use of precise definitions of geometric tolerances suitable for use in image processing, (2) the development of measurement algorithms corresponding directly to these definitions, (3) the derivation of the uncertainties in the measurement tasks, and (4) the use of this uncertainty information in the decision-making process. Initial experimental results have verified the uncertainty derivations statistically and proved that the error probabilities obtained by propagating uncertainties are lower than those obtainable without uncertainty propagation.>
Bharath R. Modayur, Linda G. Shapiro, Robert M. Haralick
CVPR3
1992 The image understanding environment program
abstract
The history of the image understanding environment (IUE) project, a five-year program to develop a common software environment for the development of algorithms and application systems, is reviewed. An overview of some of the data structures that are currently evolving as a specification for the IUE is provided. The ultimate goal of the project is to provide the basic data structures and algorithms that are required to carry state-of-the-art research in image understanding.>
Joseph L. Mundy, Thomas O. Binford, Terrance E. Boult, Allen R. Hanson, J. Ross Beveridge, Robert M. Haralick, Visvanathan Ramesh, Charles A. Kohl, Daryl T. Lawton, Doug Morgan, Keith Price, Tom Strat
CVPR6
1992 Random perturbation models and performance characterization in computer vision
abstract
It is shown how random perturbation models can be set up for a vision algorithm sequence involving edge finding, edge linking, and gap filling. By starting with an appropriate noise model for the input data, the authors derive random perturbation models for the output data at each stage of their example sequence. These random perturbation models are useful for performing model-based theoretical comparisons of the performance of vision algorithms. Parameters of these random perturbation models are related to measures of error such as the probability of misdetection of feature units, probability of false alarm, and the probability of incorrect grouping. Since the parameters of the perturbation model at the output of an algorithm are indicators of the performance of the algorithm, one could utilize these models to automate the selection of various free parameters (thresholds) of the algorithm.>
Visvanathan Ramesh, Robert M. Haralick
CVPR2
1992 Fast facet edge detection in image sequences using vector quantization
abstract
An application that uses vector quantization (VQ) to speed up the process of gradient magnitude edge detection for image sequences is presented. Because image VQ and this type of edge detection operate on block-based neighborhoods, it is possible to use VQ to perform the edge detection. The image is encoded with a VQ for which the edge/no edge decision has already been made for each block. The process of edge detection becomes a simple lookup of this information. The algorithm behaves as a trainable edge detector which has the advantage of having lower computational complexity than the facet edge detector. For a VQ with an average rate of 6 bits per vector, the method requires 55% of the multiplications and 62% of the additions of the conventional facet edge detector. It also enhances the quality of the output by rejecting low-contrast, high-frequency texture edges.>
Mysore Y. Jaisimha, Eve A. Riskin, Robert M. Haralick
ICASSP3
1992 Gray scale structuring element decomposition
abstract
Efficient implementation of morphological operations requires the decomposition of structuring elements into the dilation of smaller structuring elements. Zhuang and Haralick (1986) presented an algorithm to find optimal decompositions of structuring elements in binary morphology. In this paper the authors extend the algorithm to find optimal structuring element decomposition for gray scale morphology.>
Octavia I. Camps, Robert M. Haralick, Tapas Kanungo
ICPR (3)2
1992 Contextual decision making with degrees of belief
abstract
This paper gives a brief overview of the classical contextual pattern recognition problem. It is shown that the difficulty of this problem is really associated with the determination and use of the support of the joint prior distribution of the category labels. It is indicated how the consistent labeling framework can be used to define the support of the joint prior. It is then shown that this formulation of the problem can be generalized, and a general propositional logic framework which not only defines the support of the joint prior but also permits a calculation to be made evaluating the joint prior for any given set of joint labelings is introduced. It is shown that this formulation is indeed a formulation relating to the degree of belief. A formal system for the degree of belief in terms of an operational probability meaning is developed. The degree of belief in a proposition is exactly the probability with which the proposition can be asserted. It is then shown how the classical contextual problem can be generalized in the belief framework.>
Robert M. Haralick
ICPR (2)1
1992 Proteus: a reconfigurable computational network for computer vision
abstract
The Proteus architecture is a highly parallel MIMD, multiple instruction, multiple-data machine, optimized for large granularity tasks such as machine vision and image processing. The system can achieve 20 Giga-flops (80 Giga-flops peak). It accepts data via multiple serial links at a rate of up to 640 megabytes/second. The system employs a hierarchical reconfigurable interconnection network with the highest level being a circuit switched Enhanced Hypercube serial interconnection network for internal data transfers. The system is designed to use 256 to 1024 RISC processors. The processors use one megabyte external Read/Write Allocating Caches for reduced multiprocessor contention. The system detects, locates, and replaces faulty subsystems using redundant hardware to facilitate fault tolerance.>
Robert M. Haralick, Arun K. Somani, Craig M. Wittenbrink, Kenneth Cooper, Linda G. Shapiro, Ihsin T. Phillips, Jenq-Neng Hwang, Yung Hsi Yao, Chung-Ho Chen, Larry Yang, Brian Daugherty, Bob Lorbeski, Kent Loving, Tom Miller, Larye Parkins, Steve Soos
ICPR (4)1
1992 Morphological image processing on a token passing pyramid computer
abstract
Describes an implementation of a processor node on a Texas Instruments EVM16 microprogrammable machine. In addition, the authors give an algorithm for performing morphological image processing-a commonly used image processing tool today-on such a distributed control architecture. Finally, they give an overview of a software simulator being implemented for simulating the whole pyramid machine. A simulator is essential for testing the working of the architecture and the algorithms before building the hardware.>
Tapas Kanungo, Greg I. Chiou, Arun K. Somani, Robert M. Haralick
ICPR (4)4
1992 Object Recognition Using Prediction And Probabilistic Match
abstract
PREMIO is a CAD-based object recognition and localization system that uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the predictions to guide the search and reduce the search space. The purpose of this paper is to describe the matching algorithm and illustrative results. I Introduction Most feature-based matching schemes assume that all the features that are potentially visible in a view of an object will appear with equal probability. The resultant matching algorithms have to allow for “errors” without really understanding what the errors mean. PREMIO [2] is an object recognition/localization system that attempts to model some of the physical processes that can cause these “errors”. It uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. From these predictions, PREMIO calculates probabilities for each feature of being detected as a whole, being missed entirely, or breaking into pieces and conditional probabilities of the detection of one feature given the detection or nondetection of other features. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the probabilities to guide the search and prune the tree. The matching algorithm represents a large theoretical effort that is actually independent of the PREMIO system. The algorithm has been implemented as a C program and teated on data specifically generated to fit the abstract paradigm for the probabilistic search. The purpose of this paper is to describe the theory, the algorithm, and illustrative results.
Octavia I. Camps, Linda G. Shapiro, Robert M. Haralick
IROS3
1992 Performance characterization in image analysis: thinning, a case in point
Robert M. Haralick
Pattern Recognit. Lett.1
1992 Estimation of optimal morphological τ-opening parameters based on independent observation of signal and noise pattern spectra
Edward R. Dougherty, Robert M. Haralick, Yidong Chen 0002, Carsten Agerskov, Ulrik Jacobi, Poul Henrik Sloth
Signal Process.2
1991 Analysis and solutions of the three point perspective pose estimation problem
abstract
The major direct solutions to the three-point perspective pose estimation problems are reviewed from a unified perspective. The numerical stability of these three-point perspective solutions are discussed. It is shown that even in cases where the solution is not near the geometric unstable region considerable care must be exercised in the calculation. Depending on the order of the substitutions utilized, the relative error can change over a thousand to one. This difference is due entirely to the way the calculations are performed and not to any geometric structural instability of any problem instance. An analytical method is presented which produces a numerically stable calculation.>
Robert M. Haralick, Chung-Nan Lee, Kars Ottenburg, Michael Nölle
CVPR1
1991 Machine vision in the 1990s: Applications and how to get there
Dragutin Petkovic, J. Wilder, Masakazu Ejiri, Robert M. Haralick, Ramesh Jain 0001, Peter A. Ruetz, Jack Sklansky, C. W. Swonger
Mach. Vis. Appl.4
1991 Glossary of computer vision terms
Robert M. Haralick, Linda G. Shapiro
Pattern Recognit.1
1990 Toward the automatic generation of mathematical morphology procedures using predicate logic
abstract
A discussion is presented of the design of a system that can input a vision task specification and use its knowledge of the operations of mathematical morphology to automatically construct a procedure that can execute the task. To do this, the authors develop a predicate calculus representation to describe the essence of the states of all the images that are created during the execution of the morphological procedure and the states of the relationships among them. The authors translate the English descriptions of morphological procedures into predicate logic. In so doing they gain an understanding of the goal of each procedure and the exact conditions under which a procedure achieves its goal. With this knowledge of the operations of mathematical morphology represented in predicate logic, a search procedure can be used to automatically produce vision procedures.>
Hyonam Joo, Robert M. Haralick, Linda G. Shapiro
ICCV2
1990 Optimal affine-invariant point matching
abstract
Application of the affine-invariant point-matching scheme proposed by R. Hummel and H. Wolfson (1988) to the problem of recognizing and determining the pose of sheet metal parts is discussed. Attention is given to errors that can occur with this method due to quantization, stability, symmetry, and noise problems. These errors make the original affine-invariant matching technique unsuitable for use on the factory floor. An explicit noise model, which the Hummel and Wolfson technique lacks, is used. An optimal approach which overcomes these problems is then derived. The performance of the proposed algorithm under the influence of several distorting parameters is evaluated.>
Mauro S. Costa, Robert M. Haralick, Linda G. Shapiro
ICPR (1)2
1990 Automatic sensor and light source positioning for machine vision
abstract
The authors discuss an optimization approach to automatic sensor and light source positioning for a machine vision task where geometric measurement and/or object verification is important. The goal of the vision task is assumed to be specified in terms of edge visibility. There are two types of edge visibility: geometric edge visibility tells how much of the given edge is not occluded, and photometric visibility tells how much of the given edge has enough contrast to be detected in the image. A heuristic optimality criterion for the optimal sensor and light source position is defined in terms of these two edge visibilities. A preliminary experiment has been conducted to demonstrate the feasibility of the optimization approach. The result shows that the optimization problem formulated can be solved by mathematical programming techniques.>
Seungku Yi, Robert M. Haralick, Linda G. Shapiro
ICPR (1)2
1990 A highly robust estimator for computer vision
abstract
The authors present a highly robust estimator called an MF-estimator for general regression. It is argued that the kind of estimators needed by computer vision must be highly robust and that the classical robust estimators do not render a high robustness. It is explained that the high robustness becomes possible only through partially but completely modeling the unknown log likelihood function. Partial modeling explores a number of important heuristics implicit in the regression problem and takes place by taking them into consideration with the Bayes statistical decision rule, while maximizing the log likelihood function. Experiments with the simplest location estimation showed that the performance of the MF-estimator was superior to that of the classical M-estimator.>
Xinhua Zhuang, Robert M. Haralick
ICPR (1)2
1990 A neural net algorithm for maximum entropy image reconstruction
abstract
An open problem of whether or not there exists a solution to the maximum entropy image reconstruction problem is theoretically solved. The convergence of a maximum entropy image reconstruction algorithm proposed by X. Zhuang et al. (IEEE Trans. Acoust., Speech, Signal Processing, vol. ASSP-35, no.2, p.208-18, Feb. 1987) is proven and its fast convergence is derived.>
Xinhua Zhuang, Robert M. Haralick
ICPR (2)2
1990 Context dependent edge detection and evaluation
Robert M. Haralick, James Shih-Jong Lee
Pattern Recognit.1
1990 Multispectral image context classification using stochastic relaxation
abstract
A multispectral image context classification which is based on a stochastic relaxation algorithm and Markov-Gibbs random field is presented. The implementation of the relaxation algorithm is related to a form of optimization programming using annealing. The authors discuss the motivation for a Bayesian context-decision rule, and then use a Markov-Gibbs model to develop a contextual classification algorithm in which maximizing the posterior probability is based on stochastic relaxation. Experimental results that are based on simulated and real multispectral remote sensing images are presented to show how classification accuracy is greatly improved. The algorithm is highly parallel and exploits the equivalence between Gibbs distributions and Markov random fields.>
Mingchuan Zhang, Robert M. Haralick, James B. Campbell
IEEE Trans. Syst. Man Cybern.2
1989 Monocular vision using inverse perspective projection geometry: analytic relations
abstract
The author derives a variety of relationships, all in the reference frame of the camera, between 3-D points; 3-D lines; collections of 3-D lines; the angles between lines lying in common planes; the planes in which lines may lie; and the corresponding perspective projection of the 3-D points, lines and angles. These relationships are useful in many aspects of model-based vision and can serve as the geometric basis of a perspective-projection expert system. The derivations are outlined, and practical implications are briefly indicated.>
Robert M. Haralick
CVPR1
1989 Methodology for experimental computer vision
abstract
A key criticism in the experimental aspect in computer vision is that there are many reported experiments which illustrate results on only a few images and many experiments done all under essentially the same conditions. Such experiments are not sufficient. The author argues that the most informative kinds of experiments should state the set of controlled conditions under which a vision algorithm can be utilized and under which the vision algorithm performance exceeds some given specification. The planning document which describes the design for these experiments is called the experimental protocol. The specific elements of such a protocol are reviewed, stressing that the experimental data analysis plan must state how the hypothesis that the algorithm meets the specified requirement will be tested. The plan must be supported by theoretically developed statistical analysis which shows that an experiment carried out according to the experimental design and analyzed according to the data analysis plan will produce a statistical test itself having a given accuracy.>
Robert M. Haralick
CVPR1
1989 Binary shape recognition based on an automatic morphological shape decomposition
abstract
The authors present a technique for translation-invariant binary convex polygon shape recognition based on a morphological shape decomposition. The triangular shape primitives from the decomposition of convex shapes are used as features for shape recognition. The shape primitives are smaller and simpler than the templates of shapes; thus they are more efficient in representing shapes for discrimination. Maximum entropy reduction is used as an optimization criterion for selecting features from among the shape primitives at each node of a decision tree. Experiments on the classification of ten classes of noisy polygon shapes, where five replications per class were used for training and fifty replications per class were used for testing, achieved a recognition rate of 98.80% on the test set.>
Yunzin Zhao, Robert M. Haralick
ICASSP2
1989 A Simplex-Like Algorithm for the Relaxation Labeling Process
abstract
A simplex-like algorithm is developed for the relaxation labeling process. The algorithm is simple and has a fast convergence property which is summarized as a one-more-step theorem. The algorithm is based on fully exploiting the linearity of the variational inequality and the linear convexity of the consistent-labeling search space in a manner similar to the operation of the simplex algorithm in linear programming.>
Xinhua Zhuang, Robert M. Haralick, Hyonam Joo
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Improvement of kittler and illingworth's minimum error thresholding
Sungzoon Cho, Robert M. Haralick, Seungku Yi
Pattern Recognit.2
1989 Determining camera parameters from the perspective projection of a rectangle
Robert M. Haralick
Pattern Recognit.1
1989 Shape from shading using the facet model
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro
Pattern Recognit.2
1989 Matching topographic structures in stereo vision
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro
Pattern Recognit. Lett.2
1989 Pose estimation from corresponding point data
abstract
Solutions for four different pose estimation problems are presented. Closed-form least-squares solutions are given to the overconstrained 2D-2D and 3D-3D pose estimation problems. A globally convergent iterative technique is given for the 2D-perspective-projection-3D pose estimation problem. A simplified linear solution and a robust solution to the 2D-perspective-projection-2D-perspective-projection pose-estimation problem are also given. Simulation experiments consisting of millions of trials with varying numbers of pairs of corresponding points and varying signal-to-noise ratios (SNRs) with either Gaussian or uniform noise provide data suggesting that accurate inference of rotation and translation with noisy data may require corresponding point data sets with hundreds of corresponding point pairs when the SNR is less than 40 dB. The experimental results also show that the robust technique can suppress the blunder data which come from outliers or mismatched points.>
Robert M. Haralick, Hyonam Joo, Chung-Nan Lee, Xinhua Zhuang, Vinay G. Vaidya, Man Bae Kim
IEEE Trans. Syst. Man Cybern.1
1988 Context dependent edge detection
abstract
To simulate the edge perception ability of human eyes and detect scene edges from an image, context information must be used in the edge detection process. To accomplish the optimal use of the context, the authors introduce an edge detection scheme which uses the context of the whole image. The edge context for each pixel is the set of all row monotonically increasing paths through the pixel. The edge detector assigns a pixel that edge state having highest edge probability among all the paths. Experiments indicate the validity of the edge detector. Upon comparing the performance of the context dependent edge detector with the context free second directional derivative zero-crossing edge operator, the authors find that the context dependent edge detector is superior.>
Robert M. Haralick, James Shih-Jong Lee
CVPR1
1988 Binary morphology: working in the sampled domain
abstract
A description is given of the relationship between morphologically filtering and then sampling vs. sampling, and then morphologically filtering in the sampled domain. The authors also describe the relationship between morphologically filtering vs. sampling, morphologically filtering in the sampled domain, and then reconstructing. Unlike the standard communication sampling theory where for appropriately low-pass filtered images there is commutivity between sampling and filtering, this is not the case for appropriately morphologically simplified images. The relationship which does exist shows that the commutivity holds to within one sampling interval distance in the unsampled domain and to within two sampling intervals in the sampled domain.>
Robert M. Haralick, Xinhua Zhuang, Charlotte Lin, James Shih-Jong Lee
CVPR1
1988 From depth and optical flow to rigid body motion
abstract
The authors develop an algorithm to determine uniquely the rigid body motion from optical flow and depth, where the depth, however, does not involve any derivative information. Thus, the original assumptions made by D.H. Ballard and O.A. Kimball (1983) and by R.M. Haralick and X. Zhuang (1986) are relaxed. The proposed algorithm is appealing; in contrast to the existing linear optical flow-motion algorithms, it requires only three instead of eight optical flow image points.>
Xinhua Zhuang, Robert M. Haralick, Yunxin Zhao
CVPR2
1988 2D-3D pose estimation
abstract
A robust technique is described for solving the 3D-to-2D perspective projection pose estimation problem, given corresponding point sets. The technique has a considerable advantage over the least-squares (LS) technique in that with as many as 30% of the corresponding point pair matches completely incorrect the robust technique can determine the correct pose almost as accurately as the LS technique would if the LS technique were given only the correctly matched point pairs. The results of several hundred thousand experiments are reported that support the above conclusion.>
Robert M. Haralick, Hyonam Joo
ICPR1
1988 Context dependent edge detection
abstract
To obtain the optimal use of context, an edge detection scheme is introduced which uses the context of the whole image. The edge context for each pixel is the set of all row-monotonically-increasing paths through the pixel. The edge detector assigns to a pixel the edge state having the highest edge probability among all the paths. Experiments indicate the validity of the edge detector. A comparison with the performance of the context-free second-directional derivative zero-crossing edge operator shows that the context-dependent edge detector is superior.>
Robert M. Haralick, James Shih-Jong Lee
ICPR1
1988 Dynamic programming approach for context classification using the Markov random field
abstract
A set of multispectral image context classification techniques are discussed which are based on a recursive algorithm for optimal estimation of the state of a two-dimensional discrete Markov random field. The three recursive algorithms are forms of dynamic programming. Because the estimation equations of the recursive algorithm are quite simple, the computation complexity of the approach is low. It is shown that recursive contextual classification can improve classification performance, as compared to noncontextual classification. In addition, this algorithm has the advantage over other techniques in that it handles multispectral data naturally and simultaneously.>
Robert M. Haralick, Ming Chua Zhang, Roger W. Ehrich
ICPR1
1988 A simplified linear optic flow-motion algorithm
Xinhua Zhuang, Thomas S. Huang, Narendra Ahuja, Robert M. Haralick
Comput. Vis. Graph. Image Process.4
1988 Pipeline architectures for morphologie image analysis
A. Lynn Abbott, Robert M. Haralick, Xinhua Zhuang
Mach. Vis. Appl.2
1988 Performance assessment of near-perfect machines
Robert M. Haralick
Mach. Vis. Appl.1
1988 Gradient threshold selection using the facet model
Oscar A. Zuniga, Robert M. Haralick
Pattern Recognit.2
1988 A simple procedure to solve motion and structure from three orthographic views
abstract
In his work on motion analysis, Ullman (1979) showed that for the orthographic case, four-point correspondences over three views are sufficient to determine motion and structure of the four-point rigid configuration. However, his method is nonlinear and the conditions of uniqueness and convergence of his algorithm are not made clear. A simple linear algorithm to solve Ullman's problem of determining motion and structure from three orthographic views is presented. Necessary and sufficient conditions of a unique solution are also given.>
Xinhua Zhuang, Thomas S. Huang, Robert M. Haralick
IEEE J. Robotics Autom.3
1987 Machine vision mensuration
Robert M. Haralick
Comput. Vis. Graph. Image Process.1
1987 An Approximate Linear Time Propagate and Divide Theorem Prover for Propositional Logic
abstract
For theorem proving in the propositional calculus, we describe a semantic non-resolution based propagate and divide algorithm using the three principles of: reducing a potential theorem size by appropriate substitutions and simplifications; minimizing tree branching by using matching to carry out complete constraint propagation; and choosing alternatives which efficiently solve the subproblem currently being worked on. We provide experimental evidence based on almost 240,000 experiments indicating that this propagate and divide algorithm needs, on the average, approximately a linear amount of computation time as a function of problem size (length of well-formed formula) to prove a theorem in the propositional calculus. Furthermore, the computation time seems to depend only on the length of the well-formed formula and not on the number of distinct atoms in the well-formed formula.
Robert M. Haralick, Sheau-Huei Wu
Int. J. Pattern Recognit. Artif. Intell.1
1987 Insight: a Dataflow Language for Programming Vision Algorithms in a Reconfigurable Computational Network
abstract
Machine vision systems used in industrial applications must execute their algorithms in real time to perform such tasks as inspecting a wire bond or guiding a robot to install a part on a car body moving along a conveyer. The real time speed is achieved by employing simple-minded algorithms and by designing parallel architectures and parallel algorithms for some tasks. The majority of the work on parallel architectures has been limited to architectures that support image processing, but not mid- or high-level vision In order for more complex vision algorithms to execute in real time, a more flexible architecture is needed. Our conceptual approach to the problem is a reconfigurable computational network. Each configuration of the network implements an algorithm or class of algorithms A high-level language expresses the algorithms in a relational form that can be easily translated to the specification for a configuration. The language must be able to encode low-, mid-, and high-level vision algorithms and to efficiently handle not only pixel data, but also higher level structures. In this paper we describe a dataflow language called INSIGHT, which we have designed to meet these needs, and give several examples of parallel machine vision algorithms expressed in the language.
Linda G. Shapiro, Robert M. Haralick, Michael Goulish
Int. J. Pattern Recognit. Artif. Intell.2
1987 Image Analysis Using Mathematical Morphology
abstract
For the purposes of object or defect identification required in industrial vision applications, the operations of mathematical morphology are more useful than the convolution operations employed in signal processing because the morphological operators relate directly to shape. The tutorial provided in this paper reviews both binary morphology and gray scale morphology, covering the operations of dilation, erosion, opening, and closing and their relations. Examples are given for each morphological concept and explanations are given for many of their interrelationships.
Robert M. Haralick, Stanley R. Sternberg, Xinhua Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.1
1987 Comparison of a regular and an irregular decomposition of regions and volumes
John C. Fiala, Robert M. Haralick
Pattern Recognit.2
1987 Morphologic edge detection
abstract
Edge operators based on gray-scale morphologic operations are introduced. These operators can be efficiently implemented in near real time machine vision systems which have special hardware support for gray-scale morphologic operations. The simplest morphologic edge detectors are the dilation residue and erosion residue operators. The underlying motivation for these and some of their combinations are discussed and justified. Finally, the blur-minimum morphologic edge operator is defined. Its inherent noise sensitivity is less than the dilation or the erosion residue operators. Some experimental results are provided to show the validity of these morphologic operators. When compared with the enhancement/thresholding edge detectors and the cubic facet second derivative zero-crossing edge operator, the results show that all the edge operators have similar performance when the noise is small. However, as the noise increases, the second derivative zero-crossing edge operator and the blur-minimum morphologic edge operator have much better performance than the rest of the operators. The advantage of the blur-minimum edge operator is that it is less computationally complex than the facet edge operator.
James Shih-Jong Lee, Robert M. Haralick, Linda G. Shapiro
IEEE J. Robotics Autom.2
1987 Integrated Directional Derivative Gradient Operator
abstract
Accurate edge direction information is required in many image processing applications. A variety of operators for computing local edge direction have been proposed, many of them estimating a kind of gradient. These operators face two major problems. One problem is the inherent bias in their estimate of edge direction. The bias itself is a function of edge direction. Another problem is their sensitivity to the presence of noise in the image data. The second problem can be alleviated by an increase in the processing neighborhood size but usually at the expense of an increase in estimate bias and also inefrors in the processing of small or thin objects. An operator based on the cubic facet model is discussed, which reduces sharply both estimate bias and noise sensitivity with no increase in computational complexity. The measure of gradient strength is the maximum value of the integral of the first directional derivative taken over a rectangular or square neighborhood, the maximum being taken over all possible directions for the directional derivative. The line direction which maximizes the integral defines the new estimate of gradient direction. Experimental results show the superiority of this operator to others such as the Roberts operator, the Prewitt operator, the Sobel operator, and the standard cubic facet gradient operator for step edges and ramp edges. Under zero-noise conditions the 7× 7 integrated directional derivative gradient operator has a worst bias of less than 0.
Oscar A. Zuniga, Robert M. Haralick
IEEE Trans. Syst. Man Cybern.2
1986 A multi-threshold adaptive filtering for image enhancement
abstract
There is a compromise between noise removal and texture preservation in image enhancement. It is difficult to do an image enhancement task by using only one simple filter for a real world image which may consist of regions of various local activities. We describe a multi-threshold adaptive filter (MTA filter) for solving this problem in this paper. It uses a generalized gradient function which reflects the local contextual information as a cue to determine the nature of the filtering for each local neighborhood. In this way, several simple filters can be combined to form a more efficient and more flexible context dependent filter. As a result, specific filter is only applied to the region which is suitable for it. Thus, a balanced texture preserving and noise removal effect can be simultaneously achieved.
Jianzhong Qian, Kai-Bor Yu, Robert M. Haralick
ICASSP3
1986 From two-view motion equations to three-dimensional motion parameters and surface structure: A direct and stable algorithm
abstract
Given a set of corresponding points from two perspective projection images of a moving rigid object, this paper presents a direct and stable algorithm to solve for the parameters of the moving object. This involves determining the rotation, the translation direction and the relative depths of object points. Unlike previous algorithms (see [1]), the current algorithm does not require determining the mode of motion, i.e. if or not the motion is a pure rotation. As easily seen, it is hard or even impossible to determine the mode of motion in the presence of noise,
Xinhua Zhuang, Thomas S. Huang, Robert M. Haralick
ICRA3
1986 Computer vision theory: The lack thereof
Robert M. Haralick
Comput. Vis. Graph. Image Process.1
1986 Robot vision: Berthold Horn. MIT Press, Cambridge, MA 1986, 509 pp. $39.50
Robert M. Haralick
Comput. Vis. Graph. Image Process.1
1986 A note on "Rigid body motion from depth and optical flow"
Robert M. Haralick, Xinhua Zhuang
Comput. Vis. Graph. Image Process.1
1986 Shape from perspective: A rule-based approach
Prasanna G. Mulgaonkar, Linda G. Shapiro, Robert M. Haralick
Comput. Vis. Graph. Image Process.3
1986 Fast correlation registration method using singular value decomposition
abstract
A new, fast template-matching method using the Singular Value Decomposition (SVD) is presented. This approach involves a two-stage algorithm, which can be used to increase the speed of the matching process. In the first stage, the reference image is orthogonally separated by the SVD and then low-cost pseudo-correlation values are calculated. This reduces the number of computations to 2*N*L instead of N2L2, where L × L is the size of the reference image and N × N is the original image size. At the second stage, a small group of values near the maximum pseudo-correlation is selected. the true correlation for the small number of pixels in this group is them computed precisely in the second stage. Experimental and analytic results are presented to show how the computation complexity is greatly improved.
Mingchuan Zhang, Kai-Bor Yu, Robert M. Haralick
Int. J. Intell. Syst.3
1985 Two view motion analysis under a small perturbation
abstract
Given a set of corresponding points from a moving object is two perspective projection images. The paper completely solves the two view motion problem. We show how to use the corresponding point set to determine mode of motion, rotation, translation orientation and relative depths. Also we give a noise robust algorithm which works well under small perturbations.
Xinhua Zhuang, Robert M. Haralick
ICASSP2
1985 Comparing the laplacian zero crossing edge detector with the second directional derivative edge detector
abstract
We present evidence that the Laplacian Zero-Crossing operator does not use neighborhood information as effectively as the second directional derivative edge operator. We show that the use of a Gaussian smoother with standard deviation 5.0 for the Laplacian of Gaussian edge operator with a neighborhood size of 50 × 50 both misses and misplaces edges on an aerial image of a mobile home park. Our results of the Laplacian edge detector on a noisy test checkerboard image are also not as good as the second directional derivative edge operator. We conclude by discussing a number of open issues on edge operator evaluation.
Robert M. Haralick
ICRA1
1985 Two view motion analysis, stereo vision and a moving camera's positioning their equivalence and a new solution procedure
abstract
The three problems: Two view motion analysis, stereo vision and determining a moving camera's position are all the same problem. We explain their equivalence and introduce a noise robust procedure for solving these problems.
Xinhua Zhuang, Robert M. Haralick
ICRA2
1985 Understanding Objects with Curved Surfaces from a Single Perspective View of Boundaries
Shih Jong Lee, Robert M. Haralick, Ming Chua Zhang
Artif. Intell.2
1985 Computer Architecture for Solving Consistent Labelling Problems
abstract
Consistent labelling problems are a family of NP-complete constraint satisfaction problems such as school timetabling, for which a conventional computer may be too slow. There are a variety of techniques for reducing the elapsed time to find one or all solutions to a consistent labelling problem. In this paper we discuss and illustrate solutions consisting of special hardware to accomplish the required constraint propagation and an asynchronous network of intercommunicating computers to accomplish the tree search in parallel.
Julian R. Ullmann, Robert M. Haralick, Linda G. Shapiro
Comput. J.2
1985 Image segmentation techniques
Robert M. Haralick, Linda G. Shapiro
Comput. Vis. Graph. Image Process.1
1985 Topographic classification of digital image intensity surfaces using generalized splines and the discrete cosine transformation
Layne T. Watson, Thomas J. Laffey, Robert M. Haralick
Comput. Vis. Graph. Image Process.3
1985 Author's Reply
abstract
We present evidence that the Laplacian zero-crossing operator does not use neighborhood information as effectively as the second directional derivative edge operator. We show that the use of a Gaussian smoother with standard deviation 5.0 for the Laplacian of a Gaussian edge operator with a neighborhood size of 50 × 50 both misses and misplaces edges on an aerial image of a mobile home park. Contrary to Grimson and Hildreth's results, our results of the Laplacian edge detector on a noisy test checkerboard image are also not as good as the second directional derivative edge operator. We conclude by discussing a number of open issues on edge operator evaluation.
Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.1
1985 A Metric for Comparing Relational Descriptions
abstract
Relational models are frequently used in high-level computer vision. Finding a correspondence between a relational model and an image description is an important operation in the analysis of scenes. In this paper the process of finding the correspondence is formalized by defining a general relational distance measure that computes a numeric distance between any two relational descriptions-a model and an image description, two models, or two image descriptions. The distance measure is proved to be a metric, and is illustrated with examples of distance between object models. A variant measure used in our past studies is shown not to be a metric.
Linda G. Shapiro, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1985 Automatic inference of elevation and drainage models from a satellite image
abstract
Oblique illumination of irregular topography generates a pattern of highlighting and shadow, as the solar beam directly illuminates those slopes that face the sun while those on opposite sides of ridgelines are shadowed. On remotely sensed images these patterns appear as alternating dark and bright regions that reveal approximate positions of ridges and valleys. Knowledge of scene-specific variables (such as sun angle and elevation), general knowledge of geomorphology, atmospheric scattering, and spectral characteristics of landscapes permits reconstruction of the topography from its manifestation on the image. Raw image data record combined effects of topography, atmosphere, and diverse spectral reflections of surface materials. Our interpretation procedure isolates these several effects. From varied brightnesses caused by direct and indirect illumination, positions of ridges and valleys can be approximated. From variations in material reflectance, large rivers (channels with large areas of open water) can be detected. Finally, relative elevations can be estimated from analysis of drainage and ridge patterns using a strategy of "elevation growing" that assigns increasing elevation values to pixels as they are positioned at greater distances from rivers or other valley pixels already assigned elevations. From the estimated topographic elevations, it is possible to derive a network of drainage channels. Each stream segment in this network is labeled with information pertaining to its length, junction with other segments, direction of flow, and other properties. We then examine this network to detect logical inconsistencies in the labeling of stream segments, then apply a procedure that identifies the optimal labeling to yield the smallest error within the network.
Robert M. Haralick, James B. Campbell, Shyuan Wang
Proc. IEEE1
1985 Shape estimation from topographic primal sketch
Ting-Chuen Pong, Linda G. Shapiro, Robert M. Haralick
Pattern Recognit.3
1985 Parallel Computer Architectures and Problem Solving Strategies for the Consistent Labeling Problem
abstract
Parallel computer architectures and problem solving strategies for the consistent labeling problem are studied. Problem solving factors include: processor intercommunication methods, passing order, and selection of the initial processor to receive the problem.
Jeanette Tyler McCall, Joseph G. Tront, F. Gail Gray, Robert M. Haralick, William M. McCormack
IEEE Trans. Computers4
1984 Maximum entropy spectrum estimation from noisy correlation measurements
abstract
A computational method has been developed for solving the maximum entropy spectrum estimation with uncertainty in the correlation measurements This approach depends on first solving an unconstrained optimization problem and then an iterative search for the zero of the constraint equation along a well-defined trajectory. This trajectory is governed by a system of differential equations and the unconstrained optimal solution as the initial condition. The direction of search can be determined easily by evaluating the constraint equation at each iteration. In solving the trajectory, the Gauss-Seidel iterative algorithm is used to solve the symmetric system of equations. The choice of the previous estimate as the initial estimate greatly accelerates the convergence rate of this scheme.
Kai-Bor Yu, Xinhua Zhuang, Robert M. Haralick
ICASSP3
1984 A new approach to the solution of the maximum entropy image reconstruction problem
abstract
A new algorithm is developed for solving the maximum entropy (ME) image reconstruction problem. This approach involves solving a system of ordinary differential equations with appropriate initial conditions which can be computed easily. Instead of searching in the n+1-dimensional space as required for most ME algorithms, our approach involves solving a 1-dimensional search problem along a well-defined path. Moreover an efficient algorithm is developed to handle this search.
Xinhua Zhuang, Kai-Bor Yu, Robert M. Haralick
ICASSP3
1984 A hierarchical relational model for automated inspection tasks
abstract
An F-15 bulkhead is to be inspected by a computer system employing television cameras for vision and robot arms with tactile sensors for precise measurements. The system requires a suitable model of the object to be inspected. The model must contain very precise, accurate information for low-level vision and measurement processes and at the same time be useful to high-level vision and planning processes. This requirement makes most existing three-dimensional models useless. In this paper we define a hierarchical, relational model that we have developed to be used by the robot inspection system.
Linda G. Shapiro, Robert M. Haralick
ICRA2
1984 Experiments in segmentation using a facet model region grower
Ting-Chuen Pong, Linda G. Shapiro, Layne T. Watson, Robert M. Haralick
Comput. Vis. Graph. Image Process.4
1984 Automatic multithreshold selection
Shyuan Wang, Robert M. Haralick
Comput. Vis. Graph. Image Process.2
1984 Matching 'sticks, plates and blobs' objects using geometric and relational constraints
Prasanna G. Mulgaonkar, Linda G. Shapiro, Robert M. Haralick
Image Vis. Comput.3
1984 Digital Step Edges from Zero Crossing of Second Directional Derivatives
abstract
We use the facet model to accomplish step edge detection. The essence of the facet model is that any analysis made on the basis of the pixel values in some neighborhood has its final authoritative interpretation relative to the underlying gray tone intensity surface of which the neighborhood pixel values are observed noisy samples. With regard to edge detection, we define an edge to occur in a pixel if and only if there is some point in the pixel's area having a negatively sloped zero crossing of the second directional derivative taken in the direction of a nonzero gradient at the pixel's center. Thus, to determine whether or not a pixel should be marked as a step edge pixel, its underlying gray tone intensity surface must be estimated on the basis of the pixels in its neighborhood. For this, we use a functional form consisting of a linear combination of the tensor products of discrete orthogonal polynomials of up to degree three. The appropriate directional derivatives are easily computed from this kind of a function. Upon comparing the performance of this zero crossing of second directional derivative operator with the Prewitt gradient operator and the Marr-Hildreth zero crossing of the Laplacian operator, we find that it is the best performer; next is the Prewitt gradient operator. The Marr-Hildreth zero crossing of the Laplacian operator performs the worst.
Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.1
1984 Solving camera parameters from the perspective projection of a parameterized curve
Robert M. Haralick, Yu Hong Chu
Pattern Recognit.1
1984 Matching wire frame objects from their two dimensional perspective projections
Robert M. Haralick, Yu Hong Chu, Layne T. Watson, Linda G. Shapiro
Pattern Recognit.1
1984 Matching three-dimensional objects using a relational paradigm
Linda G. Shapiro, John D. Moriarty, Robert M. Haralick, Prasanna G. Mulgaonkar
Pattern Recognit.3
1984 Extraction of lines and regions from grey tone line drawing images
Layne T. Watson, K. Arvind, Roger W. Ehrich, Robert M. Haralick
Pattern Recognit.4
1983 Efficient Graph Automorphism by Vertex Partitioning
Glenn S. Fowler, Robert M. Haralick, F. Gail Gray, Charles Feustel, Charles M. Grinstead
Artif. Intell.2
1983 Ridges and valleys on digital images
Robert M. Haralick
Comput. Vis. Graph. Image Process.1
1983 An interpretation for probabilistic relaxation
Robert M. Haralick
Comput. Vis. Graph. Image Process.1
1983 An operating system interface for transportable image processing software
Scott Krusemark, Robert M. Haralick
Comput. Vis. Graph. Image Process.2
1983 Image random file access routines
Scott Krusemark, Robert M. Haralick
Comput. Vis. Graph. Image Process.2
1983 Decision Making in Context
abstract
From a Bayesian decision theoretic framework, we show that the reason why the usual statistical approaches do not take context into account is because of the assumptions made on the joint prior probability function and because of the simplistic loss function chosen. We illustrate how the constraints sometimes employed by artificial intelligence researchers constitute a different kind of assumption on the joint prior probability function. We discuss a couple of loss functions which do take context into account and when combined with the joint prior probability constraint create a decision problem requiring a combinatorial state space search. We also give a theory for how probabilistic relaxation works from a Bayesian point of view.
Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.1
1983 Texture analysis of aerial photographs
Ronald Lumia, Robert M. Haralick, Oscar A. Zuniga, Linda G. Shapiro, Ting-Chuen Pong, Far-Peing Wang
Pattern Recognit.2
1983 Peak noise removal by a facet model
Yoshifumi Yasuoka, Robert M. Haralick
Pattern Recognit.2
1983 The application of image analysis techniques to mineral processing
Ting-Chuen Pong, Robert M. Haralick, James R. Craig, Roe-Hoan Yoon, Woo-Zin Choi
Pattern Recognit. Lett.2
1983 Constrained Transform Coding and Surface Fitting
abstract
A constrained transform coding procedure is developed which is a combination of transform coding with differential pulse code modulation. The algorithm avoids block boundary mismatch errors, yet retains the coding efficiency of transform coding. A general theory of constrained transform coding is developed which includes the discrete cosine transformation and tensor products of splines as special cases. Results using the cosines and spines are given for two images. A complete discussion of the necessary linear algebra background is also given.
Layne T. Watson, Robert M. Haralick, Oscar A. Zuniga
IEEE Trans. Commun.2
1982 Significance of problem solving parameters on the performance of combinatorial algorithms on multi-computer parallel architectures
F. Gail Gray, William M. McCormack, Robert M. Haralick
ICPP3
1982 Understanding engineering drawings
Robert M. Haralick, David Queeney
Comput. Graph. Image Process.1
1982 Understanding engineering drawings
Robert M. Haralick, David Queeney
Comput. Graph. Image Process.1
1982 Image random file access routines
Scott Krusemark, Robert M. Haralick
Comput. Graph. Image Process.2
1982 An opening system interface for trasportable image processing software
Scott Krusemark, Robert M. Haralick
Comput. Graph. Image Process.2
1982 Organization of Relational Models for Scene Analysis
abstract
Relational models are commonly used in scene analysis systems. Most such systems are experimental and deal with only a small number of models. Unknown objects to be analyzed are usually sequentially compared to each model. In this paper, we present some ideas for organizing a large database of relational models. We define a simple relational distance measure, prove it is a metric, and using this measure, describe two organizational/access methods: clustering and binary search trees. We illustrate these methods with a set of randomly generated graphs.
Linda G. Shapiro, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1982 Design and Architectural Implications of a Spatial Information System
abstract
Image analysis, at the higher levels, works with extracted regions and line segments and their properties, not with the original raster data. Thus, a spatial information system must be able to store points, lines, and areas as well as their properties and interrelationships. In a previous paper (Shapiro and Haralick [17]), we proposed for this purpose an entity-oriented relational database system. In this paper, we describe our first experimental spatial information system which employs these concepts to store and retrieve watershed data for a portion of the state of Virginia. We describe the logical and physical design of the system and discuss the architectural implications.
Prashant D. Vaidya, Linda G. Shapiro, Robert M. Haralick, Gary J. Minden
IEEE Trans. Computers3
1981 Structural Descriptions and Inexact Matching
abstract
In this paper we formally define the structural description of an object and the concepts of exact and inexact matching of two structural descriptions. We discuss the problems associated with a brute-force backtracking tree search for inexact matching and develop several different algorithms to make the tree search more efficient. We develop the formula for the expected number of nodes in the tree for backtracking alone and with a forward checking algorithm. Finally, we present experimental results showing that forward checking is the most efficient of the algorithms tested.
Linda G. Shapiro, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1980 Sticks, Plates, and Blobs: A Three-Dimensional Object Representation for Scene Analysis
Linda G. Shapiro, John D. Moriarty, Prasanna G. Mulgaonkar, Robert M. Haralick
AAAI4
1980 Increasing Tree Search Efficiency for Constraint Satisfaction Problems
Robert M. Haralick, Gordon L. Elliott
Artif. Intell.1
1980 The Consistent Labeling Problem: Part II
abstract
In this second part of a two-part paper, we explore the power and complexity of the g=fKP and g=vKP class of look-ahead operators which can be used to speed up the tree search in the consistent labeling problem. For a specified K and P we show that the fixedpoint power of g=fKP and g=vKP is the same, that g=fKP+1 is at least as powerful as g=fKP, and that g=vK+1p is at least as powerful at g=fKP. Finally, we define a minimal compatibility relation and show how the standard tree search procedure for finding all the consistent labelings is quicker for a minimal relation. This leads to the concept of grading the complexity of compatibility relations according to how much look-ahead work it requires to reduce them to minimal relations and suggests that the reason look-ahead operators, such as Waltz filtering, work so well is that the compatibility relations used in practice are not very complex and are reducible to minimal or near minimal relations by a g=fKP or g=vKP look-ahead operator with small value for parameter P.
Robert M. Haralick, Linda G. Shapiro
IEEE Trans. Pattern Anal. Mach. Intell.1
1980 Neighbor gray levels as features in pixel classification
Narendra Ahuja, Azriel Rosenfeld, Robert M. Haralick
Pattern Recognit.3
1980 Editorial
Stephen D. Shapiro, Robert M. Haralick
Pattern Recognit.2
1980 Transportable Package Software
abstract
Abstract Package programs allow people who are not computer experts to use the power of machine computation for specialized purposes. Because the designers of scientific packages are more often experts in their scientific field than in computing, they may ignore issues of transportability and ease‐of‐use until too late, and produce a package that is difficult to use, and difficult to move to a different computer. In this paper we suggest some techniques to aid the scientific package designer. We suggest a kernel of routines that interface to the peculiar operating system of each machine, providing sophisticated but standard operating system services. This kernel makes the operating system of each computer appear identical and does not pose a difficult implementation problem. Above this interface all code can be machine‐independent, without sacrificing power and ease of use on any machine. We also suggest some novel organizations of processing routines designed to make the system easy to alter and extend. Complete independence of modules encourages centralization of tasks, which is both efficient and essential for easy extension. With this organization, the vast preponderance of package code can be written in machine‐independent ANSI FORTRAN. We suggest the use of a preprocessor like RATFOR to make this more pleasant. The paper closes with an application to an image‐processing package, in which a major problem is the flexible sequencing of processing routines.
Richard G. Hamlet, Robert M. Haralick
Softw. Pract. Exp.2
1979 Increasing Tree Search Efficiency for Constraint Satisfaction Problems
Robert M. Haralick, Gordon L. Elliott
IJCAI1
1979 The Consistent Labeling Problem: Part I
abstract
In this first part of a two-part paper we introduce a general consistent labeling problem based on a unit constraint relation T containing N-tuples of units which constrain one another, and a compatibility relation R containing N-tuples of unit-label pairs specifying which N-tuples of units are compatible with which N-tuples of labels. We show that Latin square puzzles, finding N-ary relations, graph or auto-mata homomorphisms, graph colorings, as well as determining satisfiability of propositional logic statements and solving scene and edge labeling problems, are all special cases of the general consistent labeling problem. We then discuss the various approaches that researchers have used to speed up the tree search required to find consistent labelings. Each of these approaches uses a particular look-ahead operator to help eliminate backtracking in the tree search. Finally, we define the ¿KP two-parameter class of look-ahead operators which includes, as special cases, the operators other researchers have used.
Robert M. Haralick, Linda G. Shapiro
IEEE Trans. Pattern Anal. Mach. Intell.1
1979 Decomposition of Two-Dimensional Shapes by Graph-Theoretic Clustering
abstract
This paper describes a technique for transforming a twodimensional shape into a binary relation whose clusters represent the intuitively pleasing simple parts of the shape. The binary relation can be defined on the set of boundary points of the shape or on the set of line segments of a piecewise linear approximation to the boundary. The relation includes all pairs of vertices (or segments) such that the line segment joining the pair lies entirely interior to the boundary of the shape. The graph-theoretic clustering method first determines dense regions, which are local regions of high compactness, and then forms clusters by merging together those dense regions having high enough overlap. Using this procedure on handdrawn colon shapes copied from an X-ray and on handprinted characters, the parts determined by the clustering often correspond well to decompositions that a human might make.
Linda G. Shapiro, Robert M. Haralick
IEEE Trans. Pattern Anal. Mach. Intell.2
1979 Texture pattern image generation by regular Markov chain
Ryuzo Yokoyama, Robert M. Haralick
Pattern Recognit.2
1978 Reduction operations for constraint satisfaction
Robert M. Haralick, Larry Davis 0001, Azriel Rosenfeld, David L. Milgram
Inf. Sci.1
1978 Editorial
Robert M. Haralick
Pattern Recognit.1
1978 Structural pattern recognition, homomorphisms, and arrangements
Robert M. Haralick
Pattern Recognit.1
1978 Arrangements, Homomorphisms, and Discrete Relaxation
abstract
We show how homomorphisms between arrangements, which are labeled N-ary relations, are the natural solutions to some problems requiring the integration of low-level and high-level information. Examples are given for problems in point matching, graph isomorphism, image matching, scene labeling, and spectral temporal classification of remotely sensed agricultural data. We develop characterization and representation theorems for N-ary relation homomorphisms, and we develop an algorithm consisting of a discrete relaxation method combined with a depth-first search to find such homomorphisms.
Robert M. Haralick, Jesse S. Kartus
IEEE Trans. Syst. Man Cybern.1
1977 Pattern discrimination using ellipsoidally symmetric multivariate density functions
Robert M. Haralick
Pattern Recognit.1
1977 Image Access Protocol for Image Processing Software
abstract
During the past decade a number of multiimage picture processing software packages have been put together. However, only a few of the references to picture processing systems discuss image data structure or input/output routines. This correspondence is a first step in a direction toward getting a communication process started by suggesting some specifications for a multiimage data format and standard input/output interface routines to access the image data.
Robert M. Haralick
IEEE Trans. Software Eng.1
1975 An Associative-Categorical Model of Word Meaning
Robert M. Haralick, Knut Ripken
Artif. Intell.1
1975 The pattern discrimination problem from the perspective of relation theory
Robert M. Haralick
Pattern Recognit.1
1974 The Diclique Representation and Decomposition of Binary Relations
abstract
The binary relation is often a useful mathematical structure for representing simple relationships whose essence is a directed connection. To better aid in interpreting or storing a binary relation we suggest a diclique decomposition. A diclique of a binary relation R is defined as an ordered pair ( I, O ) such that I × O ⊆ R and ( I, O ) is maximal. In this paper, an algorithm is described for determining the dicliques of a binary relation; it is proved that the set of such dicliques has a nice algebraic structure. The algebraic structure is used to show how dicliques can be coalesced, the relationship between cliques and dicliques is discussed, and an algorithm for determining cliques from dicliques is described.
Robert M. Haralick
J. ACM1
1974 A Measure for Circularity of Digital Figures
abstract
It is demonstrated that μR/σR, where R is a random variable of the distance between the center of the figure to any part of its perimeter, is a good measure for the circularity of a digital figure.
Robert M. Haralick
IEEE Trans. Syst. Man Cybern.1
1974 Comparative Study of a Discrete Linear Basis for Image Data Compression
abstract
Transform image data compression consists of dividing the image into a number of nonoverlapping subimage regions and quantizing and coding the transform of the data from each subimage. Karhunen-Loève, Hadamard, and Fourier transforms are most commonly used in transform image compression. This paper presents a new discrete linear transform for image compression which we use in conjunction with differential pulse-code modulation on spatially adjacent transformed subimage samples. For a set of thirty-three 64 × 64 images of eleven different categories, we compare the performancea of the discrete linear transform compression technique with the Karhunen-Loève and Hadamard transform techniques. Our measure of performance is the mean-squared error between the original image and the reconstructed image. We multiply the mean-squared error with a factor indicating the degree to which the error is spatially correlated. We find that for low compression rates, the Karhunen-Loève outperforms both the Hadamard and the discrete linear basis method. However, for high compression rates, the performance of the discrete transform method is very close to that of the Karhunen-Loève transform. The discrete linear transform method performs much better than the Hadamard transform method for all compression rates.
Robert M. Haralick, Karthikeyan S. Shanmugam
IEEE Trans. Syst. Man Cybern.1
1973 Glossary and index to remotely sensed image pattern recognition concepts
Robert M. Haralick
Pattern Recognit.1
1973 Textural Features for Image Classification
abstract
Texture is one of the important characteristics used in identifying objects or regions of interest in an image, whether the image be a photomicrograph, an aerial photograph, or a satellite image. This paper describes some easily computable textural features based on gray-tone spatial dependancies, and illustrates their application in category-identification tasks of three different kinds of image data: photomicrographs of five kinds of sandstones, 1:20 000 panchromatic aerial photographs of eight land-use categories, and Earth Resources Technology Satellite (ERTS) multispecial imagery containing seven land-use categories. We use two kinds of decision rules: one for which the decision regions are convex polyhedra (a piecewise linear decision rule), and one for which the decision regions are rectangular parallelpipeds (a min-max decision rule). In each experiment the data set was divided into two parts, a training set and a test set. Test set identification accuracy is 89 percent for the photomicrographs, 82 percent for the aerial photographic imagery, and 83 percent for the satellite imagery. These results indicate that the easily computable textural features probably have a general applicability for a wide variety of image-classification applications.
Robert M. Haralick, Karthikeyan S. Shanmugam, Its'hak Dinstein
IEEE Trans. Syst. Man Cybern.1
1973 A Computationally Simple Procedure for Imagery Data Compression by the Karhunen-Loeve Method
abstract
Of the several methods that have been proposed for imagery data compression, the Karhunen-Loève procedure minimizes the meansquare error between the original and reconstructed imagery data. In spite of its optimality property, the Karhunen-Loève procedure has not been widely used because of its computational complexity. The main difficulty is in the computation of the eigenvectors and the eigenvalues of the covariance matrix of the imagery data since the dimension of the covariance matrix is usually large. A computationally short procedure for calculating the eigenvalues and eigenvectors of the covariance matrix is presented. We show that the eigenvalues and eigenvectors of the N × N bisymmetric covariance matrix can be obtained from the eigenvalues and eigenvectors of two N/2 × N/2 submatrices. Since the eigenvector calculations are proportional to the third power of the matrix dimension, the proposed procedure reduces the computations by a factor of four.
Karthikeyan S. Shanmugam, Robert M. Haralick
IEEE Trans. Syst. Man Cybern.2
1971 Behavioral problems of deaf children: Clustering of variables using measures of association and similarity
Robert M. Haralick, Joy Gold Haralick
Pattern Recognit.1
1971 An Iterative Clustering Procedure
abstract
In many remote sensing applications millions of measurements can be made from a satellite at one time, and many times the data is of marginal value. In these situations clustering techniques might save much data transmission without loss of information since cluster codes may be transmitted instead of multidimensional data points. Data points within a cluster are highly similar so that interpretation of the cluster code can be meaningfully made on the basis of knowing what sort of data point is typical of those in the cluster. We introduce an iterative clustering technique; the procedure suboptimally minimizes the probability of differences between the binary reconstructions from the cluster codes and the original binary data. The iterative clustering technique was programmed for the GE 635 KANDIDATS (Kansas Digital Image Data System) and tested on two data sets. The first was a multi-image set. Twelve images of the northern part of Yellowstone Park were taken by the Michigan scanner system, and the images were reduced and run with the program. Thirty-thousand data points, each consisting of a binary vector of 25 components, were clustered into four clusters. The percentage difference between the components of the reconstructed binary data and the original binary data was 20 percent. The second data set consisted of measurements of the frequency content of the signals from lightning discharges. One hundred and thirty-four data measurements, each consisting of a binary vector of 32 components, were clustered into four clusters.
Robert M. Haralick, Its'hak Dinstein
IEEE Trans. Syst. Man Cybern.1