Jake K. Aggarwal

dblp:a/JakeKAggarwal · also J. K. Aggarwal · DBLP profile ↗
← Back
235ranked-venue papers
18as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 154 · 10 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 104 · 6 first-authorSystems, architecture and hardware · 26 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 14 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
60 papers
Video understanding and tracking · 65% 3D vision · 19% Robot navigation and mapping · 6%
Computer graphics and multimedia
30 papers
Image and video processing · 39% Audio and music processing · 32% Geometric modeling and processing · 15%
Human-computer interaction and pervasive computing
2 papers
Ubiquitous computing and smart environments · 60% Interaction techniques and input · 20% Usability and user experience research · 20%

Topics — the 30 heaviest of 151, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
activity recognition
0.742017
Multi-Type Activity Recognition from a Robot's Viewpoint · IJCAI 2017
Spatio-temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera · CVPR 2013
Stochastic Representation and Recognition of High-Level Group Activities · Int. J. Comput. Vis. 2011
Computer vision › Video understanding and tracking › activity recognition
human activity recognition
0.452011
Modeling human activities as speech · CVPR 2011
Semantic Representation and Recognition of Continued and Recursive Human Activities · Int. J. Comput. Vis. 2009
Spatio-temporal relationship match: Video structure comparison for recognition of complex human activities · ICCV 2009
Ubiquitous computing and smart environments › context recognition
activity recognition
0.212015
Robot-Centric Activity Prediction from First-Person Videos: What Will They Do to Me' · HRI 2015
Computer vision › Video understanding and tracking › action recognition › 3d action recognition
depth-based action recognition
0.212013
Spatio-temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera · CVPR 2013
Computer vision › Video understanding and tracking › action recognition
spatio-temporal interest points
0.212013
Spatio-temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera · CVPR 2013
Computer vision › Video understanding and tracking
action recognition
0.112011
A large-scale benchmark dataset for event recognition in surveillance video · CVPR 2011
Knowledge, reasoning and agents › Knowledge representation and reasoning › reasoning about action and change
action representation
0.112011
Modeling human activities as speech · CVPR 2011
Computer vision › Video understanding and tracking
event recognition
0.112011
A large-scale benchmark dataset for event recognition in surveillance video · CVPR 2011
Computer vision › Video understanding and tracking › activity recognition
group activity recognition
0.112011
Stochastic Representation and Recognition of High-Level Group Activities · Int. J. Comput. Vis. 2011
Computer vision › Video understanding and tracking › event recognition
video event recognition
0.112011
A large-scale benchmark dataset for event recognition in surveillance video · CVPR 2011
Audio and music processing › time-frequency analysis
spectrogram analysis
0.112011
Modeling human activities as speech · CVPR 2011
Computer vision › Video understanding and tracking
multi-object tracking
0.112008
Observe-and-explain: A new approach for multiple hypotheses tracking of humans and objects · CVPR 2008
Computer vision › Video understanding and tracking › multi-object tracking
multiple hypothesis tracking
0.112008
Observe-and-explain: A new approach for multiple hypotheses tracking of humans and objects · CVPR 2008
Computer vision › Video understanding and tracking › action recognition
human-object interaction recognition
0.112007
Hierarchical Recognition of Human Activities Interacting with Objects · CVPR 2007
Usability and user experience research
user assistance
0.112007
Robust Human-Computer Interaction System Guiding a User by Providing Feedback · IJCAI 2007
Biometric security › face recognition
3d face recognition
0.112007
3D Face Recognition Founded on the Structural Diversity of Human Faces · CVPR 2007
Biometric security
face recognition
0.112007
3D Face Recognition Founded on the Structural Diversity of Human Faces · CVPR 2007
Robotics › Robot navigation and mapping
localization
0.171998
Estimation of position and orientation from image sequence of a circle · ICRA 1997
Mobile robot self-location using model-image feature correspondence · IEEE Trans. Robotics Autom. 1996
Image Map Correspondence for Mobile Robot Self-Location Using Computer Graphics · IEEE Trans. Pattern Anal. Mach. Intell. 1993
Computer vision › Video understanding and tracking
egocentric video
0.112015
Robot-Centric Activity Prediction from First-Person Videos: What Will They Do to Me' · HRI 2015
Computer vision › Video understanding and tracking › activity recognition
complex activity recognition
0.112006
Recognition of Composite Human Activities through Context-Free Grammar Based Representation · CVPR (2) 2006
Computer vision › 3D vision
motion estimation
0.142007
Hierarchical Recognition of Human Activities Interacting with Objects · CVPR 2007
Moving obstacle detection from a navigating robot · IEEE Trans. Robotics Autom. 1998
Surface Correspondence and Motion Computation from a Sequence of Range Images · ICRA 1994
Computer vision › 3D vision › range sensing › depth sensing
depth camera
0.012013
Spatio-temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera · CVPR 2013
Computer vision › Image recognition and object detection
object recognition
0.032007
Hierarchical Recognition of Human Activities Interacting with Objects · CVPR 2007
Physics-based integration of multiple sensing modalities for scene interpretation · Proc. IEEE 1997
Object recognition in dense range images using a CAD system as a model base · ICRA 1990
Image and video processing
image segmentation
0.041997
A Bayesian Segmentation Framework for Textured Visual Images · CVPR 1997
The Integration of Image Segmentation Maps using Region and Edge Information · IEEE Trans. Pattern Anal. Mach. Intell. 1993
Image Interpretation Using Multiple Sensing Modalities · IEEE Trans. Pattern Anal. Mach. Intell. 1992
Computer vision › Video understanding and tracking
spatio-temporal features
0.012011
Modeling human activities as speech · CVPR 2011
Performance modeling and evaluation
benchmarking
0.012011
A large-scale benchmark dataset for event recognition in surveillance video · CVPR 2011
Robotics › Robot navigation and mapping › localization
vision-based localization
0.051997
Estimation of position and orientation from image sequence of a circle · ICRA 1997
Image Map Correspondence for Mobile Robot Self-Location Using Computer Graphics · IEEE Trans. Pattern Anal. Mach. Intell. 1993
Mobile robot self-location using model-image feature correspondence · IEEE Trans. Robotics Autom. 1996
Computer vision › 3D vision
3d reconstruction
0.081994
Generation of Architectural CAD Models Using a Mobile Robot · ICRA 1994
Integration of active and passive sensing techniques for representing three-dimensional objects · IEEE Trans. Robotics Autom. 1989
Computation of Surface Orientation and Structure of Objects Using Grid Coding · IEEE Trans. Pattern Anal. Mach. Intell. 1987
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian classification
0.021999
Automatic Tracking of Human Motion in Indoor Scenes Across Multiple Synchronized Video Streams · ICCV 1998
Tracking Human Motion in Structured Environments Using a Distributed-Camera System · IEEE Trans. Pattern Anal. Mach. Intell. 1999
Computer vision › 3D vision
stereo vision
0.051990
Stochastic Analysis of Stereo Quantization Error · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Binocular versus trinocular stereo · ICRA 1990
Quantization error in stereo imaging · CVPR 1988

Methods — techniques the papers use, named apart from their topics

onset representation · 0.4cascade histogram of time series gradients · 0.4relation history image · 0.3optimization · 0.3spatio-temporal interest features · 0.2evaluation metrics · 0.2boosted window classifiers · 0.2annotation · 0.2spatio-temporal interest point extraction · 0.2depth cuboid similarity feature · 0.2linear discriminant analysis · 0.1feedback · 0.1bayesian framework · 0.0region growing · 0.0physics-based modeling · 0.0multisensory simulation · 0.0markov random field · 0.0gabor wavelet · 0.0
YearPublicationVenuePosition
2024 Guest Editorial - Exploring the Potential of Fuzzy-Based Systems in Statistical Learning Theory
Jian Su 0001, Jake K. Aggarwal
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2017 Multi-Type Activity Recognition from a Robot's Viewpoint
abstract
The literature in computer vision is rich of works where different types of activities -- single actions, two persons interactions or ego-centric activities, to name a few -- have been analyzed. However, traditional methods treat such types of activities separately, while in real settings detecting and recognizing different types of activities simultaneously is necessary. We first design a new unified descriptor, called Relation History Image (RHI), which can be extracted from all the activity types we are interested in. We then formulate an optimization procedure to detect and recognize activities of different types. We assess our approach on a new dataset recorded from a robot-centric perspective as well as on publicly available datasets, and evaluate its quality compared to multiple baselines.
Ilaria Gori, Jake K. Aggarwal, Larry H. Matthies, Michael S. Ryoo
IJCAI2
2015 An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinement
abstract
We describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR.
Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai
AVSS26
2015 Robot-Centric Activity Prediction from First-Person Videos: What Will They Do to Me'
abstract
In this paper, we present a core technology to enable robot recognition of human activities during human-robot interactions. In particular, we propose a methodology for early recognition of activities from robot-centric videos (i.e., first-person videos) obtained from a robot's viewpoint during its interaction with humans. Early recognition, which is also known as activity prediction, is an ability to infer an ongoing activity at its early stage. We present an algorithm to recognize human activities targeting the camera from streaming videos, enabling the robot to predict intended activities of the interacting person as early as possible and take fast reactions to such activities (e.g., avoiding harmful events targeting itself before they actually occur). We introduce the novel concept of 'onset' that efficiently summarizes pre-activity observations, and design a recognition approach to consider event history in addition to visual features from first-person videos. We propose to represent an onset using a cascade histogram of time series gradients, and we describe a novel algorithmic setup to take advantage of such onset for early recognition of activities. The experimental results clearly illustrate that the proposed concept of onset enables better/earlier recognition of human activities from first-person videos collected with a robot.
Michael S. Ryoo, Thomas J. Fuchs, Jake K. Aggarwal, Larry H. Matthies
HRI4
2015 Robot-centric Activity Recognition from First-Person RGB-D Videos
abstract
We present a framework and algorithm to analyze first person RGBD videos captured from the robot while physically interacting with humans. Specifically, we explore reactions and interactions of persons facing a mobile robot from a robot centric view. This new perspective offers social awareness to the robots, enabling interesting applications. As far as we know, there is no public 3D dataset for this problem. Therefore, we record two multi-modal first-person RGBD datasets that reflect the setting we are analyzing. We use a humanoid and a non-humanoid robot equipped with a Kinect. Notably, the videos contain a high percentage of ego-motion due to the robot self-exploration as well as its reactions to the persons' interactions. We show that separating the descriptors extracted from ego-motion and independent motion areas, and using them both, allows us to achieve superior recognition results. Experiments show that our algorithm recognizes the activities effectively and outperforms other state-of-the-art methods on related tasks.
Ilaria Gori, Jake K. Aggarwal, Michael S. Ryoo
WACV3
2015 3D dynamic facial expression recognition using low-resolution videos
Jie Shao 0010, Ilaria Gori, Shaohua Wan 0002, Jake K. Aggarwal
Pattern Recognit. Lett.4
2014 Indoor Scene Recognition from RGB-D Images by Learning Scene Bases
abstract
In this paper, we propose a RGB-D indoor scene recognition method that has mainly two advantages as compared to existing methods. First, by training object detectors using RGB-D images and recognizing their spatial interrelationships, we not only achieve better object localization accuracy than using RGB images alone, but also obtain details as to how the objects are related to each other in a spatial manner, thus resulting in a more effective high-level feature representation of the scene known as the Objects and Attributes (O&A) representation. Second, we learn class-specific sub-dictionaries that capture the high-order couplings between the objects and attributes. In particular, elastic net regularization and geometric similarity constraint is imposed to increase the discriminative power of the sub-dictionaries. The proposed method is evaluated on two RGB-D datasets, the NYUD dataset and the B3DO dataset. Experiments show that superior scene recognition rate can be obtained using our method.
Shaohua Wan 0002, Changbo Hu, Jake K. Aggarwal
ICPR3
2014 Scene recognition by jointly modeling latent topics
abstract
We present a new topic model, named supervised Mixed Membership Stochastic Block Model, to recognize scene categories. In contrast to previous topic model based scene recognition, its key advantage originates from the joint modeling of the latent topics of adjacent visual words to promote the visual coherency of the latent topics. To ensure that an image is only a sparse mixture of latent topics, we use a Gini impurity based regularizer to control the freedom of a visual word taking different latent topics. We further show that the proposed model can be easily extended to incorporate the global spatial layout of the latent topics. Combined together, latent topic coherency and sparsity can rule out unlikely combinations of latent topics and guide classifier to produce more semantically meaningful interpretation of the scene. The model parameters are learned using Gibbs sampling algorithm, and the model is evaluated on three datasets, i.e. Scene-15, LabelMe, and UIUC-Sports. Experimental results demonstrate the superiority of our method over other related methods.
Shaohua Wan 0002, Jake K. Aggarwal
WACV2
2014 Spontaneous facial expression recognition: A robust metric learning approach
Shaohua Wan 0002, Jake K. Aggarwal
Pattern Recognit.2
2014 Human activity recognition from 3D data: A review
Jake K. Aggarwal
Pattern Recognit. Lett.1
2013 Spatio-temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera
abstract
Local spatio-temporal interest points (STIPs) and the resulting features from RGB videos have been proven successful at activity recognition that can handle cluttered backgrounds and partial occlusions. In this paper, we propose its counterpart in depth video and show its efficacy on activity recognition. We present a filtering method to extract STIPs from depth videos (called DSTIP) that effectively suppress the noisy measurements. Further, we build a novel depth cuboid similarity feature (DCSF) to describe the local 3D depth cuboid around the DSTIPs with an adaptable supporting size. We test this feature on activity recognition application using the public MSRAction3D, MSRDailyActivity3D datasets and our own dataset. Experimental evaluation shows that the proposed approach outperforms state-of-the-art activity recognition algorithms on depth videos, and the framework is more widely applicable than existing approaches. We also give detailed comparisons with other features and analysis of choice of parameters as a guidance for applications.
Jake K. Aggarwal
CVPR2
2013 Video event description in scene context
Changbo Hu, Qingshan Liu 0001, Jake K. Aggarwal
Neurocomputing4
2012 Toward a unified framework of motion understanding
Jake K. Aggarwal, Michael S. Ryoo
Image Vis. Comput.1
2011 AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance video
abstract
Summary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
AVSS8
2011 Modeling human activities as speech
abstract
Human activity recognition and speech recognition appear to be two loosely related research areas. However, on a careful thought, there are several analogies between activity and speech signals with regard to the way they are generated, propagated, and perceived. In this paper, we propose a novel action representation, the action spectrogram, which is inspired by a common spectrographic representation of speech. Different from sound spectrogram, an action spectrogram is a space-time-frequency representation which characterizes the short-time spectral properties of body parts' movements. While the essence of the speech signal is the variation of air pressure in time, our method models activities as the likelihood time series of action associated local interest patterns. This low-level process is realized by learning boosted window classifiers from spatially quantized spatio-temporal interest features. We have tested our algorithm on a variety of human activity datasets and achieved superior results.
Chia-Chih Chen, Jake K. Aggarwal
CVPR2
2011 A large-scale benchmark dataset for event recognition in surveillance video
abstract
We introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead.
Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai
CVPR8
2011 A Binary Stock Event Model for stock trends forecasting: Forecasting stock trends via a simple and accurate approach with machine learning
abstract
The volatile and stochastic characteristics of securities make it challenging to predict even tomorrow's stock prices. Better estimation of stock trends can be accomplished using both the significant and well-constructed set of features. Moreover, the prediction capability will gain momentum as we build the right model to capture unobservable attributes of the varying tendencies. In this paper, we propose a Binary Stock Event Model (BSEM) and generate features sets based on it in order to better predict the future trends of the stock market. We apply two learning models such as a Bayesian Naive Classifier and a Support Vector Machine to prove the efficiency of our approach in the aspects of prediction accuracy and computational cost. Our experiments demonstrate that the prediction accuracies are around 70-80% in one day predictions. In addition, our back-testing proves that our trading model outperforms well-known technical indicator based trading strategies with regards to cumulative returns by 30%-100%. As a result, this paper suggests that our BSEM based stock forecasting shows its excellence with regards to prediction accuracy and cumulative returns in a real world dataset.
Hyun Joon Jung, Jake K. Aggarwal
ISDA2
2011 Recognition of Human Activities
Jake K. Aggarwal
IWCIA1
2011 Stochastic Representation and Recognition of High-Level Group Activities
Michael S. Ryoo, Jake K. Aggarwal
Int. J. Comput. Vis.2
2010 Exploiting Geometric Restrictions in a PTZ Camera for Finding Point-orrespondences Between Configurations
abstract
A pan-tilt-zoom (PTZ) camera, fixed in location, may perform only rotational movements. There is a class of feature-based self-calibration approaches that exploit the restrictions on the camera motion in order to obtain accurate point-correspondences between two configurations of a PTZ camera. Most of these approaches require extensive computation and yet do not guarantee a satisfactory result. In this paper, we approach this problem from a different perspective. We exploit the geometric restrictions on the image planes, which are imposed by the motion restrictions on the camera. We present a simple method for estimating the camera focal length and finding the point-correspondences between two camera configurations. We compute pan-only, tilt-only and zoom-only correspondences and then combine the three to derive the geometrical relationship between any two camera configurations. We perform radial lens distortion estimation in order to calibrate distorted image coordinates. Our purely geometric approach does not require any intensive computations, feature tracking or training. However, our point-correspondence experiments show that, it still performs well-enough for most computer vision applications of PTZ cameras.
Birgi Tamersoy, Jake K. Aggarwal
AVSS2
2010 Human Shadow Removal with Unknown Light Source
abstract
In this paper, we present a shadow removal technique which effectively eliminates a human shadow cast from an unknown direction of light source. A multi-cue shadow descriptor is proposed to characterize the distinctive properties of shadows. We employ a 3-stage process to detect then remove shadows. Our algorithm improves the shadow detection accuracy by imposing the spatial constraint between the foreground subregions of human and shadow. We collect a dataset containing 81 human-shadow images for evaluation. Both descriptor ROC curves and qualitative results demonstrate the superior performance of our method.
Chia-Chih Chen, Jake K. Aggarwal
ICPR2
2010 Counting Vehicles in Highway Surveillance Videos
abstract
This paper presents a complete system for accurately and efficiently counting vehicles in a highway surveillance video. The proposed approach employs vehicle detection and tracking modules. In the detection module, an automatically trained binary classifier detects vehicles while providing robustness against view-point, poor quality videos and clutter. Efficient tracking is then achieved by a simplified multi-hypothesis approach. First an over-complete set of tracks is created considering every observed detection within a time interval. As needed, hypothesized detections are generated to force continuous tracks. Finally, a scoring function is used to separate the valid tracks in the over-complete set. Our tracking system achieved accurate results in significantly challenging highway surveillance videos.
Birgi Tamersoy, Jake K. Aggarwal
ICPR2
2010 A task-driven intelligent workspace system to provide guidance feedback
Michael S. Ryoo, Kristen Grauman, Jake K. Aggarwal
Comput. Vis. Image Underst.3
2009 Robust Vehicle Detection for Tracking in Highway Surveillance Videos Using Unsupervised Learning
abstract
This paper presents a novel approach to vehicle detection in highway surveillance videos. This method incorporates well-studied computer vision and machine learning techniques to form an unsupervised system, where vehicles are automatically ldquolearnedrdquo from video sequences. First an enhanced adaptive background mixture model is used to identify positive and negative examples. Then a classifier is trained with these examples. In the detection phase, both background subtraction and the classifier are used to achieve very accurate results while not compromising efficiency. We tested our method with very low-, medium- and high-quality, crowded and very crowded surveillance videos and got detection accuracies ranging between 90% to 96%.
Birgi Tamersoy, Jake K. Aggarwal
AVSS2
2009 Spatio-temporal relationship match: Video structure comparison for recognition of complex human activities
abstract
Human activity recognition is a challenging task, especially when its background is unknown or changing, and when scale or illumination differs in each video. Approaches utilizing spatio-temporal local features have proved that they are able to cope with such difficulties, but they mainly focused on classifying short videos of simple periodic actions. In this paper, we present a new activity recognition methodology that overcomes the limitations of the previous approaches using local features. We introduce a novel matching, spatio-temporal relationship match, which is designed to measure structural similarity between sets of features extracted from two videos. Our match hierarchically considers spatio-temporal relationships among feature points, thereby enabling detection and localization of complex non-periodic activities. In contrast to previous approaches to `classify' videos, our approach is designed to `detect and localize' all occurring activities from continuous videos where multiple actors and pedestrians are present. We implement and test our methodology on a newly-introduced dataset containing videos of multiple interacting persons and individual pedestrians. The results confirm that our system is able to recognize complex non-periodic activities (e.g. `push' and `hug') from sets of spatio-temporal features even when multiple activities are present in the scene.
Michael S. Ryoo, Jake K. Aggarwal
ICCV2
2009 Semantic labeling of track events using time series segmentation and shape analysis
abstract
This paper presents a novel framework for applying semantic labels to events within a track. A track is a two-dimensional (2D) or a three-dimensional (3D) signal in time where each point of the signal is the x and y (and z) centroid spatial coordinate of an object at a specific frame of the video. The track may be generated by the movement of a vehicle, person, or object. In the 2D case, the signal is decomposed into x and y time series for use in one-dimensional time series segmentations. Then the results of the two segmentations are combined to produce a 2D signal segmentation of the track which results in unique events to be labeled. The Procrustes measure, from shape analysis, is employed along with template matching to find the most likely trajectory of each individual event. Once each event is labeled with a semantic description from the template, we enhance the label using other basic measurements based on the track. The application of our framework on 4 vehicle tracks from original videos is shown to display the efficacy of our method.
Josh Harguess, Jake K. Aggarwal
ICIP2
2009 Patch-based face recognition from video
abstract
Face recognition from video has been extensively studied in recent years. Intuitively, video provides more information than a single image. But problems such as variation in pose and occlusion still remain. When a face is partially occluded, handling the occluded part of the face is an especially challenging task. In this paper, we propose a novel method to recognize a face from video based on face patches. First, face patches are cropped from the video frame by frame. Then, face patches are matched to an overall face model and stitched together. By accumulating the patches, a reconstructed face is built which is used in recognition. We test our method in two experiments. In the first experiment, a still face database is used by randomly occluding parts of the face and using the remaining face patches in recognition. The experiments show that our method achieves a comparable recognition rate with the recognition rate from the whole face image. In the second experiment, the method is tested on video sequences. We reach a recognition rate of 81%, while there is still missing data in the reconstructed face.
Changbo Hu, Josh Harguess, Jake K. Aggarwal
ICIP3
2009 Fusing face recognition from multiple cameras
abstract
Face recognition from video has recently received much interest. However, several challenges for such a system exist, such as resolution, occlusion (from objects or self-occlusion), motion blur, and illumination. The aim of this paper is to overcome the problem of self-occlusion by observing a person from multiple cameras with uniquely different views of the person's face and fusing the recognition results in a meaningful way. Each camera may only capture a part of the face, such as the right or left half of the face. We propose a methodology to use cylinder head models (CHMs) to track the face of a subject in multiple cameras. The problem of face recognition from video is then transformed to a still face recognition problem which has been well studied. The recognition results are fused based on the extracted pose of the face. For instance, the recognition result from a frontal face should be weighted higher than the recognition result from a face with a yaw of 30°. Eigenfaces is used for still face recognition along with the average-half-face to reduce the effect of transformation errors. Results of tracking are further aggregated to produce 100% accuracy using video taken from two cameras in our lab.
Josh Harguess, Changbo Hu, Jake K. Aggarwal
WACV3
2009 Semantic Representation and Recognition of Continued and Recursive Human Activities
Michael S. Ryoo, Jake K. Aggarwal
Int. J. Comput. Vis.2
2009 Recognizing Persons Climbing Fences
abstract
In this paper, we present two approaches to recognize a special type of human action, e.g. climbing a fence, from continuous monocular videos. Both approaches represent human activities as time series of discrete observations. In the first approach, we extract a feature vector from each frame, based on the relative configuration of human body parts against a fence. In the second approach, we propose a novel concept called "stable contact", to decompose the continuous human activities into adjacent but disjoint time intervals called primitive intervals. We then define a primitive motion unit (PMU) on each primitive interval with three attributes, thus converting each continuous activity into a sequence of PMUs. With the time series representations, we utilize Hidden Markov Model (HMM) techniques for recognition, including decoding in the first approach and searching in the second approach. Experimental results, on a set of image sequences with mixed human actions including walking and climbing a fence, show the effectiveness of the proposed methods. We compare the two approaches and discuss their applicability in different scenarios.
Elden Yu, Jake K. Aggarwal
Int. J. Pattern Recognit. Artif. Intell.2
2009 Detection of object abandonment using temporal logic
Medha Bhargava, Chia-Chih Chen, Michael S. Ryoo, Jake K. Aggarwal
Mach. Vis. Appl.4
2009 Real-Time Illegal Parking Detection in Outdoor Environments Using 1-D Transformation
abstract
With decreasing costs of high-quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in real time by applying a novel image projection that reduces the dimensionality of the data and, thus, reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require 2-D data. We evaluate the performance of our algorithm using the i-LIDS vehicle detection challenge datasets as well as videos we have taken ourselves. These videos test the algorithm in a variety of outdoor conditions, including nighttime video and instances of sudden changes in weather.
Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal
IEEE Trans. Circuits Syst. Video Technol.4
2008 Observe-and-explain: A new approach for multiple hypotheses tracking of humans and objects
abstract
This paper presents a novel approach for tracking humans and objects under severe occlusion. We introduce a new paradigm for multiple hypotheses tracking, observe-and-explain, as opposed to the previous paradigm of hypothesize-and-test. Our approach efficiently enumerates multiple possibilities of tracking by generating several likely dasiaexplanationspsila after concatenating a sufficient amount of observations. The computational advantages of our approach over the previous paradigm under severe occlusions are presented. The tracking system is implemented and tested using the i-Lids dataset, which consists of videos of humans and objects moving in a London subway station. The experimental results show that our new approach is able to track humans and objects accurately and reliably even when they are completely occluded, illustrating its advantage over previous approaches.
Michael S. Ryoo, Jake K. Aggarwal
CVPR2
2008 An adaptive background model initialization algorithm with objects moving at different depths
abstract
Background subtraction is an essential element in most object tracking and video surveillance systems. The success of this low-level processing step is highly dependent on the quality of the background model maintained. Gutchess et al. [4] proposed a novel background initialization algorithm that utilizes local optical flow information to locate the stable interval (of intensity values) which is most likely to display background. However, it is found that the accuracy of the computed background is rather sensitive to the parameters used. In addition, their algorithm is not able to handle the scenario where objects are moving at different depths. In this paper, we propose an algorithm which is adaptive to the input sequence and is able to equalize the uneven effect caused by different object depths. Our algorithm is successfully tested on complex indoor and outdoor scenes with promising results.
Chia-Chih Chen, Jake K. Aggarwal
ICIP2
2008 Recognition of box-like objects by fusing cues of shape and edges
abstract
Boxes are the universal choice for packing, storage, and transportation. In this paper we propose a template-based algorithm for recognition of box-like objects, which is invariant to scale, rotation and translation as well as robust to patterned surfaces and moderate occlusions. The algorithm first over-segments the input image to partition objects into pieces. Based on the smoothness property of surface texture, candidates for component segments of boxes are selected. Guided by a template trained linear discriminant analysis (LDA) classifier, box-like segments are reassembled from these segments of interests. For each box-like segment, we estimate its probability of being a 2D projection of a 3D box model upon the extracted contour and inner edges. Experimental results demonstrate high detection accuracy of boxes and reliable recovery of their 2D models.
Chia-Chih Chen, Jake K. Aggarwal
ICPR2
2008 3D face recognition with the average-half-face
abstract
We present a promising analysis on using the pattern of symmetry in the face to increase the accuracy of three-dimensional face recognition. We introduce the concept of the dasiaaverage-half-facepsila, motivated by the symmetry preserving singular value decomposition. We compare face recognition results using the eigenfaces face recognition algorithm with average-half-face data and full face data in several experiments on a 3D face data set of 1126 images. We show that the results from the eigenfaces face recognition system using the average-half-face is more accurate than using the full face, only the left or right half of the face or a random choice of half of the face.
Josh Harguess, Shalini Gupta, Jake K. Aggarwal
ICPR3
2008 Human activities: Handling uncertainties using fuzzy time intervals
abstract
Persons may perform an activity in many different styles, or noise may cause an identical activity to have different temporal structures. We present a robust methodology for recognition of such human activities. The recognition approach presented in this paper is able to handle person-dependent and situation-dependent uncertainties and variations of human activity executions. Our system reliably recognizes human activities with such execution variations, by semantically measuring the similarity between the observations generated by an activity execution and its optimal structure. The system detects fuzzy time intervals associated with low-level gestures of a person, and matches them hierarchically with the representation of the activity that the system is maintaining. Our system is tested for eight types of simple human interactions such as `pushing¿ and `shaking hands¿, as well as complex recursive interactions like `fighting¿ and `greeting¿. The results show that the performance of our system is superior to that of the previous systems using deterministic time intervals.
Michael S. Ryoo, Jake K. Aggarwal
ICPR2
2008 Tracking and Segmentation of Highway Vehicles in Cluttered and Crowded Scenes
abstract
Monitoring highway traffic is an important application of computer vision research. In this paper, we analyze congested highway situations where it is difficult to track individual vehicles in heavy traffic because vehicles either occlude each other or are connected together by shadow. Moreover, scenes from traffic monitoring videos are usually noisy due to weather conditions and/or video compression. We present a method that can separate occluded vehicles by tracking movements of feature points and assigning over-segmented image fragments to the motion vector that best represents the fragment's movement. Experiments were conducted on traffic videos taken from highways in Turkey, and the proposed method can successfully separate vehicles in overpopulated and cluttered scenes.
Goo Jun, Jake K. Aggarwal, Muhittin Gökmen
WACV2
2007 Detection of abandoned objects in crowded environments
abstract
With concerns about terrorism and global security on the rise, it has become vital to have in place efficient threat detection systems that can detect and recognize potentially dangerous situations, and alert the authorities to take appropriate action. Of particular significance is the case of unattended objects in mass transit areas. This paper describes a general framework that recognizes the event of someone leaving a piece of baggage unattended in forbidden areas. Our approach involves the recognition of four sub-events that characterize the activity of interest. When an unaccompanied bag is detected, the system analyzes its history to determine its most likely owner(s), where the owner is defined as the person who brought the bag into the scene before leaving it unattended. Through subsequent frames, the system keeps a lookout for the owner, whose presence in or disappearance from the scene defines the status of the bag, and decides the appropriate course of action. The system was successfully tested on the i-LIDS dataset.
Medha Bhargava, Chia-Chih Chen, Michael S. Ryoo, Jake K. Aggarwal
AVSS4
2007 Real-time detection of illegally parked vehicles using 1-D transformation
abstract
With decreasing costs of high quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in realtime by applying a novel image projection that reduces the dimensionality of the image data and thus reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require the two dimensional data. The proposed algorithm is able to successfully recognize illegally parked vehicles in real-time in the i-LIDS bag and vehicle detection challenge datasets.
Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal
AVSS4
2007 3D Face Recognition Founded on the Structural Diversity of Human Faces
abstract
We present a systematic procedure for selecting facial fiducial points associated with diverse structural characteristics of a human face. We identify such characteristics from the existing literature on anthropometric facial proportions. We also present three dimensional (3D) face recognition algorithms, which employ Euclidean/geodesic distances between these anthropometric fiducial points as features along with linear discriminant analysis classifiers. Furthermore, we show that in our algorithms, when anthropometric distances are replaced by distances between arbitrary regularly spaced facial points, their performances decrease substantially. This demonstrates that incorporating domain specific knowledge about the structural diversity of human faces significantly improves the performance of 3D human face recognition algorithms.
Shalini Gupta, Jake K. Aggarwal, Mia K. Markey, Alan C. Bovik
CVPR2
2007 Hierarchical Recognition of Human Activities Interacting with Objects
abstract
The paper presents a system that recognizes humans interacting with objects. We delineate a new framework that integrates object recognition, motion estimation, and semantic-level recognition for the reliable recognition of hierarchical human-object interactions. The framework is designed to integrate recognition decisions made by each component, and to probabilistically compensate for the failure of the components with the use of the decisions made by the other components. As a result, human-object interactions in an airport-like environment, such as 'a person carrying a baggage', 'a person leaving his/her baggage', or 'a person snatching another's baggage', are recognized. The experimental results show that not only the performance of the final activity recognition is superior to that of previous approaches, but also the accuracy of the object recognition and the motion estimation increases using feedback from the semantic layer. Several real examples illustrate the superior performance in recognition and semantic description of occurring events.
Michael S. Ryoo, Jake K. Aggarwal
CVPR2
2007 Robust Human-Computer Interaction System Guiding a User by Providing Feedback
Michael S. Ryoo, Jake K. Aggarwal
IJCAI2
2006 Recognition of Composite Human Activities through Context-Free Grammar Based Representation
abstract
This paper describes a general methodology for automated recognition of complex human activities. The methodology uses a context-free grammar (CFG) based representation scheme to represent composite actions and interactions. The CFG-based representation enables us to formally define complex human activities based on simple actions or movements. Human activities are classified into three categories: atomic action, composite action, and interaction. Our system is not only able to represent complex human activities formally, but also able to recognize represented actions and interactions with high accuracy. Image sequences are processed to extract poses and gestures. Based on gestures, the system detects actions and interactions occurring in a sequence of image frames. Our results show that the system is able to represent composite actions and interactions naturally. The system was tested to represent and recognize eight types of interactions: approach, depart, point, shake-hands, hug, punch, kick, and push. The experiments show that the system can recognize sequences of represented composite actions and interactions with a high recognition rate.
Michael S. Ryoo, Jake K. Aggarwal
CVPR (2)2
2006 Simultaneous tracking of multiple body parts of interacting persons
Sangho Park, Jake K. Aggarwal
Comput. Vis. Image Underst.2
2006 Object tracking in an outdoor environment using fusion of features and cameras
Quming Zhou, Jake K. Aggarwal
Image Vis. Comput.2
2006 Special issue on multimedia Surveillance systems: guest editorial
Jake K. Aggarwal, Rita Cucchiara
Multim. Syst.1
2005 Understanding of human motion, actions and interactions
abstract
Summary form only given. The efforts to develop computer systems able to detect humans and recognize their activities form an important area of research in computer vision today. The recognition of human activities will lead to a number of applications, including personal assistants, virtual reality, smart monitoring and surveillance systems, as well as motion analysis in sports, medicine and choreography. Motion is an important cue for the human visual system and for understanding human actions. It has been the subject of intense study in a number of fields including philosophy, psychology and neurobiology and, of course, computer vision, robotics and computer graphics. In computer vision research, motion has played an important role for the past thirty years. Prof. Aggarwal's interest in motion started with the study of motion of rigid planar objects and gradually progressed to the study of human motion. The current research includes the study of interactions at the gross (blob) level and at the detailed (head, torso, arms and legs) level. The two levels present different problems in terms of observation and analysis. For blob level analysis, we use a modified Hough transform called the temporal spatio-velocity transform to isolate pixels with similar velocity profiles. For the detailed-level analysis, we employ a multi-target, multi-assignment strategy to track blobs in consecutive frames. An event hierarchy consisting of pose, gesture, action and interaction are used to describe human-human interaction. A methodology is developed to describe the interaction at the semantic level. Professor Aggarwal's presentation will focus on the contributions from other fields leading to the study of motion in computer vision. Further, it will address the issues of interactions at the blob level and at the detailed level. In addition, it will address the directions of future research in motion and human activity recognition.
Jake K. Aggarwal
AVSS1
2005 Tracking soccer players using broadcast TV images
abstract
Tracking soccer players using broadcast TV images is a challenging problem because of players' quick movements, occlusion, and camera movements. This paper presents a framework to track multiple players using the temporal spatio-velocity (TSV) transform. The TSV transform is a combination of a windowing process and the Hough transform. We have previously introduced the TSV transform to extract velocity-stable blobs using binary image sequences. In this paper, we extend the transform to apply to grayscale or color images in order to track object textures. We also increase the computational performance of the TSV transform by setting a computational window around the blob being tracked. This operation, which we call the "local TSV transform ", maintains the computational window at the center of the tracked object and provides an alternate means of tracking when players become occluded, a common problem in broadcast TV images. The performance of the system on a soccer video is demonstrated.
Koichi Sato, Jake K. Aggarwal
AVSS2
2005 Supervised parametric and non-parametric classification of chromosome images
Mehul P. Sampat, Alan C. Bovik, Jake K. Aggarwal, Kenneth R. Castleman
Pattern Recognit.3
2004 Model-Based Human Motion Tracking and Behavior Recognition Using Hierarchical Finite State Automata
Sunghun Park, Jake K. Aggarwal
ICCSA (4)3
2004 Temporal spatio-velocity transform and its application to tracking and interaction
Koichi Sato, Jake K. Aggarwal
Comput. Vis. Image Underst.2
2004 A hierarchical Bayesian network for event recognition of human actions and interactions
Sangho Park, Jake K. Aggarwal
Multim. Syst.2
2003 Human Motion Tracking by Combining View-Based and Model-Based Methods for Monocular Video Sequences
Sangho Park, Jake K. Aggarwal
ICCSA (3)3
2003 Problems, ongoing research and future directions in motion research
Jake K. Aggarwal
Mach. Vis. Appl.1
2002 CIRES: a system for content-based retrieval in digital image libraries
abstract
This paper presents CIRES, a new online system for a content-based retrieval in digital image libraries. Content-based image retrieval systems have traditionally used color and texture analyses. These analyses have not always achieved adequate level of performance and user satisfaction. The growing need for robust image retrieval systems has led to a need for additional retrieval methodologies. CIRES addresses this issue by using image structure in addition to color and texture. The efficacy of using structure in combination with color and texture is demonstrated.
Qasim Iqbal, Jake K. Aggarwal
ICARCV2
2002 Retrieval by classification of images containing large manmade objects using perceptual grouping
Qasim Iqbal, Jake K. Aggarwal
Pattern Recognit.2
2002 Image retrieval via isotropic and anisotropic mappings
Qasim Iqbal, Jake K. Aggarwal
Pattern Recognit.2
2000 Using Head Movement to Recognize Activity
abstract
This paper presents a methodology for automatically identifying human actions in either the frontal or the lateral view. By tracking the movement of the head of the subject over successive frames of a monocular grayscale image sequence, we recognize 12 different actions. The head is segmented automatically in each frame, and the feature vectors extracted. Input sequences captured from a fixed CCD camera are matched against stored models of actions. The system uses the nearest neighbor classifier to identify the test action.
Anant Madabhushi, Jake K. Aggarwal
ICPR2
2000 Recognition of Human Interaction Using Multiple Features in Grayscale Images
abstract
This paper presents a recognition system that classifies four kinds of human interactions: shaking hands, pointing at the opposite person, standing hand-in-hand, and an intermediate/transitional state between them. Our system achieves recognition by applying the K-nearest neighbor classifier to the parametric human-interaction model, which describes the interpersonal configuration with multiple features from gray scale images (i.e., binary blob, silhouette contour, and intensity distribution). Unlike the algorithms that use temporal information about motion, our system independently classifies each frame by estimating the relative poses of the interacting persons. The system provides a tool to detect the initiation and the termination of an interaction with no parsing procedure for sequential data. Experimental results are presented and illustrated.
Sangho Park, Jake K. Aggarwal
ICPR2
2000 Bayesian recognition of targets by parts in second generation forward looking infrared images
Dinesh Nair, Jake K. Aggarwal
Image Vis. Comput.2
2000 MODEEP: a motion-based object detection and pose estimation method for airborne FLIR sequences
Alexander Strehl, Jake K. Aggarwal
Mach. Vis. Appl.2
1999 Applying Perceptual Grouping to Content-Based Image Retrieval: Building Images
abstract
This paper presents an application of perceptual grouping rules for content-based image retrieval. The semantic interrelationships between different primitive image features are exploited by perceptual grouping to detect the presence of manmade structures. A methodology based on these principles in a Bayesian framework for the retrieval of building images, and the results obtained are presented. The image database consists of monocular grayscale outdoor images taken from a ground-level camera.
Qasim Iqbal, Jake K. Aggarwal
CVPR2
1999 Hierarchical Multifeature Integration for Automatic Object Recognition in Forward Looking Infrared Images
Shishir Shah 0001, Jake K. Aggarwal
IEA/AIE2
1999 Human Motion Analysis: A Review
Jake K. Aggarwal, Quin Cai
Comput. Vis. Image Underst.1
1999 Tracking Human Motion in Structured Environments Using a Distributed-Camera System
abstract
This paper presents a comprehensive framework for tracking coarse human models from sequences of synchronized monocular grayscale images in multiple camera coordinates. It demonstrates the feasibility of an end-to-end person tracking system using a unique combination of motion analysis on 3D geometry in different camera coordinates and other existing techniques in motion detection, segmentation, and pattern recognition. The system starts with tracking from a single camera view. When the system predicts that the active camera will no longer have a good view of the subject of interest, tracking will be switched to another camera which provides a better view and requires the least switching to continue tracking. The nonrigidity of the human body is addressed by matching points of the middle line of the human image, spatially and temporally, using Bayesian classification schemes. Multivariate normal distributions are employed to model class-conditional densities of the features for tracking, such as location, intensity, and geometric features. Limited degrees of occlusion are tolerated within the system. Experimental results using a prototype system are presented and the performance of the algorithm is evaluated to demonstrate its feasibility for real time applications.
Quin Cai, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Bayesian Paradigm for Recognition of Objects - Innovative Applications
Jake K. Aggarwal, Shishir Shah 0001
ACCV (2)1
1998 Multiple Feature Integration for Robust Object Localization
abstract
This paper presents a methodology for localization of manmade objects in complex scenes by learning multiple feature models in images. The methodology is based on a modular structure consisting of multiple classi#ers, each of which solves the problem independently based on its input observations. Each classi- #er module is trained to detect manmade object regions and a higher order decision integrator collects evidencefrom each of the modules to delineate a #nal region of interest. The proposed framework is applied to the problem of Automatic Manmade Object Localization #Detection. Results obtained on the detection of vehicles in color visual and infrared imagery are presented in this paper. 1 Introduction This paper addresses the problem of object localization in complex scenes imaged by a single sensor or registered multiple sensors. The general topic of determining region of interest #ROI# and object detection is a critical step in all existing paradigms for object recognition. In...
Shishir Shah 0001, Jake K. Aggarwal
CVPR2
1998 Partial Face Recognition Using Radial Basis Function Networks
Kiminori Sato, Shishir Shah 0001, Jake K. Aggarwal
FG3
1998 Automatic Tracking of Human Motion in Indoor Scenes Across Multiple Synchronized Video Streams
abstract
This paper presents a comprehensive framework for tracking moving humans in an indoor environment from sequences of synchronized monocular grayscale images captured from multiple fixed cameras. The proposed framework consists of three main modules: Single View Tracking (SVT), Multiple View Transition Tracking (MVTT), and Automatic Camera Switching (ACS). Bayesian classification schemes based on motion analysis of human features are used to track (spatially and temporally) a subject image of interest between consecutive frames. The automatic camera switching module predicts the position of the subject along a spatial-temporal domain, and then, selects the camera which provides the best view and requires the least switching to continue tracking. Limited degrees of occlusion are tolerated within the system. Tracking is based upon the images of upper human, bodies captured from various viewing angles, and non-human moving objects are excluded using Principal Component Analysis (PCA). Experimental results are presented to evaluate the performance of the tracking system.
Quin Cai, Jake K. Aggarwal
ICCV2
1998 Feature extraction of edge by directional computation of gray-scale variation
abstract
The process of edge detection and feature extraction methods is based on converting a change of gray level between two regions of an image into a variation function that gives the difference between the gray level of each region and the gray level of the line of discontinuity. The process of computing the relative difference yields the magnitude, direction, and slope sign of the magnitude, which in turn characterize the features of an edge. By focusing on the relative variation of gray level for pixels between small regions, the filtering process of the local region identifies and smoothes uncorrelated features for a truer edge than is possible by smoothing a larger region.
Keun-Chung Kim, Jake K. Aggarwal
ICPR3
1998 A hybrid architecture for performance reasoning in classification systems
abstract
This paper presents a unified methodology for reasoning in classification systems. The methodology is based on a two-stage structure that incorporates both neural and Bayesian formulations in the first stage and a rule-based system created by extracting rules from both the classifiers in the second stage. The rule-based system provides a measure of the cause-effect relationship between the inputs and the outputs. This is a novel and useful method for reasoning about the performance of classifier systems and for representing qualitative knowledge about the causal relationship in decision-making systems. The proposed system is tested and results are reported for the problem of automatic target detection.
Shishir Shah 0001, Jake K. Aggarwal
ICPR2
1998 Nonrigid Motion Analysis: Articulated and Elastic Motion
Jake K. Aggarwal, Quin Cai, Wen-Hung Liao, Bikash Sabata
Comput. Vis. Image Underst.1
1998 Moving obstacle detection from a navigating robot
abstract
This paper presents a system that detects unexpected moving obstacles that appear in the path of a navigating robot, and estimates the relative motion of the object with respect to the robot. The system is designed for a robot navigating in a structured environment with a single wide-angle camera. The system uses polar mapping to simplify the segmentation of the moving object from the background. The polar mapping is performed with the focus of expansion as the center. A vision-based algorithm that uses the vanishing points of segments extracted from a scene in a few 3D orientations provides an accurate estimate of the robot orientation. In the transformed space qualitative estimate of moving obstacles is obtained by detecting the vertical motion of edges extracted in a few specified directions. Relative motion information about the obstacle is then obtained by computing the time to impact between the obstacles and robot from the radial component of the optical flow. The system was implemented and on an indoor mobile robot.
Dinesh Nair, Jake K. Aggarwal
IEEE Trans. Robotics Autom.2
1997 A Bayesian Segmentation Framework for Textured Visual Images
abstract
This paper presents a new framework for segmentation of textured visual imagery. The proposed method consists of a Bayesian formulation for labeling similar regions. Similarity is defined via texture features obtained by Gabor Wavelets. Multivariate Gaussian distributions are employed to model the feature class-conditional densities, while the Markov process is used to characterize the distributions of the region labeling due to each feature. A coarse nearest neighbor clustering is performed over the feature space to estimate the initial labelings. An iterative solution to the Maximum A Posteriori (MAP) estimation is developed, where the parameters of the prior distribution of region labels are estimated using the Expectation-Maximization (EM) algorithm. Finally, for man-made object segmentation, a region-growing procedure is used to analyze the classified texture regions by incorporating measures of local shape characteristics to obtain smooth boundaries and region homogeneity. Results of the developed algorithm on real scene images are presented.
Shishir Shah 0001, Jake K. Aggarwal
CVPR2
1997 Multisensor Integration for Scene Classifiction: An Experiment in Human Form Detection
abstract
This paper presents a system for classification of scenes using a multisensor integration framework. Indoor scenes are imaged using a visual and an infrared sensor and the images processed in three stages to perform classification of sensed objects into two classes: human and background. Finally, information from individual classifiers is integrated in order to obtain an improved classification performance. Details of feature extraction and classification using neural network combining a multi-Bayesian framework are presented. Segmentation of the imaged scene is performed using existing techniques such as texture analysis and histogram modeling. Classification results on real-world data are presented. The system represents a first step in the development of improved, robust classifiers based on the concepts of neural networks and multisensor integration.
Shishir Shah 0001, Jake K. Aggarwal, Jayan Eledath, Joydeep Ghosh
ICIP (2)2
1997 Estimation of position and orientation from image sequence of a circle
abstract
A method to estimate the relative position and orientation of a camera to a circle having unknown radius is presented. A possible application of this problem is the automatic landing of a helicopter because the method can replace human pilots who estimate their position and orientation relative to the landing site by observing the circle marked in heliports. The method is formulated on the framework of the recursive estimation using an image sequence of a circle. The results of the experiment both on the synthetic data and the real image show the proposed method works adequately.
Machiko Sato, Jake K. Aggarwal
ICRA2
1997 Line Correspondences from Cooperating Spatial and Temporal Grouping Processes for a Sequence of Images
Yuh-Lin Chang, Jake K. Aggarwal
Comput. Vis. Image Underst.2
1997 The reconstruction of dynamic 3D structure of biological objects using stereo microscope images
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
Mach. Vis. Appl.3
1997 Mobile robot navigation and scene modeling using stereo fish-eye lens system
Shishir Shah 0001, Jake K. Aggarwal
Mach. Vis. Appl.2
1997 Physics-based integration of multiple sensing modalities for scene interpretation
abstract
The fusion of multiple imaging modalities offers many advantages over the analysis, separately, of the individual sensory modalities. In this paper we present a unique approach to the integrated analysis of disparate sources of imagery for object recognition. The approach is based on physics-based modeling of the image generation mechanisms. Such models make possible features that are physically meaningful and have an improved capacity to differentiate between multiple classes of objects. We illustrate the use of physics-based approach to develop multisensory vision systems for different object recognition application domains. The paper discusses the integration of different suites of sensors, the integration of image-derived information with model-derived information and the physics-based simulation of multisensory imagery.
Nagaraj Nandhakumar, Jake K. Aggarwal
Proc. IEEE2
1996 A Focused Target Segmentation Paradigm
Dinesh Nair, Jake K. Aggarwal
ECCV (1)2
1996 Tracking human motion using multiple cameras
abstract
Presents a framework for tracking human motion in an indoor environment from sequences of monocular grayscale images obtained from multiple fixed cameras. Multivariate Gaussian models are applied to find the most likely matches of human subjects between consecutive frames taken by cameras mounted in various locations. Experimental results from real data show the robustness of the algorithm and its potential for real time applications.
Quin Cai, Jake K. Aggarwal
ICPR2
1996 Curve and surface interpolation using rational radial basis functions
abstract
This paper addresses the problem of reconstructing two-dimensional curves and three-dimensional surfaces from scattered, sparse measurements. We extend the rational Gaussian (RaG) functions introduced by Goshtasby (1993) to general rational radial basis functions and develop a method to compute the smoothness parameters for the shape model by considering the adjacency relation of the control points. Experimental results demonstrate substantial improvements over the original RaG-based method when the input data is sparse and the distribution of the control points is highly nonuniform.
Wen-Hung Liao, Jake K. Aggarwal
ICPR2
1996 Hierarchical, modular architectures for object recognition by parts
abstract
We present a methodology for object recognition by parts. The methodology is based on a hierarchical, modular structure for object recognition. Recognition is performed at different levels in the hierarchy, and the type of recognition performed differs from level to level. Each level is made up of modules, where each module is an expert on a particular part of an object, that is, each module is specifically trained to recognize one part of an object. We present a Bayesian system in which the expert modules represent the probability density functions of each part, modeled as a mixture of densities to incorporate different views (aspects) of each part. Results obtained for object recognition in second generation forward looking infrared (FLIR) images are also presented in this paper.
Dinesh Nair, Jake K. Aggarwal
ICPR2
1996 Bayesian range segmentation using focus cues
abstract
The objective of range segmentation is to partition a scene into regions with different depth ranges. We first perform range classification by combining two paradigms: focus cues and Bayesian estimation. A criterion function from focus cues provides a basic rule for measuring the ranges of a region in images. A Bayesian estimation is obtained by modeling the class field as a Markov random field (MRF). To combine these two paradigms, we define a combined energy function in terms of the cost function using the criterion function values for focus measure and the energy function of the Gibbs distribution of the class field. Then the combined energy function is minimized by a modified simulated annealing method to obtain range classification. The range classification is based on quantized ranges, and it provides an initial range segmentation. For range segmentation, we obtain interpolated range values, and perform a merging process by modeling the field of ranges as a Gaussian Markov random field. The range segmentation result gives a description of the 3-D structure of a scene.
Changhoon Yim, Alan C. Bovik, Jake K. Aggarwal
ICPR3
1996 Robust automatic target recognition in second generation FLIR images
abstract
In this paper we present a system for the detection and recognition of targets in second generation forward looking infrared (FLIR) images. The system uses new algorithms for target detection and segmentation of the targets. Recognition is based on a methodology far target recognition by parts. A diffusion based approach for determining the parts of a target is also presented here. Experimental results on a large database of FLIR images validate the robustness of the system, and its applicability to FLIR imagery obtained from real scenes.
Dinesh Nair, Jake K. Aggarwal
WACV2
1996 Surface Correspondence and Motion Computation from a Pair of Range Images
Bikash Sabata, Jake K. Aggarwal
Comput. Vis. Image Underst.2
1996 Pattern Category Assignment by Neural Networks and Nearest Neighbors Rules: A Synopsis and A Characterization
abstract
The purpose of this paper is two-fold: to give a synoptic description of favored neural networks and to characterize the potency of these neural networks as pattern classifiers, against the background of the familiar nearest neighbors classification. We limit the study to those neural network structures most commonly used for pattern classification: the multilayer perceptron, the Kohonen associative memory, and the Carpenter–Grossberg clustering network, for which we give a tutorial description with the aim of making the driving concepts apparent. The nearest neighbors rule is presented with improved nearest neighbor search and reference data sample pruning. To gain some familiarity with the classifiers, we expound the sequence of computations implicated in pattern category assignment by each classifier. A characterization of the classifiers is drawn from observed and expected properties and from experiments in automatic target recognition and optical character recognition as summarized in comparative tables of performance. This characterization supports the suggestion that nearest neighbors classification always be considered before endorsing alternative pattern classifiers such as neural networks.
Amar Mitiche, Jake K. Aggarwal
Int. J. Pattern Recognit. Artif. Intell.2
1996 Intrinsic parameter calibration procedure for a (high-distortion) fish-eye lens camera with distortion model and accuracy estimation
Shishir Shah 0001, Jake K. Aggarwal
Pattern Recognit.2
1996 Mobile robot self-location using model-image feature correspondence
abstract
The problem of establishing reliable and accurate correspondence between a stored 3-D model and a 2-D image of it is important in many computer vision tasks, including model-based object recognition, autonomous navigation, pose estimation, airborne surveillance, and reconnaissance. This paper presents an approach to solving this problem in the context of autonomous navigation of a mobile robot in an outdoor urban, man-made environment. The robot's environment is assumed consist of polyhedral buildings. The 3-D descriptions of the lines constituting the buildings' rooftops is assumed to be given as the world model. The robot's position and pose are estimated by establishing correspondence between the straight line features extracted from the images acquired by the robot and the model features. The correspondence problem is formulated as a two-stage constrained search problem. Geometric visibility constraints are used to reduce the search space of possible model-image feature correspondences. Techniques for effectively deriving and capturing these visibility constraints from the given world model are presented. The position estimation technique presented is shown to be robust and accurate even in the presence of errors in the feature detection, incomplete model description, and occlusions. Experimental results of testing this approach using a model of an airport scene are presented.
Raj Talluri, Jake K. Aggarwal
IEEE Trans. Robotics Autom.2
1995 Analysis of Left Ventricular Motion
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
ACCV3
1995 Modeling Structured Environments Using Robot Vision
Shishir Shah 0001, Jake K. Aggarwal
ACCV2
1995 Tracking human motion in an indoor environment
abstract
This paper presents an approach to tracking human motion in a sequence of monocular images. The process consists of detecting motion, segmenting moving subjects by recovering the background and, finally, tracking the subject of interest. The usual assumptions of small image motion, a fixed viewing system and constant velocity are systematically relaxed. Two cases are studied: (1) a viewing system with negligible motion, and (2) a viewing system with non-negligible motion.
Quin Cai, Amar Mitiche, Jake K. Aggarwal
ICIP3
1995 On comparing the performance of object recognition systems
abstract
We give a methodology for the evaluation and comparison of object recognition systems. The methodology is based on indicators of two kinds: (1) statistical and (2) algorithmic. Statistical indicators measure the significance of the performance difference between different systems and rank the performances when the difference is significant. Of the various statistical indicators, we use the Kruskal-Wallis H test. Algorithmic indicators include the usual space and time complexity measures, and various performance curves in variables of error and test data sample size. We illustrate the methodology to evaluate a number of pattern classifiers.
Dinesh Nair, Amar Mitiche, Jake K. Aggarwal
ICIP3
1995 Stereo Matching in the Presence of Narrow Occluding Objects Using Dynamic Disparity Search
abstract
Most contemporary stereo correspondence algorithms impose global consistency among candidate match-points using spatial hierarchy mechanism (SHM) based techniques that rely on either the local support within a 2D neighborhood in the image plane and/or cooperative processes between multiple levels of a pixel-resolution or structural-description hierarchy. We analyze the stereo matching failures in SHM-based techniques in the presence of narrow occluding objects and propose the dynamic disparity search (DDS) framework to reduce false-positive matches. Experiments with indoor and outdoor scenes demonstrate a significant reduction in the false-positive match rates of a DDS-based stereo algorithm as compared to those of two existing algorithms.>
Umesh R. Dhond, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Representing and estimating 3-D lines
Yuh-Lin Chang, Jake K. Aggarwal
Pattern Recognit.2
1995 Robot self-location using visual reasoning relative to a single target object
Michael J. Magee, Jake K. Aggarwal
Pattern Recognit.2
1994 Reconstruction of Cynamic 3-D Structures of Biological Objects Using Stereo Microscopy
abstract
The authors address the analysis of three dimensional shape and shape change in nonrigid biological objects imaged via a stereo light microscope (SLM). Most existing stereo or motion analysis techniques cannot be applied to microscopic biological images because they usually lack salient features. The authors propose an integrated approach for the reconstruction of 3D structures and motion analysis for scenes where only a few informative features are available. The key components of this framework are: (1) image registration, (2) region-of-interest extraction, and (3) stereo and motion analysis using a cooperative spatial and temporal matching process. The authors describe these three stages of processing and illustrate the efficacy of the proposed approach using real images of a live frog's ventricle. The reconstructed dynamic 3-D structures of the ventricle are demonstrated in the authors' experimental results.>
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
ICIP (3)3
1994 Detecting Unexpected Moving Obstacles that Appear in the Path of a Navigating Robot
abstract
Presents a system that detects unexpected moving obstacles that appear in the path of a navigating robot, and that estimates the relative motion of the object with respect to the robot. The system is designed for a robot navigating in a structured environment with a single wide-angle camera. The system uses polar mapping to simplify the segmentation of the moving object from the background. The polar mapping is performed with the focus of expansion (FOE) as the center, obtained from the vanishing point of the significant lines in the environment. A method for determining a qualitative estimate of the motion of the object, coupled with an efficient clustering algorithm, is used to segment the moving obstacle from the background. The system uses an area based feature matching method to compute the optical flow in the segmented regions of the image sequence, and time-to-impact is computed from the radial component of the optical flow in the image plane. A framework for a real time implementation of this system on a robot is also presented.>
Dinesh Nair, Jake K. Aggarwal
ICIP (2)2
1994 Depth Estimation using Stereo Fish-Eye Lenses
abstract
This paper presents the estimation of depth in an indoor, structured environment based on a stereo setup consisting of two fish-eye lenses, with parallel optical axes, mounted on a robot platform. The use of fish-eye lenses provides for a large field of view to estimate better the depth of features very close to the lens. To extract significant information from the fish-eye lens images, we first correct for the distortion before using a special line detector, based on vanishing points, to extract significant features. We use a relaxation procedure to achieve correspondence between features in the left and right images. The process of prediction and recursive verification of the hypotheses is utilized to find a one-to-one correspondence. Experimental results obtained on several stereo images are presented, and an accuracy analysis is performed. Further, the algorithm is tested using a pair of wide-angle lenses, and the accuracy and difference in the spatial information obtained are compared.>
Shishir Shah 0001, Jake K. Aggarwal
ICIP (2)2
1994 Generation of Architectural CAD Models Using a Mobile Robot
abstract
This paper describes new algorithms for automatically constructing a computer aided design (CAD) model of a structured scene as imaged by a single camera on a mobile robot. The scene to be modeled is assumed to be composed mostly of linear edges with particular orientations in 3-D. This is the case for most indoor scenes as well as some outdoor urban scenes. The orientation data is used by a motion stereo algorithm to estimate the 3-D structure using a sequence of images. The algorithm assumes that the linear edges are the boundaries of opaque planar patches, such as the floor, the ceiling and the walls. The resulting 3-D description is a CAD model of the scene. Applications of this technique include CAD modeling for architecture and computer graphics, and robot navigation. This paper completes earlier publications and therefore concentrates on the latter parts of processing, including automatically tracking segments, and using the resulting models in different applications.>
Xavier Lebègue, Jake K. Aggarwal
ICRA2
1994 Surface Correspondence and Motion Computation from a Sequence of Range Images
abstract
The estimation of the motion transformation of a moving object from a sequence of images is of prime interest in computer vision. In this paper, the issues in estimating the motion parameters from a sequence of range images are addressed. The motion estimation task involves: (1) extracting the surfaces and establish the correspondence of the surfaces over the frames in the sequence of range images, and (2) computing the motion transformation using these surface correspondences. A novel procedure based on a hierarchical hypergraph representation of range images is presented for finding surface correspondence. Motion transformation between image frames is computed using the planar and the quadric surface pairings. The authors' solution strategy computes the motion by extracting unique linear features from the quadric surfaces and using them to compute the motion transformation.>
Bikash Sabata, Jake K. Aggarwal
ICRA2
1994 A Simple Calibration Procedure for Fish-Eye (High-Distortion) Lens Camera
abstract
Presents a new algorithm for the geometric camera calibration of a fish-eye lens (a high distortion lens) mounted on a CCD TV camera. The algorithm determines a mapping between points in the world coordinate system and their corresponding point locations in the image plane. The parameters to be calibrated are effective focal length, one-pixel width on the image plane, image distortion center, and distortion coefficients. A simple calibration pattern consisting of equally spaced dots is introduced as a reference for calibration. Some parameters to be calibrated are eliminated by setting up the calibration pattern precisely and assuming negligible distortion at the image distortion center. Thus, the number of unknown parameters to be calibrated is drastically reduced, enabling simple and useful calibration. The method employs a polynomial transformation between points in the world coordinate system and their corresponding image plane locations. The coefficients of the polynomial are determined using the Lagrangian estimation. Furthermore, the effectiveness of the proposed calibration method is confirmed by experimentation.>
Shishir Shah 0001, Jake K. Aggarwal
ICRA2
1994 Manipulations of Octrees and Quadtrees on Multiprocessors
abstract
Octrees offer a powerful means for representing and manipulating 3-D objects. This paper presents an implementation of octree manipulations using a new approach on a shared memory architecture. Octrees are hierarchical data structures used to model 3-D objects. The manipulation of these data structures involves performing independent computations on each node of the octree. Octrees are much easier to deal with than other forms of representations used to model 3-D objects especially where extensive manipulations are involved. When these operations are distributed among multiple processing elements (PEs) and executed simultaneously, a significant speedup may be achieved. Manipulations such as a complement, a union, an intersection and other operations such as finding the volume and centroid which this paper describes are implemented on the Sequent Balance multiprocessor. In this approach the PEs are allocated dynamically, resulting in a uniform load balancing among them. The experimental results presented illustrate the feasibility of the approach. Although this evaluation has been originally done for shared memory machines, it will provide insight for the evaluation of other architectures.
Vipin Chaudhary, K. Kumari, Prakash Arunachalam, Jake K. Aggarwal
Int. J. Pattern Recognit. Artif. Intell.4
1994 Unified modeling of non-homogeneous 3D objects for thermal and visual image synthesis
Nagaraj Nandhakumar, Sankaran Karthik, Jake K. Aggarwal
Pattern Recognit.3
1993 The Integration of Image Segmentation Maps using Region and Edge Information
abstract
We present an algorithm that integrates multiple region segmentation maps and edge maps. It operates independently of image sources and specific region-segmentation or edge-detection techniques. User-specified weights and the arbitrary mixing of region/edge maps are allowed. The integration algorithm enables multiple edge detection/region segmentation modules to work in parallel as front ends. The solution procedure consists of three steps. A maximum likelihood estimator provides initial solutions to the positions of edge pixels from various inputs. An iterative procedure using only local information (without edge tracing) then minimizes the contour curvature. Finally, regions are merged to guarantee that each region is large and compact. The channel-resolution width controls the spatial scope of the initial estimation and contour smoothing to facilitate multiscale processing. Experimental results are demonstrated using data from different types of sensors and processing techniques. The results show an improvement over individual inputs and a strong resemblance to human-generated segmentation.>
Chen-Chau Chu, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Image Map Correspondence for Mobile Robot Self-Location Using Computer Graphics
abstract
The use of computer graphics in estimating the position of an autonomous mobile robot navigating in an outdoor mountainous environment is discussed. A digital elevation map (DEM) of the area in which the robot is to navigate is given, and the robot is equipped with a camera that can be panned and tilted, a compass, and an altimeter. The position of the robot is estimated by establishing a correspondence between the images acquired by the camera on the robot (actual images) and the images generated from the DEM (predicted images) using computer graphics techniques. Features are extracted from the predicted images and the actual images that are used in establishing the correspondence. The features used are the horizon line contours (HLCs) in the images. To reduce the search space a constrained search paradigm is used. Geometric constraints help prune the search space significantly.>
Raj Talluri, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Calibrating a mobile camera's parameters
Yuh-Lin Chang, Xavier Lebègue, Jake K. Aggarwal
Pattern Recognit.3
1993 A Generalized Scheme for Mapping Parallel Algorithms
abstract
A generalized mapping strategy that uses a combination of graph theory, mathematical programming, and heuristics is proposed. The authors use the knowledge from the given algorithm and the architecture to guide the mapping. The approach begins with a graphical representation of the parallel algorithm (problem graph) and the parallel computer (host graph). Using these representations, the authors generate a new graphical representation (extended host graph) on which the problem graph is mapped. An accurate characterization of the communication overhead is used in the objective functions to evaluate the optimality of the mapping. An efficient mapping scheme is developed which uses two levels of optimization procedures. The objective functions include minimizing the communication overhead and minimizing the total execution time which includes both computation and communication times. The mapping scheme is tested by simulation and further confirmed by mapping a real world application onto actual distributed environments.>
Vipin Chaudhary, Jake K. Aggarwal
IEEE Trans. Parallel Distributed Syst.2
1993 A Sliding Memory Plane Array Processor
abstract
A mesh-connected single-input multiple-data (SIMD) architecture called a sliding memory plane (SliM) array processor is proposed. Differing from existing mesh-connected SIMD architectures, SliM has several salient features such as a sliding memory plane that provides inter-PE communication during computation. Two I/O planes provide an I/O overlapping capability. Thus, inter-PE communication and I/O overhead can be overlapped with computation. Inter-PE communication time is invisible in most image processing tasks because the computation time is larger than the communication time on SliM. The ability to overlap inter-PE communication with computation, regardless of window size and shape and without using a coprocessor or an on-chip DMA controller is unique to SliM.>
Myung Hoon Sunwoo, Jake K. Aggarwal
IEEE Trans. Parallel Distributed Syst.2
1993 Significant line segments for an indoor mobile robot
abstract
New algorithms for detecting and interpreting linear features of a real scene as imaged by a single camera on a mobile robot are described. The low-level processing stages are specifically designed to increase the usefulness and the quality of the extracted features for indoor scene understanding. In order to derive 3-D information from a 2-D image, we consider only lines with particular orientation in 3-D. The detection and interpretation processes provide a 3-D orientation hypothesis for each 2-D segment. This in turn is used to estimate the robot's orientation and relative position in the environment. Next, the orientation data is used by a motion stereo algorithm to fully estimate the 3-D structure when a sequence of images becomes available. From detection to 3-D estimation, a strong emphasis is placed on real-world applications and very fast processing with conventional hardware. Results of experimentation with a mobile robot under realistic conditions are given and discussed.>
Xavier Lebègue, Jake K. Aggarwal
IEEE Trans. Robotics Autom.2
1992 Computing stereo correspondences in the presence of narrow occluding objects
abstract
Contemporary stereo correspondence algorithms disambiguate multiple candidate matches by using a spatial hierarchy mechanism. Narrow occluding objects in 3-D scenes cannot be handled by the spatial hierarchy mechanism alone. The authors propose a dynamic disparity search (DDS) framework that combines the spatial hierarchy mechanism with a new disparity hierarchy mechanism, to reduce the stereo matching errors caused by narrow occluding objects. They demonstrate the merits of the DDS approach on real stereo images.>
Umesh R. Dhond, Jake K. Aggarwal
CVPR2
1992 Detecting 3-D Parallel Lines for Perceptual Organization
Xavier Lebègue, Jake K. Aggarwal
ECCV2
1992 Analysis of the stereo correspondence process in scenes with narrow occluding objects
abstract
Most contemporary stereo correspondence algorithms impose global consistency among candidate match-points by solely using a spatial hierarchy mechanism. This paper analyzes the stereo matching failures caused by the spatial hierarchy mechanism in the presence of narrow occluding objects. The lessons learned from this occlusion analysis are used to formulate a new global matching framework-dynamic disparity search. The authors present experimental results showing a significant decrease in the false positive match-rate with a dynamic disparity search-based stereo algorithm as compared to a spatial hierarchy mechanism-based algorithm.>
Umesh R. Dhond, Jake K. Aggarwal
ICPR (1)2
1992 Extraction and interpretation of semantically significant line segments for a mobile robot
abstract
The authors describe algorithms for detecting and interpreting linear features of a real scene as images by a single camera on a mobile robot. The low-level processing stages were specifically designed to increase the usefulness and the quality of the extracted features for a semantic interpretation. The detection and interpretation processes provided a 3-D orientation hypothesis for each 2-D segment. This, in turn, was used to estimate the robot's orientation and relative position in the environment and to delimit the free space visible in the image. The orientation data was used by a motion stereo algorithm to fully estimate the 3-D structure when a sequence of images becomes available. From detection to 3-D estimation, an emphasis was placed on real-world applications and very fast processing with conventional hardware.>
Xavier Lebègue, Jake K. Aggarwal
ICRA2
1992 Range image understanding
Jake K. Aggarwal, Baba C. Vemuri
Image Vis. Comput.1
1992 A flexibility coupled hypercube multiprocessor for high level vision
Myung Hoon Sunwoo, Jake K. Aggarwal
Mach. Vis. Appl.2
1992 Image Interpretation Using Multiple Sensing Modalities
abstract
The AIMS (automatic interpretation using multiple sensors) system, which uses registered laser radar and thermal imagers, is discussed. Its objective is to detect and recognize man-made objects at kilometer range in outdoor scenes. The multisensor fusion approach is applied to four sensing modalities (range, intensity, velocity, and thermal) to improve both image segmentation and interpretation. Low-level attributes of image segments (regions) are computed by the segmentation modules and then converted to the KEE format. The knowledge-based interpretation modules are constructed using KEE and Lisp. AIMS applies forward chaining in a bottom-up fashion to derive object-level interpretations from databases generated by the low-level processing modules. The efficiency of the interpretaton process is enhanced by transferring nonsymbolic processing tasks to a concurrent service manager (program). A parallel implementation of the interpretation module is reported. Experimental results using real data are presented.>
Chen-Chau Chu, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 Applying perceptual organization to the detection of man-made objects in non-urban scenes
H. Q. Lu, Jake K. Aggarwal
Pattern Recognit.2
1992 Position estimation for an autonomous mobile robot in an outdoor environment
abstract
The position estimation problem for an autonomous land vehicle navigating in an unstructured mountainous terrain is solved. A digital elevation map (DEM) of the area is assumed to be given. It is also assumed that the robot is equipped with a camera that can be panned and tilted, a compass, and an altimeter. No recognizable landmarks are assumed to be present, and the robot is not assumed to have an initial estimate of its position. The solution structures the problem as a constrained search paradigm by searching the DEM for the possible robot location. The shape and position of the horizon line in the image plane and the known camera geometry of the perspective projection are used as parameters. Geometric constraints are used to prune the search space significantly. The algorithm is made robust to errors in the imaging process by accounting for worst case errors. The approach is tested using real terrain data of areas in Colorado and Texas.>
Raj Talluri, Jake K. Aggarwal
IEEE Trans. Robotics Autom.2
1991 Positional estimation of a mobile robot using edge visibility regions
abstract
A technique is presented for estimating the position and pose of an autonomous mobile robot navigating in an outdoor, urban environment consisting of polyhedral buildings. The 3-D description of the roof tops of the buildings are assumed to be given as a world model. A discussion is presented of the uses of an intermediate representation of the environment, called the edge visibility regions (EVRs), in the positional estimation process.>
Raj Talluri, Jake K. Aggarwal
CVPR2
1991 On the Complexity of Parallel Image Component Labeling
Vipin Chaudhary, Jake K. Aggarwal
ICPP (3)2
1991 The VEDIC Network for Multicomputers
Vipin Chaudhary, Bikash Sabata, Jake K. Aggarwal
ICPP (1)3
1991 Mobile robot self-location using constrained search
abstract
Presents a technique for estimating the position and pose of an autonomous mobile robot navigating in an outdoor, urban environment consisting of polyhedral buildings. The 3D descriptions of the rooftops of the buildings are assumed to be given as a world model. The robot is assumed to be equipped with a visual camera. The position and pose are estimated by establishing a correspondence between the lines that constitute the rooftops of the buildings (world model features) and their images. A constrained search paradigm is used to isolate a consistent set of correspondences. The geometric constraints between the world model features and their images are used to prune the search space. To effectively capture the geometric relations between the world model features with respect to their visibility from various positions of the robot in the plane in which it navigates (visibility plane), the free space of the robot is partitioned into a set of distinct, nonoverlapping regions called the edge visibility regions (EVRs). The use of these EVRs in isolating a consistent set of correspondences between the world model and the image features is discussed.>
Raj Talluri, Jake K. Aggarwal
IROS2
1991 Estimation of motion from a pair of range images: A review
Bikash Sabata, Jake K. Aggarwal
CVGIP Image Underst.2
1991 A cost-benefit analysis of a third camera for stereo correspondence
Umesh R. Dhond, Jake K. Aggarwal
Int. J. Comput. Vis.2
1991 The interpretation of laser radar images by a knowledge-based system
Chen-Chau Chu, Jake K. Aggarwal
Mach. Vis. Appl.2
1991 Analysis of video image sequences using point and line correspondences
Yuan-Fang Wang, Nitin Karandikar, Jake K. Aggarwal
Pattern Recognit.3
1990 Reconstructing 3D lines from a sequence of 2D projections: representation and estimation
abstract
The reconstruction of a 3-D line from a sequence of its 2-D projections is addressed. The authors first consider the problem of 3-D line representation and then the recursive estimation algorithm. They point out the problems with all previous 3-D line representation models, and suggest a novel approach based on simple geometrical observations. They then derive the corresponding recursive estimation algorithm for the novel representation, based on simple linear algebra. The representation model and the recursive algorithm are implemented on an IBM RT PC and simulation results are presented.>
Yuh-Lin Chang, Jake K. Aggarwal
ICCV2
1990 The integration of region and edge-based segmentation
abstract
An algorithm is presented that integrates segmentation maps using both region and edge segmentation maps as input. The result of integration is a region map in which each region is large and compact. The operation is efficient and independent of image sources as well as segmentation techniques. The proposed algorithm allows multiple input maps and applies user-selected weights on various information sources. The scope of integration is parametrically controlled for the desired spatial resolution. A maximum likelihood estimator provides initial solutions of edge positions and strengths from multiple inputs. An iterative procedure is then used to smooth the resultant edge patterns. The edge map is converted to a region map using closed edge contours if desired. Finally, regions are merged to ensure that every region has the required properties. Experimental results are demonstrated using various segmentation techniques and real data from laser radar and thermal sensors.>
Chen-Chau Chu, Jake K. Aggarwal
ICCV2
1990 Terrain matching by analysis of aerial images
abstract
A terrain matching algorithm has been developed for use in a passive aircraft navigation system. A sequence of aerial, optical images is matched to a reference digital map of the three-dimensional terrain. Stereo analysis of successive images results in the recovery of an elevation map of the observed terrain. A 'cliff map' is then used as a novel, compact representation of the terrain surface. The position and heading of the aircraft are determined via a terrain matching algorithm that locates the observed cliff map within the reference cliff map. Novel results using real terrain data and on implementation of a real stereo algorithm provide an important extension to previous results using simulated stereo.>
Jake K. Aggarwal
ICCV2
1990 Segmentation of 3D range images using pyramidal data structures
abstract
Given a 3D range image of a scene containing multiple arbitrarily shaped objects, the authors segment the scene into homogeneous surface patches. A novel modular framework for the segmentation task is proposed. In the first module, over-segmentation is achieved using zeroth and first order local surface properties. The segmentation is then refined in the second module using high order surface representations dictated by the high level vision tasks. The procedure has been applied successfully to many range images, five of which are presented.>
Bikash Sabata, Farshid Arman, Jake K. Aggarwal
ICCV3
1990 Generalized Mapping of Parallel Algorithms Onto Parallel Architectures
Vipin Chaudhary, Jake K. Aggarwal
ICPP (2)2
1990 A two-stage hybrid approach to the correspondence problem via forward-searching and backward-correcting
abstract
A two-stage solution to the problem of the correspondence of points for motion estimation in computer vision is presented. The first stage of the algorithm is a sequential forward-searching algorithm (FSA), which extends all the survivor trajectories. The second stage of the algorithm is a batch-type, rule-based, backward-correcting algorithm (BCA). Under the simple error assumption (no chain errors), seven rules are sufficient to handle all the possible errors made by the FSA. The BCA takes the last four frames of points as input and rearranges the correspondence among them according to these rules. The FSA and BCA are applied alternatively. This algorithm is able to establish the correspondence of a sequence of frames of points without assuming that the numbers of points in all frames are equal or that the correspondence of the first two frames has been established. Experiments illustrate the robustness of the algorithm on sequences of synthetic data as well as on real images.>
C.-L. Cheng, Jake K. Aggarwal
ICPR (1)2
1990 A sliding memory array processor for low level vision
abstract
A mesh-connected architecture, called a Sliding Memory Plane (SliM) array processor, for low-level vision tasks is described. SliM alleviates several disadvantages of existing mesh-connected architectures. It is a fine-grained massively parallel single-instruction multiple-data architecture. The inter-PE communication can occur without interrupting PEs. Two I/O planes can provide a buffering capability. The communication and I/O overheads can each be overlapped with the computation, and both are therefore greatly diminished. Bit-serial communication and bit-parallel computation are used. Various types of communication and virtual interconnections for communication between diagonal-direction PEs are provided. Each PE can operate separately, based on some degree of local autonomy. SliM shows thorough applicability with these configurations. Parallel algorithms for low-level vision tasks on SliM having a zero or an O(1) communication complexity are proposed. The performance of SliM is illustrated with several examples of low-level computer vision tasks.>
Myung Hoon Sunwoo, Jake K. Aggarwal
ICPR (2)2
1990 Vista for a general purpose computer vision system
abstract
An integrated vision triarchitecture, Vista, for a general-purpose computer vision system is described. Vista consists of three parallel architectures. A fine-grained mesh-connected single-instruction multiple-data (SIMD) architecture is proposed to minimize the communication overhead between processing elements in existing mesh-connected architectures. A medium-grained multiple-instruction multiple-data (MIMD) architecture is proposed to alleviate the communication overhead in loosely coupled multiprocessors and the memory contention in tightly coupled multiprocessors. A coarse-grained MIMD architecture is proposed to reduce the communication and input/output (I/O) overheads in existing hypercube multiprocessors. These architectures are exploited for low-, intermediate-, and high-level computer vision, respectively. Each architecture shows a performance improvement in each class. The three architectures are pipelined into Vista for a general-purpose computer vision system.>
Myung Hoon Sunwoo, Jake K. Aggarwal
ICPR (2)2
1990 Object recognition in dense range images using a CAD system as a model base
abstract
A model-based vision system is proposed in which a commercial CAD system has been used for object modeling. Assuming that the model is known, the corresponding object in the scene is located. Given the CAD model of an object, certain features of the model are extracted, while others are precalculated and stored. The given dense 3-D range image is segmented into a set of homogeneous surface patches using a segmentation procedure. Properties such as curvature, surface normal, and surface area are approximated for each surface patch. For each extracted surface patch, three filters are applied to the previously obtained model features to find the best match. Then, a global consistency filter is applied to remove ambiguities and to find the best matched model.>
Farshid Arman, Jake K. Aggarwal
ICRA2
1990 Binocular versus trinocular stereo
abstract
Consideration is given to the twin issues of the gain in accuracy of stereo correspondence and the accompanying increase in computational cost due to the use of a third camera for stereo analysis. The relative performance of binocular and trinocular stereo algorithms on stereo images with known ground truth is evaluated. A technology for comparing the matching performance of stereo algorithms by using real digital elevation maps is developed. Trinocular local matching reduces the percentage of mismatches having large disparity errors by more than one half when compared to binocular matching, while trinocular stereopsis increases the computational cost of local matching over binocular stereopsis by only about one fourth.>
Umesh R. Dhond, Jake K. Aggarwal
ICRA2
1990 Flexibly Coupled Multiprocessors for Image Processing
Myung Hoon Sunwoo, Jake K. Aggarwal
J. Parallel Distributed Comput.2
1990 A System Design/Scheduling Strategy for Parallel Image Processing
abstract
A system design/scheduling strategy is described for a real-time parallel image processing system. A parallel image processing model based on the spatial and temporal parallelisms extractable in image processing tasks is formulated. This model consists of linear pipeline stages, each of which is a multiprocessing module. The strategy is discussed for two cases: static and dynamic. The description of the detailed hardware structure of each module is not attempted because of its dependency on a specific task. In the static case, where the processing time is constant, the processing times of all the stages are adjusted to be identical. In the dynamic case, where the processing time varies with each image, the strategy needs to be modified to achieve the maximum possible processing speed. The strategy is demonstrated for the static case by implementing image processing tasks on two different multiprocessor systems.>
Soo-Young Lee, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Stochastic Analysis of Stereo Quantization Error
abstract
The probability density function of the range estimation error and the expected value of the range error magnitude are derived in terms of the various design parameters of a stereo imaging system. In addition, the relative range error is proposed as a better way of quantifying the range resolution of a stereo imaging system than the percent range error when the depths in the scene lie within a narrow range.>
Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Matching Aerial Images to 3-D Terrain Maps
abstract
A terrain-matching algorithm is presented for use in a passive aircraft navigation system. A sequence of aerial images is matched to a reference digital map of the 3-D terrain. Stereo analysis of successive images results in a recovered elevation map. A cliff map is then used as a novel compact representation of the 3-D surfaces. The position and heading of the aircraft are determined with a terrain-matching algorithm that locates the unknown cliff map within the reference cliff map. The robustness of the matching algorithm is demonstrated by experimental results using real terrain data.>
Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Pyramid-based image segmentation using multisensory data
H. Asar, Nagaraj Nandhakumar, Jake K. Aggarwal
Pattern Recognit.3
1990 Image segmentation using laser radar data
Chen-Chau Chu, Nagaraj Nandhakumar, Jake K. Aggarwal
Pattern Recognit.3
1989 Integrated modelling of thermal and visual image generation
abstract
A unified approach for modeling objects which are imaged by thermal (infrared) and visual cameras is presented. The model supports the generation of both infrared (8 mu m-12 mu m wavelength) images and monochrome visual images under different viewing and ambient-scene conditions. A modified octree data structure is used for object modeling. The octree serves two different purposes: surface information encoded in boundary nodes and efficient tree-traversal algorithms facilitate the generation of monochrome visual images; and the compact volumetric representation facilitates simulation of heat flow in the object which gives rise to surface temperature variation, which in turn is used to synthesize the thermal image. The detailed object model allows for more accurate prediction of thermal and visual images of objects. It also predicts the values of discriminatory features used in classification. The model developed is designed to be used in a model-based vision system which uses a hypothesize-and-verify strategy to interpret thermal and visual images of scenes. Several blocks-world examples are presented to show typical images generated by the approach.>
Chanhee Oh, Nagaraj Nandhakumar, Jake K. Aggarwal
CVPR3
1989 Hierarchical segmentation of 3-D range images
abstract
The authors present a novel approach for segmentation of dense three-dimensional range images. In this approach, four local properties, namely the 3-D coordinate, the surface normal, the Gaussian curvature, and the mean curvature of each data point, are combined in a hierarchical data structure to segment a given 3-D dense range map into surface patches. This algorithm is applicable to planar as well as curved surfaces; several examples of segmentation of such surfaces are presented.>
Farshid Arman, Bikash Sabata, Jake K. Aggarwal
SMC3
1989 Counting straight lines
Amar Mitiche, Olivier D. Faugeras, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.3
1989 Recent progress in object recognition from range data
J. P. Brady, Nagaraj Nandhakumar, Jake K. Aggarwal
Image Vis. Comput.3
1989 Model Construction and Shape Recognition from Occluding Contours
abstract
A technique is presented for recognizing a 3D object (a model in an image library) from a single 2D silhouette using information such as corners (points with high positive curvatures) and occluding contours, rather than straight line segments. The silhouette is assumed to be a parallel projection of the object. Each model is stored as a set of the principal quadtrees, from which the volume/surface octree of the model is generated. Feature points (i.e. corners) are extracted to guide the recognition process. Four-point correspondences between the 2D feature points of the observed object and 3D feature points of each model are hypothesized, and then verified by applying a variety of constraints to their associated viewing parameters. The result of the hypothesis and verification process is further validated by 2D contour matching. This approach allows for a method of handling both planar and curved objects in a uniform manner, and provides a solution to the recognition of multiple objects with occlusion as demonstrated by the experimental results.>
Chiun-Hong Chien, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Obtaining a solid model from optical serial sections
Fernando Macías-Garza, Kenneth R. Diller, Alan C. Bovik, Shanti J. Aggarwal, Jake K. Aggarwal
Pattern Recognit.5
1989 Integration of active and passive sensing techniques for representing three-dimensional objects
abstract
An algorithm is introduced for inferring the structure of three-dimensional (3-D) objects from multiple viewing directions using an integration of active and passive sensing. The structural description of a 3-D object is constructed in two stages: (i) the surface orientation and partial structure are first inferred from a set of single views, and (ii) the visible surface structures inferred from different viewpoints are integrated to complete the description of the 3-D object. In the first stage, an active stripe coding technique is used for recovering visible surface orientation and partial structure. In the second stage, an iterative construction/refinement scheme is used which exploits both passive and active sensing for representing the object surfaces. The final surface structure is recorded in a data structure where the surface contours in a set of parallel planar cross sections are stored. The system construction is inexpensive, and the algorithms introduced are adaptive, versatile, and suitable for applications in dynamic environments.>
Yuan-Fang Wang, Jake K. Aggarwal
IEEE Trans. Robotics Autom.2
1989 Structure from stereo-a review
abstract
The authors review major recent developments in establishing stereo correspondence for the extraction of the 3D structure of a scene. Broad categories of stereo algorithms are identified on the basis of differences in imaging geometry, matching primitives, and the computational structure used. The performance of these stereo techniques on various classes of test images is reviewed, and possible directions of future research are indicated.>
Umesh R. Dhond, Jake K. Aggarwal
IEEE Trans. Syst. Man Cybern.2
1988 Generation of volume/surface octree from range data
abstract
The authors propose a scheme to generate the volume/surface octree structure from range data. The scheme is similar to that of the quadtree generation algorithm. However, in this case, each node in the quadtree is a binary tree corresponding to a range data point. Consequently, the octree of the viewed object can be generated efficiently by merging the neighboring binary trees recursively. Surface normals can be computed directly from the range image. They are encoded into associated binary trees and subsequently propagated to the corresponding octree nodes during the merging process. Since 3-D information of the viewed object is available in each range image, the proposed scheme is capable of capturing the concave structures in objects, which cannot be detected from intensity model construction. Furthermore, since the algorithms developed in this research are essentially recursive tree traversal procedures, they are suitable for parallel implementation.>
Chium-Hong Chien, Y. B. Sim, Jake K. Aggarwal
CVPR3
1988 Quantization error in stereo imaging
abstract
In the design of a stereo imaging system, one often chooses the parameters to meet a desired error level. The probability density function of the range estimation error and the expected value of the range error magnitude are derived in terms of the various design parameters. The relative range error is proposed as a better way of quantifying the range resolution of a stereo imaging system.>
Jake K. Aggarwal
CVPR2
1988 Localization of objects from range data
abstract
A technique is presented for determining the orientation of an object from single-view range data. The objects and models are represented by regions that are a collection of surface patches homogeneous in curvature-based properties. Given exactly one point on the surface of the unknown view of an object and the corresponding point on the surface of the model of the object, the principal vectors (if unique) can be used to determine the 3-D rotation required to bring the model into the same orientation as that of the object. To determine this one-point correspondence, curves of constant principal curvature are extracted from the surfaces (corresponding to regions possessing the same sign of the principal curvatures) of the model and the unknown view of an object. Then, the local maxima in the curvature of these curves are utilized as a means to establish the one-point correspondence.>
Baba C. Vemuri, Jake K. Aggarwal
CVPR2
1988 The missing cone problem and low-pass distortion in optical serial sectioning microscopy
abstract
In optical serial sectioning, the 3-D structure of a microscopic specimen is observed by incrementing the focusing plane of a light microscope through the specimen. If the depth of field of the microscope is infinitesimal, the image obtained from each focusing plane is an in-focus slice of the optical density of the specimen. The authors show that the finite aperture of any practical microscope inevitably results in the loss of a biconic region of frequencies in the 3-D Fourier spectrum of the optical density, oriented in the direction of the optical axis. Thus, the resolution along this axis is severely reduced. Outside the missing cone of frequencies, the spectrum is distorted by a strong low-pass effect. A closed form expression is obtained for the overall distortion function using principles of geometric optics, and by assuming that the absorption of the specimen is linear and nondiffractive. Methods for restoring the 3-D images obtained through optical serial sectioning are considered, and several examples are provided.>
Fernando Macías-Garza, Alan C. Bovik, Kenneth R. Diller, Shanti J. Aggarwal, Jake K. Aggarwal
ICASSP5
1988 Flexibly Coupled Multiprocessors for Image Processing
Myung Hoon Sunwoo, Jake K. Aggarwal
ICPP (1)2
1988 Recent progress in the recognition of objects from range data
abstract
After a brief summary of range acquisition techniques, some of the progress made with three-dimensional data in the field of computer vision is examined. The present state of three-dimensional object recognition using range data is surveyed. Some work involving three-dimensional object recognition using intensity images is also included when it is applicable to, or extendable to, recognition with range data.>
J. P. Brady, Nagaraj Nandhakumar, Jake K. Aggarwal
ICPR3
1988 Counting the line
abstract
A brief review is given of the current formulations of the problem of recovering 3D structure and motion from image line correspondences, and their minimum requirement is formally established in terms of the number of views and lines. Only general rigid motion is considered. Computation of the rotational component and translation component of general motion can be separated. When the rotational component is determined, computation of the translation component is straightforward.>
Amar Mitiche, Olivier D. Faugeras, Jake K. Aggarwal
ICPR3
1988 Thermal and visual information fusion for outdoor scene perception
abstract
A novel technique is presented for automated image analysis. Information from thermal and visual imagery is fused for classifying objects in outdoor scenes. Pixel-level information fusion yields a feature based on the lumped thermal capacitance of the objects. Region-level fusion using a decision tree classifier categorizes imaged objects as being either vegetation, building, pavement, or a vehicle.>
Nagaraj Nandhakumar, Jake K. Aggarwal
ICRA2
1988 Analyzing orthographic projection of multiple 3D velocity vector fields in optical flow
Hideo Tsukune, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1988 On Smoothness of a Vector Field-Application to Optical Flow
abstract
Two measures of smoothness of a vector field over a domain are suggested, together with associated smoothing operators that satisfy given constrains. The technique is illustrated using operators that smooth the invariants of optical flow curl and divergence, estimated from intensity profiles of moving nondeformable objects. The operators have been applied with satisfactory results to derived flow estimates with high noise levels.>
Amar Mitiche, R. Grisell, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1988 Using Constancy of Distance to Estimate Position and Displacement in Space
abstract
The problem of computing structure and motion from the observation of points in two distinct images of a scene is considered. The proposed method explicitly utilizes the principle of conservation of distance during rigid-body motion. The formulation is such that it separates the problem of estimating object position from that of determining motion parameters. The equations of invariance of distance for a rigid body are solved for the points' position in space. When these coordinates in space are known, motion parameters are computed in a simple and straightforward manner. Examples are given to illustrate the efficiency of the algorithm.>
Amar Mitiche, Steven Seida, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1988 Integrated Analysis of Thermal and Visual Images for Scene Interpretation
abstract
An approach for computer perception of outdoor scenes is presented. The approach is based on integrating information extracted from thermal images and visual images, which provides information not available by processing either type of image alone. The thermal image is analyzed to provide estimates of surface temperature. The visual image provides surface absorptivity and relative orientation. These parameters are used together to provide estimates of heat fluxes at the surfaces of viewed objects. The thermal behavior of scene objects is described in terms of surface heat fluxes. Features based on estimated values of surface heat fluxes are shown to be more meaningful and specific in distinguishing scene components.>
Nagaraj Nandhakumar, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1988 On the computation of motion from sequences of images-A review
abstract
Recent developments are reviewed in the computation of motion and structure of objects in a scene from a sequence of images. Two distinct paradigms are highlighted: (i) the feature-based approach and (ii) the optical-flow-based approach. The comparative merits/demerits of these approaches are discussed. The current status of research in these areas is reviewed and future research directions are indicated.>
Jake K. Aggarwal, Nagaraj Nandhakumar
Proc. IEEE1
1987 Analysis of a sequence of images using point and line correspondences
abstract
In this paper, a feature based approach for analyzing time-varying imagery is presented. Instead of restricting the analysis to be based on points or lines, we propose to use points and lines simultaneously. A special case where four point and one line correspondences are combined for computing the structure and motion of 3D objects from a given image sequence is discussed in detail. Computer simulation results are provided to demonstrate the validity of the algorithm.
Jake K. Aggarwal, Y. F. Wang
ICRA1
1987 On modelling 3-D objects using multiple sensory data
abstract
In this paper, we introduce a new algorithm for modelling 3-D objects using information gathered from both active and passive sensing mechanisms. Construction of the structural description of a 3-D object is composed of two stages: (i) The visible surface orientation and partial structure are first inferred from a set of single views, and (ii) the partial surface structures inferred from different viewpoints are integrated to complete the 3-D structural description of the object. In the first stage, an active stripe coding technique is used which projects spatially modulated patterns to encode the objects surfaces for analysis. The visible surface orientation is inferred using a constraint satisfaction process based upon the observed orientation of the projected patterns. The visible surface structure is recovered through integrating a dense orientation map. In the second stage, an iterative construction/refinement scheme is used which exploits both passive and active sensing for representing the object surfaces. The bounding volume description of the object is first constructed using multiple occluding contours which are acquired through passive sensing. The bounding volume is then refined using the partial surface structures inferred from active sensing. The final surface structure is recorded in a memory efficient data structure where the surface contours in a set of parallel planar cross sections are recorded. We expect this approach to be widely applicable in the field of robotics, geometric modeling and factory automation.
Yuan-Fang Wang, Jake K. Aggarwal
ICRA2
1987 Parallel 2-D Convolution on a Mesh Connected Array Processor
abstract
In this correspondence, a parallel 2-D convolution scheme is presented. The processing structure is a mesh connected array processor consisting of the same number of simple processing elements as the number of pixels in the image. For most windows considered, the number of computation steps required is the same as that of the coefficients of a convolution window. The proposed scheme can be easily extended to convolution windows of arbitrary size and shape. The basic idea of the proposed scheme is to apply the 1-D systolic concept to 2-D convolution on a mesh structure. The computation is carried out along a path called a convolution path in a systolic manner. The efficiency of the scheme is analyzed for windows of various shapes. The ideal convolution path is a Hamiltonian path ending at the center of the window, the length of which is equal to the number of window coefficients. The simple architecture and control strategy make the proposed scheme suitable for VLSI implementation.
Soo-Young Lee, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1987 Computation of Surface Orientation and Structure of Objects Using Grid Coding
abstract
In this correspondence, algorithms are introduced to infer surface orientation and structure of visible object surfaces using grid coding. We adopt the active lighting technique to spatially ``encode'' the scene for analysis. The observed objects, which can have surfaces of arbitrary shape, are assumed to rest on a plane (base plane) in a scene which is ``encoded'' with light cast through a grid plane. Two orthogonal grid patterns are used, where each pattern is obtained with a set of equally spaced stripes marked on a glass pane. The scene is observed through a camera and the object surface orientation is determined using the projected patterns on the object surface. If the surfaces under consideration obey certain smoothness constraints, a dense orientation map can be obtained through proper interpolation. The surface structure can then be recovered given this dense orientation map. Both planar and curved surfaces can be handled in a uniform manner. The algorithms we propose yield reasonably accurate results and are relatively tolerant to noise, especially when compared to shape-from-shading techniques. In contrast to other grid coding techniques reported which match the grid junctions for depth reconstruction under the stereopsis principle, our techniques use the direction of the projected stripes to infer local surface orientation and do not require any correspondence relationship between either the grid lines or the grid junctions to be specified. The algorithm has the ability to register images and can therefore be embedded in a system which integrates knowledge from multiple views.
Yuan-Fang Wang, Amar Mitiche, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1987 Parallel image normalization on a mesh connected array processor
Sudhakar Yalamanchili, Jake K. Aggarwal
Pattern Recognit.3
1987 Experiments in computing optical flow with the gradient-based, multiconstraint method
Amar Mitiche, Yuan-Fang Wang, Jake K. Aggarwal
Pattern Recognit.3
1987 Multiple resolution imagery and texture analysis
S. J. Roan, Jake K. Aggarwal, Worthy N. Martin
Pattern Recognit.2
1987 A Mapping Strategy for Parallel Processing
abstract
This paper presents a mapping strategy for parallel processing using an accurate characterization of the communication overhead. A set of objective functions is formulated to evaluate the optimality of mapping a problem graph onto a system graph. One of them is especially suitable for real-time applications of parallel processing. These objective functions are different from the conventional objective functions in that the edges in the problem graph are weighted and the actual distance rather than the nominal distance for the edges in the system graph is employed. This facilitates a more accurate quantification of the communication overhead. An efficient mapping scheme has been developed for the objective functions, where two levels of assignment optimization procedures are employed: initial assignment and pairwise exchange. The mapping scheme has been tested using the hypercube as a system graph.
Soo-Young Lee, Jake K. Aggarwal
IEEE Trans. Computers2
1987 A Characterization and Analysis of Parallel Processor Interconnection Networks
abstract
The permuting properties of various interconnection networks have been extensively studied. However, not too much attention has been focused on how the permuting properties interact with the mapping of tasks to processors in realizing the communication requirements between tasks. In this paper we focus on characterizing the abilities of some interconnection networks in realizing intertask communication that can be specified as permutations of the task names. From the point of view of the intertask communications requirements, the perceived permuting capabilities may depend upon the specific assignment of tasks to processors. Distinct network permutations may actually result in equivalent intertask communication patterns depending upon the mapping of tasks to processors. Characterizations of networks are developed based upon the theory of permutation groups. A number of properties as well as limitations of these networks become evident from this characterization. Finally, a class of switching networks is identified, that possess many useful properties that make them preferable to multistage interconnection networks in specific applications.
Sudhakar Yalamanchili, Jake K. Aggarwal
IEEE Trans. Computers2
1987 Positioning three-dimensional objects using stereo images
abstract
A simple stereo algorithm is presented for determining the three-dimensional position of object points in a scene. In this algorithm two independent measures of similarity, the zero-crossing pattern and the intensity gradient, are combined to improve the matching process. Zero-crossing neighborhoods are classified into 16 possible patterns according to their local connectivity. In the matching process a relaxation method is used to find the best matches. Three constraints are incorporated into a single relaxation process, namely, disparity continuity, figural continuity, and smoothness of the probability of matching.
Yeon Kim, Jake K. Aggarwal
IEEE J. Robotics Autom.2
1987 Determining object motion in a sequence of stereo images
abstract
The motion of a three-dimensional object is determined from a sequence of stereo images by extracting three-dimensional features, establishing correspondences between these features, and finally, computing the rigid motion parameters. Three-dimensional features are extracted from the depth map of a scene. A two-pass relaxation method is developed for matching features extracted from successive depth maps. In each iteration, geometrical relationships between a feature and its neighbors in one map are compared to those between a candidate in the other map and its neighbors to update the matching probability of the candidate. The comparison of the geometrical relationship is based on the principle of conservation of distance and angle between features during rigid motion. The use of three-dimensional features allows one to find the rotation and translation components of motion separately via solving linear equations. Experimental results using several sets of real data are presented to illustrate results and difficulties.
Yeon Kim, Jake K. Aggarwal
IEEE J. Robotics Autom.2
1986 Volume/surface octrees for the representation of three-dimensional objects
Chiun-Hong Chien, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1986 Identification of 3D objects from multiple silhouettes using quadtrees/octrees
Chiun-Hong Chien, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1986 Computer vision and image processing research at the University of Texas at Austin
Jake K. Aggarwal, Kenneth R. Diller, Alan C. Bovik
Image Vis. Comput.1
1986 Curvature-based representation of objects from range data
Baba C. Vemuri, Amar Mitiche, Jake K. Aggarwal
Image Vis. Comput.3
1986 Determining motion parameters using intensity guided range sensing
Jake K. Aggarwal, Michael J. Magee
Pattern Recognit.1
1986 Surface reconstruction and representation of 3-D scenes
Yuan-Fang Wang, Jake K. Aggarwal
Pattern Recognit.2
1986 Rectangular parallelepiped coding: A volumetric representation of three-dimensional objects
abstract
A new three-dimensional object representation scheme called rectangular parallelepiped coding is presented. Rectangular parallelepiped coding is an extension to three-dimensional space of the two-dimensional rectangular coding scheme. It is a volume-based representation of three-dimensional objects constructed from three orthogonal views (silhouettes) of the object using volume intersection method and is coded as a list of rectangular parallelepipeds. The representation is versatile and easily edited. For the objects coded in this representation, the operations of translation, rotation, and scaling are easily performed and the properties of volume, surface, and moments are easily computed. Examples are given for which processing time and storage requirements are examined.
Yeon Kim, Jake K. Aggarwal
IEEE J. Robotics Autom.2
1985 Using multisensory images to derive the structure of three-dimensional objects - A review
Michael J. Magee, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1985 Image segmentation by conventional and information-integrating techniques: a synopsis
Amar Mitiche, Jake K. Aggarwal
Image Vis. Comput.2
1985 Experiments in Intensity Guided Range Sensing Recognition of Three-Dimensional Objects
abstract
With the advent of devices that can directly sense and determine the coordinates of points in space, the goal of constructing and recognizing descriptors of three-dimensional (3-D) objects is attracting the attention of many researchers in the image processing community. Unfortunately, the time required to fully sense a range image is large relative to the time required to sense an intensity image. Conversely, a single intensity image lacks the depth information required to construct 3-D object descriptors. This paper presents a method of combining the two sensory sources, intensity and range, such that the time required for range sensing is considerably reduced. The approach is to extract potential points of interest from the intensity image and then selectively sense range at these feature points. After the range information is known at these points, a graph structure representing the object in the scene is constructed. This structure is compared to the stored graph models using an algorithm for partial matching. The results of applying the method to both synthetic data and real intensity/range images are presented.
Michael J. Magee, Brian A. Boyter, Chiun-Hong Chien, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.4
1985 The artificial intelligence approach to pattern recognition--a perspective and an overview
Nagaraj Nandhakumar, Jake K. Aggarwal
Pattern Recognit.2
1985 Analysis of a model for parallel image processing
Sudhakar Yalamanchili, Jake K. Aggarwal
Pattern Recognit.2
1985 A system organization for parallel image processing
Sudhakar Yalamanchili, Jake K. Aggarwal
Pattern Recognit.2
1985 Nonlinear systems: Stability analysis
abstract
This work, which is part of the benchmark series in electrical engineering and computer science, is a compilation of the research papers representing the major advances in the area of nonlinear systems stability. It contains a total of 28 papers that pertain to both Lyapunov-like approach applied to ordinary differential equations and functional-analytic approach applied to feedback systems. The papers are categorized into eight parts, listed below.
Jake K. Aggarwal, Mathukumalli Vidyasagar
IEEE Trans. Syst. Man Cybern.1
1984 Algebraic Properties of some Parallel Processor Interconnection Networks
abstract
Interconnection networks form an integral part of parallel processing systems, providing the facility for communication between several concurrently executing tasks. This paper presents an algebraic characterization of the sets of communication paths that a network can simultaneously establish between the processing elements of the system. This characterization provides a uniform framework for the analysis of interconnection networks and is shown to yield answers to several important questions concerning the capabilities, limitations and operation of these networks. The networks considered include one way and two way shift register rings, near neighbor meshes, single stage switching networks and briefly, multistage switching networks.
Sudhakar Yalamanchili, Jake K. Aggarwal
ICDE2
1984 Determining the position of a robot using a single calibration object
abstract
A procedure for determining the position uniquely, of a mobile robot in a three-dimensional space is presented. The method consists of viewing a single sphere with horizontal and vertical calibration great circles and computing distance and elevation and azimuth angles with respect to the sphere. The method is simple and provides good results as long as the sphere is projected onto a large portion of the image plane. Results using simulated and actual data are presented.
Michael J. Magee, Jake K. Aggarwal
ICRA2
1984 A normalized quadtree representation
Chiun-Hong Chien, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1984 Determining vanishing points from perspective images
Michael J. Magee, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1984 Matching Three-Dimensional Objects Using Silhouettes
abstract
A method for matching three-dimensional objects against a library of models from an observed sequence of silhouettes is presented in this correspondence. Based upon the observed silhouettes, the three-dimensional structure of the object is constructed and refined. The principal moments and three primary silhouettes are computed for the constructed three-dimensional objects to represent the aggregate and detailed structure parameters. The adaptive matching technique requires that sufficient silhouettes be added to modify the structure of the unknown object until consistent and steady matching results are obtained. The library for matching is based on three primary silhouettes of the model objects. Experiments conducted show a fast convergence to a consistent result may be achieved provided that a reasonable choice of silhouettes is made.
Yuan-Fang Wang, Michael J. Magee, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1984 Robot guidance using computer vision
J. W. Courtney, Michael J. Magee, Jake K. Aggarwal
Pattern Recognit.3
1984 A model for characterizing the motion of the solid-liquid interface in freezing solutions
Baba C. Vemuri, Kenneth R. Diller, Jake K. Aggarwal
Pattern Recognit.3
1984 Formulation of parallel image processing tasks
Sudhakar Yalamanchili, Jake K. Aggarwal
Pattern Recognit. Lett.2
1983 Editorial
Jake K. Aggarwal
Comput. Vis. Graph. Image Process.1
1983 Editorial
Jake K. Aggarwal
Comput. Vis. Graph. Image Process.1
1983 Experiments in combining intensity and range edge maps
Baldemar Gil, Amar Mitiche, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.3
1983 Contour registration by shape-specific points for shape matching
Amar Mitiche, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1983 Shape and correspondence
Jon A. Webb, Jake K. Aggarwal
Comput. Vis. Graph. Image Process.2
1983 Volumetric Descriptions of Objects from Multiple Views
abstract
Occluding contours from an image sequence with view-point specifications determine a bounding volume approximating the object generating the contours. The initial creation and continual refinement of the approximation requires a volumetric representation that facilitates modification yet is descriptive of surface detail. The ``volume segment'' representation presented in this paper is one such representation.
Worthy N. Martin, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1983 Detection of Edges Using Range Information
abstract
Range data provide an important source of 3-D shape information. This information can be used to extract jump boundaries which correspond to occluding boundaries of objects in a scene and ``edges'' which correspond to points lying between significantly different regions on the surface of objects. We are mainly interested in range data obtained from sensors such as lasers. The main problem with this type of range finder is the fact that the accuracy of the measurements depends on the power of the signal that reaches the receiver. This study describes how a range edge detection procedure can be designed that has low sensitivity to noise and imbeds all the knowledge available on the range measurement accuracy.
Amar Mitiche, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1983 Image analysis of solid-liquid interface morphology in freezing solutions
Baba C. Vemuri, Kenneth R. Diller, Larry Davis 0001, Jake K. Aggarwal
Pattern Recognit.4
1982 A comparsion between time- and frequency-domain techniques for time-varying signal processing
abstract
This paper compares the performances of time-varying digital filters designed by time- and frequency-domain techniques. The filter performances are compared based on the mean-square difference between the actual and the desired output sequences. Theoretically, such a comparison would favor the conventional time-domain technique which, however, requires much more computation time. In this paper, numerical examples are presented to illustrate the tradeoffs between the computation time and the filter performance. We show that the frequency-domain technique is efficient but yields suboptimal results.
Nian-Chyi Huang, Jake K. Aggarwal
ICASSP2
1982 Dynamic scenes and object descriptions
abstract
In this paper we will address several fundamental concepts in the analysis of dynamic scenes and will describe a system that derives three-dimensional object representations from the multiple views provided by dynamic images. The two fundamental concepts are that of the correspondence problem (1) and occlusion analysis (2). The former involves the process which is necessary to span the discontinuity inherent in image sequence representations of dynamic scenes. A detailed description of a dynamic scene analysis system (3) is presented with special emphasis on the two major goals pursued in the work, namely, to lessen the dependence on feature point measurements in a structure from motion system and to develop a descriptive three-dimensional object representation that is suitable for dynamic scene analysis systems. The results are a structure from occluding contours system and a "volume segment" representation scheme.
Worthy N. Martin, Jake K. Aggarwal
ICASSP2
1982 Detection of edges using range information
abstract
Range data provide an important source of 3-D shape information. This information can be used to extract jump boundaries which correspond to occluding boundaries of objects in a scene and "edges" which correspond to points lying between significantly different regions on the surface of objects. We are mainly interested in range data obtained from sensors such as lasers. The main problem with this type of range finders is the fact that the accuracy of the measurements depends on the power of the signal that reaches the receiver. This study describes how a range edge detection procedure can be designed that has low sensitivity to noise and imbeds all the knowledge available on the range measurement accuracy.
Amar Mitiche, Jake K. Aggarwal
ICASSP2
1982 Structure from Motion of Rigid and Jointed Objects
Jon A. Webb, Jake K. Aggarwal
Artif. Intell.2
1982 Experiments in combining intensity and range-edge maps
Baldemar Gil, Amar Mitiche, Jake K. Aggarwal
Comput. Graph. Image Process.3
1982 Contour registration by shape-specific points for shape matching
Amar Mitiche, Jake K. Aggarwal
Comput. Graph. Image Process.2
1982 Shape and correspondence
Jon A. Webb, Jake K. Aggarwal
Comput. Graph. Image Process.2
1982 Extraction of moving object descriptions via differencing
Sudhakar Yalamanchili, Worthy N. Martin, Jake K. Aggarwal
Comput. Graph. Image Process.3
1982 Volumetric descriptions from dynamic scenes
Worthy N. Martin, Jake K. Aggarwal
Pattern Recognit. Lett.2
1982 On combining range and intensity data
Amar Mitiche, Baldemar Gil, Jake K. Aggarwal
Pattern Recognit. Lett.3
1981 Implementation of two-dimensional semicausal recursive digital filters
abstract
The use of semicausal or half-plane filters in image processing generally requires a large amount of extra grid points to produce the output image compared with the case employing causal or quarter-plane filters. This is due to the inherent nature of semicausal filters. In this paper, a practical aspect of implementing 2D semicausal filters for image processing is investigated in terms of state-space realization, and a method to reduce the total number of spatial grid points necessary to evaluate the output image, termed simplified implementation, is presented.
Hyokang Chang, Jake K. Aggarwal
ICASSP2
1981 Spectral modifications using linear shift-variant digital filters
abstract
The present paper develops a framework for the analysis and synthesis of linear shift-variant (LSV) digital filters in the frequency domain. First, LSV digital filters are modeled by the successive use of linear shift-invariant (LSI) filters. Further, shift-variant digital filtering is discussed in relation to the notions of the short-time spectrum and the generalized frequency function. In addition, we propose an efficient implementation procedure which reduces the number of filter coefficients and the amount of computation. The effectiveness of LSV digital filters in processing time-varying signals is demonstrated by experimental verification.
Nian-Chyi Huang, Jake K. Aggarwal
ICASSP2
1981 Structure from Motion of Rigid and Jointed Objects
Jon A. Webb, Jake K. Aggarwal
IJCAI2
1981 An Empirical Evaluation of Generalized Cooccurrence Matrices
abstract
A comparative study of generalized cooccurrence texture analysis tools is presented. A generalized cooccurrence matrix (GCM) reflects the shape, size, and spatial arrangement of texture features. The particular texture features considered in this paper are 1) pixel-intensity, for which generalized cooccurrence reduces to traditional cooccurrence; 2) edge-pixel; and 3) extended-edges. Three experiments are discussed-the first based on a nearest neighbor classifier, the second on a linear discriminant classifier, and the third on the Battacharyya distance figure of merit.
Larry Davis 0001, M. Clearman, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1981 Segmentation of chromatic images
Alireza Sarabi, Jake K. Aggarwal
Pattern Recognit.2
1979 Stabilization of two-dimensional recursive filters
abstract
In this paper we present a new stabilization procedure based on the spectral factorization capability of PLSI (planar least square inverse) polynomials of semicausal form. It is not involved with an intermediate PLSI polynomial since stabilization of an unstable filter is directly obtained from its autocorrelation function. We also introduce a measure of amplitude distortion due to stabilization. The new procedure offers a remedy for flaws in Shanks' original procedure for stabilization.
Hyokang Chang, Jake K. Aggarwal
ICASSP2
1979 Extraction of Moving Object Images Through Change Detection
Ramesh Jain 0001, Worthy N. Martin, Jake K. Aggarwal
IJCAI3
1979 Texture Analysis Using Generalized Co-Occurrence Matrices
abstract
We present a new approach to texture analysis based on the spatial distribution of local features in unsegmented textures. The textures are described using features derived from generalized co-occurrence matrices (GCM). A GCM is determined by a spatial constraint predicate F and a set of local features P = {(Xi, Yi, di), i = 1,..., m} where (Xi, Yi) is the location of the ith feature, and di is a description of the ith feature. The GCM of P under F, GF, is defined by GF(i, j) = number of pairs, pk, pl such that F(pk, pl) is true and di and dj are the descriptions of pk and pl, respectively. We discuss features derived from GCM's and present an experimental study using natural textures.
Larry Davis 0001, Steven A. Johns, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.3
1979 Computer Tracking of Objects Moving in Space
abstract
A method is developed to represent movement of convex blocks in three-dimensional space from a sequence of two-dimensional camera images. The goals are to determine the objects' movement toward or away from the camera as well as left/right and up/down movement in the image plane and to build models of the blocks. The movement information is used as part of a hierarchical matching process that determines the correspondence of blocks between scenes.
John W. Roach, Jake K. Aggarwal
IEEE Trans. Pattern Anal. Mach. Intell.2
1979 Computer analysis of dynamic scenes containing curvilinear figures
Worthy N. Martin, Jake K. Aggarwal
Pattern Recognit.2
1978 Design of semicausal two-dimensional recursive filters
abstract
A procedure for the design of 2D recursive filters is developed using the Semicausal or half-plane filters. For recursive filters, semicausal or half-plane filters are more general than causal or quarter-plane filters to approximate the arbitrary magnitude characteristics. In approximating the given characteristics, the concept of 2D spectral factorization is utilized. A stabilization procedure is also presented and incorporated to produce the finite order stable filters. The transfer function of the resulting filter consists of both numerator and denominator polynomials.
Hyokang Chang, Jake K. Aggarwal
ICASSP2
1977 Computer Analysis of Planar Curvilinear Moving Images
abstract
The present correspondence develops an algorithm for tracking the motion of planar curvilinear moving objects from a sequence of scenes produced by a television camera. The shape of the objects is unrestricted except that holes may not be present. The shape of the objects may vary slightly from scene to scene. The linear and angular velocities of the objects are computed, and a model is created and updated each time a new scene is processed. Furthermore, the algorithm generates a prediction of each successive input scene. The algorithm is implemented in the computer language Lisp, and experimental results are presented.
W. K. Chow, Jake K. Aggarwal
IEEE Trans. Computers2
1977 Computer Recognition of Partial Views of Curved Objects
abstract
The use of a variation of chain encoding of binary boundary images is described in the context of a scene analysis system. The system can learn single views of the three-dimensional curved objects and can later recognize various partial views of the same projections of the objects. The system is completely automatic except for specifying whether the object is to be learned or recognized.
James W. McKee, Jake K. Aggarwal
IEEE Trans. Computers2
1976 Error Analysis of Two-Dimensional Recursive Digital Filters Employing Floating-Point Arithmetic
abstract
In this correspondence, error analysis for two-dimensional direct form recursive digital filters employing floating-point arithmetic is carried out. A systematic way of estimating the mean-squared errors due to roundoff, coefficient, and input quantizations is discussed. Norm error bounds are also given. Extensive numerical experiments have shown that the present approach leads to satisfactory results.
Ming-Duenn Ni, Jake K. Aggarwal
IEEE Trans. Computers2
1975 Finding the edges of the surfaces of three-dimensional curved objects by computer
James W. McKee, Jake K. Aggarwal
Pattern Recognit.2
1975 Computer Analysis of Moving Polygonal Images
abstract
A general mathematical model is developed as an idealization of the problem of determining cloud motions from satellite pictures. The model consists of superimposed planes of rigid moving polygons. The problem is to determine from a sequence of scenes the linear and angular velocities of the figures, and to decompose the scene into its component figures. Study of the model reveals a number of fundamental relations that form the basis for an analysis program. In particular, a systematic anaylsis is given of the topological changes that can occur when overlapping figures move together or apart. A computer program based on these results is described, and experimental results are presented.
Jake K. Aggarwal, Richard O. Duda
IEEE Trans. Computers1
1974 Two-Dimensional Digital Filtering and its Error Analysis
abstract
The concept of a two-dimensional recursive digital filter is introduced and block diagram representations are given. Expressions for error bounds and mean-square, error for roundoff error accumulation are derived. Other sources of error are described. Simulation of the implementation of digital filters is discussed, and corresponding to this simulation, calculation of error is performed. Numerical examples are given. The derived analytic results are shown to be in good agreement with the simulation results.
Ming-Duenn Ni, Jake K. Aggarwal
IEEE Trans. Computers2
1971 Optimal flow in networks with gains and costs
abstract
Abstract The problem of routing a desired flow at minimum cost through a network in which associated with each arc is a capacity, a linear cost, a fixed cost and a flow multiplier is considered. The equivalence of the problems on a general graph and a bipartite graph is shown; feasibility is discussed and properties of solutions are investigated.
Manu Malek-Zavarei, Jake K. Aggarwal
Networks2