Michael A. Greenspan

dblp:g/MichaelAGreenspan · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-6054-8770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 9 since 2021Systems, architecture and hardware · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Deep Learning Framework for Vision-Based Intraocular Pressure Monitoring Using Microfluidic Contact Lenses
Justin Jacob, Tanzila Afrin, Yong Jun Lai, Michael A. Greenspan
AIME (2)4
2026 CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
Nishq Poorav Desai, Michael A. Greenspan, Ali Etemad
ICPR (2)2
2025 Federated Unsupervised Domain Generalization Using Global and Local Alignment of Gradients
abstract
We address the problem of federated domain generalization in an unsupervised setting for the first time. We first theoretically establish a connection between domain shift and alignment of gradients in unsupervised federated learning and show that aligning the gradients at both client and server levels can facilitate the generalization of the model to new (target) domains. Building on this insight, we propose a novel method named FedGaLA, which performs gradient alignment at the client level to encourage clients to learn domain-invariant features, as well as global gradient alignment at the server to obtain a more generalized aggregated model. To empirically evaluate our method, we perform various experiments on four commonly used multi-domain datasets, PACS, OfficeHome, DomainNet, and TerraInc. The results demonstrate the effectiveness of our method which outperforms comparable baselines. Ablation and sensitivity studies demonstrate the impact of different components and parameters in our approach.
Farhad Pourpanah, Mahdiyar Molahasani, Milad Soltany, Michael A. Greenspan, Ali Etemad
AAAI4
2025 Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training
abstract
In this paper, we propose a novel approach, Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training (FedSB), to address the challenges of data heterogeneity within a federated learning framework. FedSB utilizes label smoothing at the client level to prevent overfitting to domain-specific features, thereby enhancing generalization capabilities across diverse domains when aggregating local models into a global model. Additionally, FedSB incorporates a decentralized budgeting mechanism which balances training among clients, which is shown to improve the performance of the aggregated global model. Extensive experiments on four commonly used multi-domain datasets, PACS, VLCS, OfficeHome, and TerraInc, demonstrate that FedSB outperforms competing methods, achieving state-of-the-art results on three out of four datasets, indicating the effectiveness of FedSB in addressing data heterogeneity.
Milad Soltany, Farhad Pourpanah, Mahdiyar Molahasani, Michael A. Greenspan, Ali Etemad
ICASSP4
2025 PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
Mahdiyar Molahasani, Azadeh Motamedi, Michael A. Greenspan, Il-Min Kim 0001, Ali Etemad
ICCV3
2025 Socially-Informed Reconstruction for Pedestrian Trajectory Forecasting
abstract
Pedestrian trajectory prediction remains a challenge for autonomous systems, particularly due to the intricate dynamics of social interactions. Accurate forecasting requires a comprehensive understanding not only of each pedestrian's previous trajectory but also of their interaction with the surrounding environment, an important part of which are other pedestrians moving dynamically in the scene. To learn effective socially-informed representations, we propose a model that uses a reconstructor alongside a conditional variational autoencoder-based trajectory forecasting module. This module generates pseudo-trajectories, which we use as augmentations throughout the training process. To further guide the model towards social awareness, we propose a novel social loss that aids in forecasting of more stable trajectories. We validate our approach through extensive experiments, demonstrating strong performances in comparison to state-of-the-art methods on the ETH/UCY and SDD benchmarks.
Haleh Damirchi, Ali Etemad, Michael A. Greenspan
WACV3
2025 CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
abstract
Self-driving research often underrepresents cyclist collisions and safety. To address this, we present CycleCrash, a novel dataset consisting of 3,000 dashcam videos with 436,347 frames that capture cyclists in a range of critical situations, from collisions to safe interactions. This dataset enables 9 different cyclist collision prediction and classification tasks focusing on potentially hazardous conditions for cyclists and is annotated with collision-related, cyclist-related, and scene-related labels. Next, we propose Vid-NeXt, a novel method that leverages a ConvNeXt spatial encoder and a non-stationary transformer to capture the temporal dynamics of videos for the tasks defined in our dataset. To demonstrate the effectiveness of our method and create additional baselines on CycleCrash, we apply and compare 7 models along with a detailed ablation. We release the dataset and code at https://github.com/DeSinister/CycleCrash/.
Nishq Poorav Desai, Ali Etemad, Michael A. Greenspan
WACV3
2024 Pseudo-Keypoint RKHS Learning for Self-supervised 6DoF Pose Estimation
Yangzheng Wu, Michael A. Greenspan
ECCV (60)2
2024 Learning Better Keypoints for Multi-Object 6DoF Pose Estimation
abstract
We address the problem of keypoint selection, and find that the performance of 6DoF pose estimation methods can be improved when pre-defined keypoint locations are learned, rather than being heuristically selected as has been the standard approach. We found that accuracy and efficiency can be improved by training a graph network to select a set of disperse keypoints with similarly distributed votes. These votes, learned by a regression network to accumulate evidence for the keypoint locations, can be regressed more accurately compared to previous heuristic keypoint algorithms. The proposed KeyGNet, supervised by a combined loss measuring both Wasserstein distance and dispersion, learns the color and geometry features of the target objects to estimate optimal keypoint locations. Experiments demonstrate the keypoints selected by KeyGNet improved the accuracy for all evaluation metrics of all seven datasets tested, for three keypoint voting methods. The challenging Occlusion LINEMOD dataset notably improved ADD(S) by +16.4% on PVN3D, and all core BOP datasets showed an AR improvement for all objects, of between +1% and +21.5%. There was also a notable increase in performance when transitioning from single object to multiple object training using KeyGNet keypoints, essentially eliminating the SISO-MIMO gap for Occlusion LINEMOD.
Yangzheng Wu, Michael A. Greenspan
WACV2
2023 Context-Aware Pedestrian Trajectory Prediction with Multimodal Transformer
abstract
We propose a novel solution for predicting future trajectories of pedestrians. Our method uses a multimodal encoder-decoder transformer architecture, which takes as input both pedestrian locations and ego-vehicle speeds. Notably, our decoder predicts the entire future trajectory in a single-pass and does not perform one-step-ahead prediction, which makes the method effective for embedded edge deployment. We perform detailed experiments and evaluate our method on two popular datasets, PIE and JAAD. Quantitative results demonstrate the superiority of our proposed model over the current state-of-the-art, which consistently achieves the lowest error for 3 time horizons of 0.5, 1.0 and 1.5 seconds. Moreover, the proposed method is significantly faster than the state-of-the-art for the two datasets of PIE and JAAD. Lastly, ablation experiments demonstrate the impact of the key multimodal configuration of our method.
Haleh Damirchi, Michael A. Greenspan, Ali Etemad
ICIP2
2023 Continual Learning for Out-of-Distribution Pedestrian Detection
abstract
A continual learning solution is proposed to address the out-of-distribution generalization problem for pedestrian detection. While recent pedestrian detection models have achieved impressive performance on various datasets, they remain sensitive to shifts in the distribution of the inference data. Our method adopts and modifies Elastic Weight Consolidation to a backbone object detection network, in order to penalize the changes in the model weights based on their importance towards the initially learned task. We show that when trained with one dataset and fine-tuned on another, our solution learns the new distribution and maintains its performance on the previous one, avoiding catastrophic forgetting. We use two popular datasets, CrowdHuman and CityPersons for our cross-dataset experiments, and show considerable improvements over standard fine-tuning, with a 9% and 18% miss rate percent reduction improvement in the CrowdHuman and CityPersons datasets, respectively.
Mahdiyar Molahasani, Ali Etemad, Michael A. Greenspan
ICIP3
2023 Flow-Based Spatio-Temporal Structured Prediction of Motion Dynamics
abstract
Conditional Normalizing Flows (CNFs) are flexible generative models capable of representing complicated distributions with high dimensionality and large interdimensional correlations, making them appealing for structured output learning. Their effectiveness in modelling multivariates spatio-temporal structured data has yet to be completely investigated. We propose MotionFlow as a novel normalizing flows approach that autoregressively conditions the output distributions on the spatio-temporal input features. It combines deterministic and stochastic representations with CNFs to create a probabilistic neural generative approach that can model the variability seen in high-dimensional structured spatio-temporal data. We specifically propose to use conditional priors to factorize the latent space for the time dependent modeling. We also exploit the use of masked convolutions as autoregressive conditionals in CNFs. As a result, our method is able to define arbitrarily expressive output probability distributions under temporal dynamics in multivariate prediction tasks. We apply our method to different tasks, including trajectory prediction, motion prediction, time series forecasting, and binary segmentation, and demonstrate that our model is able to leverage normalizing flows to learn complicated time dependent conditional distributions.
Mohsen Zand, Ali Etemad, Michael A. Greenspan
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Keypoint Cascade Voting for Point Cloud Based 6DoF Pose Estimation
abstract
We propose a novel keypoint voting 6DoF object pose estimation method, which takes pure unordered point cloud geometry as input without RGB information. The proposed cascaded keypoint voting method, called RCVPose3D, is based upon a novel architecture which separates the task of semantic segmentation from that of keypoint regression, thereby increasing the effectiveness of both and improving the ultimate performance. The method also introduces a pairwise constraint in between different keypoints to the loss function when regressing the quantity for keypoint estimation, which is shown to be effective, as well as a novel Voter Confident Score which enhances both the learning and inference stages. Our proposed RCVPose3D achieves state-of-the-art performance on the Occlusion LINEMOD (74.5%) and YCB-Video (96.9%) datasets, outperforming existing pure RGB and RGB-D based methods, as well as being competitive with RGB plus point cloud methods.
Yangzheng Wu, Alireza Javaheri, Mohsen Zand, Michael A. Greenspan
3DV4
2022 Vote from the Center: 6 DoF Pose Estimation in RGB-D Images by Radial Keypoint Voting
Yangzheng Wu, Mohsen Zand, Ali Etemad, Michael A. Greenspan
ECCV (10)4
2022 ObjectBox: From Centers to Boxes for Anchor-Free Object Detection
Mohsen Zand, Ali Etemad, Michael A. Greenspan
ECCV (10)3
2022 Multiscale Crowd Counting and Localization By Multitask Point Supervision
abstract
We propose a multitask approach for crowd counting and person localization in a unified framework. As the detection and localization tasks are well-correlated and can be jointly tackled, our model benefits from a multitask solution by learning multiscale representations of encoded crowd images, and subsequently fusing them. In contrast to the relatively more popular density-based methods, our model uses point supervision to allow for crowd locations to be accurately identified. We test our model on two popular crowd counting datasets, ShanghaiTech A and B, and demonstrate that our method achieves strong results on both counting and localization tasks, with MSE measures of 110.7 and 15.0 for crowd counting and AP measures of 0.71 and 0.75 for localization, on ShanghaiTech A and B respectively. Our detailed ablation experiments show the impact of our multiscale approach as well as the effectiveness of the fusion module embedded in our network. Our code is available at: https://github.com/RCVLab-AiimLab/crowdcounting
Mohsen Zand, Haleh Damirchi, Andrew Farley, Mahdiyar Molahasani, Michael A. Greenspan, Ali Etemad
ICASSP5
2022 Neuro-symbolic Natural Logic with Introspective Revision for Natural Language Inference
abstract
Abstract We introduce a neuro-symbolic natural logic framework based on reinforcement learning with introspective revision. The model samples and rewards specific reasoning paths through policy gradient, in which the introspective revision algorithm modifies intermediate symbolic reasoning steps to discover reward-earning operations as well as leverages external knowledge to alleviate spurious reasoning and training inefficiency. The framework is supported by properly designed local relation models to avoid input entangling, which helps ensure the interpretability of the proof paths. The proposed model has built-in interpretability and shows superior capability in monotonicity inference, systematic generalization, and interpretability, compared with previous models on the existing datasets.
Xiaoyu Yang 0002, Xiaodan Zhu 0001, Michael A. Greenspan
Trans. Assoc. Comput. Linguistics4
2022 Oriented Bounding Boxes for Small and Freely Rotated Objects
abstract
A novel object detection method is presented that handles freely rotated objects of arbitrary sizes, including tiny objects as small as$2 \times 2$pixels. Such tiny objects appear frequently in remotely sensed images, and present a challenge to recent object detection algorithms. More importantly, current object detection methods have been designed originally to accommodate axis-aligned bounding box detection, and therefore fail to accurately localize oriented boxes that best describe freely rotated objects. In contrast, the proposed convolutional neural network (CNN) -based approach uses potential pixel information at multiple scale levels without the need for any external resources, such as anchor boxes. The method encodes the precise location and orientation of features of the target objects at grid cell locations. Unlike existing methods that regress the bounding box location and dimension, the proposed method learns all the required information by classification, which has the added benefit of enabling oriented bounding box detection without any extra computation. It thus infers the bounding boxes only at inference time by finding the minimum surrounding box for every set of the same predicted class labels. Moreover, a rotation-invariant feature representation is applied to each scale, which imposes a regularization constraint to enforce covering the 360° range of in-plane rotation of the training samples to share similar features. Evaluations on the xView and dataset for object detection in aerial images (DOTA) data sets show that the proposed method uniformly improves performance over existing state-of-the-art methods.
Mohsen Zand, Ali Etemad, Michael A. Greenspan
IEEE Trans. Geosci. Remote. Sens.3
2021 Teacher-Student Adversarial Depth Hallucination to Improve Face Recognition
abstract
We present the Teacher-Student Generative Adversarial Network (TS-GAN) to generate depth images from single RGB images in order to boost the performance of face recognition systems. For our method to generalize well across unseen datasets, we design two components in the architecture, a teacher and a student. The teacher, which itself consists of a generator and a discriminator, learns a latent mapping between input RGB and paired depth images in a supervised fashion. The student, which consists of two generators (one shared with the teacher) and a discriminator, learns from new RGB data with no available paired depth information, for improved generalization. The fully trained shared generator can then be used in runtime to hallucinate depth from RGB for downstream applications such as face recognition. We perform rigorous experiments to show the superiority of TS-GAN over other methods in generating synthetic depth images. Moreover, face recognition experiments demonstrate that our hallucinated depth along with the input RGB images boost performance across various architectures when compared to a single RGB modality by average values of +1.2%, +2.6%, and +2.6% for IIIT-D, EURECOM, and LFW datasets respectively. We make our implementation public at: https://github.com/hardik-uppal/teacher-student-gan.git.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
ICCV3
2021 Multistream Validnet: Improving 6D Object Pose Estimation by Automatic Multistream Validation
abstract
This work presents a novel approach to improve the results of pose estimation by detecting and distinguishing between the occurrence of True and False Positive results. It achieves this by training a binary classifier on the output of an arbitrary pose estimation algorithm, and returns a binary label indicating the validity of the result. We demonstrate that our approach improves upon a state-of-the-art pose estimation result on the Siléane dataset, outperforming a variation of the alternative CullNet method by 4.15% in average class accuracy and 0.73% in overall accuracy at validation. Applying our method can also improve the pose estimation average precision results of Op-Net by 6.06% on average.
Joy Mazumder, Mohsen Zand, Michael A. Greenspan
ICIP3
2021 Depth as Attention for Face Representation Learning
abstract
Face representation learning solutions have recently achieved great success for various applications such as verification and identification. However, face recognition approaches that are based purely on RGB images rely solely on intensity information, and therefore are more sensitive to facial variations, notably pose, occlusions, and environmental changes such as illumination and background. A novel depth-guided attention mechanism is proposed for deep multi-modal face recognition using low-cost RGB-D sensors. Our novel attention mechanism directs the deep network “where to look” for visual features in the RGB image by focusing the attention of the network using depth features extracted by a Convolution Neural Network (CNN). The depth features help the network focus on regions of the face in the RGB image that contain more prominent person-specific information. Our attention mechanism then uses this correlation to generate an attention map for RGB images from the depth features extracted by the CNN. We test our network on four public datasets, showing that the features obtained by our proposed solution yield better results on the Lock3DFace, CurtinFaces, IIIT-D RGB-D, and KaspAROV datasets which include challenging variations in pose, occlusion, illumination, expression, and time lapse. Our solution achieves average (increased) accuracies of 87.3% (+5.0%), 99.1% (+0.9%), 99.7% (+0.6%) and 95.3%(+0.5%) for the four datasets respectively, thereby improving the state-of-the-art. We also perform additional experiments with thermal images, instead of depth images, showing the high generalization ability of our solution when adopting other modalities for guiding the attention mechanism instead of depth information.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
IEEE Trans. Inf. Forensics Secur.3
2020 Exploring End-to-End Differentiable Natural Logic Modeling
abstract
We explore end-to-end trained differentiable models that integrate natural logic with neural networks, aiming to keep the backbone of natural language reasoning based on the natural logic formalism while introducing subsymbolic vector representations and neural components.The proposed model adapts module networks to model natural logic operations, which is enhanced with a memory component to model contextual information.Experiments show that the proposed framework can effectively model monotonicity-based reasoning, compared to the baseline neural network models without built-in inductive bias for monotonicity-based reasoning.Our proposed model shows to be robust when transferred from upward to downward inference.We perform further analyses on the performance of the proposed model on aggregation, showing the effectiveness of the proposed subcomponents on helping achieve better intermediate aggregation performance.
Zi'ou Zheng, Quan Liu 0003, Michael A. Greenspan, Xiaodan Zhu 0001
COLING4
2020 One Shot Radial Distortion Correction by Direct Linear Transformation
abstract
A novel method is proposed to estimate and correct image radial distortion. The solution is based upon an algebraic expansion of the homographic relationship between a planar pattern and its distorted projection into a single image, which is solved using the Direct Linear Transformation. The method requires ten point correspondences, is fully automatic and estimates both the first two parameters of the division model, and the center of distortion. Experimental results show the method to be more accurate than other recent one- and two-parameter point-based and plumb-line-based approaches.
Sheikh Ziauddin, Mohsen Zand, Michael A. Greenspan
ICIP4
2020 Multi-Modal Fusion With Observation Points For Skeleton Action Recognition
abstract
Current methods for skeleton-based action recognition compute features based on the given skeleton joint information. We show that introducing new observation points in skeleton motion sequences and using them to create fused representations from multiple modalities such as joints and bones, can enhance the discriminative power of the original modalities. Moreover, such representations can be used to create new streams in multi-stream networks that fuse constructively with other streams trained on the original modalities, effectively exhibiting a dual behaviour and collectively boosting the performance of the network even further. We present one possible multi-modal fusion system with a single observation point that can easily be incorporated in existing networks and improves state-of-the-art results on the two popular J-HMDB and Kinetics-Skeleton action recognition datasets.
Iqbal Singh, Xiaodan Zhu 0001, Michael A. Greenspan
ICIP3
2020 Two-Level Attention-based Fusion Learning for RGB-D Face Recognition
abstract
With recent advances in RGB-D sensing technologies as well as improvements in machine learning and fusion techniques, RGB-D facial recognition has become an active area of research. A novel attention aware method is proposed to fuse two image modalities, RGB and depth, for enhanced RGB-D facial recognition. The proposed method first extracts features from both modalities using a convolutional feature extractor. These features are then fused using a two layer attention mechanism. The first layer focuses on the fused feature maps generated by the feature extractor, exploiting the relationship between feature maps using LSTM recurrent learning. The second layer focuses on the spatial features of those maps using convolution. The training database is preprocessed and augmented through a set of geometric transformations, and the learning process is further aided using transfer learning from a pure 2D RGB image training process. Comparative evaluations demonstrate that the proposed method outperforms other state-of-the-art approaches, including both traditional and deep neural network-based methods, on the challenging CurtinFaces and IIIT-D RGB-D benchmark databases, achieving classification accuracies over 98.2% and 99.3% respectively. The proposed attention mechanism is also compared with other attention mechanisms, demonstrating more accurate results.
Hardik Uppal, Alireza Sepas-Moghaddam, Michael A. Greenspan, Ali Etemad
ICPR3
2020 Inverse Rectification for Efficient Procam Pattern Correspondence
abstract
A method called inverse rectification, is proposed which facilitates the establishment of correspondences across a projected pattern and an acquired image. A pattern of features comprising vertical dashes is warped by the inverse of the rectifying homography of the projector-camera pair, prior to projection. This warping imparts upon the system the property that projected features will fall on distinct conjugate epipolar lines of the rectified projector and acquired camera images. This reduces the correspondence search to a trivial constant-time table lookup once a feature is found in the camera image, and leads to robust, accurate, and extremely efficient disparity calculations. A projectorcamera range sensor is developed based on this method, and is shown experimentally to be effective, with bandwidth exceeding some existing consumer-level range sensors.
Yubo Qiu, Jonathon Malcolm, Abhay Vatoo, Sheikh Ziauddin, Michael A. Greenspan
WACV5
2017 Point Cloud Registration with Virtual Interest Points from Implicit Quadric Surface Intersections
abstract
A novel method is presented to robustly and efficiently register two partially overlapping point clouds. Following segmentation, the regions are represented as implicit quadric surfaces using polynomials of degree two in three variables. The registration establishes correspondences of virtual interest points, which do not exist in the original point cloud data, and are defined by the intersection of three implicit quadric surfaces extracted from the point cloud regions. Implicit quadric surfaces exist in abundance in both natural and architectural scenes, and can be used to identify stable regions in the data, which in turn leads to repeatable virtual interest points. Large regions in a point cloud can be represented by a few implicit surfaces, which reduces the computational cost of registration and also makes the algorithm robust to noise and data density variations. Experiments were performed on seven data sets from various sensors. The proposed method outperformed most of the feature based registration and non-feature based registration methods for computational efficiency and convergence.
Mirza Tahir Ahmed, Joshua A. Marshall, Michael A. Greenspan
3DV3
2015 Super Generalized 4PCS for 3D Registration
abstract
The 4-Points Congruent Sets (4PCS) Algorithm is an established approach to registering two overlapping 3D point sets with partial overlap and arbitrary initial poses. 4PCS performs the registration efficiently using a special set of 4 points, also known as a base, formed by two co-planar pairs of points within a RANSAC framework. The SUPER 4PCS algorithm uses intelligent indexing to reduce the complexity of the original 4PCS algorithm. Although SUPER 4PCS is efficient, we show in this work that one can gain significant practical improvements in runtime by reducing the number of congruent 4-point bases across the two 3D point sets. We accomplish this by using a generalized 4-point base which considers non-coplanar 4-point bases as well as planar ones. We show through experimentation that the number of 4-point bases decreases, sometimes exponentially, with a non-coplanar base. Using this property, we propose the Super Generalized 4PCS algorithm which can exhibit a significant speed-up of up to 6.5x over the Super 4PCS algorithm as demonstrated experimentally.
Mustafa Mohamad, Mirza Tahir Ahmed, David Rappaport, Michael A. Greenspan
3DV4
2015 Visual indoor positioning with a single camera using PnP
abstract
This paper introduces an accurate and inexpensive method for localizing a calibrated monocular camera in 3D indoor environments. The objective of this work is to localize in 6 degrees-of-freedom (6 DOF) in the presence of a 3D map that contains 3D point clouds co-registered with intensity information. This is done by solving the Perspective-n-Point (PnP) problem to accurately compute the camera location in 6 DOF. An efficient data structure is used to store a large set of point clouds co-registered with intensity information, image features, and transformations between the images. This data structure, referred to as the feature database, is implemented such that it retrieves a match for a query image efficiently. Thus the overall process of localization in 6 DOF becomes a real-time process with high efficiency and accuracy. Our technique was tested with two ground truth data sets of indoor environments, an office and a laboratory. The experimental results show the accuracy and the efficiency of our technique, with an average localization error of less than 10 mm from the ground truth in both environments. In addition, localization results on query images obtained using two different cameras in four different environments are presented. This demonstrates that any type of monocular camera may be used during localization, as long as a sufficient number of environmental features can be extracted from the query images.
Edith Deretey, Mirza Tahir Ahmed, Joshua A. Marshall, Michael A. Greenspan
IPIN4
2014 Generalized 4-Points Congruent Sets for 3D Registration
abstract
The 4-Points Congruent Sets (4PCS) algorithm is a state-of-the-art RANSAC-based algorithm for registering two partially overlapping 3D point sets using raw points. Unlike other RANSAC-based algorithms, which try to achieve registration by searching for matching 3-point bases, it uses a base of two coplanar pairs of points to reduce the search space matching bases. In this work, we first generalize the algorithm by allowing the two pairs to fall on two different planes which have an arbitrary distance, i.e. Degree of separation, between them. Furthermore, we show that increasing the degree of separation exponentially decreases the search space of matching bases. Using this property, we show that using the new generalized base allows for more efficient registration than the original 4PCS base type. We achieve a maximum run-time improvement of 83.10% for 3D registration.
Mustafa Mohamad, David Rappaport, Michael A. Greenspan
3DV3
2014 Scene Dynamics Estimation for Parameter Adjustment of Gaussian Mixture Models
abstract
The scene dynamics can provide useful statistical information for adjusting parameters of Gaussian mixture models (GMMs) in video surveillance. The contributions of this paper are twofold. First, an adaptive scene dynamics estimation approach is proposed. Second, we propose a scene-dynamics based method to adjust two types of GMMs' parameters, i.e., the learning rates and number of Gaussian components. For the learning rates, the scene dynamics are integrated into different kinds of pixel-type feedback schemes to control different kinds of learning rates. Experimental results demonstrate that the proposed method can effectively improve the performance of GMMs in surveillance scenes with complex dynamic backgrounds.
Weiguo Gong, Victor Grzeda, Andrew Yaworski, Michael A. Greenspan
IEEE Signal Process. Lett.5
2013 Object Class Recognition in Mobile Urban Lidar Data Using Global Shape Descriptors
abstract
A method is presented to automatically classify objects that lie within the vicinity of streets in 3D point clouds of urban environments. The system first successfully segments objects of interest from the scene through a combination of ground segmentation and road extraction using a Kalman filtering approach, and cluster extraction region growing. Those clusters that fall close to the road are then passed to a classification phase, where they are compared against a labelled dataBase of such clusters. The comparison of clusters is Based upon Variable Dimensional Global Shape Descriptors, which encode the geometry of the objects into multidimensional histograms, the similarities of which are measured against the dataBase clusters using a variety of metrics including Earth Mover's Distance and Bhattacharya similarity. The method was applied to dense data acquired from central New York City, covering an area of 78,000 m2. On a test set containing 101 objects partitioned into 5 classes, the method had an average successful recognition rate of 94.5% for a rich set of vehicles, pedestrians, and street furniture such as fire hydrants, street signs, and poles.
Salar Awan, Mustafa Muhamad, Kresimir Kusevic, Paul Mrstik, Michael A. Greenspan
3DV5
2013 3D Object Recognition by Surface Registration of Interest Segments
abstract
An object recognition system Based on registering repeatable interest segments from 3D surfaces is presented. The strength of this approach lies in its independence of local features, which can be unreliable when corrupted by noise, and indistinct for certain objects and surfaces. The proposed framework is Based on recent advances in segmenting 3D data into repeatable interest segments, followed by efficient surface registration of model and scene segments, where pose clustering returns the best pose candidates. A quality measure Based on reprojection of the model points and pose refinement are then used to select the best pose. The proposed method is demonstrated experimentally to be both accurate and robust when tested against a variety of partially occluded free-form objects in cluttered scenes, achieving an average accuracy of 93% on an accurate and high resolution LiDAR data set, and 81% on a noisy and low resolution Kinect data set.
Joseph Lam, Michael A. Greenspan
3DV2
2013 Automatic Rail Extraction in Terrestrial and Airborne LiDAR Data
abstract
Datasets of stretches of railway tracks are collected using both Airborne and Terrestrial LiDAR scanners having varying density, resolution and provide different views of the railway track. Manual feature extraction from such datasets is tedious and labour intensive. Therefore, automatic extraction of desired features is highly desirable. In this work, we propose a technique to extract the rails from these two types of datasets. Our rail extraction technique models the a railway track as a dynamic system of local pairs of parallel line segments and uses the Kalman filter to predict and monitor the state of the system. The system's state is composed of the two centroids of the parallel line segments as well as their common direction. Additionally, we augment the Kalman filter process to deal with special cases such as missing railway track segments, sensor noise, and data sparseness. Our technique is effective on both types of data sets as we achieve a precision of 97% and a recall of 78% on a high resolution the Terrestrial dataset and a precision of 95% and a recall of 83% on the Airborne dataset.
Mustafa Mohamad, Kresimir Kusevic, Paul Mrstik, Michael A. Greenspan
3DV4
2012 Probabilistic shape parsing for view-based object recognition
Diego Macrini, Chris Whiten, Robert Laganière, Michael A. Greenspan
ICPR4
2012 Nonparametric on-line background generation for surveillance video
Weiguo Gong, Andrew Yaworski, Michael A. Greenspan
ICPR4
2011 Local shape descriptor selection for object recognition in range data
Babak Taati, Michael A. Greenspan
Comput. Vis. Image Underst.2
2011 Model-based segmentation and recognition of dynamic gestures in continuous video streams
Michael A. Greenspan
Pattern Recognit.2
2010 Real-time Object Recognition in Sparse Range Images Using Error Surface Embedding
Limin Shang, Michael A. Greenspan
Int. J. Comput. Vis.2
2007 Variable Dimensional Local Shape Descriptors for Object Recognition in Range Data
abstract
We propose a new set of highly descriptive local shape descriptors (LSDs) for model-based object recognition and pose determination in input range data. Object recognition is performed in three phases: point matching, where point correspondences are established between range data and the complete model using local shape descriptors; pose recovery, where a computationally robust algorithm generates a rough alignment between the model and its instance in the scene, if such an instance is present; and pose refinement. While previously developed LSDs take a minimalist approach, in that they try to construct low dimensional and compact descriptors, we use high (up to 9) dimensional descriptors as the key to more accurate and robust point correspondence. Our strategy significantly simplifies the computational burden of the pose recovery phase by investing more time in the point matching phase. Experiments with Lidar and dense stereo range data illustrate the effectiveness of the approach by providing a higher percentage of correct matches in the candidate point matches list than a leading minimalist technique. Consequently, the number of RANSAC iterations required for recognition and pose determination is drastically smaller in our approach.
Babak Taati, Michel Bondy, Piotr Jasiobedzki, Michael A. Greenspan
ICCV4
2007 Segmentation and Recognition of Continuous Gestures
abstract
A novel method is introduced to segment and recognize time-varying human gestures from continuous video streams. Motion is represented by a 3D spatio-temporal surface based upon the evolution of a contour over time. The warping paths between the input signal and a set of gesture models are obtained using continuous dynamic programming and the boundary of a gesture is located by analyzing all possible gesture candidates during a specific period of time. Correlation and mutual information are employed to select the best candidate when more than one gesture is recognized at the same time period. The system has been implemented and tested on continuous gesture sequences containing 8 different gestures performed by 4 subjects. The results demonstrate that the proposed method is very effective, achieving a recognition rate of 95.9%.
Michael A. Greenspan
ICIP (1)2
2007 An iterative algebraic approach to TCF matrix estimation
abstract
One of the greatest challenges in developing eye- in-hand visual servoing systems is accurate calibration. The central problem of eye-in-hand calibration is to solve the unknown extrinsic parameter X in the homogeneous equation AX = XB. A new formulation using algebraic rearrangement in combination with singular value decomposition is introduced to solve this classical problem. In addition, an iterative algorithm that reduces the error in the initial estimate is used to further improve the results. Experimental results both in simulation and with real data show the method to be easy to use, efficient and accurate. One of the greatest challenges in developing eye-in-hand visual servoing systems is accurate calibration. The central problem of eye-in-hand calibration is to solve the unknown extrinsic parameter X in the homogeneous equation AX = XB. A new formulation using algebraic rearrangement in combination with singular value decomposition is introduced to solve this classical problem. In addition, an iterative algorithm that reduces the error in the initial estimate is used to further improve the results. Experimental results both in simulation and with real data show the method to be easy to use, efficient and accurate.
Joseph Lam, Michael A. Greenspan
IROS2
2007 Model-Based Tracking by Classification in a Tiny Discrete Pose Space
abstract
A method is presented for tracking 3D objects as they transform rigidly in space within a sparse range image sequence. The method operates in discrete space and exploits the coherence across image frames that results from the relationship between known bounds on the object's velocity and the sensor frame rate. These motion bounds allow the interframe transformation space to be reduced to a reasonable and indeed tiny size, comprising only tens or hundreds of possible states. The tracking problem is in this way cast into a classification framework, effectively trading off localization precision for runtime efficiency and robustness. The method has been implemented and tested extensively on a variety of freeform objects within a sparse range data stream comprising only a few hundred points per image. It has been shown to compare favorably against continuous domain Iterative Closest Point (ICP) tracking methods, performing both more efficiently and more robustly. A hybrid method has also been implemented that executes a small number of ICP iterations following the initial discrete classification phase. This hybrid method is both more efficient than the ICP alone and more robust than either the discrete classification method or the ICP separately.
Limin Shang, Piotr Jasiobedzki, Michael A. Greenspan
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Validation of bone segmentation and improved 3-D registration using contour coherency in CT data
abstract
A method is presented to validate the segmentation of computed tomography (CT) image sequences, and improve the accuracy and efficiency of the subsequent registration of the three-dimensional surfaces that are reconstructed from the segmented slices. The method compares the shapes of contours extracted from neighborhoods of slices in CT stacks of tibias. The bone is first segmented by an automatic segmentation technique, and the bone contour for each slice is parameterized as a one-dimensional function of normalized arc length versus inscribed angle. These functions are represented as vectors within a K-dimensional space comprising the first K amplitude coefficients of their Fourier Descriptors. The similarity or coherency of neighboring contours is measured by comparing statistical properties of their vector representations within this space. Experimentation has demonstrated this technique to be very effective at identifying low-coherency segmentations. Compared with experienced human operators, in a set of 23 CT stacks (1,633 slices), the method correctly detected 87.5% and 80% of the low-coherency and 97.7% and 95.5% of the high coherency segmentations, respectively from two different automatic segmentation techniques. Removal of the automatically detected low-coherency segmentations also significantly improved the accuracy and time efficiency of the registration of 3-D bone surface models. The registration error was reduced by over 500% (i.e., a factor of 5) and 280%, and the computational performance was improved by 540% and 791% for the two respective segmentation methods.
Liping Ingrid Wang, Michael A. Greenspan, Randy E. Ellis
IEEE Trans. Medical Imaging2
2005 Multi-Scale Gesture Recognition from Time-Varying Contours
abstract
A novel method is introduced to recognize and estimate the scale of time-varying human gestures. It exploits the changes in contours along spatiotemporal directions. Each contour is first parameterized as a 2D function of radius vs. cumulative contour length, and a 3D surface is composed from a sequence of such functions. In a two-phase recognition process, dynamic time warping is employed to rule out significantly different gesture models, and then mutual information (MI) is applied for matching the remaining models. The system has been tested on 8 gestures performed by 5 subjects with varied time scales. The two-phase process is compared against exhaustively testing three similarity measures based upon MI, correlation, and nonparametric kernel density estimation. Experimental results demonstrate that the exhaustive application of MI is the most robust with a recognition rate of 90.6%, however, the two-phase approach is much more computationally efficient with a comparable recognition rate of 90.0%.
Michael A. Greenspan
ICCV2
2005 An Estimation/Correction Algorithm for Detecting Bone Edges in CT Images
abstract
The normal direction of the bone contour in computed tomography (CT) images provides important anatomical information and can guide segmentation algorithms. Since various bones in CT images have different sizes, and the intensity values of bone pixels are generally nonuniform and noisy, estimation of the normal direction using a single scale is not reliable. We propose a multiscale approach to estimate the normal direction of bone edges. The reliability of the estimation is calculated from the estimated results and, after re-scaling, the reliability is used to further correct the normal direction. The optimal scale at each point is obtained while estimating the normal direction; this scale is then used in a simple edge detector. Our experimental results have shown that use of this estimated/corrected normal direction improves the segmentation quality by decreasing the number of unexpected edges and discontinuities (gaps) of real contours. The corrected normal direction could also be used in postprocessing to delete false edges. Our segmentation algorithm is automatic, and its performance is evaluated on CT images of the human pelvis, leg, and wrist.
Weiguang Yao, Purang Abolmaesumi, Michael A. Greenspan, Randy E. Ellis
IEEE Trans. Medical Imaging3
2004 Efficient Tracking with the Bounded Hough Transform
Michael A. Greenspan, Limin Shang, Piotr Jasiobedzki
CVPR (1)1
2004 Robotic pool: an experiment in automatic potting
abstract
A robotic system is presented which automatically pots (i.e., sinks) pool balls. A homography is estimated that relates the gantry robot coordinate frame to the overhead (global) camera coordinate frame. This homography is computed by first calculating the mapping between the camera frame and a projection of the robot frame, and then solving the pool table plane equation in the robot frame. A measurement technique has been developed which is based upon a local camera attached to the robot end effector. This local camera allows the robot to be positioned accurately over circular targets placed on the table. The homography and table plane equation are then estimated by establishing correspondences between at least 4 measured target positions in the global camera and robot frames. The resulting homography allows the gantry to be positioned to within an average of 0.6 mm of a global camera frame position over the extent of a full sized pool table. The system has been used to pot a ball with 67% accuracy over the extent of the table, with a high repeatability.
Johan Herland, Marie-Christine Tessier, Darryl Naulls, Andrew Roth, Gerhard Roth, Michael A. Greenspan
IROS7
2002 Geometric Probing of Dense Range Data
abstract
A new method is presented for the efficient and reliable pose determination of 3D objects in dense range image data. The method is based upon a minimalistic Geometric Probing strategy that hypothesizes the intersection of the object with some selected image point, and searches for additional surface data at locations relative to that point. The strategy is implemented in the discrete domain as a binary decision tree classifier. The tree leaf nodes represent individual voxel templates of the model, with one template per distinct model pose. The internal nodes represent the union of the templates of their descendant leaf nodes. The union of all leaf node templates is the complete template set of the model over its discrete pose space. Each internal node also encodes a single voxel which is the most common element of its child node templates. Traversing the free is equivalent to efficiently matching the large set of templates at a selected image seed location. The method was implemented and extensive experiments were conducted for a variety of combinations of tree designs and traversals under isolated, cluttered, and occluded scene conditions. The results demonstrated a tradeoff between efficiency and reliability. It was concluded that there exist combinations of tree design and traversal which are both highly efficient and reliable.
Michael A. Greenspan
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Beyond Range Sensing: XYZ-RGB Digitizing and Modeling
abstract
Discusses the progress and the evolution of the development of range sensing techniques at the NRC laboratories. Essentially a 3D imaging project at the beginning, it has evolved to a new media project which requires the development of new tools for 3D modeling, editing, database searching and visualization. Generic applications related to documentation, inspection, target tracking and visual communication will be discussed.
Marc Rioux, François Blais, J.-Angelo Beraldin, Guy Godin, Pierre Boulanger, Michael A. Greenspan
ICRA6
1998 The Sample Tree: A Sequential Hypothesis Testing Approach to 3D Object Recognition
abstract
A method is presented for efficient and reliable object recognition within noisy, cluttered, and occluded range images. The method is based on a strategy which hypothesizes the intersection of the object with some selected image point, and searches for additional surface data at locations relative to that point. At each increment, the image is queried for the existence of surface data at a specific spatial location, and the set of possible object poses is further restricted. Eventually, either the object is identified and localized, or the initial hypothesis is refuted. The strategy is implemented in the discrete domain as a binary decision tree classifier. The tree leaf nodes represent individual voxel templates of the model. The internal tree nodes represent the union of the templates of their descendant leaf nodes. The union of all leaf node templates is the complete template set of the model over its discrete pose space. Each internal node also references a single voxel which is the most common element of its child node templates. Traversing the tree is equivalent to efficiently matching the large set of templates at a selected image seed location. The process is approximately 3 orders of magnitude more efficient than brute-force template matching. Experimental results are presented in which objects are reliably recognized and localized in 6 dimensions in less than 60 seconds within noisy and significantly occluded range images.
Michael A. Greenspan
CVPR1
1997 Sticky and slippery collision avoidance for tele-excavation
abstract
Two new modes of real-time collision avoidance are presented which aid in the effectiveness of remote tele-operation. Sticky collision mode improves system safety by disallowing contact with modelled obstacles. Slippery mode is used do improve the efficiency of the process by reducing the required level of operator skill. Both modes execute in real-time and have been implemented and tested on a remote excavator system. It was found that precise geometries could be excavated using extremely simplified cutting strategies that required very little fine control on the part of the operator, thereby improving the effectiveness of remote tele-excavation.
Michael A. Greenspan, John Ballantyne, Mike Lipsett
IROS1
1996 Obstacle count independent real-time collision avoidance
abstract
Robotic manipulator real-time collision avoidance is a safety critical mode of teleoperation where motion commands which would result in a collision are disallowed. To achieve real-time performance, it is necessary to efficiently detect impending collisions between the manipulator and the workspace obstacles. A collision detection method is presented which is based upon two representations. The dynamic elements, such as the manipulator links, are modelled as sets of spheres. The static elements, such as the workspace obstacles, are represented as a weighted voxel map, in which the value of any voxel is indicative of its distance to the nearest obstacle. Combining these two representations results in a collision detection method which is obstacle count independent, i.e. independent of the number of obstacles in the workspace. This property is desirable for operation in cluttered environments with many obstacles, where the total number of calculations in the alternative collision detection paradigm of pairwise comparison will prohibit real-time performance. The method is efficient enough to satisfy a hard real-time constraint Novel algorithms are described to generate the voxel map and spherical model representations, and an implementation is described which uses the collision detection method for real-time teleoperated collision avoidance and online path planning of a Puma 560 manipulator.
Michael A. Greenspan, Nestor Burtnyk
ICRA1