Mohsen Zand

dblp:160/2516 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
8since 2021 · last 2023
0000-0001-8177-6000ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 Moving Object Detection by Low-Rank Analysis of Region-Based Correlated Motion Fields
abstract
This paper proposes a novel approach for moving object detection in video sequences captured by nonstationary cameras. The approach, called RCMFD, uses region-based correlated motion fields decomposition, which exploits the sparsity of foreground motions against the low-rank structured background motion. The method uses spatial correlations of region-based features to boost accurate change detection for motion estimation, and motion features across object boundaries are used to exploit moving objects. A dense optical field, which is robust to illumination changes and noise, is established using cross-correlation of region-based features, and a robust principal component analysis (RPCA) model is applied to partition exploited motions into background and foreground motions. Experiments demonstrate the robustness of the proposed method on real video sequences.
Bahareh Kalantar, Naonori Ueda, Mohsen Zand, Husam A. H. Al-Najjar
IGARSS3
2023 Flow-Based Spatio-Temporal Structured Prediction of Motion Dynamics
abstract
Conditional Normalizing Flows (CNFs) are flexible generative models capable of representing complicated distributions with high dimensionality and large interdimensional correlations, making them appealing for structured output learning. Their effectiveness in modelling multivariates spatio-temporal structured data has yet to be completely investigated. We propose MotionFlow as a novel normalizing flows approach that autoregressively conditions the output distributions on the spatio-temporal input features. It combines deterministic and stochastic representations with CNFs to create a probabilistic neural generative approach that can model the variability seen in high-dimensional structured spatio-temporal data. We specifically propose to use conditional priors to factorize the latent space for the time dependent modeling. We also exploit the use of masked convolutions as autoregressive conditionals in CNFs. As a result, our method is able to define arbitrarily expressive output probability distributions under temporal dynamics in multivariate prediction tasks. We apply our method to different tasks, including trajectory prediction, motion prediction, time series forecasting, and binary segmentation, and demonstrate that our model is able to leverage normalizing flows to learn complicated time dependent conditional distributions.
Mohsen Zand, Ali Etemad, Michael A. Greenspan
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Keypoint Cascade Voting for Point Cloud Based 6DoF Pose Estimation
abstract
We propose a novel keypoint voting 6DoF object pose estimation method, which takes pure unordered point cloud geometry as input without RGB information. The proposed cascaded keypoint voting method, called RCVPose3D, is based upon a novel architecture which separates the task of semantic segmentation from that of keypoint regression, thereby increasing the effectiveness of both and improving the ultimate performance. The method also introduces a pairwise constraint in between different keypoints to the loss function when regressing the quantity for keypoint estimation, which is shown to be effective, as well as a novel Voter Confident Score which enhances both the learning and inference stages. Our proposed RCVPose3D achieves state-of-the-art performance on the Occlusion LINEMOD (74.5%) and YCB-Video (96.9%) datasets, outperforming existing pure RGB and RGB-D based methods, as well as being competitive with RGB plus point cloud methods.
Yangzheng Wu, Alireza Javaheri, Mohsen Zand, Michael A. Greenspan
3DV3
2022 Vote from the Center: 6 DoF Pose Estimation in RGB-D Images by Radial Keypoint Voting
Yangzheng Wu, Mohsen Zand, Ali Etemad, Michael A. Greenspan
ECCV (10)2
2022 ObjectBox: From Centers to Boxes for Anchor-Free Object Detection
Mohsen Zand, Ali Etemad, Michael A. Greenspan
ECCV (10)1
2022 Multiscale Crowd Counting and Localization By Multitask Point Supervision
abstract
We propose a multitask approach for crowd counting and person localization in a unified framework. As the detection and localization tasks are well-correlated and can be jointly tackled, our model benefits from a multitask solution by learning multiscale representations of encoded crowd images, and subsequently fusing them. In contrast to the relatively more popular density-based methods, our model uses point supervision to allow for crowd locations to be accurately identified. We test our model on two popular crowd counting datasets, ShanghaiTech A and B, and demonstrate that our method achieves strong results on both counting and localization tasks, with MSE measures of 110.7 and 15.0 for crowd counting and AP measures of 0.71 and 0.75 for localization, on ShanghaiTech A and B respectively. Our detailed ablation experiments show the impact of our multiscale approach as well as the effectiveness of the fusion module embedded in our network. Our code is available at: https://github.com/RCVLab-AiimLab/crowdcounting
Mohsen Zand, Haleh Damirchi, Andrew Farley, Mahdiyar Molahasani, Michael A. Greenspan, Ali Etemad
ICASSP1
2022 Oriented Bounding Boxes for Small and Freely Rotated Objects
abstract
A novel object detection method is presented that handles freely rotated objects of arbitrary sizes, including tiny objects as small as$2 \times 2$pixels. Such tiny objects appear frequently in remotely sensed images, and present a challenge to recent object detection algorithms. More importantly, current object detection methods have been designed originally to accommodate axis-aligned bounding box detection, and therefore fail to accurately localize oriented boxes that best describe freely rotated objects. In contrast, the proposed convolutional neural network (CNN) -based approach uses potential pixel information at multiple scale levels without the need for any external resources, such as anchor boxes. The method encodes the precise location and orientation of features of the target objects at grid cell locations. Unlike existing methods that regress the bounding box location and dimension, the proposed method learns all the required information by classification, which has the added benefit of enabling oriented bounding box detection without any extra computation. It thus infers the bounding boxes only at inference time by finding the minimum surrounding box for every set of the same predicted class labels. Moreover, a rotation-invariant feature representation is applied to each scale, which imposes a regularization constraint to enforce covering the 360° range of in-plane rotation of the training samples to share similar features. Evaluations on the xView and dataset for object detection in aerial images (DOTA) data sets show that the proposed method uniformly improves performance over existing state-of-the-art methods.
Mohsen Zand, Ali Etemad, Michael A. Greenspan
IEEE Trans. Geosci. Remote. Sens.1
2021 Multistream Validnet: Improving 6D Object Pose Estimation by Automatic Multistream Validation
abstract
This work presents a novel approach to improve the results of pose estimation by detecting and distinguishing between the occurrence of True and False Positive results. It achieves this by training a binary classifier on the output of an arbitrary pose estimation algorithm, and returns a binary label indicating the validity of the result. We demonstrate that our approach improves upon a state-of-the-art pose estimation result on the Siléane dataset, outperforming a variation of the alternative CullNet method by 4.15% in average class accuracy and 0.73% in overall accuracy at validation. Applying our method can also improve the pose estimation average precision results of Op-Net by 6.06% on average.
Joy Mazumder, Mohsen Zand, Michael A. Greenspan
ICIP2
2020 One Shot Radial Distortion Correction by Direct Linear Transformation
abstract
A novel method is proposed to estimate and correct image radial distortion. The solution is based upon an algebraic expansion of the homographic relationship between a planar pattern and its distorted projection into a single image, which is solved using the Direct Linear Transformation. The method requires ten point correspondences, is fully automatic and estimates both the first two parameters of the division model, and the center of distortion. Experimental results show the method to be more accurate than other recent one- and two-parameter point-based and plumb-line-based approaches.
Sheikh Ziauddin, Mohsen Zand, Michael A. Greenspan
ICIP3
2017 Visual and semantic context modeling for scene-centric image annotation
Mohsen Zand, Shyamala C. Doraisamy, Alfian Abdul Halin, Mas Rina Mustaffa
Multim. Tools Appl.1
2017 Multiple Moving Object Detection From UAV Videos Using Trajectories of Matched Regional Adjacency Graphs
abstract
Image registration has been long used as a basis for the detection of moving objects. Registration techniques attempt to discover correspondences between consecutive frame pairs based on image appearances under rigid and affine transformations. However, spatial information is often ignored, and different motions from multiple moving objects cannot be efficiently modeled. Moreover, image registration is not well suited to handle occlusion that can result in potential object misses. This paper proposes a novel approach to address these problems. First, segmented video frames from unmanned aerial vehicle captured video sequences are represented using region adjacency graphs of visual appearance and geometric properties. Correspondence matching (for visible and occluded regions) is then performed between graph sequences by using multigraph matching. After matching, region labeling is achieved by a proposed graph coloring algorithm which assigns a background or foreground label to the respective region. The intuition of the algorithm is that background scene and foreground moving objects exhibit different motion characteristics in a sequence, and hence, their spatial distances are expected to be varying with time. Experiments conducted on several DARPA VIVID video sequences as well as self-captured videos show that the proposed method is robust to unknown transformations, with significant improvements in overall precision and recall compared to existing works.
Bahareh Kalantar, Shattri Mansor, Alfian Abdul Halin, Helmi Z. M. Shafri, Mohsen Zand
IEEE Trans. Geosci. Remote. Sens.5
2016 Ontology-Based Semantic Image Segmentation Using Mixture Models and Multiple CRFs
abstract
Semantic image segmentation is a fundamental yet challenging problem, which can be viewed as an extension of the conventional object detection with close relation to image segmentation and classification. It aims to partition images into non-overlapping regions that are assigned predefined semantic labels. Most of the existing approaches utilize and integrate low-level local features and high-level contextual cues, which are fed into an inference framework such as, the conditional random field (CRF). However, the lack of meaning in the primitives (i.e., pixels or superpixels) and the cues provides low discriminatory capabilities, since they are rarely object-consistent. Moreover, blind combinations of heterogeneous features and contextual cues exploitation through limited neighborhood relations in the CRFs tend to degrade the labeling performance. This paper proposes an ontology-based semantic image segmentation (OBSIS) approach that jointly models image segmentation and object detection. In particular, a Dirichlet process mixture model transforms the low-level visual space into an intermediate semantic space, which drastically reduces the feature dimensionality. These features are then individually weighed and independently learned within the context, using multiple CRFs. The segmentation of images into object parts is hence reduced to a classification task, where object inference is passed to an ontology model. This model resembles the way by which humans understand the images through the combination of different cues, context models, and rule-based learning of the ontologies. Experimental evaluations using the MSRC-21 and PASCAL VOC'2010 data sets show promising results.
Mohsen Zand, Shyamala C. Doraisamy, Alfian Abdul Halin, Mas Rina Mustaffa
IEEE Trans. Image Process.1
2015 Texture classification and discrimination for region-based image retrieval
abstract
In RBIR, texture features are crucial in determining the class a region belongs to since they can overcome the limitations of color and shape features. Two robust approaches to model texture features are Gabor and curvelet features. Although both features are close to human visual perception, sufficient information needs to be extracted from their sub-bands for effective texture classification. Moreover, shape irregularity can be a problem since Gabor and curvelet transforms can only be applied on the regular shapes. In this paper, we propose an approach that uses both the Gabor wavelet and the curvelet transforms on the transferred regular shapes of the image regions. We also apply a fitting method to encode the sub-bands’ information in the polynomial coefficients to create a texture feature vector with the maximum power of discrimination. Experiments on texture classification task with ImageCLEF and Outex databases demonstrate the effectiveness of the proposed approach.
Mohsen Zand, Shyamala C. Doraisamy, Alfian Abdul Halin, Mas Rina Mustaffa
J. Vis. Commun. Image Represent.1