Maria A. Amer

dblp:172/9646 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-6821-3032ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Ranking of Visual Trackers Using Robust Error Norms
abstract
Object trackers are typically ranked by the average of averages, that is, a performance measure averaged over all frames of a video and then averaged over the entire dataset. The average is not a robust estimator. We propose to rank trackers based on robust error norms: we divide the performances of a set of trackers for a video, sorted from best to worst, into outliers (edge trackers) and inliers (trackers with similar performances); we propose an edge-stopping function that assigns the highest score to the highest-performance (top) tracker and scores other trackers accordingly. Our edge-stopping function stops at edge trackers (outliers) using a robust scale defined using the difference (error) between the performances of the top tracker and neighboring trackers. Our method is not a new performance measure but an approach to rank trackers robustly and systematically. We test our methods using five video datasets and 20 trackers. We show that the proposed score is more robust and representative of a tracker’s performance than the widely-used average of averages.
Julien Valognes, Maria A. Amer
ICASSP2
2024 Separable Self and Mixed Attention Transformers for Efficient Object Tracking
abstract
The deployment of transformers for visual object tracking has shown state-of-the-art results on several benchmarks. However, the transformer-based models are underutilized for Siamese lightweight tracking due to the computational complexity of their attention blocks. This paper proposes an efficient self and mixed attention transformer-based architecture for lightweight tracking. The proposed backbone utilizes the separable mixed attention transformers to fuse the template and search regions during feature extraction to generate superior feature encoding. Our prediction head performs global contextual modeling of the encoded features by leveraging efficient self-attention blocks for robust target state estimation. With these contributions, the proposed lightweight tracker deploys a transformer-based backbone and head module concurrently for the first time. Our ablation study testifies to the effectiveness of the proposed combination of backbone and head modules. Simulations show that our Separable Self and Mixed Attention-based Tracker, SMAT, surpasses the performance of related lightweight trackers on GOT10k, TrackingNet, LaSOT, NfS30, UAV123, and AVisT datasets, while running at 37 fps on CPU, 158 fps on GPU, and having 3.8M parameters. For example, it significantly surpasses the closely related trackers E.T.Track and MixFormerV2-S on GOT10ktest by a margin of 7.9% and 5.8%, respectively, in the AO metric. The tracker code and model is available at https://github.com/goutamyg/SMAT.
Goutam Yelluru Gopal, Maria A. Amer
WACV2
2024 Defending Object Detection Models against Image Distortions
abstract
Image distortions pose a significant challenge to object detection. To address this issue, our paper introduces a novel data augmentation method that generates new samples resembling the original training images. The new sample exhibits randomly altered pixels based on a pixel distribution obtained from multiple image distortions using kernel density estimation (KDE). The main steps of our method, GSES, are generating distorted versions of each pixel of an original training image, selecting a set of pixels in each version, and then, for each selected pixel, estimating its distribution using KDE and then sampling one pixel from this distribution. By employing this approach, the new samples possess distorted pixels while maintaining a certain degree of similarity to the original image. This degree of similarity is essential to balance the accuracy of object detection models under distorted and clean images. Our approach improves the accuracy of different object detection models under 15 image distortions, such as motion blur, fog, and noise. For example, the average accuracy of YOLOv4 improves by 9.19% and 9.54 % across all 15 distortions added to the COCO and PASCAL datasets, respectively. Our method surpasses other defence methods to combat image distortions. Our ablation and stability studies show why our method performs well. Moreover, we also show that our method can be well used to improve the accuracy of image classification under 15 distortions and cross-domains. Our code is available at https://github.com/moforio/GSES/.
Mark Ofori-Oduro, Maria A. Amer
WACV2
2024 Artificial immune systems for data augmentation
Mark Ofori-Oduro, Maria A. Amer
Image Vis. Comput.2
2024 Reliable interconnected channels for dynamic DCF based visual tracking
Goutam Yelluru Gopal, Maria A. Amer
Multim. Tools Appl.2
2023 Mobile Vision Transformer-based Visual Object Tracking
Goutam Yelluru Gopal, Maria A. Amer
BMVC2
2021 Realistic Augmentation For Effective 2d Human Pose Estimation Under Occlusion
abstract
Occlusion is a major challenge for effective human pose estimation, occurring naturally in a high percentage of real-world images. Handling occlusion has been a difficult challenge in literature due to a lack of a proper dataset with an actual focus on occlusion, prompting researchers to create artificial datasets as a means of data augmentation. However, all of these datasets lack the features of a real-world occlusion. In this work, we introduce a new realistic data augmentation approach built on top of a base dataset (here the Human3.6m) that tackles this issue, creating realistic samples similar to those found in the wild. Arguing that CNN models pay higher attention to local as opposed to global features, we define occlusion levels, select many to-occlude objects from different categories, and blend those within the original image from the base dataset. We, then, test top-performing 2D human pose estimation models with and without this occlusion-augmented dataset (called RealOcc) to display the drop in performance under occlusion and then train them on the new dataset to show the increase in the accuracy of the model under occlusion, without any change to the models themselves.
Amin Ansarian, Maria A. Amer
ICIP2
2020 Drift Detection and Correction Post-Tracking
abstract
Accurate object tracking is a challenging problem due to numerous factors, that may cause the tracker to drift away from the target object. Typically, the output of a tracker is a bounding box (BB); such BB may not well discriminate the object from its background and may not be centered correctly around the object. This paper proposes a method that first detects, at each frame, if a tracker tends to drift by analyzing saliency features of the output BB of a tracker, and then applies automatic seeded object segmentation on the BB to correct the drift once detected. Such segmentation is meant to relocate (recenter) the BB adaptive to the object segmented. As seeds, we propose to use SIFT and salient points conditioned they are non-background pixels. Different than related work, our approach thus models drift external to a base tracker by examining its output BB at each and corrects drift, as needed, by updating that BB adaptive to segmentation. We show the ability of the proposed method to significantly improve the tracking quality of base trackers. We also show that the proposed method outperforms by far segmentation-based trackers.
Tarek Ghoniemy, Maria A. Amer
ICASSP2
2020 Dynamic Channel Pruning For Correlation Filter Based Object Tracking
abstract
Fusion of multi-channel representations has played a crucial role in the success of correlation filter (CF) based trackers. But, all channels do not contain useful information for target localization at every frame. During challenging scenarios, ambiguous responses of non-discriminative or unreliable channels lead to erroneous results and cause tracker drift. To mitigate this problem, we propose a method for dynamic channel pruning through online (i.e., at every frame) learning of channel weights. Our method uses estimated reliability scores to compute channel weights, to nullify the impact of highly unreliable channels. The proposed method for learning of channel weights is modeled as a non-smooth convex optimization problem. We then propose an algorithm to solve the resulting problem efficiently compared to off-the-shelf solvers. Results on VOT2018 and TC128 datasets show that proposed method improves the performance of baseline CF trackers.
Goutam Yelluru Gopal, Maria A. Amer
ICASSP2
2020 Optimization Using Artificial Immune Systems Applied To Object Tracking And Segmentation
abstract
This paper proposes the use of an artificial immune systems (AIS) to obtain the values of hyperparameters of networks such as the kernel parameter of the support vector machines (SVM) in object tracking and weighting factor of the loss term in object segmentation. The proposed iterative AIS method is generic to extend to other image processing tasks by formulating a corresponding objective function (fitness). We verify our method on the STRUCK method that uses SVM to track objects. Depending on feature variations between video frames, our AIS approach incorporates a complementary SVM model to select the SVM parameters for the main SVM model, where our AIS stopping criteria are classification accuracy and number of iterations. We then apply our AIS method to find the parameters that simultaneously minimize both false positives and false negatives of the object segmentation method Graph-Cut. Our results show that our AIS approach achieves significant enhancement of Graph-Cut segmentation accuracy and of STRUCK tracking quality.
Tarek Ghoniemy, Maria A. Amer
ICIP2
2020 Reliable Temporally Consistent Feature Adaptation for Visual Object Tracking
abstract
Correlation Filter (CF) based trackers have been the frontiers on various object tracking benchmarks. Use of multiple features and sophisticated learning methods have increased the accuracy of tracking results. However, the contribution of features are often fixed throughout the video sequence. Unreliable features lead to erroneous target localization and result in tracking failures. To alleviate this problem, we propose a method for online adaptation of feature weights based on their reliability. Our method also includes the notion of temporal consistency, to handle noisy reliability estimates. The two objectives are coupled to model a convex optimization problem for robust learning of feature weights. We also propose an algorithm to efficiently solve the resulting optimization problem, without hindering tracking speed. Results on VOT2018, TC128 and NfS30 datasets show that proposed method improves the performance of baseline CF trackers.
Goutam Yelluru Gopal, Maria A. Amer
ICIP2
2020 Data Augmentation Using Artificial Immune Systems For Noise-Robust CNN Models
abstract
CNN based models are the state-of-the-art in object detection. The effect of noise on their performances has not been extensively examined. In this paper, we examine the models SSD, Yolo, and Faster RCNN, and show that their performance dropped greatly under white noise by an average mAP of 10.33 on the PASCAL-VOC dataset. We propose mitigating this issue by augmenting the training dataset with antibodies generated using Artificial Immune Systems (AIS). We then test the CNN models under different noise levels and show that our data augmentation approach significantly improves their performance under noise by more than 55%, i.e., the average mAP drop reduced to 4.37, without altering their speed. We also show that training of said CNN models under noise does improve their performance but interestingly less than when using our data augmentation AIS approach.
Mark Ofori-Oduro, Maria A. Amer
ICIP2
2018 Robust Scoring and Ranking of Object Tracking Techniques
abstract
Object tracking is an active research area and numerous techniques have been proposed recently. To evaluate a new tracker, its performance is compared against existing ones typically by averaging its quality based on a performance measure, over all test video sequences. Such averaging is, however, not representative as it does not account for outliers (or similarities) between trackers. This paper presents a framework for scoring and ranking of trackers using uncorrelated quality metrics (overlap ratio and failure rate), coupled with a robust estimator (median absolute deviation) against outliers. Ten different performing trackers are scored and ranked using the proposed framework on a public benchmark of 100 sequences. The obtained results show that our framework well highlights and distinguishes the relative performance of each tracker.
Tarek Ghoniemy, Julien Valognes, Maria A. Amer
ICIP3
2018 Low-Frequency Image Noise Removal Using White Noise Filter
abstract
Image noise filters usually assume noise as white Gaussian. However, in a capturing pipeline, noise often becomes spatially correlated due to in-camera processing that aims to suppress the noise and increase the compression rate. Mostly, only high-frequency noise components are suppressed since the image signal is more likely to appear in the low-frequency components of the captured image. As a result, noise emerges as coarse grain which makes white (all-pass) noise filters ineffective, especially when the resolution of the target display is lower than the captured image. Denoising of image approximation in coarse scale has the advantage of removing low-frequency noise, however, lack of spatial resolution degrades the image quality. This paper presents an approach for a coarse-grain removal. Our approach utilizes existing white Gaussian noise filters to address low-frequency component of spatially correlated noises, employing pixel decoupling, local shrinkage, and soft thresholding. Subjective and objective results show that the proposed approach better handles low-frequency noise compared to related work.
Meisam Rakhshanfar, Maria A. Amer
ICIP2
2018 Deep 3D Human Pose Estimation Under Partial Body Presence
abstract
This paper addresses the problem of 3D human pose estimation when not all body parts are present in the input image, i.e., when some body joints are present while other joints are fully absent (we exclude self-occlusion). State-of-the-art is not designed and thus not effective for such cases. We propose a deep CNN to regress the human pose directly from an input image; we design and train this network to work under partial body presence. Parallel to this, we train a detection network to classify the presence or absence of each of the main body joints in the input image. The outputs of our detection and regression networks are a) joints that are present and b) joints that are absent. With these outputs, our method reconstructs the full body skeleton. Evaluations on the Hu-man3.6M dataset yield promising results compared to related work.
Saeid Vosoughi, Maria A. Amer
ICIP2
2016 Estimation of Gaussian, Poissonian-Gaussian, and Processed Visual Noise and Its Level Function
abstract
We propose a method for estimating the image and video noises of different types: white Gaussian (signal-independent), mixed Poissonian-Gaussian (signal-dependent), or processed (non-white). Our method also estimates the noise level function (NLF) of these types. We do so by classifying image patches based on their intensity and variance in order to find homogeneous regions that best represent the noise. We assume that the noise variance is a piecewise linear function of intensity in each intensity class. To find noise representative regions, noisy (signal-free) patches are first nominated in each intensity class. Next, clusters of connected patches are weighted, where the weights are calculated based on the degree of similarity to the noise model. The highest ranked cluster defines the peak noise variance, and other selected clusters are used to approximate the NLF. The more information we incorporate, such as temporal data and camera settings, the more reliable the estimation becomes. To account for the processed noise, (i.e., remaining after in-camera processing), we consider the ratio of low-to-high-frequency energies. We address noise variations along video signals using a temporal stabilization of the estimated noise. Objective and subjective simulations demonstrate that the proposed method outperforms other noise estimation techniques, both in accuracy and speed.
Meisam Rakhshanfar, Maria A. Amer
IEEE Trans. Image Process.2
2015 No-reference image quality assessment for removal of processed and unprocessed noise
abstract
We present a fast no-reference quality measurement method based on the entropy analysis of the image structure. The information that an images carries is represented not only by intensity values but also their position in the image. We examine the behavior of the entropy by altering the positions to distinguish the image structure, noise and blur. We propose a method to rearrange the pixel placements using a random pixel scattering transform. Using this transform we firstly determine the local power of image structure relative to noise and then we obtain the quality index by combining this information with our proposed blur factor. We show that this method is useful to select denoising parameters automatically in both unprocessed (white Gaussian) and processed (frequency-dependent) noise when the reference image is not available. Our method is easy to implement and yet rivals state-of-the-art quality measurement approaches.
Meisam Rakhshanfar, Maria A. Amer
ICIP2
2015 Object tracking with adaptive motion modeling of particle filter and support vector machines
abstract
In this paper, we propose an approach for tracking arbitrary objects based on dynamic motion model and support vector machines (SVMs). When the motion of target is large or abrupt, modeling target motion is crucial for robust object tracking, however, recent advanced trackers ignore this component. In our proposed approach, we represent target motion as a random stochastic process. We use Kernelized Harmonic Means to predict the next state of target motion using few prior state vectors, and then utilize Particle filter to further optimize the predicted state. Because SVMs possess good generalization ability, while being robust against noise, we adopt online ker-nelized SVMs to the tracking problem. Our approach learns the appearance of the target during tracking, and thus the proposed method is able to adapt online to target appearance changes and its surrounding background. In addition, we incorporate our motion model within the online kernelized SVMs framework as an energy map to assign higher energy for the support vectors that are closer to the location predicted by the proposed motion model. This allows reliably maintaining smooth trajectories without unnatural jittering artifacts, which is important for long-term object tracking in the presence of occlusion and noise. Taking the dynamic model into account also improves the computational efficiency as it reduces the dense search space required for localizing the target candidate. Experimentally, we demonstrate that the proposed method outperforms state-of-the-art trackers on particularly challenging standard datasets.
Kumara Ratnayake, Maria A. Amer
ICIP2