Umar Asif

dblp:11/10094 · DBLP profile ↗
← Back
13ranked-venue papers
13as first author
1since 2021 · last 2023
0000-0001-5209-7084ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-authorSystems, architecture and hardware · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Robot manipulation · 46% Image recognition and object detection · 28% Efficient and distributed learning · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › grasping
grasp detection
1.032019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded Forests · IEEE Trans. Robotics 2017
Robotics › Robot manipulation
grasping
1.032019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded Forests · IEEE Trans. Robotics 2017
Computer vision › Image recognition and object detection › object recognition › multimodal object recognition
RGB-D object recognition
0.622018
A Multi-Modal, Discriminative and Spatially Invariant CNN for RGB-D Object Labeling · IEEE Trans. Pattern Anal. Mach. Intell. 2018
RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded Forests · IEEE Trans. Robotics 2017
Machine learning › Efficient and distributed learning › efficient neural network design
efficient CNN architecture
0.312018
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Computer vision › Image recognition and object detection
object recognition
0.312017
RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded Forests · IEEE Trans. Robotics 2017
Computer vision › 3D vision › 3d scene reconstruction
dense scene reconstruction
0.212016
Simultaneous dense scene reconstruction and object labeling · ICRA 2016
Computer vision › Image recognition and object detection › image classification
object classification
0.212015
Efficient RGB-D object categorization using cascaded ensembles of randomized decision trees · ICRA 2015
Robotics › Robot manipulation › grasping › grasp planning
grasp selection
0.212014
Model-Free Segmentation and Grasp Selection of Unknown Stacked Objects · ECCV (5) 2014
Computer vision › Image recognition and object detection › object detection
object proposal generation
0.112019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112018
A Multi-Modal, Discriminative and Spatially Invariant CNN for RGB-D Object Labeling · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Image recognition and object detection
object labeling
0.112016
Simultaneous dense scene reconstruction and object labeling · ICRA 2016
Computer vision › 3D vision › multimodal perception
RGB-D perception
0.112015
Efficient RGB-D object categorization using cascaded ensembles of randomized decision trees · ICRA 2015

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.7dilated convolution · 0.7dense residual connection · 0.7layer-wise feature fusion · 0.4fully convolutional network · 0.4spatial transformer network · 0.3fisher encoding · 0.3conditional random field · 0.3uncertainty minimization · 0.3hierarchical cascaded forests · 0.3
YearPublicationVenuePosition
2023 DeepActsNet: A deep ensemble framework combining features from face, hands, and body for action recognition
Umar Asif, Deval Mehta 0001, Stefan von Cavallar, Jianbin Tang, Stefan Harrer
Pattern Recognit.1
2020 SSHFD: Single Shot Human Fall Detection with Occluded Joints Resilience
abstract
Falling can have fatal consequences for elderly people especially if the fallen person is unable to call for help due to loss of consciousness or any injury. Automatic fall detection systems can assist through prompt fall alarms and by minimizing the fear of falling when living independently at home. Existing vision-based fall detection systems lack generalization to unseen environments due to challenges such as variations in physical appearances, different camera viewpoints, occlusions, and background clutter. In this paper, we explore ways to overcome the above challenges and present Single Shot Human Fall Detector (SSHFD), a deep learning based framework for automatic fall detection from a single image. This is achieved through two key innovations. First, we present a human pose based fall representation which is invariant to appearance characteristics. Second, we present neural network models for 3d pose estimation and fall recognition which are resilient to missing joints due to occluded body parts. Experiments on public fall datasets show that our framework successfully transfers knowledge of 3d pose estimation and fall recognition learnt purely from synthetic data to unseen real-world data, showcasing its generalization capability for accurate fall detection in real-world scenarios.
Umar Asif, Stefan von Cavallar, Jianbin Tang, Stefan Harrer
ECAI1
2020 Ensemble Knowledge Distillation for Learning Improved and Efficient Networks
abstract
Ensemble models comprising of deep Convolutional Neural Networks (CNN) have shown significant improvements in model generalization but at the cost of large computation and memory requirements. In this paper, we present a framework for learning compact CNN models with improved classification performance and model generalization. For this, we propose a CNN architecture of a compact student model with parallel branches which are trained using ground truth labels and information from high capacity teacher networks in an ensemble learning fashion. Our framework provides two main benefits: i) Distilling knowledge from different teachers into the student network promotes heterogeneity in learning features at different branches of the student network and enables the network to learn diverse solutions to the target problem. ii) Coupling the branches of the student network through ensembling encourages collaboration and improves the quality of the final predictions by reducing variance in the network outputs. Experiments on the well established CIFAR-10 and CIFAR-100 datasets show that our Ensemble Knowledge Distillation (EKD) improves classification accuracy and model generalization especially in situations with limited training data. Experiments also show that our EKD based compact networks outperform in terms of mean accuracy on the test datasets compared to other knowledge distillation based methods.
Umar Asif, Jianbin Tang, Stefan Harrer
ECAI1
2019 Densely Supervised Grasp Detector (DSGD)
abstract
This paper presents Densely Supervised Grasp Detector (DSGD), a deep learning framework which combines CNN structures with layer-wise feature fusion and produces grasps and their confidence scores at different levels of the image hierarchy (i.e., global-, region-, and pixel-levels). Specifically, at the global-level, DSGD uses the entire image information to predict a grasp. At the region-level, DSGD uses a region proposal network to identify salient regions in the image and uses a grasp prediction network to generate segmentations and their corresponding grasp poses of the salient regions. At the pixel-level, DSGD uses a fully convolutional network and predicts a grasp and its confidence at every pixel. During inference, DSGD selects the most confident grasp as the output. This selection from hierarchically generated grasp candidates overcomes limitations of the individual models. DSGD outperforms state-of-the-art methods on the Cornell grasp dataset in terms of grasp accuracy. Evaluation on a multi-object dataset and real-world robotic grasping experiments show that DSGD produces highly stable grasps on a set of unseen objects in new environments. It achieves 97% grasp detection accuracy and 90% robotic grasping success rate with real-time inference speed.
Umar Asif, Jianbin Tang, Stefan Harrer
AAAI1
2018 EnsembleNet: Improving Grasp Detection using an Ensemble of Convolutional Neural Networks
Umar Asif, Jianbin Tang, Stefan Harrer
BMVC1
2018 GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices
abstract
Recent research on grasp detection has focused on improving accuracy through deep CNN models, but at the cost of large memory and computational resources. In this paper, we propose an efficient CNN architecture which produces high grasp detection accuracy in real-time while maintaining a compact model design. To achieve this, we introduce a CNN architecture termed GraspNet which has two main branches: i) An encoder branch which downsamples an input image using our novel Dilated Dense Fire (DDF) modules - squeeze and dilated convolutions with dense residual connections. ii) A decoder branch which upsamples the output of the encoder branch to the original image size using deconvolutions and fuse connections. We evaluated GraspNet for grasp detection using offline datasets and a real-world robotic grasping setup. In experiments, we show that GraspNet achieves superior grasp detection accuracy compared to the stateof-the-art computation-efficient CNN models with real-time inference speed on embedded GPU hardware (Nvidia Jetson TX1), making it suitable for low-powered devices.
Umar Asif, Jianbin Tang, Stefan Harrer
IJCAI1
2018 A Multi-Modal, Discriminative and Spatially Invariant CNN for RGB-D Object Labeling
abstract
While deep convolutional neural networks have shown a remarkable success in image classification, the problems of inter-class similarities, intra-class variances, the effective combination of multi-modal data, and the spatial variability in images of objects remain to be major challenges. To address these problems, this paper proposes a novel framework to learn a discriminative and spatially invariant classification model for object and indoor scene recognition using multi-modal RGB-D imagery. This is achieved through three postulates: 1) spatial invariance $-$ this is achieved by combining a spatial transformer network with a deep convolutional neural network to learn features which are invariant to spatial translations, rotations, and scale changes, 2) high discriminative capability $-$ this is achieved by introducing Fisher encoding within the CNN architecture to learn features which have small inter-class similarities and large intra-class compactness, and 3) multi-modal hierarchical fusion$-$ this is achieved through the regularization of semantic segmentation to a multi-modal CNN architecture, where class probabilities are estimated at different hierarchical levels (i.e., image- and pixel-levels), and fused into a Conditional Random Field (CRF)-based inference hypothesis, the optimization of which produces consistent class labels in RGB-D images. Extensive experimental evaluations on RGB-D object and scene datasets, and live video streams (acquired from Kinect) show that our framework produces superior object and scene classification results compared to the state-of-the-art methods.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded Forests
abstract
This paper presents an efficient framework to perform recognition and grasp detection of objects from RGB-D images of real scenes. The framework uses a novel architecture of hierarchical cascaded forests, in which object-class and grasp-pose probabilities are computed at different levels of an image hierarchy (e.g., patch and object levels) and fused to infer the class and the grasp of unseen objects. We introduce a novel training objective function that minimizes the uncertainties of the class labels and the grasp ground truths at the leaves of the forests, thereby enabling the framework to perform the recognition and grasp detection of objects. Our objective function is learned from features that are extracted from RGB-D point clouds of the objects. For that, we propose a novel method to encode an RGB-D point cloud into a representation that facilitates the use of large convolution neural networks to extract discriminative features from RGB-D images. We evaluate our framework on challenging object datasets, where we demonstrate that our framework outperforms the state-of-the-art methods in terms of object-recognition and grasp-detection accuracies. We also show experiments by using live video streams from a Kinect mounted on our in-house robotic platform.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
IEEE Trans. Robotics1
2016 Simultaneous dense scene reconstruction and object labeling
abstract
This paper presents an efficient system for simultaneous dense scene reconstruction and object labeling in real-world environments (captured with an RGB-D sensor). The proposed system starts with the generation of object proposals in the scene. It then tracks spatio-temporally consistent object proposals across multiple frames and produces a dense reconstruction of the scene. In parallel, the proposed system uses an efficient inference algorithm, where object class probabilities are computed at an object-level and fused into a voxel-based prediction hypothesis modeled on the voxels of the reconstructed scene. Our extensive experiments using challenging RGB-D object and scene datasets, and live video streams from Microsoft Kinect show that the proposed system achieved competitive 3D scene reconstruction and object labeling results compared to the state-of-the-art methods.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
ICRA1
2015 Efficient RGB-D object categorization using cascaded ensembles of randomized decision trees
abstract
This paper presents an efficient framework for the categorization of objects in real-world scenes (captured with an RGB-D sensor). The proposed framework uses ensembles of randomized decision trees in a hierarchical cascaded architecture to compute consistent object-class inferences of unseen objects. Specifically, the proposed framework computes object-class probabilities at three levels of an image hierarchy (i.e., pixel-, surfel-, and object-levels) using Random Forest classifiers. Next, these probabilities are fused together to compute a cumulative probabilistic output which is used to infer object categories. This fusion results in an improved object categorization performance compared with the state-of-the-art methods.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
ICRA1
2015 Discriminative feature learning for efficient RGB-D object recognition
abstract
This paper presents an efficient approach to recognize objects captured with an RGB-D sensor. The proposed approach uses a Bag-of-Words (BOW) model to learn feature representations from raw RGB-D point clouds in a weakly supervised manner. To this end, we introduce a novel method based on randomized clustering trees to learn visual vocabularies which are fast to compute and more discriminative compared to the vocabularies generated by classical methods such as k-means. We show that, when combined with standard spatial pooling strategies, our proposed approach yields a powerful feature representation for RGB-D object recognition. Our extensive experimental evaluation on two challenging RGB-D object datasets and live video streams from Kinect shows that our learned features result in superior object recognition accuracies compared with the state-of-the-art methods.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
IROS1
2014 Model-Free Segmentation and Grasp Selection of Unknown Stacked Objects
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
ECCV (5)1
2014 A model-free approach for the segmentation of unknown objects
abstract
We address the problem of object segmentation from depth images of highly complex indoor scenes. We propose a model-free segmentation approach, which robustly separates unknown stacked objects in real-world scenes. Our approach constructs geometrically constrained 3D clusters known as salient-regions, which are subsequently merged into high-level object hypotheses by analyzing the local geometrical characteristics (such as local shape and homogeneity) of the area of their shared boundaries. We tested our approach using depth images from live Kinect video streams and publicly available RGB-D datasets. Our approach is highly efficient and achieves superior performance compared to state-of-the-art techniques.
Umar Asif, Mohammed Bennamoun, Ferdous Sohel
IROS1