Jianbin Tang

dblp:99/11070 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0001-5440-0796ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Robot manipulation · 65% Efficient and distributed learning · 30% Image recognition and object detection · 5%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 92% Embedded and real-time systems · 8%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › grasping
grasp detection
0.722019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Robotics › Robot manipulation
grasping
0.722019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Machine learning › Efficient and distributed learning › efficient neural network design
efficient CNN architecture
0.312018
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices · IJCAI 2018
Emerging computing paradigms
neuromorphic computing
0.312017
Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic Chips · IJCAI 2017
Emerging computing paradigms
neuromorphic hardware
0.312017
Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic Chips · IJCAI 2017
Emerging computing paradigms › neuromorphic computing
neuromorphic hardware deployment
0.312017
Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic Chips · IJCAI 2017
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.312017
Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic Chips · IJCAI 2017
Computer vision › Image recognition and object detection › object detection
object proposal generation
0.112019
Densely Supervised Grasp Detector (DSGD) · AAAI 2019

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.0dilated convolution · 0.7dense residual connection · 0.7layer-wise feature fusion · 0.4fully convolutional network · 0.4spiking neuron model · 0.3binary crossbar training · 0.3
YearPublicationVenuePosition
2023 DeepActsNet: A deep ensemble framework combining features from face, hands, and body for action recognition
Umar Asif, Deval Mehta 0001, Stefan von Cavallar, Jianbin Tang, Stefan Harrer
Pattern Recognit.4
2020 SSHFD: Single Shot Human Fall Detection with Occluded Joints Resilience
abstract
Falling can have fatal consequences for elderly people especially if the fallen person is unable to call for help due to loss of consciousness or any injury. Automatic fall detection systems can assist through prompt fall alarms and by minimizing the fear of falling when living independently at home. Existing vision-based fall detection systems lack generalization to unseen environments due to challenges such as variations in physical appearances, different camera viewpoints, occlusions, and background clutter. In this paper, we explore ways to overcome the above challenges and present Single Shot Human Fall Detector (SSHFD), a deep learning based framework for automatic fall detection from a single image. This is achieved through two key innovations. First, we present a human pose based fall representation which is invariant to appearance characteristics. Second, we present neural network models for 3d pose estimation and fall recognition which are resilient to missing joints due to occluded body parts. Experiments on public fall datasets show that our framework successfully transfers knowledge of 3d pose estimation and fall recognition learnt purely from synthetic data to unseen real-world data, showcasing its generalization capability for accurate fall detection in real-world scenarios.
Umar Asif, Stefan von Cavallar, Jianbin Tang, Stefan Harrer
ECAI3
2020 Ensemble Knowledge Distillation for Learning Improved and Efficient Networks
abstract
Ensemble models comprising of deep Convolutional Neural Networks (CNN) have shown significant improvements in model generalization but at the cost of large computation and memory requirements. In this paper, we present a framework for learning compact CNN models with improved classification performance and model generalization. For this, we propose a CNN architecture of a compact student model with parallel branches which are trained using ground truth labels and information from high capacity teacher networks in an ensemble learning fashion. Our framework provides two main benefits: i) Distilling knowledge from different teachers into the student network promotes heterogeneity in learning features at different branches of the student network and enables the network to learn diverse solutions to the target problem. ii) Coupling the branches of the student network through ensembling encourages collaboration and improves the quality of the final predictions by reducing variance in the network outputs. Experiments on the well established CIFAR-10 and CIFAR-100 datasets show that our Ensemble Knowledge Distillation (EKD) improves classification accuracy and model generalization especially in situations with limited training data. Experiments also show that our EKD based compact networks outperform in terms of mean accuracy on the test datasets compared to other knowledge distillation based methods.
Umar Asif, Jianbin Tang, Stefan Harrer
ECAI2
2019 Densely Supervised Grasp Detector (DSGD)
abstract
This paper presents Densely Supervised Grasp Detector (DSGD), a deep learning framework which combines CNN structures with layer-wise feature fusion and produces grasps and their confidence scores at different levels of the image hierarchy (i.e., global-, region-, and pixel-levels). Specifically, at the global-level, DSGD uses the entire image information to predict a grasp. At the region-level, DSGD uses a region proposal network to identify salient regions in the image and uses a grasp prediction network to generate segmentations and their corresponding grasp poses of the salient regions. At the pixel-level, DSGD uses a fully convolutional network and predicts a grasp and its confidence at every pixel. During inference, DSGD selects the most confident grasp as the output. This selection from hierarchically generated grasp candidates overcomes limitations of the individual models. DSGD outperforms state-of-the-art methods on the Cornell grasp dataset in terms of grasp accuracy. Evaluation on a multi-object dataset and real-world robotic grasping experiments show that DSGD produces highly stable grasps on a set of unseen objects in new environments. It achieves 97% grasp detection accuracy and 90% robotic grasping success rate with real-time inference speed.
Umar Asif, Jianbin Tang, Stefan Harrer
AAAI2
2019 PubLayNet: Largest Dataset Ever for Document Layout Analysis
abstract
Recognizing the layout of unstructured digital documents is an important step when parsing the documents into structured machine-readable format for downstream applications. Deep neural networks that are developed for computer vision have been proven to be an effective method to analyze layout of document images. However, document layout datasets that are currently publicly available are several magnitudes smaller than established computing vision datasets. Models have to be trained by transfer learning from a base model that is pre-trained on a traditional computer vision dataset. In this paper, we develop the PubLayNet dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central. The size of the dataset is comparable to established computer vision datasets, containing over 360 thousand document images, where typical document layout elements are annotated. The experiments demonstrate that deep neural networks trained on PubLayNet accurately recognize the layout of scientific articles. The pre-trained models are also a more effective base mode for transfer learning on a different document domain. We release the dataset (https://github.com/ibm-aur-nlp/PubLayNet) to support development and evaluation of more advanced models for document layout analysis.
Xu Zhong, Jianbin Tang, Antonio Jimeno-Yepes
ICDAR2
2018 EnsembleNet: Improving Grasp Detection using an Ensemble of Convolutional Neural Networks
Umar Asif, Jianbin Tang, Stefan Harrer
BMVC2
2018 GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices
abstract
Recent research on grasp detection has focused on improving accuracy through deep CNN models, but at the cost of large memory and computational resources. In this paper, we propose an efficient CNN architecture which produces high grasp detection accuracy in real-time while maintaining a compact model design. To achieve this, we introduce a CNN architecture termed GraspNet which has two main branches: i) An encoder branch which downsamples an input image using our novel Dilated Dense Fire (DDF) modules - squeeze and dilated convolutions with dense residual connections. ii) A decoder branch which upsamples the output of the encoder branch to the original image size using deconvolutions and fuse connections. We evaluated GraspNet for grasp detection using offline datasets and a real-world robotic grasping setup. In experiments, we show that GraspNet achieves superior grasp detection accuracy compared to the stateof-the-art computation-efficient CNN models with real-time inference speed on embedded GPU hardware (Nvidia Jetson TX1), making it suitable for low-powered devices.
Umar Asif, Jianbin Tang, Stefan Harrer
IJCAI2
2018 Semantic Labeling Using a Low-Power Neuromorphic Platform
abstract
Deep learning is a powerful technique for the analysis of remote sensing imagery. For applications that require real-time processing on mobile platforms, a low power consumption processing unit is advantageous. The human brain is remarkably powerful at image recognition tasks while operating at very low power consumption levels. Neuromorphic computing designs aim to achieve energy efficiency through the use of spiking neurons and low-precision synapses to perform data processing. We demonstrate here the classification of red, green, blue and depth and hyperspectral data sets using a neuromorphic processing unit (IBM TrueNorth Neurosynaptic System). The convolutional neural-network architecture of the classifier network has been adapted to fit the neuromorphic architecture. The results on overhead imagery and hyperspectral imagery data show that neuromorphic platforms can achieve the state-of-theart performance in semantic labeling with significantly (≈1000×) lower power consumption than traditional GPU-based solutions.
Jianbin Tang, Benjamin S. Mashford, Antonio Jimeno-Yepes
IEEE Geosci. Remote. Sens. Lett.1
2017 Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic Chips
abstract
Deep Neural Networks (DNN) achieve human level performance in many image analytics tasks but DNNs are mostly deployed to GPU platforms that consume a considerable amount of power. New hardware platforms using lower precision arithmetic achieve drastic reductions in power consumption. More recently, brain-inspired spiking neuromorphic chips have achieved even lower power consumption, on the order of milliwatts, while still offering real-time processing. However, for deploying DNNs to energy efficient neuromorphic chips the incompatibility between continuous neurons and synaptic weights of traditional DNNs, discrete spiking neurons and synapses of neuromorphic chips need to be overcome. Previous work has achieved this by training a network to learn continuous probabilities, before it is deployed to a neuromorphic architecture, such as IBM TrueNorth Neurosynaptic System, by random sampling these probabilities. The main contribution of this paper is a new learning algorithm that learns a TrueNorth configuration ready for deployment. We achieve this by training directly a binary hardware crossbar that accommodates the TrueNorth axon configuration constrains and we propose a different neuron model. Results of our approach trained on electroencephalogram (EEG) data show a significant improvement with previous work (76% vs 86% accuracy) while maintaining state of the art performance on the MNIST handwritten data set.
Antonio Jimeno-Yepes, Jianbin Tang, Benjamin S. Mashford
IJCAI2
2016 Weighted Population Code for low power neuromorphic image classification
abstract
Recent digital spiking neuromorphic chips can perform complex computations in real-time with very low power consumption. The input data to such systems needs to first be converted into spikes using a spike encoding scheme. Current examples of such schemes include rate codes and population codes. The selected coding scheme might heavily impact the system's energy consumption, communication bandwidth, processing frame-rate, and computation accuracy. Hence it is important to make an educated decision when selecting the most appropriate spike coding scheme for a given task. To this end, we present a novel spike coding scheme named Weighted Population Code (WPC). WPC is compared to existing coding schemes to transduce images for classification using the TrueNorth chip. Extensive on-chip experimentation with the MNIST and the Flickr-LOGOS32 datasets sheds light on the trade-offs between accuracy, bandwidth, frame rate, network size and energy consumption for image classification, showing the advantages of WPC when high dynamic range and accuracy are needed.
Antonio Jimeno-Yepes, Jianbin Tang, Shreya Saxena, Tobias Brosch, Arnon Amir
IJCNN2