Tanveer Hussain 0001

dblp:205/7268-1 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
23since 2021 · last 2025
0000-0003-4861-8347ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Big Data Analysis for Industrial Activity Recognition Using Attention-Inspired Sequential Temporal Convolution Network
abstract
Deep-learning-based human activity recognition (HAR) methods have significantly transformed a wide range of domains over recent years. However, the adoption of Big Data techniques in industrial applications remains challenging due to issues such as generalized weight optimization, diverse viewpoints, and the complex spatiotemporal features of videos. To address these challenges, this work presents an industrial HAR framework consisting of two main phases. First, a squeeze bottleneck attention block (SBAB) is introduced to enhance the learning capabilities of the backbone model for contextual learning, which allows for the selection and refinement of an optimal feature vector. In the second phase, we propose an effective sequential temporal convolutional network (STCN), which is designed in parallel fashion to mitigate the issues of exploding and vanishing gradients associated with sequence learning. The high-dimensional spatiotemporal feature vectors from the STCN undergo further refinement through our proposed SBAB in a sequential manner, to optimize the features for HAR and enhance the overall performance. The efficacy of the proposed framework is validated through extensive experiments on six datasets, including data from industrial and general activities.
Altaf Hussain 0002, Tanveer Hussain 0001, Waseem Ullah, Samee Ullah Khan, Min Je Kim, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik
IEEE Trans. Big Data2
2024 AllWeather-Net: Unified Image Enhancement for Autonomous Driving Under Adverse Weather and Low-Light Conditions
Chenghao Qian, Mahdi Rezaei 0001, Saeed Anwar, Wenjing Li 0005, Tanveer Hussain 0001, Mohsen Azarmi, Wei Wang 0335
ICPR (30)5
2024 An intelligent correlation learning system for person Re-identification
Samee Ullah Khan, Noman Khan, Tanveer Hussain 0001, Sung Wook Baik
Eng. Appl. Artif. Intell.3
2024 Deep multi-scale pyramidal features network for supervised video summarization
Habib Khan, Tanveer Hussain 0001, Samee Ullah Khan, Zulfiqar Ahmad Khan 0002, Sung Wook Baik
Expert Syst. Appl.2
2024 A modified vision transformer architecture with scratch learning capabilities for effective fire detection
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Tanveer Hussain 0001, Sung Wook Baik
Expert Syst. Appl.3
2024 Deep-ReID: deep features and autoencoder assisted image patching strategy for person re-identification in smart cities surveillance
Samee Ullah Khan, Tanveer Hussain 0001, Amin Ullah, Sung Wook Baik
Multim. Tools Appl.2
2024 A Trapezoid Attention Mechanism for Power Generation and Consumption Forecasting
abstract
Effective operation of smart grids relies on accurate forecasting models for renewable power generation (RPG) and power consumption. The intermittent and unpredictable nature of RPG, coupled with diverse consumption patterns, underscores the importance of robust forecasting approaches. Existing models often employ stacked layers, integrating direct features into fully connected layers, yielding suboptimal results with limited generalization capabilities. Addressing these limitations, we propose a novel two-stream architecture for RPG and power consumption forecasting. The first stream leverages dilated causal convolutional layers to capture intricate patterns, while the second stream focuses on temporal information extraction. Importantly, we fine-tune the hyperparameters of both streams using Bayesian algorithms to optimize the learning process. The outputs from these two streams are then intelligently combined and channeled into our innovative trapezoid attention module (TAM) for feature refinement, resulting in superior pattern representation. The TAM incorporates three distinct dimensions (spatial, temporal, and spatiotemporal) and enriches feature maps by integrating a skip connection from the pre-TAM features. The output post-TAM features are then employed for final forecasting. Our approach showcases remarkable performance in short-term forecasting across a spectrum of datasets, including RPG, regional, residential, and industrial power consumption. By addressing the shortcomings of existing forecasting models, our research contributes to the advancement of smart grid technologies, ensuring more reliable and efficient energy management.
Zulfiqar Ahmad Khan 0002, Tanveer Hussain 0001, Waseem Ullah, Sung Wook Baik
IEEE Trans. Ind. Informatics2
2023 TransCNN: Hybrid CNN and transformer mechanism for surveillance anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Mi Young Lee, Sung Wook Baik
Eng. Appl. Artif. Intell.2
2023 AD-Graph: Weakly Supervised Anomaly Detection Graph Neural Network
abstract
The main challenge faced by video‐based real‐world anomaly detection systems is the accurate learning of unusual events that are irregular, complicated, diverse, and heterogeneous in nature. Several techniques utilizing deep learning have been created to detect anomalies, yet their effectiveness on real‐world data is often limited due to the insufficient incorporation of motion patterns. To address these problems and enhance the traditional functionality of anomaly detection systems for surveillance video data, we propose a weakly supervised graph neural‐network‐assisted video anomaly detection framework called AD‐Graph. To identify temporal information from a series of frames, we extract 3D visual and motion features and represent these in a language‐based knowledge graph format. Next, a robust clustering strategy is applied to group together meaningful neighbourhoods of the graph with similar vertices. Furthermore, spectral filters are applied to these graphs, and spectral graph theory is used to generate graph signals and detect anomalous events. Extensive experimental results over two challenging datasets, UCF‐Crime and ShanghaiTech, show improvements of 0.35% and 0.78% against a state‐of‐the‐art model.
Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Khan Muhammad 0001, Mahmoud Hassaballah, Joel J. P. C. Rodrigues, Sung Wook Baik, Victor Hugo C. de Albuquerque
Int. J. Intell. Syst.2
2023 Vision transformer attention with multi-reservoir echo state network for anomaly recognition
Waseem Ullah, Tanveer Hussain 0001, Sung Wook Baik
Inf. Process. Manag.2
2023 Deep Learning Assists Surveillance Experts: Toward Video Data Prioritization
abstract
Video summarization (VS) suppresses high-dimensional (HD) video data by only extracting the important information. However, prior research has not focused on the need for surveillance VS, that is used for many applications to assist video surveillance experts, including video retrieval and data storage. In addition, mainstream techniques commonly use two-dimensional (2-D) deep models for VS, ignoring event occurrences. Accordingly, we present a two-fold 3-D deep learning-assisted VS framework. First, we employ an inflated 3-D ConvNet model to extract temporal features; these features are optimized using a proposed encoder mechanism. The input video is temporally segmented using a feature comparison technique for selecting a single frame from each video segment. The segmented shots are evaluated using our novel shot segmentation evaluation scheme and are input into a saliency computation mechanism for keyframe selection in a second fold. Qualitative and quantitative analyses over VS benchmarks and surveillance videos demonstrate the superior performance of our framework, with 0.3- and 4.2-unit increases in the F1 scores for YouTube and title-based video summarization datasets, respectively. Along with accurate VS, a key contribution of our study is the novel shot segmentation criterion prior to VS, which can be used as a benchmark in future research to effectively prioritize HD visual data.
Tanveer Hussain 0001, Fath U Min Ullah, Samee Ullah Khan, Amin Ullah, Umair Haroon, Khan Muhammad 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics1
2022 Artificial Intelligence of Things-assisted two-stream neural network for anomaly detection in surveillance Big Video Data
Waseem Ullah, Amin Ullah, Tanveer Hussain 0001, Khan Muhammad 0001, Ali Asghar Heidari, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque
Future Gener. Comput. Syst.3
2022 Intelligent dual stream CNN and echo state network for anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Zulfiqar Ahmad Khan 0002, Umair Haroon, Sung Wook Baik
Knowl. Based Syst.2
2022 Clustering Feature Constraint Multiscale Attention Network for Shadow Extraction From Remote Sensing Images
abstract
Shadow extraction is an important and challenging task in remote sensing image analysis because the presence of shadows not only reduces radiation information but also affects the interpretation of remote sensing images. In this article, a clustering feature constraint multiscale attention network for shadow extraction from remote sensing images is proposed. First, in addition to the pixel-level description of the traditional neural network, our method focuses on the clustering relationships between pixel pairs to obtain the pixel group features of shadows. The feature extraction capability of the network is improved with a reweighting mechanism at the pixel level and pixel group features. Second, we employ a feature fusion algorithm by considering contextual information to improve the network’s attention toward shadow areas and enhance the nonlinear expression ability during the encoding and decoding layers. Furthermore, considering the most prominent multiscale features of shadows in remote sensing images, a deep multiscale feature aggregation structure is established to better fit the multiscale feature expression of shadows. Finally, we construct a shadow extraction dataset to verify the proposed approach. We compare our method with the results of state-of-the-art deep learning models. The results show that the intersection over union (IOU) of our method is improved by 0.85%–9.51% and that the$F1$-score is improved by 0.73–6.48. In addition, the test results for images with different resolutions prove that the proposed approach is more robust than the other methods.
Yakun Xie, Dejun Feng, Yangge Liu, Jun Zhu 0007, Tanveer Hussain 0001, Sung Wook Baik
IEEE Trans. Geosci. Remote. Sens.6
2022 A Multi-Stream Sequence Learning Framework for Human Interaction Recognition
abstract
Human interaction recognition (HIR) is challenging due to multiple humans’ involvement and their mutual interaction in a single frame, generated from their movements. Mainstream literature is based on three-dimensional (3-D) convolutional neural networks (CNNs), processing only visual frames, where human joints data play a vital role in accurate interaction recognition. Therefore, this article proposes a multistream network for HIR that intelligently learns from skeletons’ key points and spatiotemporal visual representations. The first stream localises the joints of the human body using a pose estimation model and transmits them to a 1-D CNN and bidirectional long short-term memory to efficiently extract the features of the dynamic movements of each human skeleton. The second stream feeds the series of visual frames to a 3-D convolutional neural network to extract the discriminative spatiotemporal features. Finally, the outputs of both streams are integrated via fully connected layers that precisely classify the ongoing interactions between humans. To validate the performance of the proposed network, we conducted a comprehensive set of experiments on two benchmark datasets, UT-interaction and TV human interaction, and found 1.15% and 10.0% improvement in the accuracy.
Umair Haroon, Amin Ullah, Tanveer Hussain 0001, Waseem Ullah, Khan Muhammad 0001, Mi Young Lee, Sung Wook Baik
IEEE Trans. Hum. Mach. Syst.3
2022 Optimized Dual Fire Attention Network and Medium-Scale Fire Classification Benchmark
abstract
Vision-based fire detection systems have been significantly improved by deep models; however, higher numbers of false alarms and a slow inference speed still hinder their practical applicability in real-world scenarios. For a balanced trade-off between computational cost and accuracy, we introduce dual fire attention network (DFAN) to achieve effective yet efficient fire detection. The first attention mechanism highlights the most important channels from the features of an existing backbone model, yielding significantly emphasized feature maps. Then, a modified spatial attention mechanism is employed to capture spatial details and enhance the discrimination potential of fire and non-fire objects. We further optimize the DFAN for real-world applications by discarding a significant number of extra parameters using a meta-heuristic approach, which yields around 50% higher FPS values. Finally, we contribute a medium-scale challenging fire classification dataset by considering extremely diverse, highly similar fire/non-fire images and imbalanced classes, among many other complexities. The proposed dataset advances the traditional fire detection datasets by considering multiple classes to answer the following question: what is on fire? We perform experiments on four widely used fire detection datasets, and the DFAN provides the best results compared to 21 state-of-the-art methods. Consequently, our research provides a baseline for fire detection over edge devices with higher accuracy and better FPS values, and the proposed dataset extension provides indoor fire classes and a greater number of outdoor fire classes; these contributions can be used in significant future research. Our codes and dataset will be publicly available at https://github.com/tanveer-hussain/DFAN.
Hikmat Yar, Tanveer Hussain 0001, Mohit Agarwal 0003, Zulfiqar Ahmad Khan 0002, Suneet K. Gupta 0001, Sung Wook Baik
IEEE Trans. Image Process.2
2022 Vision-Based Semantic Segmentation in Scene Understanding for Autonomous Driving: Recent Achievements, Challenges, and Outlooks
abstract
Scene understanding plays a crucial role in autonomous driving by utilizing sensory data for contextual information extraction and decision making. Beyond modeling advances, the enabler for vehicles to become aware of their surroundings is the availability of visual sensory data, which expand the vehicular perception and realizes vehicular contextual awareness in real-world environments. Research directions for scene understanding pursued by related studies include person/vehicle detection and segmentation, their transition analysis, lane change, and turns detection, among many others. Unfortunately, these tasks seem insufficient to completely develop fully-autonomous vehicles i.e., achieving level-5 autonomy, travelling just like human-controlled cars. This latter statement is among the conclusions drawn from this review paper: scene understanding for autonomous driving cars using vision sensors still requires significant improvements. With this motivation, this survey defines, analyzes, and reviews the current achievements of the scene understanding research area that mostly rely on computationally complex deep learning models. Furthermore, it covers the generic scene understanding pipeline, investigates the performance reported by the state-of-the-art, informs about the time complexity analysis of avant garde modeling choices, and highlights major triumphs and noted limitations encountered by current research efforts. The survey also includes a comprehensive discussion on the available datasets, and the challenges that, even if lately confronted by researchers, still remain open to date. Finally, our work outlines future research directions to welcome researchers and practitioners to this exciting domain.
Khan Muhammad 0001, Tanveer Hussain 0001, Hayat Ullah, Javier Del Ser, Mahdi Rezaei 0001, Neeraj Kumar 0001, Mohammad Hijji, Paolo Bellavista, Victor Hugo C. de Albuquerque
IEEE Trans. Intell. Transp. Syst.2
2022 PMAL: A Proxy Model Active Learning Approach for Vision Based Industrial Applications
abstract
Deep Learning models’ performance strongly correlate with availability of annotated data; however, massive data labeling is laborious, expensive, and error-prone when performed by human experts. Active Learning (AL) effectively handles this challenge by selecting the uncertain samples from unlabeled data collection, but the existing AL approaches involve repetitive human feedback for labeling uncertain samples, thus rendering these techniques infeasible to be deployed in industry related real-world applications. In the proposed Proxy Model based Active Learning technique (PMAL) , this issue is addressed by replacing human oracle with a deep learning model, where human expertise is reduced to label only two small subsets of data for training proxy model and initializing the AL loop. In the PMAL technique, firstly, proxy model is trained with a small subset of labeled data, which subsequently acts as an oracle for annotating uncertain samples. Secondly, active model's training, uncertain samples extraction via uncertainty sampling, and annotation through proxy model is carried out until predefined iterations to achieve higher accuracy and labeled data. Finally, the active model is evaluated using testing data to verify the effectiveness of our technique for practical applications. The correct annotations by the proxy model are ensured by employing the potentials of explainable artificial intelligence. Similarly, emerging vision transformer is used as an active model to achieve maximum accuracy. Experimental results reveal that the proposed method outperforms the state-of-the-art in terms of minimum labeled data usage and improves the accuracy with 2.2%, 2.6%, and 1.35% on Caltech-101, Caltech-256, and CIFAR-10 datasets, respectively. Since the proposed technique offers a highly reasonable solution to exploit huge multimedia data, it can be widely used in different evolutionary industrial domains.
Ijaz Ul Haq, Tanveer Hussain 0001, Khan Muhammad 0001, Mohammad Hijji, Victor Hugo C. de Albuquerque, Sung Wook Baik
ACM Trans. Multim. Comput. Commun. Appl.3
2021 DeepSmoke: Deep learning model for smoke detection and segmentation in outdoor environments
Salman Khan 0004, Khan Muhammad 0001, Tanveer Hussain 0001, Javier Del Ser, Fabio Cuzzolin, Siddhartha Bhattacharyya 0001, Zahid Akhtar, Victor Hugo C. de Albuquerque
Expert Syst. Appl.3
2021 Conflux LSTMs Network: A Novel Approach for Multi-View Action Recognition
abstract
Multi-view action recognition (MVAR) is an optimal technique to acquire numerous clues from different views data for effective action recognition, however, it is not well explored yet. There exist several challenges to MVAR domain such as divergence in viewpoints, invisible regions, and different scales of appearance in each view require better solutions for real world applications. In this paper, we present a conflux long short-term memory (LSTMs) network to recognize actions from multi-view cameras. The proposed framework has four major steps; 1) frame level feature extraction, 2) its propagation through conflux LSTMs network for view self-reliant patterns learning, 3) view inter-reliant patterns learning and correlation computation, and 4) action classification. First, we extract deep features from a sequence of frames using a pre-trained VGG19 CNN model for each view. Second, we forward the extracted features to conflux LSTMs network to learn the view self-reliant patterns. In the next step, we compute the inter-view correlations using the pairwise dot product from output of the LSTMs network corresponding to different views to learn the view inter-reliant patterns. In the final step, we use flatten layers followed by SoftMax classifier for action recognition. Experimental results over benchmark datasets compared to state-of-the-art report an increase of 3% and 2% on northwestern-UCLA and MCAD datasets, respectively.
Amin Ullah, Khan Muhammad 0001, Tanveer Hussain 0001, Sung Wook Baik
Neurocomputing3
2021 An Efficient Deep Learning Framework for Intelligent Energy Management in IoT Networks
abstract
Green energy management is an economical solution for better energy usage, but the employed literature lacks focusing on the potentials of edge intelligence in controllable Internet of Things (IoT). Therefore, in this article, we focus on the requirements of todays' smart grids, homes, and industries to propose a deep-learning-based framework for intelligent energy management. We predict future energy consumption for short intervals of time as well as provide an efficient way of communication between energy distributors and consumers. The key contributions include edge devices-based real-time energy management via common cloud-based data supervising server, optimal normalization technique selection, and a novel sequence learning-based energy forecasting mechanism with reduced time complexity and lowest error rates. In the proposed framework, edge devices relate to a common cloud server in an IoT network that communicates with the associated smart grids to effectively continue the energy demand and response phenomenon. We apply several preprocessing techniques to deal with the diverse nature of electricity data, followed by an efficient decision-making algorithm for short-term forecasting and implement it over resource-constrained devices. We perform extensive experiments and witness 0.15 and 3.77 units reduced mean-square error (MSE) and root MSE (RMSE) for residential and commercial datasets, respectively.
Tao Han 0004, Khan Muhammad 0001, Tanveer Hussain 0001, Jaime Lloret Mauri, Sung Wook Baik
IEEE Internet Things J.3
2021 Multiview Summarization and Activity Recognition Meet Edge Computing in IoT Environments
abstract
Multiview video summarization (MVS) has not received much attention from the research community due to inter-view correlations and views' overlapping, etc. The majority of previous MVS works are offline, relying on only summary, and require additional communication bandwidth and transmission time, with no focus on foggy environments. We propose an edge intelligence-based MVS and activity recognition framework that combines artificial intelligence with Internet of Things (IoT) devices. In our framework, resource-constrained devices with cameras use a lightweight CNN-based object detection model to segment multiview videos into shots, followed by mutual information computation that helps in a summary generation. Our system does not rely solely on a summary, but encodes and transmits it to a master device using a neural computing stick for inter-view correlations computation and efficient activity recognition, an approach which saves computation resources, communication bandwidth, and transmission time. Experiments show an increase of 0.4 unit in F-measure on an MVS Office dateset and 0.2% and 2% improved accuracy for UCF-50 and YouTube 11 datesets, respectively, with lower storage and transmission times. The processing time is reduced from 1.23 to 0.45 s for a single frame and optimally 0.75 seconds faster MVS. A new dateset is constructed by synthetically adding fog to an MVS dateset to show the adaptability of our system for both certain and uncertain IoT surveillance environments.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Javier Del Ser, Amir Hossein Gandomi, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.1
2021 A comprehensive survey of multi-view video summarization
Tanveer Hussain 0001, Khan Muhammad 0001, Weiping Ding 0001, Jaime Lloret Mauri, Sung Wook Baik, Victor Hugo C. de Albuquerque
Pattern Recognit.1
2020 Cost-Effective Video Summarization Using Deep CNN With Hierarchical Weighted Fusion for IoT Surveillance Networks
abstract
Video summarization (VS) has attracted intense attention recently due to its enormous applications in various computer vision domains, such as video retrieval, indexing, and browsing. Traditional VS researches mostly target at the effectiveness of the VS algorithms by introducing the high quality of features and clusters for selecting representative visual elements. Due to the increased density of vision sensors network, there is a tradeoff between the processing time of the VS methods with reasonable and representative quality of the generated summaries. It is a challenging task to generate a video summary of significant importance while fulfilling the needs of Internet of Things (IoT) surveillance networks with constrained resources. This article addresses this problem by proposing a new computationally effective solution through designing a deep CNN framework with hierarchical weighted fusion for the summarization of surveillance videos captured in IoT settings. The first stage of our framework designs discriminative rich features extracted from deep CNNs for shot segmentation. Then, we employ image memorability predicted from a fine-tuned CNN model in the framework, along with aesthetic and entropy features to maintain the interestingness and diversity of the summary. Third, a hierarchical weighted fusion mechanism is proposed to produce an aggregated score for the effective computation of the extracted features. Finally, an attention curve is constituted using the aggregated score for deciding outstanding keyframes for the final video summary. Experiments are conducted using benchmark data sets for validating the importance and effectiveness of our framework, which outperforms the other state-of-the-art schemes.
Khan Muhammad 0001, Tanveer Hussain 0001, Muhammad Tanveer 0001, Giovanna Sannino, Victor Hugo C. de Albuquerque
IEEE Internet Things J.2
2020 Efficient CNN based summarization of surveillance videos for resource-constrained devices
Khan Muhammad 0001, Tanveer Hussain 0001, Sung Wook Baik
Pattern Recognit. Lett.2
2020 Intelligent Embedded Vision for Summarization of Multiview Videos in IIoT
abstract
Nowadays, video sensors are used on a large scale for various applications, including security monitoring and smart transportation. However, the limited communication bandwidth and storage constraints make it challenging to process such heterogeneous nature of Big Data in real time. Multiview video summarization (MVS) enables us to suppress redundant data in distributed video sensors settings. The existing MVS approaches process video data in offline manner by transmitting them to the local or cloud server for analysis, which requires extra streaming to conduct summarization, huge bandwidth, and are not applicable for integration with industrial Internet of Things (IIoT). This article presents a light-weight convolutional neural network (CNN) and IIoT-based computationally intelligent (CI) MVS framework. Our method uses an IIoT network containing smart devices, Raspberry Pi (RPi) (clients and master) with embedded cameras to capture multiview video data. Each client RPi detects target in frames via light-weight CNN model, analyzes these targets for traffic and crowd density, and searches for suspicious objects to generate alert in the IIoT network. The frames of each client RPi are encoded and transmitted with approximately 17.02% smaller size of each frame to master RPi for final MVS. Empirical analysis shows that our proposed framework can be used in industrial environments for various applications such as security and smart transportation and can be proved beneficial for saving resources.11[Online]. Available: https://github.com/tanveer-hussain/Embedded-Vision-for-MVS.
Tanveer Hussain 0001, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics1
2020 Cloud-Assisted Multiview Video Summarization Using CNN and Bidirectional LSTM
abstract
The massive amount of video data produced by surveillance networks in industries instigate various challenges in exploring these videos for many applications, such as video summarization (VS), analysis, indexing, and retrieval. The task of multiview video summarization (MVS) is very challenging due to the gigantic size of data, redundancy, overlapping in views, light variations, and interview correlations. To address these challenges, various low-level features and clustering-based soft computing techniques are proposed that cannot fully exploit MVS. In this article, we achieve MVS by integrating deep neural network based soft computing techniques in a two-tier framework. The first online tier performs target-appearance-based shots segmentation and stores them in a lookup table that is transmitted to cloud for further processing. The second tier extracts deep features from each frame of a sequence in the lookup table and pass them to deep bidirectional long short-term memory (DB-LSTM) to acquire probabilities of informativeness and generates a summary. Experimental evaluation on benchmark dataset and industrial surveillance data from YouTube confirms the better performance of our system compared to the state-of-the-art MVS methods.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Zehong Cao, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics1
2020 DeepReS: A Deep Learning-Based Video Summarization Strategy for Resource-Constrained Industrial Surveillance Scenarios
abstract
The exponential growth in the production of video contents in different industries causes an urgent need for effective video summarization (VS) techniques, in order to get an optimal storage and preservation of key information in the video. Compared to other domains, industrial videos are more challenging to process, as they usually contain diverse and complex events, which make their online processing a difficult task. In this article, we introduce an online system for intelligent video capturing, coarse and fine redundancy removal, and summary generation. First, we capture video data through resource-constrained devices in an industrial Internet of Things network, equipped with vision sensors and apply coarse redundancy removal through the comparison of low-level features. Second, we transmit the resulting frames to the cloud for detailed analysis, where sequential features are extracted for the selection of candidate keyframes. Finally, we refine the candidate keyframes in order to discriminate those with maximum information as part of the summary. The key contributions of this article include the coarse and fine refining of video data implemented over resource-restricted devices and the presentation of important data in the form of a summary. Experiments11[Online]. Available: https://github.com/tanveer-hussain/DeepRes-Video-Summarization. over publicly available datasets evince a 0.3-unit increase in the F1 score when compared to state-of-the-art and with reduced time complexity. Furthermore, we provide convincing results on our newly created dataset in an industrial environment, which is made publicly available for the research community along with its labeled ground truth.
Khan Muhammad 0001, Tanveer Hussain 0001, Javier Del Ser, Vasile Palade, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics2
2019 CNN-based anti-spoofing two-tier multi-factor authentication system
Salman Khan 0004, Tanveer Hussain 0001, Khan Muhammad 0001, Arun Kumar Sangaiah, Aniello Castiglione, Christian Esposito 0001, Sung Wook Baik
Pattern Recognit. Lett.3