Sung Wook Baik

dblp:b/SungWookBaik · also Sung Baik · DBLP profile ↗
← Back
111ranked-venue papers
9as first author
47since 2021 · last 2026
0000-0002-6678-7788ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 7 since 2021Computer networks · 8 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Differential attention transformer-enhanced graph neural network for accurate material property prediction
Altaf Hussain 0002, Muhammad Munsif, Kee-Sun Sohn, Sung Wook Baik
Eng. Appl. Artif. Intell.5
2026 Pseudo-labeling driven refinement of benchmark object detection datasets via analysis of learning patterns
abstract
Benchmark Object Detection (OD) datasets are essential for advancing computer vision in areas such as autonomous driving, robotics, and surveillance. MS-COCO has become the standard benchmark due to its object diversity and scene complexity. Despite the considerable time since its release, MS-COCO remains widely used, as many OD algorithms have been evaluated on it. However, it suffers from issues including missing or incorrect labels, inaccurate bounding boxes, duplicates, and inconsistent group labeling. These data quality problems are recognized as critical bottlenecks, motivating a data-centric shift that prioritizes improving datasets over simply tuning models. To address these challenges, we propose a comprehensive data refinement framework and present MJ-COCO, a newly re-annotated version of MS-COCO. Our approach combines loss and gradient-based error detection with a four-stage pseudo-labeling process: (1) bounding box generation via invertible transformations, (2) IoU-based duplicate removal and confidence merging, (3) class consistency verification through an expert recognizer, and (4) spatial adjustment using region activation maps. This pipeline enables scalable, accurate correction of annotation errors without manual re-labeling. Extensive experiments with one-stage (RetinaNet, YOLOv3, YOLOX) and two-stage (Faster R-CNN, Libra R-CNN) models across MS-COCO, Sama COCO, Objects365, and PASCAL VOC show that MJ-COCO consistently outperforms MS-COCO, particularly on high-quality validation sets, with notable improvements in Average Precision (AP) and AP S . Annotation coverage also improved significantly, with over 200,000 more small object annotations. These results confirm MJ-COCO as a more accurate, robust, and scalable alternative for modern OD tasks. Our annotations are publicly available at https://www.kaggle.com/datasets/mjcoco2025/mj-coco-2025 .
Min Je Kim, Muhammad Munsif, Altaf Hussain 0002, Hikmat Yar, Sung Wook Baik
Neurocomputing5
2026 Class-incremental learning network for real-time anomaly recognition in surveillance environments
Adnan Hussain, Waseem Ullah, Noman Khan, Zulfiqar Ahmad Khan 0002, Hikmat Yar, Sung Wook Baik
Pattern Recognit.6
2026 Bootstrapping for fuzzy mediation and moderated-mediation analysis using Fuzzy Lease Squares Estimation (FLSE) and Fuzzy Least Absolute Deviations (FLAD) with evolutionary algorithms
Da Jeong Kang, Hyeon Gu Kang, Zong Woo Geem, Geon Hee Lee, Sung Wook Baik, Jin Hee Yoon
Soft Comput.5
2026 Deep hybrid network with additive attention for accurate population forecasting
Hikmat Yar, Adnan Hussain, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik
Soft Comput.5
2025 Hierarchical attention-based framework for enhanced prediction and optimization of organic and inorganic material synthesis
Muhammad Munsif, Altaf Hussain 0002, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik
Adv. Eng. Informatics5
2025 Edge-assisted framework for instant anomaly detection and cloud-based anomaly recognition in smart surveillance
Adnan Hussain, Noman Khan, Zulfiqar Ahmad Khan 0002, Hikmat Yar, Sung Wook Baik
Eng. Appl. Artif. Intell.5
2025 Action understanding in low-light and pitch-dark conditions: A comprehensive survey
Muhammad Munsif, Samee Ullah Khan, Noman Khan, Altaf Hussain 0002, Sung Wook Baik
Eng. Appl. Artif. Intell.5
2025 EFNet-CSM: EfficientNet with a modified attention mechanism for effective fire detection
Hikmat Yar, Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik
Knowl. Based Syst.5
2025 Dual stream deep attention networks for annual population projection
Adnan Hussain, Hikmat Yar, Noman Khan, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik
Pattern Anal. Appl.6
2025 Optimized cross-module attention network and medium-scale dataset for effective fire detection
Zulfiqar Ahmad Khan 0002, Fath U Min Ullah, Hikmat Yar, Waseem Ullah, Noman Khan, Min Je Kim, Sung Wook Baik
Pattern Recognit.7
2025 Big Data Analysis for Industrial Activity Recognition Using Attention-Inspired Sequential Temporal Convolution Network
abstract
Deep-learning-based human activity recognition (HAR) methods have significantly transformed a wide range of domains over recent years. However, the adoption of Big Data techniques in industrial applications remains challenging due to issues such as generalized weight optimization, diverse viewpoints, and the complex spatiotemporal features of videos. To address these challenges, this work presents an industrial HAR framework consisting of two main phases. First, a squeeze bottleneck attention block (SBAB) is introduced to enhance the learning capabilities of the backbone model for contextual learning, which allows for the selection and refinement of an optimal feature vector. In the second phase, we propose an effective sequential temporal convolutional network (STCN), which is designed in parallel fashion to mitigate the issues of exploding and vanishing gradients associated with sequence learning. The high-dimensional spatiotemporal feature vectors from the STCN undergo further refinement through our proposed SBAB in a sequential manner, to optimize the features for HAR and enhance the overall performance. The efficacy of the proposed framework is validated through extensive experiments on six datasets, including data from industrial and general activities.
Altaf Hussain 0002, Tanveer Hussain 0001, Waseem Ullah, Samee Ullah Khan, Min Je Kim, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik
IEEE Trans. Big Data8
2024 Dual Deep Learning Network for Abnormal Action Detection
abstract
Neural networks have demonstrated remarkable effectiveness in solving distinct real-world vision problems pertaining to activity recognition and violence detection in surveillance scenarios. The broad reliance on practicing a single network for spatial and motion information collection has made them less effective for long-term dependency analysis in video snippets. Our work solves this issue through a multi-network fusion strategy suitable for real-world surveillance. Initially, the spatial information is accessed from a compound coefficient strategy inspired by a robust convolutional neural network (ConvNet). Next, the pyramidal convolutional features from two consecutive frames are obtained through LiteFlowNet. The output from both the networks (ConvNet and LiteFlowNet) is separately passed into a deep-gated recurrent Unit (GRU) that is assembled for a skip connection. The latter obtained from each GRU is fused and further propagated to the dense layer for final decision. The results on the datasets and the ablation study confirm our method’s efficiency, outperforming the state-of-the-art methods. (Code: GitHub)
Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Sung Wook Baik, Estefanía Talavera, Saeed Anwar, Khan Muhammad 0001
AVSS3
2024 AI-driven behavior biometrics framework for robust human activity recognition in surveillance systems
Altaf Hussain 0002, Samee Ullah Khan, Noman Khan, Mohammad Shabaz, Sung Wook Baik
Eng. Appl. Artif. Intell.5
2024 An intelligent correlation learning system for person Re-identification
Samee Ullah Khan, Noman Khan, Tanveer Hussain 0001, Sung Wook Baik
Eng. Appl. Artif. Intell.4
2024 TDS-Net: Transformer enhanced dual-stream network for video Anomaly Detection
Adnan Hussain, Waseem Ullah, Noman Khan, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik
Expert Syst. Appl.6
2024 Deep multi-scale pyramidal features network for supervised video summarization
Habib Khan, Tanveer Hussain 0001, Samee Ullah Khan, Zulfiqar Ahmad Khan 0002, Sung Wook Baik
Expert Syst. Appl.5
2024 A modified vision transformer architecture with scratch learning capabilities for effective fire detection
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Tanveer Hussain 0001, Sung Wook Baik
Expert Syst. Appl.4
2024 An efficient deep learning architecture for effective fire detection in smart surveillance
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Imad Rida, Waseem Ullah, Min Je Kim, Sung Wook Baik
Image Vis. Comput.6
2024 Contextual visual and motion salient fusion framework for action recognition in dark environments
Muhammad Munsif, Samee Ullah Khan, Noman Khan, Altaf Hussain 0002, Min Je Kim, Sung Wook Baik
Knowl. Based Syst.6
2024 Deep-ReID: deep features and autoencoder assisted image patching strategy for person re-identification in smart cities surveillance
Samee Ullah Khan, Tanveer Hussain 0001, Amin Ullah, Sung Wook Baik
Multim. Tools Appl.4
2024 A Trapezoid Attention Mechanism for Power Generation and Consumption Forecasting
abstract
Effective operation of smart grids relies on accurate forecasting models for renewable power generation (RPG) and power consumption. The intermittent and unpredictable nature of RPG, coupled with diverse consumption patterns, underscores the importance of robust forecasting approaches. Existing models often employ stacked layers, integrating direct features into fully connected layers, yielding suboptimal results with limited generalization capabilities. Addressing these limitations, we propose a novel two-stream architecture for RPG and power consumption forecasting. The first stream leverages dilated causal convolutional layers to capture intricate patterns, while the second stream focuses on temporal information extraction. Importantly, we fine-tune the hyperparameters of both streams using Bayesian algorithms to optimize the learning process. The outputs from these two streams are then intelligently combined and channeled into our innovative trapezoid attention module (TAM) for feature refinement, resulting in superior pattern representation. The TAM incorporates three distinct dimensions (spatial, temporal, and spatiotemporal) and enriches feature maps by integrating a skip connection from the pre-TAM features. The output post-TAM features are then employed for final forecasting. Our approach showcases remarkable performance in short-term forecasting across a spectrum of datasets, including RPG, regional, residential, and industrial power consumption. By addressing the shortcomings of existing forecasting models, our research contributes to the advancement of smart grid technologies, ensuring more reliable and efficient energy management.
Zulfiqar Ahmad Khan 0002, Tanveer Hussain 0001, Waseem Ullah, Sung Wook Baik
IEEE Trans. Ind. Informatics4
2024 Darkness-Adaptive Action Recognition: Leveraging Efficient Tubelet Slow-Fast Network for Industrial Applications
abstract
Infrared (IR) technology has emerged as a solution for monitoring dark environments. It offers resilience to shifting illumination, appearance changes, and shadows, with applications spanning self-driving cars, robotics, nighttime security, and many other fields. While existing state-of-the-art RGB-based human action recognition (AR) models exhibit limitations in scalability for action understanding under uncertain, low-light, or dark conditions. Integrating these with IR data faces challenges due to changes in modality, high resource demands, and strict latency requirements. Such issues hinder the deployment of these technologies in real-world settings. To overcome these challenges, we introduce a novel slow-fast tubelet (SFT) processing framework designed for efficient and accurate AR in IR-based scenarios. The SFT framework comprises three modules: tubelet preprocessing (TPP), feature extraction, and the feature lateral connection and recognition module (FELCM). The TPP module refines IR streams by extracting the region of interest, filtering detected objects, removing noise, and generating tubelets of refined frames. The FELCM processes refined tubelet through two pathways, where the fast tubelet path operates at a high rate and the slow tubelet path operates at a slow rate. These pathways interconnect through lateral connections, facilitating mutual updates, and enhancing the prediction efficiency. We conducted extensive experiments on benchmark datasets, including NTURGB-D 120 and infrared action recognition (InfAR). The results demonstrate that our proposed SFT framework surpasses state-of-the-art approaches in terms of accuracy (2.7% and 3.3% improvement, respectively), computational cost, and inference latency while maintaining the competitive recognition performance. Our framework's promising results underscore its potential for direct deployment in real-world applications.
Muhammad Munsif, Noman Khan, Altaf Hussain 0002, Min Je Kim, Sung Wook Baik
IEEE Trans. Ind. Informatics5
2023 TransCNN: Hybrid CNN and transformer mechanism for surveillance anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Mi Young Lee, Sung Wook Baik
Eng. Appl. Artif. Intell.5
2023 Sequential attention mechanism for weakly supervised video anomaly detection
Waseem Ullah, Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Sung Wook Baik
Expert Syst. Appl.4
2023 A modified YOLOv5 architecture for efficient fire detection in smart cities
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Fath U Min Ullah, Waseem Ullah, Sung Wook Baik
Expert Syst. Appl.5
2023 AD-Graph: Weakly Supervised Anomaly Detection Graph Neural Network
abstract
The main challenge faced by video‐based real‐world anomaly detection systems is the accurate learning of unusual events that are irregular, complicated, diverse, and heterogeneous in nature. Several techniques utilizing deep learning have been created to detect anomalies, yet their effectiveness on real‐world data is often limited due to the insufficient incorporation of motion patterns. To address these problems and enhance the traditional functionality of anomaly detection systems for surveillance video data, we propose a weakly supervised graph neural‐network‐assisted video anomaly detection framework called AD‐Graph. To identify temporal information from a series of frames, we extract 3D visual and motion features and represent these in a language‐based knowledge graph format. Next, a robust clustering strategy is applied to group together meaningful neighbourhoods of the graph with similar vertices. Furthermore, spectral filters are applied to these graphs, and spectral graph theory is used to generate graph signals and detect anomalous events. Extensive experimental results over two challenging datasets, UCF‐Crime and ShanghaiTech, show improvements of 0.35% and 0.78% against a state‐of‐the‐art model.
Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Khan Muhammad 0001, Mahmoud Hassaballah, Joel J. P. C. Rodrigues, Sung Wook Baik, Victor Hugo C. de Albuquerque
Int. J. Intell. Syst.7
2023 Efficient Person Reidentification for IoT-Assisted Cyber-Physical Systems
abstract
The main objective of this study is to propose a cyber–physical system (CPS)-based person reidentification (P-ReID) framework for smart surveillance. The Internet of Things (IoT)-based interconnected vision sensors in smart cities are considered essential elements of a CPS, and contribute significantly to urban security. However, the reidentification of targeted persons using emerging edge AI techniques still faces certain challenges. To improve efficiency at the edge and overcome the traditional sensing of video cameras, we employed an AI-based P-ReID framework for CPS that is functional in IoT environments. In addition, we present dual attention dilated network (DADNet), which integrates an energy-efficient convolutional neural network (CNN) with a self-attention module to substantially improve the person matching probability. Furthermore, we applied dual feature fusion to intelligently integrate discriminative and robust features using early and late fusion strategies that allow DADNet to significantly consider the foreground and marginally utilize the background information. Furthermore, we impose diversity orthogonality regularization over several CNN layers, which boosts the performance of DADNet, resulting in an appropriate usage over IoT networks. A comprehensive set of ablation studies, comparison with other state-of-the-art approaches, and a time complexity analysis confirm the strength of our DADNet for reidentification tasks in AI-enabled IoT settings that are well suited for a CPS.
Samee Ullah Khan, Ijaz Ul Haq, Noman Khan, Amin Ullah, Khan Muhammad 0001, Huiling Chen 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.7
2023 AI-Assisted Hybrid Approach for Energy Management in IoT-Based Smart Microgrid
abstract
Power generation (PG) prediction from renewable energy sources (RESs) plays a vital role in effective energy management in smart cities. However, harnessing the potential of edge intelligence in well-controlled Internet of Things (IoT) networks poses significant challenges. To address this, we propose an IoT-based framework for intelligent and efficient PG prediction in smart microgrids. The framework begins by acquiring data from various RESs, including wind and solar. Before the training process, the data undergoes cleaning and normalization steps that use denoising and cleansing filters. For forecasting renewable energy (RE), we introduce a hybrid model that integrates a multi-head attention (MHA)-based deep autoencoder (AE) with extreme gradient boosting (XGB) algorithm. The AE’s encoder component extracts discriminative features from the cleaned data sequence, which are then learned by XGB to provide a final PG forecast. This edge computing layer facilitates information sharing through fog computing, which ensures power balancing between suppliers and consumers. Furthermore, the framework also incorporates various power consumption (PC) sectors and entities within smart cities, such as transportation and healthcare, to ensure efficient management. We evaluate the proposed hybrid model using publicly accessible benchmarks and locally gathered data sets, demonstrating state-of-the-art performance in terms of error metrics. The computational complexity of the proposed model is also suitable for resource-constrained IoT devices connected to a shared IoT-Fog setup, enabling seamless communication with smart microgrids for effective power management.
Noman Khan, Samee Ullah Khan, Fath U Min Ullah, Mi Young Lee, Sung Wook Baik
IEEE Internet Things J.5
2023 Vision transformer attention with multi-reservoir echo state network for anomaly recognition
Waseem Ullah, Tanveer Hussain 0001, Sung Wook Baik
Inf. Process. Manag.3
2023 Deep Learning Assists Surveillance Experts: Toward Video Data Prioritization
abstract
Video summarization (VS) suppresses high-dimensional (HD) video data by only extracting the important information. However, prior research has not focused on the need for surveillance VS, that is used for many applications to assist video surveillance experts, including video retrieval and data storage. In addition, mainstream techniques commonly use two-dimensional (2-D) deep models for VS, ignoring event occurrences. Accordingly, we present a two-fold 3-D deep learning-assisted VS framework. First, we employ an inflated 3-D ConvNet model to extract temporal features; these features are optimized using a proposed encoder mechanism. The input video is temporally segmented using a feature comparison technique for selecting a single frame from each video segment. The segmented shots are evaluated using our novel shot segmentation evaluation scheme and are input into a saliency computation mechanism for keyframe selection in a second fold. Qualitative and quantitative analyses over VS benchmarks and surveillance videos demonstrate the superior performance of our framework, with 0.3- and 4.2-unit increases in the F1 scores for YouTube and title-based video summarization datasets, respectively. Along with accurate VS, a key contribution of our study is the novel shot segmentation criterion prior to VS, which can be used as a benchmark in future research to effectively prioritize HD visual data.
Tanveer Hussain 0001, Fath U Min Ullah, Samee Ullah Khan, Amin Ullah, Umair Haroon, Khan Muhammad 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics7
2022 Artificial Intelligence of Things-assisted two-stream neural network for anomaly detection in surveillance Big Video Data
Waseem Ullah, Amin Ullah, Tanveer Hussain 0001, Khan Muhammad 0001, Ali Asghar Heidari, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque
Future Gener. Comput. Syst.7
2022 Learning to rank: An intelligent system for person reidentification
abstract
Person reidentification (P-Reid) is an emerging research domain in the field of information retrieval that has gained exponential growth due to its wide range of applications in pedestrian tracking and crime prevention. The primary goal of P-Reid is to recognize a person based on previous appearance in multiview surveillance videos. The mainstream approaches apply fully supervised learning techniques that have poor scalability when deployed in complex real-world scenes, due to the overfitting problem, caused by the lack of sufficient annotated data. Further, optimization of these models for unlabeled data in real-time surveillance is a challenging task. To tackle these issues, an intelligent framework (LR-Net) is proposed, consisting of three tiers including fine-tuning (FT), siamese network (SN), and fusion strategy (FS). In the first tier, a deep learning model is fine-tuned for P-Reid that can handle both labeled and unlabeled data. Next, with the assistance of transfer learning, an SN is proposed that has a strong discriminative capability in terms of similarity between a pair of images. Finally, a learning-to-rank strategy is applied to optimize the learning capability of the SN, in which a triplet network extracts spatial-temporal patterns from unlabeled samples. In addition, a bayesian fusion model (BFM) is introduced to integrate the spatiotemporal and visual features, which yields 4.4%, 9.3%, and 0.8% improvement in the matching score over Market-1501, DukeMCMT-reID, and CUHK03 data sets, respectively. The conducted experiments and ablation study on the benchmark data sets empirically validate the proposed system, which obtains a high Rank-1 score as compared with the state-of-the-art (SOTA) methods.
Samee Ullah Khan, Ijaz Ul Haq, Noman Khan, Khan Muhammad 0001, Mohammad Hijji, Sung Wook Baik
Int. J. Intell. Syst.6
2022 An intelligent system for complex violence pattern analysis and detection
abstract
Video surveillance has shown encouraging outcomes to monitor human activities and prevent crimes in real time. To this extent, violence detection (VD) has received substantial attention from the research community due to its vast applications, such as ensuring security over public areas and industrial settings through smart machine intelligence. However, because of changing illumination, complex background and low resolution, the analysis of violence patterns remains challenging in the industrial video surveillance domain. In this paper, we propose a computationally intelligent VD approach to precisely detect violent scenes through deep analysis of surveillance video sequential patterns. First, the video stream acquired through the vision sensor is processed by a lightweight convolutional neural network (CNN) for the segmentation of important shots. Next, temporal optical flow features are extracted from the informative shots via a residential optical flow CNN. These are concatenated with appearance-invariant features extracted from a Darknet CNN model. Finally, a multilayer long short-term memory network is plugged to generate the final feature map for learning the violence patterns in a sequence of frames. In addition, we contribute to the existing surveillance VD data set by considering its indoor and outdoor scenarios separately for the proposed method's evaluation, achieving a 2% increase in accuracy over surveillance fight data set. Experiments also show encouraging results over the state of the art on other challenging benchmark data sets.
Fath U Min Ullah, Mohammad S. Obaidat, Khan Muhammad 0001, Amin Ullah, Sung Wook Baik, Fabio Cuzzolin, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque
Int. J. Intell. Syst.5
2022 Enabling automation and edge intelligence over resource constraint IoT devices for smart home
Mansoor Nasir, Khan Muhammad 0001, Amin Ullah, Jamil Ahmad 0003, Sung Wook Baik
Neurocomputing5
2022 Intelligent dual stream CNN and echo state network for anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Zulfiqar Ahmad Khan 0002, Umair Haroon, Sung Wook Baik
Knowl. Based Syst.5
2022 Damaged Building Detection From Post-Earthquake Remote Sensing Imagery Considering Heterogeneity Characteristics
abstract
Damaged building detection from remote sensing imagery helps to quickly and rapidly assess losses after an earthquake. In recent years, deep learning technology has become a favorable tool for remote sensing image information detection. Based on the characteristics of damaged buildings in remote sensing images, in this paper, a framework for damaged building detection that considers heterogeneity characteristics is proposed. First, a local-global context attention module is proposed to improve the feature detection ability of the network, which can extract the features of damaged buildings from different directions and effectively aggregate global and local features. In addition, the module takes the correlation between feature maps at different scales into account while extracting information. Second, a feature fusion module with self-attention is established to replace the simple connection between the encoding and decoding processes, which improves the detail feature recovery ability of the network during the upsampling process. Finally, to fully aggregate semantic and detail features at different scales, a multibranch auxiliary classifier is established by adding two separate branches in the prediction stage. The effectiveness of the proposed approach is verified based on data from the 2010 Haiti earthquake, and comparisons with 3 object-oriented methods and 16 existing excellent deep learning models are performed. The IOU increase of 0.03%-7.39% is achieved using the proposed approach compared with excellent deep learning models.
Yakun Xie, Dejun Feng, Hongyu Chen 0004, Wenfei Mao, Jun Zhu 0007, Ya Hu, Sung Wook Baik
IEEE Trans. Geosci. Remote. Sens.8
2022 Clustering Feature Constraint Multiscale Attention Network for Shadow Extraction From Remote Sensing Images
abstract
Shadow extraction is an important and challenging task in remote sensing image analysis because the presence of shadows not only reduces radiation information but also affects the interpretation of remote sensing images. In this article, a clustering feature constraint multiscale attention network for shadow extraction from remote sensing images is proposed. First, in addition to the pixel-level description of the traditional neural network, our method focuses on the clustering relationships between pixel pairs to obtain the pixel group features of shadows. The feature extraction capability of the network is improved with a reweighting mechanism at the pixel level and pixel group features. Second, we employ a feature fusion algorithm by considering contextual information to improve the network’s attention toward shadow areas and enhance the nonlinear expression ability during the encoding and decoding layers. Furthermore, considering the most prominent multiscale features of shadows in remote sensing images, a deep multiscale feature aggregation structure is established to better fit the multiscale feature expression of shadows. Finally, we construct a shadow extraction dataset to verify the proposed approach. We compare our method with the results of state-of-the-art deep learning models. The results show that the intersection over union (IOU) of our method is improved by 0.85%–9.51% and that the$F1$-score is improved by 0.73–6.48. In addition, the test results for images with different resolutions prove that the proposed approach is more robust than the other methods.
Yakun Xie, Dejun Feng, Yangge Liu, Jun Zhu 0007, Tanveer Hussain 0001, Sung Wook Baik
IEEE Trans. Geosci. Remote. Sens.7
2022 A Multi-Stream Sequence Learning Framework for Human Interaction Recognition
abstract
Human interaction recognition (HIR) is challenging due to multiple humans’ involvement and their mutual interaction in a single frame, generated from their movements. Mainstream literature is based on three-dimensional (3-D) convolutional neural networks (CNNs), processing only visual frames, where human joints data play a vital role in accurate interaction recognition. Therefore, this article proposes a multistream network for HIR that intelligently learns from skeletons’ key points and spatiotemporal visual representations. The first stream localises the joints of the human body using a pose estimation model and transmits them to a 1-D CNN and bidirectional long short-term memory to efficiently extract the features of the dynamic movements of each human skeleton. The second stream feeds the series of visual frames to a 3-D convolutional neural network to extract the discriminative spatiotemporal features. Finally, the outputs of both streams are integrated via fully connected layers that precisely classify the ongoing interactions between humans. To validate the performance of the proposed network, we conducted a comprehensive set of experiments on two benchmark datasets, UT-interaction and TV human interaction, and found 1.15% and 10.0% improvement in the accuracy.
Umair Haroon, Amin Ullah, Tanveer Hussain 0001, Waseem Ullah, Khan Muhammad 0001, Mi Young Lee, Sung Wook Baik
IEEE Trans. Hum. Mach. Syst.8
2022 AI-Assisted Edge Vision for Violence Detection in IoT-Based Industrial Surveillance Networks
abstract
Analyzing surveillance videos is mandatory for the public and industrial security. Overwhelming growth in computer vision fields has been made to automate the surveillance system in terms of human activity recognition, such as behavior analysis and violence detection (VD). However, it is challenging to detect and analyze the violent scenes intelligently to fulfill the notion of Industrial Internet of Things (IIoT)-based surveillance buoyed by constrained resources to reduce computational power. To tackle this challenge, in this article, an artificial intelligence enabled IIoT-based framework with VD-Network (VD-Net) is proposed. First, the input video frames are passed to light-weight convolutional neural network model for important information collection including humans or suspicious objects such as knives/guns. Upon suspicious object detection, an alert is generated as an earlier VD in IIoT network while the information is shared with concern departments. Only the frames with objects are forwarded to cloud for detail investigation where features are extracted using convolutional long short-term memory (ConvLSTM). The latter from ConvLSTM is propagated to gated recurrent unit for final VD. The conducted experiments and ablation study on the existing surveillance and nonsurveillance datasets empirically validate the effectiveness of the proposed VD-Net by improving 3.9% increase in the accuracy compared with the state-of-the-art VD methods.
Fath U Min Ullah, Khan Muhammad 0001, Ijaz Ul Haq, Noman Khan, Ali Asghar Heidari, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics6
2022 Optimized Dual Fire Attention Network and Medium-Scale Fire Classification Benchmark
abstract
Vision-based fire detection systems have been significantly improved by deep models; however, higher numbers of false alarms and a slow inference speed still hinder their practical applicability in real-world scenarios. For a balanced trade-off between computational cost and accuracy, we introduce dual fire attention network (DFAN) to achieve effective yet efficient fire detection. The first attention mechanism highlights the most important channels from the features of an existing backbone model, yielding significantly emphasized feature maps. Then, a modified spatial attention mechanism is employed to capture spatial details and enhance the discrimination potential of fire and non-fire objects. We further optimize the DFAN for real-world applications by discarding a significant number of extra parameters using a meta-heuristic approach, which yields around 50% higher FPS values. Finally, we contribute a medium-scale challenging fire classification dataset by considering extremely diverse, highly similar fire/non-fire images and imbalanced classes, among many other complexities. The proposed dataset advances the traditional fire detection datasets by considering multiple classes to answer the following question: what is on fire? We perform experiments on four widely used fire detection datasets, and the DFAN provides the best results compared to 21 state-of-the-art methods. Consequently, our research provides a baseline for fire detection over edge devices with higher accuracy and better FPS values, and the proposed dataset extension provides indoor fire classes and a greater number of outdoor fire classes; these contributions can be used in significant future research. Our codes and dataset will be publicly available at https://github.com/tanveer-hussain/DFAN.
Hikmat Yar, Tanveer Hussain 0001, Mohit Agarwal 0003, Zulfiqar Ahmad Khan 0002, Suneet K. Gupta 0001, Sung Wook Baik
IEEE Trans. Image Process.6
2022 PMAL: A Proxy Model Active Learning Approach for Vision Based Industrial Applications
abstract
Deep Learning models’ performance strongly correlate with availability of annotated data; however, massive data labeling is laborious, expensive, and error-prone when performed by human experts. Active Learning (AL) effectively handles this challenge by selecting the uncertain samples from unlabeled data collection, but the existing AL approaches involve repetitive human feedback for labeling uncertain samples, thus rendering these techniques infeasible to be deployed in industry related real-world applications. In the proposed Proxy Model based Active Learning technique (PMAL) , this issue is addressed by replacing human oracle with a deep learning model, where human expertise is reduced to label only two small subsets of data for training proxy model and initializing the AL loop. In the PMAL technique, firstly, proxy model is trained with a small subset of labeled data, which subsequently acts as an oracle for annotating uncertain samples. Secondly, active model's training, uncertain samples extraction via uncertainty sampling, and annotation through proxy model is carried out until predefined iterations to achieve higher accuracy and labeled data. Finally, the active model is evaluated using testing data to verify the effectiveness of our technique for practical applications. The correct annotations by the proxy model are ensured by employing the potentials of explainable artificial intelligence. Similarly, emerging vision transformer is used as an active model to achieve maximum accuracy. Experimental results reveal that the proposed method outperforms the state-of-the-art in terms of minimum labeled data usage and improves the accuracy with 2.2%, 2.6%, and 1.35% on Caltech-101, Caltech-256, and CIFAR-10 datasets, respectively. Since the proposed technique offers a highly reasonable solution to exploit huge multimedia data, it can be widely used in different evolutionary industrial domains.
Ijaz Ul Haq, Tanveer Hussain 0001, Khan Muhammad 0001, Mohammad Hijji, Victor Hugo C. de Albuquerque, Sung Wook Baik
ACM Trans. Multim. Comput. Commun. Appl.8
2021 Conflux LSTMs Network: A Novel Approach for Multi-View Action Recognition
abstract
Multi-view action recognition (MVAR) is an optimal technique to acquire numerous clues from different views data for effective action recognition, however, it is not well explored yet. There exist several challenges to MVAR domain such as divergence in viewpoints, invisible regions, and different scales of appearance in each view require better solutions for real world applications. In this paper, we present a conflux long short-term memory (LSTMs) network to recognize actions from multi-view cameras. The proposed framework has four major steps; 1) frame level feature extraction, 2) its propagation through conflux LSTMs network for view self-reliant patterns learning, 3) view inter-reliant patterns learning and correlation computation, and 4) action classification. First, we extract deep features from a sequence of frames using a pre-trained VGG19 CNN model for each view. Second, we forward the extracted features to conflux LSTMs network to learn the view self-reliant patterns. In the next step, we compute the inter-view correlations using the pairwise dot product from output of the LSTMs network corresponding to different views to learn the view inter-reliant patterns. In the final step, we use flatten layers followed by SoftMax classifier for action recognition. Experimental results over benchmark datasets compared to state-of-the-art report an increase of 3% and 2% on northwestern-UCLA and MCAD datasets, respectively.
Amin Ullah, Khan Muhammad 0001, Tanveer Hussain 0001, Sung Wook Baik
Neurocomputing4
2021 An Efficient Deep Learning Framework for Intelligent Energy Management in IoT Networks
abstract
Green energy management is an economical solution for better energy usage, but the employed literature lacks focusing on the potentials of edge intelligence in controllable Internet of Things (IoT). Therefore, in this article, we focus on the requirements of todays' smart grids, homes, and industries to propose a deep-learning-based framework for intelligent energy management. We predict future energy consumption for short intervals of time as well as provide an efficient way of communication between energy distributors and consumers. The key contributions include edge devices-based real-time energy management via common cloud-based data supervising server, optimal normalization technique selection, and a novel sequence learning-based energy forecasting mechanism with reduced time complexity and lowest error rates. In the proposed framework, edge devices relate to a common cloud server in an IoT network that communicates with the associated smart grids to effectively continue the energy demand and response phenomenon. We apply several preprocessing techniques to deal with the diverse nature of electricity data, followed by an efficient decision-making algorithm for short-term forecasting and implement it over resource-constrained devices. We perform extensive experiments and witness 0.15 and 3.77 units reduced mean-square error (MSE) and root MSE (RMSE) for residential and commercial datasets, respectively.
Tao Han 0004, Khan Muhammad 0001, Tanveer Hussain 0001, Jaime Lloret Mauri, Sung Wook Baik
IEEE Internet Things J.5
2021 Multiview Summarization and Activity Recognition Meet Edge Computing in IoT Environments
abstract
Multiview video summarization (MVS) has not received much attention from the research community due to inter-view correlations and views' overlapping, etc. The majority of previous MVS works are offline, relying on only summary, and require additional communication bandwidth and transmission time, with no focus on foggy environments. We propose an edge intelligence-based MVS and activity recognition framework that combines artificial intelligence with Internet of Things (IoT) devices. In our framework, resource-constrained devices with cameras use a lightweight CNN-based object detection model to segment multiview videos into shots, followed by mutual information computation that helps in a summary generation. Our system does not rely solely on a summary, but encodes and transmits it to a master device using a neural computing stick for inter-view correlations computation and efficient activity recognition, an approach which saves computation resources, communication bandwidth, and transmission time. Experiments show an increase of 0.4 unit in F-measure on an MVS Office dateset and 0.2% and 2% improved accuracy for UCF-50 and YouTube 11 datesets, respectively, with lower storage and transmission times. The processing time is reduced from 1.23 to 0.45 s for a single frame and optimally 0.75 seconds faster MVS. A new dateset is constructed by synthetically adding fog to an MVS dateset to show the adaptability of our system for both certain and uncertain IoT surveillance environments.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Javier Del Ser, Amir Hossein Gandomi, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.7
2021 CNN features with bi-directional LSTM for real-time anomaly detection in surveillance networks
Waseem Ullah, Amin Ullah, Ijaz Ul Haq, Khan Muhammad 0001, Sung Wook Baik
Multim. Tools Appl.6
2021 A comprehensive survey of multi-view video summarization
Tanveer Hussain 0001, Khan Muhammad 0001, Weiping Ding 0001, Jaime Lloret Mauri, Sung Wook Baik, Victor Hugo C. de Albuquerque
Pattern Recognit.5
2020 One-Shot Learning for Surveillance Anomaly Recognition using Siamese 3D CNN
abstract
One-shot image recognition has been explored for many applications in computer vision community. However, its applications in video analytics is not deeply investigated yet. For instance, surveillance anomaly recognition is an open challenging problem and one of its hurdles is the lack of accurate temporally annotated data. This paper addresses the lack of data issue using one-shot learning strategy and proposes an anomaly recognition framework which exploits a 3D CNN siamese network that yields the similarity between two anomaly sequences. This paper also investigates the existing 3D CNNs for this task and then proposes a lightweight 3D CNN model that efficiently handles one-shot anomaly recognition. Once our network is trained, then we can use the powerful discriminative 3D CNN features to predict anomalies not only for the new data but also for entirely new classes. The proposed model is trained using temporally annotated test set of UCF Crime dataset. Finally, the trained model is used to recognize the anomalies and produce temporal automatic labels for the video level weakly annotated training set of the dataset.
Amin Ullah, Khan Muhammad 0001, Kilichbek Haydarov, Ijaz Ul Haq, Mi Young Lee, Sung Wook Baik
IJCNN6
2020 Mining top-k frequent patterns from uncertain databases
Tuong Le, Bay Vo, Van-Nam Huynh, Ngoc Thanh Nguyen 0001, Sung Wook Baik
Appl. Intell.5
2020 Raspberry Pi assisted face recognition framework for enhanced law-enforcement services in smart cities
Mansoor Nasir, Khan Muhammad 0001, Siraj Khan, Zahoor Jan, Arun Kumar Sangaiah, Mohamed Elhoseny, Sung Wook Baik
Future Gener. Comput. Syst.8
2020 Egocentric visual scene description based on human-object interaction and deep spatial relations among objects
Gulraiz Khan, Muhammad Usman Ghani Khan, Aiman Siddiqi, Zahoor-Ur Rehman, Sung Wook Baik, Irfan Mehmood
Multim. Tools Appl.6
2020 Efficient CNN based summarization of surveillance videos for resource-constrained devices
Khan Muhammad 0001, Tanveer Hussain 0001, Sung Wook Baik
Pattern Recognit. Lett.3
2020 Intelligent Embedded Vision for Summarization of Multiview Videos in IIoT
abstract
Nowadays, video sensors are used on a large scale for various applications, including security monitoring and smart transportation. However, the limited communication bandwidth and storage constraints make it challenging to process such heterogeneous nature of Big Data in real time. Multiview video summarization (MVS) enables us to suppress redundant data in distributed video sensors settings. The existing MVS approaches process video data in offline manner by transmitting them to the local or cloud server for analysis, which requires extra streaming to conduct summarization, huge bandwidth, and are not applicable for integration with industrial Internet of Things (IIoT). This article presents a light-weight convolutional neural network (CNN) and IIoT-based computationally intelligent (CI) MVS framework. Our method uses an IIoT network containing smart devices, Raspberry Pi (RPi) (clients and master) with embedded cameras to capture multiview video data. Each client RPi detects target in frames via light-weight CNN model, analyzes these targets for traffic and crowd density, and searches for suspicious objects to generate alert in the IIoT network. The frames of each client RPi are encoded and transmitted with approximately 17.02% smaller size of each frame to master RPi for final MVS. Empirical analysis shows that our proposed framework can be used in industrial environments for various applications such as security and smart transportation and can be proved beneficial for saving resources.11[Online]. Available: https://github.com/tanveer-hussain/Embedded-Vision-for-MVS.
Tanveer Hussain 0001, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics4
2020 Cloud-Assisted Multiview Video Summarization Using CNN and Bidirectional LSTM
abstract
The massive amount of video data produced by surveillance networks in industries instigate various challenges in exploring these videos for many applications, such as video summarization (VS), analysis, indexing, and retrieval. The task of multiview video summarization (MVS) is very challenging due to the gigantic size of data, redundancy, overlapping in views, light variations, and interview correlations. To address these challenges, various low-level features and clustering-based soft computing techniques are proposed that cannot fully exploit MVS. In this article, we achieve MVS by integrating deep neural network based soft computing techniques in a two-tier framework. The first online tier performs target-appearance-based shots segmentation and stores them in a lookup table that is transmitted to cloud for further processing. The second tier extracts deep features from each frame of a sequence in the lookup table and pass them to deep bidirectional long short-term memory (DB-LSTM) to acquire probabilities of informativeness and generates a summary. Experimental evaluation on benchmark dataset and industrial surveillance data from YouTube confirms the better performance of our system compared to the state-of-the-art MVS methods.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Zehong Cao, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics5
2019 SPPC: a new tree structure for mining erasable patterns in data streams
Tuong Le, Bay Vo, Philippe Fournier-Viger, Mi Young Lee, Sung Wook Baik
Appl. Intell.5
2019 Action recognition using optimized deep autoencoder and CNN for surveillance data streams of non-stationary environments
Amin Ullah, Khan Muhammad 0001, Ijaz Ul Haq, Sung Wook Baik
Future Gener. Comput. Syst.4
2019 Energy-Efficient Deep CNN for Smoke Detection in Foggy IoT Environment
abstract
Smoke detection in Internet of Things (IoT) environment is a primary component of early disaster-related event detection in smart cities. Recently, several smoke and fire detection methods are presented with reasonable accuracy and running time for normal IoT environment. However, these methods are unable to detect smoke in foggy IoT environment, which is a challenging task. In this paper, we propose an energy-efficient system based on deep convolutional neural networks for early smoke detection in both normal and foggy IoT environments. Our method takes advantage of VGG-16 architecture, considering its sensible stability between the accuracy and time efficiency for smoke detection compared to the other computationally expensive networks, such as GoogleNet and AlexNet. Experiments performed on benchmark smoke detection datasets and their results in terms of accuracy, false alarms rate, and efficiency reveal the better performance of our technique compared to state-of-the-art and verifies its applicability in smart cities for early detection of smoke in normal and foggy IoT environments.
Salman Khan 0004, Khan Muhammad 0001, Shahid Mumtaz, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.4
2019 A fast and accurate approach for bankruptcy forecasting using squared logistics loss with GPU-based extreme gradient boosting
Tuong Le, Bay Vo, Hamido Fujita, Ngoc Thanh Nguyen 0001, Sung Wook Baik
Inf. Sci.5
2019 Raspberry Pi assisted facial expression recognition framework for smart security in law-enforcement services
Mansoor Nasir, Fath U Min Ullah, Khan Muhammad 0001, Arun Kumar Sangaiah, Sung Wook Baik
Inf. Sci.6
2019 Deep features-based speech emotion recognition for smart affective services
Abdul Malik Badshah, Nasir Rahim, Noor Ullah, Jamil Ahmad 0003, Khan Muhammad 0001, Mi Young Lee, Soonil Kwon, Sung Wook Baik
Multim. Tools Appl.8
2019 An efficient computerized decision support system for the analysis and 3D visualization of brain tumor
Irfan Mehmood, Khan Muhammad 0001, Syed Inayat Ali Shah, Arun Kumar Sangaiah, Sung Wook Baik
Multim. Tools Appl.7
2019 CNN-based anti-spoofing two-tier multi-factor authentication system
Salman Khan 0004, Tanveer Hussain 0001, Khan Muhammad 0001, Arun Kumar Sangaiah, Aniello Castiglione, Christian Esposito 0001, Sung Wook Baik
Pattern Recognit. Lett.8
2019 Efficient Fire Detection for Uncertain Surveillance Environment
abstract
Tactile Internet can combine multiple technologies by enabling intelligence via mobile edge computing and data transmission over a 5G network. Recently, several convolutional neural networks (CNN) based methods via edge intelligence are utilized for fire detection in certain environment with reasonable accuracy and running time. However, these methods fail to detect fire in uncertain Internet of Things (IoT) environment having smoke, fog, and snow. Furthermore, achieving good accuracy with reduced running time and model size is challenging for resource constrained devices. Therefore, in this paper, we propose an efficient CNN based system for fire detection in videos captured in uncertain surveillance scenarios. Our approach uses light-weight deep neural networks with no dense fully connected layers, making it computationally inexpensive. Experiments are conducted on benchmark fire datasets and the results reveal the better performance of our approach compared to state-of-the-art. Considering the accuracy, false alarms, size, and running time of our system, we believe that it is a suitable candidate for fire detection in uncertain IoT environment for mobile and embedded vision applications during surveillance.
Khan Muhammad 0001, Salman Khan 0004, Mohamed Elhoseny, Syed Hassan Ahmed, Sung Wook Baik
IEEE Trans. Ind. Informatics5
2019 Efficient Deep CNN-Based Fire Detection and Localization in Video Surveillance Applications
abstract
Convolutional neural networks (CNNs) have yielded state-of-the-art performance in image classification and other computer vision tasks. Their application in fire detection systems will substantially improve detection accuracy, which will eventually minimize fire disasters and reduce the ecological and social ramifications. However, the major concern with CNN-based fire detection systems is their implementation in real-world surveillance networks, due to their high memory and computational requirements for inference. In this paper, we propose an original, energy-friendly, and computationally efficient CNN architecture, inspired by the SqueezeNet architecture for fire detection, localization, and semantic understanding of the scene of the fire. It uses smaller convolutional kernels and contains no dense, fully connected layers, which helps keep the computational requirements to a minimum. Despite its low computational needs, the experimental results demonstrate that our proposed solution achieves accuracies that are comparable to other, more complex models, mainly due to its increased depth. Moreover, this paper shows how a tradeoff can be reached between fire detection accuracy and efficiency, by considering the specific characteristics of the problem of interest and the variety of fire data.
Khan Muhammad 0001, Jamil Ahmad 0003, Zhihan Lyu, Paolo Bellavista, Po Yang 0001, Sung Wook Baik
IEEE Trans. Syst. Man Cybern. Syst.6
2018 Privacy-preserving image retrieval for mobile devices with deep features on the cloud
Nasir Rahim, Jamil Ahmad 0003, Khan Muhammad 0001, Arun Kumar Sangaiah, Sung Wook Baik
Comput. Commun.5
2018 Efficient algorithms for mining top-rank-k erasable patterns using pruning strategies and the subsume concept
Tuong Le, Bay Vo, Sung Wook Baik
Eng. Appl. Artif. Intell.3
2018 Object-oriented convolutional features for fine-grained image retrieval in large surveillance datasets
Jamil Ahmad 0003, Khan Muhammad 0001, Sambit Bakshi, Sung Wook Baik
Future Gener. Comput. Syst.4
2018 Image steganography using uncorrelated color space and its application for security of visual contents in online social networks
Khan Muhammad 0001, Irfan Mehmood, Seungmin Rho, Sung Wook Baik
Future Gener. Comput. Syst.5
2018 Early fire detection using convolutional neural networks during surveillance for effective disaster management
Khan Muhammad 0001, Jamil Ahmad 0003, Sung Wook Baik
Neurocomputing3
2018 ARM-AMO: An efficient association rule mining algorithm based on animal migration optimization
Le Hoang Son, Francisco Chiclana, Raghvendra Kumar 0001, Mamta Mittal, Manju Khari, Jyotir Moy Chatterjee, Sung Wook Baik
Knowl. Based Syst.7
2018 A Route Optimized Distributed IP-Based Mobility Management Protocol for Seamless Handoff across Wireless Mesh Networks
Peer Azmat Shah, Khalid M. Awan, Zahoor-Ur Rehman, Khalid Iqbal, Farhan Aadil, Khan Muhammad 0001, Irfan Mehmood, Sung Wook Baik
Mob. Networks Appl.8
2018 Determining speaker attributes from stress-affected speech in emergency situations with hybrid SVM-DNN architecture
Jamil Ahmad 0003, Seungmin Rho, Soonil Kwon, Mi Young Lee, Sung Wook Baik
Multim. Tools Appl.6
2018 Integrating salient colors with rotational invariant texture features for image representation in retrieval systems
Amin Ullah, Jamil Ahmad 0003, Naveed Abbas, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.6
2018 Efficient Conversion of Deep Features to Compact Binary Codes Using Fourier Decomposition for Multimedia Big Data
abstract
Exponential growth of multimedia data has been witnessed in recent years from various industries, such as e-commerce, health, transportation, and social networks, etc. Access to desired data in such gigantic datasets require sophisticated and efficient retrieval methods. In the last few years, neuronal activations generated by a pretrained convolutional neural network (CNN) have served as generic descriptors for various tasks including image classification, object detection and segmentation, and image retrieval. They perform incredibly well compared to hand-crafted features. However, these features are usually high dimensional, requiring a lot of memory and computations for indexing and retrieval. For very large datasets, utilization of these high dimensional features in raw form becomes infeasible. In this paper, a highly efficient method is proposed to transform high dimensional deep features into compact binary codes using bidirectional Fourier decomposition. This compact bit code saves memory and eases computations during retrieval. Further, these codes can also serve as hash codes, allowing very efficient access to images in large datasets using approximate nearest neighbor (ANN) search techniques. Our method does not require any training and achieves considerable retrieval accuracy with short length codes. It has been tested on features extracted from fully connected layers of a pretrained CNN. Experiments conducted with several large datasets reveal the effectiveness of our approach for a wide variety of datasets.
Jamil Ahmad 0003, Khan Muhammad 0001, Jaime Lloret Mauri, Sung Wook Baik
IEEE Trans. Ind. Informatics4
2018 Secure Surveillance Framework for IoT Systems Using Probabilistic Image Encryption
abstract
This paper proposes a secure surveillance framework for Internet of things (IoT) systems by intelligent integration of video summarization and image encryption. First, an efficient video summarization method is used to extract the informative frames using the processing capabilities of visual sensors. When an event is detected from keyframes, an alert is sent to the concerned authority autonomously. As the final decision about an event mainly depends on the extracted keyframes, their modification during transmission by attackers can result in severe losses. To tackle this issue, we propose a fast probabilistic and lightweight algorithm for the encryption of keyframes prior to transmission, considering the memory and processing requirements of constrained devices that increase its suitability for IoT systems. Our experimental results verify the effectiveness of the proposed method in terms of robustness, execution time, and security compared to other image encryption algorithms. Furthermore, our framework can reduce the bandwidth, storage, transmission cost, and the time required for analysts to browse large volumes of surveillance data and make decisions about abnormal events, such as suspicious activity detection and fire detection in surveillance applications.
Khan Muhammad 0001, Rafik Hamza, Jamil Ahmad 0003, Jaime Lloret Mauri, Haoxiang Wang 0001, Sung Wook Baik
IEEE Trans. Ind. Informatics6
2017 Mining erasable itemsets with subset and superset itemset constraints
Bay Vo, Tuong Le, Witold Pedrycz, Sung Wook Baik
Expert Syst. Appl.5
2017 Efficient object-based surveillance image search using spatial pooling of convolutional features
Jamil Ahmad 0003, Irfan Mehmood, Sung Wook Baik
J. Vis. Commun. Image Represent.3
2017 Analysis of interaction trace maps for active authentication on smart devices
Jamil Ahmad 0003, Zahoor Jan, Irfan Mehmood, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.6
2017 Image steganography for authenticity of visual contents in social networks
Khan Muhammad 0001, Jamil Ahmad 0003, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.4
2017 Mobile-cloud assisted framework for selective encryption of medical images with steganography for resource-constrained devices
Khan Muhammad 0001, Sung Wook Baik, Seungmin Rho, Zahoor Jan, Sang-Soo Yeo, Irfan Mehmood
Multim. Tools Appl.3
2016 Divide-and-conquer based summarization framework for extracting affective video content
Irfan Mehmood, Seungmin Rho, Sung Wook Baik
Neurocomputing4
2016 Multi-scale local structure patterns histogram for describing visual contents in social image retrieval systems
Jamil Ahmad 0003, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.4
2016 A novel magic LSB substitution method (M-LSB-SM) using multi-level encryption and achromatic component of an image
Khan Muhammad 0001, Irfan Mehmood, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.5
2015 Image super-resolution using sparse coding over redundant dictionary based on effective image representations
Irfan Mehmood, Sung Wook Baik
J. Vis. Commun. Image Represent.3
2015 Digital image super-resolution using adaptive interpolation based on Gaussian function
Naveed Ejaz, Irfan Mehmood, Sung Wook Baik
Multim. Tools Appl.4
2014 Adaptive image data hiding using transformation and error replacement
Naveed Ejaz, Jamil Anwar, Muhammad Ishtiaq, Sung Wook Baik
Multim. Tools Appl.4
2014 Multi-kernel based adaptive interpolation for image super-resolution
Naveed Ejaz, Sung Wook Baik
Multim. Tools Appl.3
2013 Efficient visual attention based framework for extracting key frames from videos
Naveed Ejaz, Irfan Mehmood, Sung Wook Baik
Signal Process. Image Commun.3
2012 Adaptive key frame extraction for video summarization using an aggregation mechanism
Naveed Ejaz, Tayyab Bin Tariq, Sung Wook Baik
J. Vis. Commun. Image Represent.3
2012 Video summarization using a network of radial basis functions
Naveed Ejaz, Sung Wook Baik
Multim. Syst.2
2011 Neural network modeling of inter-characteristics of silicon nitride film deposited by using a plasma-enhanced chemical vapor deposition
Su Jin Lee, Byungwhan Kim, Sung Wook Baik
Expert Syst. Appl.3
2010 Distributed Data Mining System Based on Multi-agent Communication Mechanism
Sung Gook Kim, Kyeong Deok Woo, Jerzy W. Bala, Sung Wook Baik
KES-AMSTA (2)4
2010 Immersive modeling system (IMMS) for personal electronic products using a multi-modal interface
Yong-Gu Lee, Hyungjun Park, Woontack Woo, Jeha Ryu, Hong Kook Kim, Sung Wook Baik, Kwang Hee Ko, Han Kyun Choi, Sun-Uk Hwang, Duck Bong Kim, Hyun Soo Kim, Kwan H. Lee
Comput. Aided Des.6
2007 A Robust Localization Method for Mobile Robots Based on Ceiling Landmarks
Viet Thang Nguyen, Moon Seok Jeong, Sung Mahn Ahn, Seungbin Moon, Sung Wook Baik
MDAI5
2006 Aerial Photo Image Retrieval Using Adaptive Image Classification
Sung Wook Baik, Moon Seok Jeong, Ran Baik
KES (3)1
2005 Background Modeling Using Color, Disparity, and Motion Information
Jong-Weon Lee 0002, Hyo Sung Jeon, Sung Min Moon, Sung Wook Baik
ACIVS4
2005 Neural Network Based Adult Image Classification
Wonil Kim, Han-Ku Lee, Seong-Joon Yoo, Sung Wook Baik
ICANN (1)4
2005 Minimal RBF Networks by Gaussian Mixture Model
Sung Mahn Ahn, Sung Wook Baik
ICIC (1)2
2005 Distributed Visual Interfaces for Collaborative Exploration of Data Spaces
Sung Wook Baik, Jerzy W. Bala, Yung Jo
KES (1)1
2005 Aerial Photograph Image Retrieval Using the MPEG-7 Texture Descriptors
Sang Kim, Sung Wook Baik, Yung Jo, Seungbin Moon, Dae Woong Rhee
KES (2)2
2004 Adaptive Texture Recognition in Image Sequences with Prediction through Features Interpolation
Sung Wook Baik, Ran Baik
ICCSA (3)1
2004 A Decision Tree Algorithm for Distributed Data Mining: Towards Network Intrusion Detection
Sung Wook Baik, Jerzy W. Bala
ICCSA (4)1
2004 Visualizing Predictive Models in Decision Tree Generation
Sung Wook Baik, Jerzy W. Bala, Sung Ahn
ICCSA (4)1
2004 Texture Feature Extraction and Selection for Classification of Images in a Sequence
Khin K. Win, Sung Wook Baik, Ran Baik, Sung Ahn, Sang Kim, Yung Jo
IWCIA2
2004 Agent Based Distributed Data Mining
Sung Wook Baik, Jerzy W. Bala, Ju Sang Cho
PDCAT1
2003 Adaptive Segmentation of Remote-Sensing Images for Aerial Surveillance
Sung Wook Baik, Sung Mahn Ahn, Jong-Weon Lee 0002, Khin K. Win
CAIP1
2002 Online model modification for adaptive texture recognition in image sequences
abstract
This paper presents and validates a method for adaptive texture recognition in image sequences under dynamic perceptual conditions and, consequently, under changing texture characteristics. The approach builds a closed-loop interaction between texture recognition and model modification systems. Texture recognition applies a modified radial-basis function (RBF) classifier to a current image of a sequence. The feedback reinforcement generation mechanism evaluates the classification results when compared to the previous images and activates classifier modification, if needed. Classifier modification selects a strategy and employs four behaviors in adapting the classifier's structure and parameters. These behaviors include accommodation, translation, generation, and extinction applied to selected classifier components. Accommodation modifies the component's boundary/spread. Translation shifts a given component over the feature space. Generation creates a new component of the RBF classifier. Extinction eliminates components that are no longer in use. The evolved RBF model is verified in order to confirm applied model modifications. Experimental results are presented for indoor and outdoor image sequences. The approach is validated and compared with traditional nonadaptive methods for texture recognition.
Sung Wook Baik, Peter W. Pachowicz
IEEE Trans. Syst. Man Cybern. Part A1
2000 On-Line Model Modification for Adaptive Object Recognition
Peter W. Pachowicz, Sung Wook Baik
ECAI2
2000 Adaptive RBF Classifier for Object Recognition in Image Sequences
abstract
This paper presents an adaptive NN-RBF classifier developed for object recognition under continuously time-varying perceptual conditions. The classifier is a hybrid of a neural net and a control environment. Adaptability of the classifier involves processes of image analysis, reinforcement generation, and classifier modification. An NN-RBF classifier is applied to a single image of a sequence. A feedback reinforcement generation mechanism evaluates the classification results when compared to the previous images and activates classifier modification, if needed. Classifier modification selects a strategy and employs four behaviors in adapting the classifier's structure and parameters. The developed approach is tested on indoor and outdoor image sequences.
Peter W. Pachowicz, Sung Wook Baik
IJCNN (6)2
1999 A cluster analysis based on a regularization method
abstract
The paper shows a clustering application of a density estimation method that utilizes the Gaussian mixture model and the regularization theory. We define a "closeness measure" as a clustering criterion to see how close two Gaussian components are. The closeness measure is defined as the ratio of log likelihood between two Gaussian components. According to simulations using artificial data, the clustering algorithm turned out to be very powerful in that it can correctly determine clusters in complex situations, and very flexible in that it can produce different sizes of clusters based on different threshold values.
Sung Mahn Ahn, Sung Wook Baik
IJCNN2
1999 Adaptive object recognition based on the radial basis function paradigm
abstract
This paper presents an online model modification methodology for the recognition of objects. The methodology deals with the problem of continuously changing object characteristics used by the recognition processes. This characteristics change is due to the dynamics of scene-viewer relationship such as resolution and lighting. The approach applies a modified radial-basis function paradigm for model-based object modeling and recognition. Online adaptation of these models is developed and works in a closed loop with the object recognition system to perceive discrepancies between object models and varying object characteristics. Developed methodology is experimentally validated through experiments with a sequence of outdoor images.
Sung Wook Baik, Peter W. Pachowicz
IJCNN1