Li Zhang 0013

dblp:89/5992-13 · DBLP profile ↗
← Back
88ranked-venue papers
18as first author
38since 2021 · last 2026
0000-0001-6674-692XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 66 · 14 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 14 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 TSD-Rec: Metric semantic noise enhanced diffusion based contrastive learning with topology prior for multi-behavior recommendation
Rong Gao 0001, Yabo Guo, Yonghong Yu, Zhiwei Ye, Li Zhang 0013, Lingyu Yan
Expert Syst. Appl.5
2026 Wasserstein distance-based graph contrastive learning for recommendation
Yonghong Yu, Yujie Liao, Li Zhang 0013, Rong Gao 0001
Expert Syst. Appl.4
2026 Deep learning-based emotion recognition using unimodal facial expressions or physiological signals: A review
abstract
ABSTRACT Emotion recognition has become a key component of intelligent systems, enabling improved human–computer interaction across domains such as healthcare, education, and robotics. Progress has been achieved using facial expressions and physiological signals, particularly Electroencephalography (EEG) and Electrocardiogram (ECG), supported by advances in deep learning. This paper presents a comprehensive review of unimodal emotion recognition based on facial expressions or physiological signals, with each modality analysed independently to understand signal-specific characteristics and modelling strategies. In addition to commonly studied modalities, this review covers physiological signals including Galvanic Skin Response (GSR), Photoplethysmography (PPG), Electrooculography (EOG), Electromyography (EMG), Respiration Rate (RR), Skin Temperature (SKT), and functional near-infrared spectroscopy (fNIRS). The review emphasises the design and evaluation of deep learning architectures, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM), Gated Recurrent Units (GRU), and emerging approaches such as transformer-based models, Vision Transformers (ViTs), and transfer learning techniques. Studies are analysed using a structured evaluation framework considering model design, preprocessing strategies, datasets, evaluation protocols, and performance. Unlike existing surveys, this work provides a critical analysis of prior studies, highlighting strengths, limitations, and trade-offs, with emphasis on generalisation capability and evaluation strategies. The review identifies risks associated with improper data partitioning, where data leakage can lead to overestimated performance. Key challenges include limited dataset sizes, lack of standardised evaluation protocols, and inconsistencies in performance reporting. This study provides a structured understanding of current research trends and outlines future directions for developing more robust, reliable, and generalisable emotion recognition systems.
Mohsen Golafrouz, Houshyar Asadi, Mohammad Anwar Hosen, Mohammad Reza Chalak Qazani, Seyed Amin Khatami, Mojgan Fayyazi, Li Zhang 0013, Siamak Pedrammehr, Lei Wei 0002, Cp Lim, Saeid Nahavandi
Knowl. Based Syst.7
2026 Enhancing Data Efficiency With a Trustworthy Counterfactual Generative Model
abstract
Leveraging limited data to synthesize an additional training set is essential for robotic vision, particularly in dynamic environments where collecting large datasets is impractical. Traditional robotic vision systems rely on extensive training data for object recognition and scene understanding but struggle to generalize to real-world variations, such as lighting conditions, occlusions, and sensor noise. This article proposes causal diffuse variational autoencoder (causal DiffuseVAE), a novel method integrating causal inference with high-fidelity image synthesis to generate counterfactual images. By combining the disentanglement properties of variational autoencoders (VAEs) with the generative capabilities of diffusion models, causal DiffuseVAE produces realistic, interpretable simulations of variations, such as shadows and occlusions. This combination enables data-efficient generative modeling by learning from small subsets and synthesizing missing or unseen samples. In addition, causal inference ensures that generated data follow real-world dependencies, making it robust and interpretable for deployment in unpredictable environments. Four baseline approaches are evaluated across six different datasets, demonstrating that causal DiffuseVAE consistently outperforms the four baseline approaches.
Zhaoan Ye, Dezong Zhao, Li Zhang 0013, Xidong Yan, Qinglin Bi, David Flynn
IEEE Trans. Ind. Informatics3
2025 Enhanced U-Net Models with Hybrid Attention Elements for Medical Image Segmentation
abstract
This research presents an enhanced deep learning-based approach for medical image segmentation by integrating multiple attention mechanisms into the U-Net architecture. Specifically, the proposed models incorporate seven advanced attention mechanisms, including Convolutional Block Attention Module, Attention Gate, Squeeze-and-Excitation, Halo, Coordinate, Spatial, and Triplet Attention strategies, to improve the model’s capability to capture both channel and spatial contextual information within medical images. These attention-enhanced U-Net variants are rigorously evaluated on multiple medical imaging datasets, demonstrating significant improvements in segmentation accuracy compared to benchmark models such as U-Net and DeepLabV3+. The results indicate superior performance, especially in complex scenarios with overlapping structures and fine organ/lesion details, showcasing the effectiveness of attention mechanisms for improving segmentation in medical imaging.
Piyushkumar Banugariya, Li Zhang 0013, Yonghong Yu, Vivian Sedov
IJCNN2
2025 Novel Convolutional Neural Networks with Multi-layer Feature Aggregation and Pooling Permutations for Sound Classification
abstract
Existing Convolutional Neural Networks (CNNs) suffer from limitations such as purely using features from the final layers for decision making, without leveraging important information obtained from middle layers. To tackle such limitations, we propose a novel CNN variant with multi-layer feature aggregation and pooling permutations for sound classification. We also introduce a method reshaping intermediate features from different layers to summarize in time-axis without losing important information. The proposed model results in significant boosting on model performance over the original networks for diverse sound classification problems.
Hyosun Choi, Li Zhang 0013, Chris Watkins, Arjun Panesar
SMC2
2025 Deep Fusion Networks with Hybrid Attention Mechanisms for Voice Diabetes Detection
abstract
About 50% of Type 2 diabetes patients develop diabetic neuropathy, which can damage nerves throughout the body, including those controlling the vocal cords, leading to issues like vocal fold paralysis, hoarseness, or vocal strain. Therefore, this research pioneers automated diabetes diagnosis using voice and speech recordings, with the attempt to provide a non-invasive, easily accessible tool for early detection of diabetes and prediabetes conditions. Firstly, five speech/voice datasets have been generated, including audio recordings of vowel letters (‘A’, ‘E’, ‘O’) and high and low speed countings of digital numbers (1-20), provided by participants with and without diabetes conditions. Initially, Mel-frequency Cepstral Coefficient (MFCC) features from various voice clips of diabetic and non-diabetic participants are extracted. 1D Convolutional Neural Network (1D CNN) and Bidirectional Long Short-Term Memory (BiLSTM) incorporating various hybrid attention mechanisms are proposed for diabetes classification. Specifically, attention methods, i.e. Squeeze and Excitation Block (SE), Convolutional Block Attention Module (CBAM), and Global Context (GC) blocks, as well as their hybrid strategies, i.e. SE+CBAM and SE+CBAM+GC, have been proposed to extract the most important acoustic features and incorporated with BiLSTM and 1D CNN, respectively. Moreover, within and cross-dataset hybrid attention feature maps of diverse resulting networks are constructed to further emphasize the most crucial disease-indicative patterns from a global perspective, leading to the best accuracy rate of 82%. The empirical studies suggest that speech analysis could be a viable method for the preliminary diagnosis of diabetes.
Abhishek Gangani, Li Zhang 0013, Arjun Panesar
SMC2
2025 Beta distribution-based monogamous pairs genetic algorithm for knowledge transfer in many-task optimization
Ting Yee Lim, Choo Jun Tan, Yi Wen Kerk, Li Zhang 0013, Chee Peng Lim
Knowl. Based Syst.4
2025 Cluster search optimisation of deep neural networks for audio emotion classification
abstract
Automated patient monitoring solutions greatly benefit from audio emotion classification, although the considerable variance in individual expression and interpretation of emotions poses a challenge. Current approaches often employ standard Audio Spectrogram Transformer (AST) and deep learning models such as Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN)-based networks. However, their performance can be enhanced by integrating neural architecture search techniques using swarm optimisation algorithms. In this research, we explore AST with hyperparameter optimisation for speech emotion recognition. Three deep learning architectures with optimisable τ b -block structures and variable filter numbers, i.e. 1DCNN, bidirectional LSTM (BiLSTM) and CNN-BiLSTM, are also proposed, enabling the optimisation of network depth and width. A novel Cluster Search Optimisation (CSO) algorithm is introduced. It incorporates Cluster Centroid Search, a Cluster Distance Improvement metric and reinforcement learning to dispatch different search actions based on clustering convergence and Q -learning strategies, respectively. A novel Noise Tempered K-means (NTKM) clustering model is also proposed with the integration of Gaussian-based noise insertion and cluster compactness-separation measurement, to further fine-tune the cluster centriods obtained using OPTICS clustering. CSO is used for hyperparameter and architecture search for AST and aforementioned deep networks. Attention mechanisms are also integrated with CSO-optimised networks to further enhance feature learning. We evaluate the resulting models against those devised by other optimisation algorithms across the EMO-DB, SAVEE, and TESS datasets. The empirical results demonstrate that CSO-optimised AST and CNN-BiLSTM with attention mechanisms outperform other architectures and yield favourable comparison results against those from existing state-of-the-art audio emotion classification methods. • Evolving transformer and deep networks are devised for audio emotion recognition. • A Cluster Search Optimisation algorithm is proposed to adapt hyperparameters. • It incorporates Noise Tempered K-means clustering and Cluster Distance Improvement. • The Q-learning algorithm is used to optimise search behaviours. • Our study indicates CSO-optimised deep networks’ effectiveness across datasets.
Sam Slade, Li Zhang 0013, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Arjun Panesar, Philip Fei Wu, Rong Gao 0001
Knowl. Based Syst.2
2025 Self-supervised extracted contrast network for facial expression recognition
Lingyu Yan, Jinquan Yang, Jinyao Xia, Rong Gao 0001, Li Zhang 0013, Yuan Yan Tang
Multim. Tools Appl.5
2025 Audio-Visual Emotion Classification Using Reinforcement Learning-Enhanced Particle Swarm Optimisation
abstract
The extraction of fine-grained spatial-temporal characteristics for emotion classification is a challenging task owing to the subtlety and ambiguity of emotional expressions through video and audio channels. In this research, we propose an audio-visual ensemble model, comprising a two-stream 3D Convolutional Neural Network (CNN) architecture with RGB and optical flow as inputs for video emotion classification, as well as a variant of Wav2Vec2 for audio emotion recognition. The Wav2Vec2 variant integrates additional recurrent and attention layers with each transformer block to extract long- and short-term dependencies. A new Particle Swarm Optimisation (PSO) algorithm is proposed to fine-tune hyper-parameters of 3D CNNs and the enhanced Wav2Vec2, and formulate audio-visual ensemble models with the smallest sizes. It integrates a reinforcement learning (RL) algorithm, i.e. Asynchronous Advantage Actor-Critic (A3C), for search parameter and hybrid leader construction, and another RL algorithm, Proximal Policy Optimisation (PPO), for search action selection, as well as hypotrochoid and super formula-based search operations. Evaluated using audio-visual emotion datasets, our evolving ensemble model outperforms those devised by other search methods and existing state-of-the-art deep networks, significantly.
Karolis Kondrotas, Li Zhang 0013, Chee Peng Lim, Houshyar Asadi, Yonghong Yu
IEEE Trans. Affect. Comput.2
2025 Contrastive Translation With Dynamical Temperature for Sequential Recommendation
abstract
Contrastive learning is a promising solution to the problem of data sparsity in the field of recommendation system since it is able to extract self-supervised signals from raw data. The traditional contrastive learning-based sequential recommendation algorithms generate augmentations of original item sequences by utilizing crop, mask and reorder operations. However, those augmentation schemes destroy the underlying semantics of item sequences, resulting in difficulty in accurately defining positive and negative samples. To address this issue, we propose a contrastive translation based sequential recommendation algorithm, namely, CT4Rec. Specifically, CT4Rec generates augmented views of item sequences by injecting noises into embeddings of users and items, which is able to guarantee that the underlying semantics of augmented views are consistent with those of original item sequence. Hence, CT4Rec is able to effectively learn the invariances among the augmented views. In addition, the personalized translation operations are utilized to model the third-order relationships among entities. Moreover, it is difficult for contrastive learning-based recommendation algorithms with static temperature to simultaneously capture the differences among individual users/items and among the clusters of users/items. Hence, we utilize a dynamic temperature strategy to enhance CT4Rec, which endows CT4Rec with the capabilities of group-wise discrimination and instance discrimination. Our validation on five benchmark datasets shows that CT4Rec outperforms SOTA sequential recommendation methods. Our code is released athttps://github.com/zar123123/CT4Rec.
Aoran Zhang 0001, Yonghong Yu, Li Zhang 0013, Rong Gao 0001, Hongzhi Yin
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Hyperbolic Adversarial Learning for Personalized Item Recommendation
Aoran Zhang 0001, Yonghong Yu, Gongyou Xu, Rong Gao 0001, Li Zhang 0013, Hongzhi Yin
DASFAA (3)5
2024 Plant Species Classification Using Evolving Ensemble and Siamese Networks
abstract
Image-based dried plant specimen identification poses a significant challenge due to the large number of possible classes and the extreme scarcity of labelled training samples. To tackle these limitations and mitigate classification biases, this research proposes a Particle Swarm Optimisation (PSO)-based weighted evolving ensemble model as well as a Siamese network for plant species classification. Specifically, we first diversify the base classifier pool by employing three networks, i.e. ResNet50, Xception, and VGG19, fine-tuned using the specimen samples. Besides the adoption of a mean average ensemble model, a weighted ensemble scheme with PSO-based optimal weighting factor generation is also utilised to integrate the outputs of the three base networks for tackling classification variances. In addition, to further tackle species classification with extremely imbalanced data, a Siamese network with ResNet50 as the backbone is utilised. Evaluated using a challenging FGVC6 data set with Melastomataceae images, the PSO-based weighted ensemble model is able to assign more influence to the best performing base networks for ensemble prediction and outperforms the traditional mean average ensemble method. Moreover, the Siamese network also obtains competitive performance for solving imbalanced specimen classification by performing comparing similarity scores between image embeddings.
Jed Arno, Olwen Grace, Isabel Larridon, Li Zhang 0013
SMC4
2024 Auxiliary Generative Adversarial Networks with Iliustration2Vec and Q-Learning based Hyperparameter Optimisation for Anime Image Synthesis
abstract
Harnessing the power of Generative Adversarial Networks (GANs) for the specialised task of anime face generation, this study introduces enhanced models of Auxiliary Classifier GAN (AC-GAN) and Wasserstein Auxiliary Classifier GAN (WAC-GAN) with modified network architectures and reinforcement learning-based hyperparameter optimisation. These models are uniquely adapted to handle the distinct nuances of anime-style imagery, a domain where conventional GANs often stumble due to complex stylistic variations and a heightened risk of mode collapse. Novel elements of our approach include, (1) modification of existing generator and discriminator architectures of both AC-GAN and WAC-GAN, (2) Q-learning based optimal hyperparameter selection, and (3) Illustration2Vec (I2V)-based automated attribute label extraction. Specifically, the Q-learning method is employed for hyperparameter search which effectively explores the search space of key network configurations by fulfilling the principles of Bellman optimality. Besides that, a deep learning-based 12V's method is utilised to generate attribute class labels and latent vectors to inform the generation process. Furthermore, we augment AC-GAN and WAC-GAN with additional layers to enhance their feature learning and generative capabilities. The insertion of these additional layers is calibrated based on the optimised network learning settings as well as the class labels derived from 12V, to fine-tune model scalability and diversity. Our experimental studies indicate that the conjunction of these techniques has led to a significant improvement in generating high-fidelity anime faces, adeptly handling the diverse and complex attributes inherent in anime-style imagery. The proposed strategies also showcase the potential of our customised AC-GAN and WAC-GAN models to master the nuanced art of anime face generation.
Vivian Sedov, Li Zhang 0013
SMC2
2024 Genetic Algorithm with Reinforcement Learning based Parameter Optimisation
abstract
Optimisation of parameters in Genetic algorithms (GA) can improve the speed and accuracy of the solution produced, but well optimised parameters are dependant on the problem being solved, and the substantial additional cost of spending time pre-computing good parameters can offset the benefit. This research investigates the use of reinforcement learning algorithms to optimise the parameters of the GA during its runtime. Specifically, we propose a variant of the GA method which embeds the Q-learning algorithm to select an optimal mutation rate at each iteration. Evaluating with a set of benchmark functions, the proposed GA model with Q-learning shows promising performance with lower mean scores than those of the original GA for most test functions. In particular, the Q-learning algorithm shows a promising emergent behaviour, i.e. selecting a high mutation rate when the population variance is low to increase swarm and search diversity. Evaluated using diverse unimodal and multimodal numerical optimisation problems, the proposed model outperforms several baseline GAs with a statistical significance.
Alexander Woodcock, Li Zhang 0013
SMC2
2024 Application of artificial intelligence in cognitive load analysis using functional near-infrared spectroscopy: A systematic review
abstract
Cognitive load theory suggests that overloading of working memory may negatively affect the performance of human in cognitively demanding tasks. Evaluation of cognitive load is a difficult task; it is often assessed through feedback and evaluation from experts. Cognitive load classification based on Functional Near-InfraRed Spectroscopy (fNIRS) is now one of the key research areas in recent years, due to its resistance of artefacts, cost-effectiveness, and portability. To make fNIRS more practical in various applications, it is necessary to develop robust algorithms that can automatically classify fNIRS signals and less reliant on trained signals. Many of the analytical tools used in cognitive sciences have used Deep Learning (DL) modalities to uncover relevant information for mental workload classification. This review investigates the research questions on the design and overall effectiveness of DL as well as its key characteristics. We have identified 38 studies published between 2011 and 2022, that specifically proposed Machine Learning (ML) models for classifying cognitive load using data obtained from fNIRS devices. Those studies were analyzed based on type of feature selection methods, input, and DL model architectures. Most of the existing cognitive load studies are based on ML algorithms, which follow signal filtration and hand-crafted features. It is observed that hybrid DL architectures that integrate convolution and LSTM operators performed significantly better in comparison with other models. However, DL models especially hybrid models have not been extensively investigated for the classification of cognitive load captured by fNIRS devices. The current trends and challenges are highlighted to provide directions for the development of DL models pertaining to fNIRS research.
Mehshan Ahmed Khan, Houshyar Asadi, Li Zhang 0013, Mohammad Reza Chalak Qazani, Sam Oladazimi, Chu Kiong Loo, Chee Peng Lim, Saeid Nahavandi
Expert Syst. Appl.3
2024 Video Deepfake classification using particle swarm optimization-based evolving ensemble models
abstract
The recent breakthrough of deep learning based generative models has led to the escalated generation of photo-realistic synthetic videos with significant visual quality. Automated reliable detection of such forged videos requires the extraction of fine-grained discriminative spatial-temporal cues. To tackle such challenges, we propose weighted and evolving ensemble models comprising 3D Convolutional Neural Networks (CNNs) and CNN-Recurrent Neural Networks (RNNs) with Particle Swarm Optimization (PSO) based network topology and hyper-parameter optimization for video authenticity classification. A new PSO algorithm is proposed, which embeds Muller's method and fixed-point iteration based leader enhancement, reinforcement learning-based optimal search action selection, a petal spiral simulated search mechanism, and cross-breed elite signal generation based on adaptive geometric surfaces. The PSO variant optimizes the RNN topologies in CNN-RNN, as well as key learning configurations of 3D CNNs, with the attempt to extract effective discriminative spatial-temporal cues. Both weighted and evolving ensemble strategies are used for ensemble formulation with aforementioned optimized networks as base classifiers. In particular, the proposed PSO algorithm is used to identify optimal subsets of optimized base networks for dynamic ensemble generation to balance between ensemble complexity and performance. Evaluated using several well-known synthetic video datasets, our approach outperforms existing studies and various ensemble models devised by other search methods with statistical significance for video authenticity classification. The proposed PSO model also illustrates statistical superiority over a number of search methods for solving optimization problems pertaining to a variety of artificial landscapes with diverse geometrical layouts.
Li Zhang 0013, Dezong Zhao, Chee Peng Lim, Houshyar Asadi, Haoqian Huang, Yonghong Yu, Rong Gao 0001
Knowl. Based Syst.1
2024 Video deepfake detection using Particle Swarm Optimization improved deep neural networks
abstract
Abstract As complexity and capabilities of Artificial Intelligence technologies increase, so does its potential for misuse. Deepfake videos are an example. They are created with generative models which produce media that replicates the voices and faces of real people. Deepfake videos may be entertaining, but they may also put privacy and security at risk. A criminal may forge a video of a politician or another notable person in order to affect public opinions or deceive others. Approaches for detecting and protecting against these types of forgery must evolve as well as the methods of generation to ensure that proper information is supplied and to mitigate the risks associated with the fast evolution of deepfakes. This research exploits the effectiveness of deepfake detection algorithms with the application of a Particle Swarm Optimization (PSO) variant for hyperparameter selection. Since Convolutional Neural Networks excel in recognizing objects and patterns in visual data while Recurrent Neural Networks are proficient at handling sequential data, in this research, we propose a hybrid EfficientNet-Gated Recurrent Unit (GRU) network as well as EfficientNet-B0-based transfer learning for video forgery classification. A new PSO algorithm is proposed for hyperparameter search, which incorporates composite leaders and reinforcement learning-based search strategy allocation to mitigate premature convergence. To assess whether an image or a video is manipulated, both models are trained on datasets containing deepfake and genuine photographs and videos. The empirical results indicate that the proposed PSO-based EfficientNet-GRU and EfficientNet-B0 networks outperform the counterparts with manual and optimal learning configurations yielded by other search methods for several deepfake datasets.
Leandro Cunha, Li Zhang 0013, Bilal Sowan, Chee Peng Lim, Yinghui Kong
Neural Comput. Appl.2
2024 Hyperbolic Translation-Based Sequential Recommendation
abstract
The goal of sequential recommendation algorithms is to predict personalized sequential behaviors of users (i.e., next-item recommendation). Learning representations of entities (i.e., users and items) from sparse interaction behaviors and capturing the relationships between entities are the main challenges for sequential recommendation. However, most sequential recommendation algorithms model relationships among entities in Euclidean space, where it is difficult to capture hierarchical relationships among entities. Moreover, most of them utilize independent components to model the user preferences and the sequential behaviors, ignoring the correlation between them. To simultaneously capture the hierarchical structure relationships and model the user preferences and the sequential behaviors in a unified framework, we propose a general hyperbolic translation-based sequential recommendation framework, namely HTSR. Specifically, we first measure the distance between entities in hyperbolic space. Then, we utilize personalized hyperbolic translation operations to model the third-order relationships among a user, his/her latest visited item, and the next item to consume. In addition, we instantiate two hyperbolic translation-based sequential recommendation models, namely Poincaré translation-based sequential recommendation (PoTSR) and Lorentzian translation-based sequential recommendation (LoTSR). PoTSR and LoTSR utilize the Poincaré distance and Lorentzian distance to measure similarities between entities, respectively. Moreover, we utilize the tangent space optimization method to determine optimal model parameters. Experimental results on five real-world datasets show that our proposed hyperbolic translation-based sequential recommendation methods outperform the state-of-the-art sequential recommendation algorithms.
Yonghong Yu, Aoran Zhang 0001, Li Zhang 0013, Rong Gao 0001, Hongzhi Yin
IEEE Trans. Comput. Soc. Syst.3
2024 Neural Inference Search for Multiloss Segmentation Models
abstract
Semantic segmentation is vital for many emerging surveillance applications, but current models cannot be relied upon to meet the required tolerance, particularly in complex tasks that involve multiple classes and varied environments. To improve performance, we propose a novel algorithm, neural inference search (NIS), for hyperparameter optimization pertaining to established deep learning segmentation models in conjunction with a new multiloss function. It incorporates three novel search behaviors, i.e., Maximized Standard Deviation Velocity Prediction, Local Best Velocity Prediction, and n -dimensional Whirlpool Search. The first two behaviors are exploratory, leveraging long short-term memory (LSTM)-convolutional neural network (CNN)-based velocity predictions, while the third employs n -dimensional matrix rotation for local exploitation. A scheduling mechanism is also introduced in NIS to manage the contributions of these three novel search behaviors in stages. NIS optimizes learning and multiloss parameters simultaneously. Compared with state-of-the-art segmentation methods and those optimized with other well-known search algorithms, NIS-optimized models show significant improvements across multiple performance metrics on five segmentation datasets. NIS also reliably yields better solutions as compared with a variety of search methods for solving numerical benchmark functions.
Sam Slade, Li Zhang 0013, Haoqian Huang, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Hanhe Lin, Rong Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Human Action Recognition Using Multi-Stream Fusion and Hybrid Deep Neural Networks
abstract
Action Recognition in videos is a topic of interest in the area of computer vision, due to potential applications such as multimedia indexing and surveillance in public areas. In this research, we first propose spatial and temporal Convolutional Neural Network (CNNs), based on transfer learning using ResNetl0l, GoogleNet and VGG16, for undertaking human action recognition. Besides that, hybrid networks such as CNN-Recurrent Neural Network (RNN) models are also exploited as encoder-decoder architectures for video action classification. In particular, different types of RNNs such as Long Short-Term Memory (LSTM), Bidirectional-LSTM (BiLSTM), Gated Recurrent Unit (GRU), and Bidirectional-GRU (BiGRU), are exploited as the decoders for action recognition. To further enhance performance, diverse aggregation networks of CNN and CNN-RNN models are implemented. Specifically, an Average Fusion method is used to integrate spatial and temporal CNN s trained on images, as well as CNN - RNN trained on videos, where the final classification is formed by combining Softmax scores of these models via a late fusion. A total of 22 models (1 motion CNN, 3 spatial CNNs, 12 CNN-RNNs and 6 fusion networks) are implemented which are evaluated using UCF11, UCFSO, and UCF10l datasets for performance comparison. The empirical results indicate the significant efficiency of Average Fusion of multiple Spatial-CNNs with one Motion-CNN, and ResNet101-BiGRU, among all the networks for undertaking realistic video action recognition.
Saurabh Chopra, Li Zhang 0013, Ming Jiang 0020
SMC2
2023 Mask R-CNN Transfer Learning Variants for Multi-Organ Medical Image Segmentation
abstract
Medical abdomen image segmentation is a challenging task owing to discernible characteristics of the tumour against other organs. As an effective image segmenter, Mask R-CNN has been employed in many medical imaging applications, e.g. for segmenting nucleus from cytoplasm for leukaemia diagnosis and skin lesion segmentation. Motivated by such existing studies, this research takes advantage of the strengths of Mask R-CNN in leveraging on pre-trained CNN architectures such as ResNet and proposes three variants of Mask R-CNN for multi-organ medical image segmentation. Specifically, we propose three variants of the Mask R-CNN transfer learning model successively, each with a set of configurations modified from the one preceding. To be specific, the three variants are (1) the traditional transfer learning with customized loss functions with comparatively more weightage on the segmentation performance, (2) transfer learning based on Mask R-CNN with deepened re-trained layers instead of only the last two/three layers as in traditional transfer learning, and (3) the fine-tuning of Mask R-CNN with expansion of the Region of Interest pooling sizes. Evaluating using Beyond-the-Cranial-Vault (BTCV) abdominal dataset, a well-established benchmark for multi-organ medical image segmentation, the three proposed variants of Mask R-CNN obtain promising performances. In particular, the empirical results indicate the effectiveness of the proposed adapted loss functions, the deepened transfer learning process, as well as the expansion of the RoI pooling sizes. Such variations account for the great efficiency of the proposed transfer learning variant schemes for undertaking multi-organ image segmentation tasks.
Hongjian Lem, Li Zhang 0013
SMC2
2023 Medical Image Classification Using Transfer Learning and Network Pruning Algorithms
abstract
Deep neural networks show great advancement in recent decades in classifying medical images (such as CT-scans) with high precision to aid disease diagnosis. However, the training of deep neural networks requires significant sample sizes for learning enriched discriminative spatial features. Building a high quality dataset large enough to satisfy model training requirement is a challenging task due to limited disease sample cases, and various data privacy constraints. Therefore in this research, we perform medical image classification using transfer learning based on several well-known deep networks, i.e. GoogLeNet, Resnet and EfficientNet. To tackle data sparsity issues, a Wasserstein Generative Adversarial Network (WGAN) is used to generate new medical image samples to increase the numbers of training instances of the minority classes. The transfer learning process itself also allows the building of strong classifiers by transferring knowledge from the pre-trained image domain to a new medical domain using a small sample size. Moreover, the lottery ticket hypothesis is also used to prune each transfer learning network trained using the new target image data sets. Specifically, the L1 norm unstructured pruning technique is used for network reduction. Hyper-parameter fine-tuning is also performed to identify optimal settings of key network hyper-parameters such as learning rate, batch size and weight decay. A total of 20 trials are used for optimal hyper-parameter selection. Evaluated using multi-class lung X-ray images for pneumonia conditions and brain tumor CT-scans, the fine-tuned EfficientNet model obtains the best brain tumor classification accuracy rate of 96% and a fine-tuned GoogLeNet model with pruning has the highest pneumonia classification accuracy rate of 81.5%.
Luca Saleh, Li Zhang 0013
SMC2
2023 Enhanced bare-bones particle swarm optimization based evolving deep neural networks
abstract
In this research, we propose a variant of the Bare-Bones Particle Swarm Optimization (BBPSO) algorithm for hyper-parameter selection and deep architecture generation for image, audio and video classification tasks. Since the search process of the original BBPSO model is guided by a single leader and the particles’ personal best experiences, there is a lack of interactions pertaining to the neighbouring elite solutions. To overcome this limitation, we propose a versatile search process for a modified BBPSO model that incorporates a number of effective components and operations. These include the neighbouring and global best signals, search actions with Cauchy/Levy scale factors, sub-dimension operations guided by the local and global elite solutions, and a Levy-driven local search mechanism. Moreover, root-finding algorithms are employed which use informative mathematical principles to estimate new root offspring for leader/particle enhancement. A reinforcement learning algorithm is subsequently used to identify the optimal sequential deployment of these numerical analysis methods to increase robustness. Several medical imaging data sets, i.e., ISIC 2017, PH2 and Dermofit skin lesion databases, the ALL-IDB2 microscopic blood image data set, the MURA musculoskeletal radiographic database, the CK + facial expression data set, as well as the Coswara respiratory audio data set and UCF101 video action data set, are employed for evaluation. The proposed BBPSO-optimized Convolutional Neural Network (CNN), bidirectional Long Short-Term Memory (BiLSTM) with attention mechanism, and CNN-BiLSTM models outperform those devised by other PSO and BBPSO variants, as well as state-of-the-art existing studies, significantly, for image, audio respiratory abnormality and realistic video action recognition.
Li Zhang 0013, Chee Peng Lim, Chengyu Liu 0001
Expert Syst. Appl.1
2023 Semantic segmentation using Firefly Algorithm-based evolving ensemble deep neural networks
abstract
Automatic segmentation of salient objects in real-world images has gained increasing interests owing to its popularity in diverse real-world applications, such as autonomous driving, medical diagnosis, aviation security, and underwater surveillance. In this research, we propose Firefly Algorithm (FA)-enhanced evolving ensemble deep networks for semantic segmentation and visual saliency prediction. An improved FA model is proposed to optimize network hyper-parameters. Specifically, it employs mutation operators and a neighbouring search strategy with granular search steps to establish search intensification. It also emphasizes search diversification by adopting multiple dynamic hybrid leaders and diverse adaptive sine and cosine search trajectories in full and randomly selected sub-dimensions to overcome stagnation. Because of its competent segmentation performance, DeepLabV3+ is fine-tuned using transfer learning with FA-based hyper-parameter identification. We optimize the learning rate, momentum and weight decay of the transfer learning network. A number of optimized DeepLabV3+ networks with distinguishing learning configurations are yielded. An ensemble model is subsequently constructed by incorporating three optimized base networks to further strengthen segmentation performance. Evaluated using diverse challenging semantic segmentation and saliency prediction tasks using underwater and medical image data sets, our evolving ensemble deep network illustrates significant superiority over other state-of-the-art deep networks and existing studies. The proposed FA model also outperforms other search methods in solving diverse mathematical landscapes with statistical significance.
Li Zhang 0013, Sam Slade, Chee Peng Lim, Houshyar Asadi, Saeid Nahavandi, Haoqian Huang
Knowl. Based Syst.1
2023 Dynamics simulation-based deep residual neural networks to detect flexible shafting faults
Haimin Zhu, Qingzhang Chen, Li Zhang 0013, Miaomiao Li 0002, Rupeng Zhu
Knowl. Based Syst.3
2023 Hybrid PSO feature selection-based association classification approach for breast cancer detection
Bilal Sowan, Mohammed Eshtay, Keshav P. Dahal, Hazem Qattous, Li Zhang 0013
Neural Comput. Appl.5
2023 Personalized tag recommendation via denoising auto-encoder
Weibin Zhao, Lin Shang 0001, Yonghong Yu, Li Zhang 0013, Can Wang 0004, Jiajun Chen 0001
World Wide Web (WWW)4
2022 Hyperbolic Personalized Tag Recommendation
Weibin Zhao, Aoran Zhang 0001, Lin Shang 0001, Yonghong Yu, Li Zhang 0013, Can Wang 0004, Jiajun Chen 0001, Hongzhi Yin
DASFAA (2)5
2022 Deep Learning Based Short-Term Total Cloud Cover Forecasting
abstract
In this research, we conduct deep learning based Total Cloud Cover (TCC) forecasting using satellite images. The proposed system employs the Otsu's method for cloud segmentation and Long Short-Term Memory (LSTM) variant models for TCC prediction. Specifically, a region-based Otsu's method is used to segment clouds from satellite images. A time-series dataset is generated using the TCC information extracted from each image in image sequences using a new feature extraction method. The generated time series data are subsequently used to train several LSTM variant models, i.e. LSTM, bi-directional LSTM and Convolutional Neural Network (CNN)-LSTM, for future TCC forecasting. Our approach achieves impressive average RMSE scores with multi-step forecasting, i.e. 0.0543 and 0.0823, with respect to both the first half of daytime and full daytime TCC forecasting on a given day, using the generated dataset.
Ishara Bandara, Li Zhang 0013, Kamlesh Mistry
IJCNN2
2022 Failure Mode Identification of Elastomer for Well Completion Systems using Mask R-CNN
abstract
Sealing mechanisms are at the heart of any drilling, completion, or production system and are the primary components on which the functional success and longevity of the system rest. Element extrusion and cracking are the earliest failure modes to indicate leakage on the elastomers. With the trend towards higher operating temperatures and pressures of the well completion systems, more accurate and reliable methods are needed to identify the failure modes of the elastomer. Currently, elastomer failure type tests are carried out by a trained human supervisor. This method is time consuming and inefficient in inspecting a large quantity of elastomers. Furthermore, this manual approach depends on the human examiner's knowledge and experience, which is subjective and may be prone to errors. To overcome these issues, we employ Mask R-CNN with a modified loss function to identify the failure types and localize failure regions of elastomers of the well completion systems. To train Mask R-CNN for crack and extrusion detection, a total of 70 elastomer images with different types of failure conditions are collected from real-world environments by the oil and gas company. The VGG annotator is used to annotate the labels of the elastomer failure types and positions with polygonal regions in the images. Diverse image augmentation techniques are also used to further enhance performance. The mean average precision results are used to indicate model performance with different datasets. The empirical results indicate the robustness and efficiency of Mask R-CNN in performing failure detection and segmentation of elastomers.
Weihang Chen, Li Zhang 0013, Ming Jiang 0020
IJCNN2
2022 Human Action Recognition Using Hybrid Deep Evolving Neural Networks
abstract
Human action recognition can be applied in a multitude of fully diversified domains such as active large-scale surveillance, threat detection, personal safety in hazardous environments, human assistance, health monitoring, and intelligent robotics. Owing to its high demands in real-world applications, it has drawn significant attention. In this research, we propose hybrid deep neural networks, i.e. Convolutional Long Short-Term Memory (ConvLSTM) Networks, Long-term Recurrent Convolutional Networks (LRCN), for tackling video action classification. In particular, for the LRCN model, different CNN encoder architectures such as VGG16, ResNet50, DenseNet121 and MobileNet, as well as several Long Short-Term Memory (LSTM) variant decoder architectures, such as LSTM, bidirectional LSTM (BiLSTM) and Gated Recurrent Unit (GRU), are used for spatial-temporal feature extraction to test model performance. We adopt diverse experimental settings including using different numbers of frames per video and learning configurations to optimize performance. The empirical results indicate the superiority of MobileNet in combination with a BiLSTM network over other hybrid network settings for the action classification using the UCF50 dataset. Owing to the lightweight MobileNet encoder, this LRCN model also achieves a better trade-off between performance and training and inference computational costs, while outperforming existing state-of-the-art methods.
Pavan Dasari, Li Zhang 0013, Yonghong Yu, Haoqian Huang, Rong Gao 0001
IJCNN2
2022 An evolving ensemble model of multi-stream convolutional neural networks for human action recognition in still images
abstract
Abstract Still image human action recognition (HAR) is a challenging problem owing to limited sources of information and large intra-class and small inter-class variations which requires highly discriminative features. Transfer learning offers the necessary capabilities in producing such features by preserving prior knowledge while learning new representations. However, optimally identifying dynamic numbers of re-trainable layers in the transfer learning process poses a challenge. In this study, we aim to automate the process of optimal configuration identification. Specifically, we propose a novel particle swarm optimisation (PSO) variant, denoted as EnvPSO, for optimal hyper-parameter selection in the transfer learning process with respect to HAR tasks with still images. It incorporates Gaussian fitness surface prediction and exponential search coefficients to overcome stagnation. It optimises the learning rate, batch size, and number of re-trained layers of a pre-trained convolutional neural network (CNN). To overcome bias of single optimised networks, an ensemble model with three optimised CNN streams is introduced. The first and second streams employ raw images and segmentation masks yielded by mask R-CNN as inputs, while the third stream fuses a pair of networks with raw image and saliency maps as inputs, respectively. The final prediction results are obtained by computing the average of class predictions from all three streams. By leveraging differences between learned representations within optimised streams, our ensemble model outperforms counterparts devised by PSO and other state-of-the-art methods for HAR. In addition, evaluated using diverse artificial landscape functions, EnvPSO performs better than other search methods with statistically significant difference in performance.
Sam Slade, Li Zhang 0013, Yonghong Yu, Chee Peng Lim
Neural Comput. Appl.2
2022 A New Prepositioning Technique of a Motion Simulator Platform Using Nonlinear Model Predictive Control and Recurrent Neural Network
abstract
The motion cueing algorithm (MCA) is the main algorithm in motion simulators in charge of generating vehicle motions within the platform’s constraints. The classical washout filter is one of the popular types of MCA, which is used in air and land vehicle motion simulators. The fixed home position of the simulator platform is always cogitated in the MCA to washout the motion simulator after generating each motion. Unfortunately, considering the fixed home position reduces the efficient consumption of the workspace in the linear directions. The linear motion of the motion simulator is due to the production of the high-pass frequency part of the motion scenarios. Prepositioning is used to tackle this assumption by varying the home position rather than the fixed position. The linear motion limitations of the motion simulator can virtually be enlarged using the prepositioning method. The efficient regeneration of the high-pass motion cues using a new propositioning technique is the main goal of this study to increase the motion realism of the simulator and remove any false motion cues due to the platform limitations. The proposed model utilised the recurrent neural network (RNN) to estimate the motion scenario along the prediction horizon. The nonlinear model predictive control (MPC) uses the estimated motion signals to extract the best optimal off-centre position of the motion simulator platform. The newly developed prepositioning technique is developed in the simulation environment of MATLAB to validate the proposed technique in terms of efficiency and applicability. The outcomes prove the capability of the proposed technique against the recently developed prepositioning technique using fuzzy logic and RNN.
Mohammad Reza Chalak Qazani, Houshyar Asadi, Li Zhang 0013, Farzin Tabarsinezhad, Shady M. K. Mohamed, Chee Peng Lim, Saeid Nahavandi
IEEE Trans. Intell. Transp. Syst.3
2021 An Interactive Evolution Strategy based Deep Convolutional Generative Adversarial Network for 2D Video Game Level Procedural Content Generation
abstract
The generation of desirable video game contents has been a challenge of games level design and production. In this research, we propose a game player flow experience driven interactive latent variable evolution strategy incorporated with a Deep Convolutional Generative Adversarial Network (DCGAN) for undertaking game content generation with respect to a 2D Super Mario video game. Since the Generative Adversarial Network (GAN) models tend to capture the high-level style of the input images by learning the latent vectors, they are used to generate game scenarios and context images in this research. However, as GANs employ arbitrary inputs for game image generation without taking specific features into account, they generate game level images in an incoherent manner without the specific playable game level properties, such as a broken pipe in the Mario game level image. In order to overcome such drawbacks, we propose a game player flow experience driven optimised mechanism with human intervention, to guide the game level content generation process so that only plausible and even enjoyable images will be generated as the candidates for the final game design and production.
Ming Jiang 0020, Li Zhang 0013
IJCNN2
2021 Deep Recurrent Neural Networks with Attention Mechanisms for Respiratory Anomaly Classification
abstract
In recent years, a variety of deep learning techniques and methods have been adopted to provide AI solutions to issues within the medical field, with one specific area being audio-based classification of medical datasets. This research aims to create a novel deep learning architecture for this purpose, with a variety of different layer structures implemented for undertaking audio classification. Specifically, bidirectional Long Short-Term Memory (BiLSTM) and Gated Recurrent Units (GRU) networks in conjunction with an attention mechanism, are implemented in this research for chronic and non-chronic lung disease and COVID-19 diagnosis. We employ two audio datasets, i.e. the Respiratory Sound and the Coswara datasets, to evaluate the proposed model architectures pertaining to lung disease classification. The Respiratory Sound Database contains audio data with respect to lung conditions such as Chronic Obstructive Pulmonary Disease (COPD) and asthma, while the Coswara dataset contains coughing audio samples associated with COVID-19. After a comprehensive evaluation and experimentation process, as the most performant architecture, the proposed attention BiLSTM network (A-BiLSTM) achieves accuracy rates of 96.2% and 96.8% for the Respiratory Sound and the Coswara datasets, respectively. Our research indicates that the implementation of the BiLSTM and attention mechanism was effective in improving performance for undertaking audio classification with respect to various lung condition diagnoses.
Conor Wall, Li Zhang 0013, Yonghong Yu, Kamlesh Mistry
IJCNN2
2021 Intelligent human action recognition using an ensemble model of evolving deep networks with swarm-based optimization
Li Zhang 0013, Chee Peng Lim, Yonghong Yu
Knowl. Based Syst.1
2020 Neural Pairwise Ranking Factorization Machine for Item Recommendation
Lihong Jiao, Yonghong Yu, Ningning Zhou, Li Zhang 0013, Hongzhi Yin
DASFAA (1)4
2020 Graph Neural Networks Boosted Personalized Tag Recommendation Algorithm
abstract
Personalized tag recommender systems recommend a set of tags for items based on users' historical behaviors, and play an important role in the collaborative tagging systems. However, traditional personalized tag recommendation methods cannot guarantee that the collaborative signal hidden in the interactions among entities is effectively encoded in the process of learning the representations of entities, resulting in insufficient expressive capacity for characterizing the preferences or attributes of entities. In this paper, we proposed a graph neural networks boosted personalized tag recommendation model, which integrates the graph neural networks into the pairwise interaction tensor factorization model. Specifically, we consider two types of interaction graph (i.e. the user-tag interaction graph and the item-tag interaction graph) that is derived from the tag assignments. For each interaction graph, we exploit the graph neural networks to capture the collaborative signal that is encoded in the interaction graph and integrate the collaborative signal into the learning of representations of entities by transmitting and assembling the representations of entity neighbors along the interaction graphs. In this way, we explicitly capture the collaborative signal, resulting in rich and meaningful representations of entities. Experimental results on real world datasets show that our proposed graph neural networks boosted personalized tag recommendation model outperforms the traditional tag recommendation models.
Yonghong Yu, Fengyixin Jiang, Li Zhang 0013, Rong Gao 0001, Haiyan Gao
IJCNN4
2020 A Multi-Population FA for Automatic Facial Emotion Recognition
abstract
Automatic facial emotion recognition system is popular in various domains such as health care, surveillance and human-robot interaction. In this paper we present a novel multi-population FA for automatic facial emotion recognition. The overall system is equipped with horizontal vertical neighborhood local binary patterns (hvnLBP) for feature extraction, a novel multi-population FA for feature selection and diverse classifiers for emotion recognition. First, we extract features using hvnLBP, which are robust to illumination changes, scaling and rotation variations. Then, a novel FA variant is proposed to further select most important and emotion specific features. These selected features are used as input to the classifier to further classify seven basic emotions. The proposed system is evaluated with multiple facial expression datasets and also compared with other state-of-the-art models.
Kamlesh Mistry, Baqar Rizvi, Chris Rook, Sadaf Iqbal, Li Zhang 0013, Colin Paul Joy
IJCNN5
2020 Adaptive melanoma diagnosis using evolving clustering, ensemble and deep neural networks
Teck Yan Tan, Li Zhang 0013, Chee Peng Lim
Knowl. Based Syst.2
2020 Enhanced factorization machine via neural pairwise ranking and attention networks
abstract
The factorization machine models attract significant attention nowadays since they improve recommendation performance by incorporating context information into recommendation modeling. However, traditional factorization machine models often adopt the point-wise learning method for model parameter learning, as well as only model the linear interactions between features. They substantially fail to capture the complex interactions among features, which degrades the performance of factorization machine models. In this research, we propose a neural pairwise ranking factorization machine for item recommendation, namely NPRFM, which integrates the multi-layer perceptual neural networks into the pairwise ranking factorization machine model. Specifically, to capture the high-order and nonlinear interactions among features, we stack a multi-layer perceptual neural network over the bi-interaction layer, which encodes the second-order interactions between features. Moreover, instead of the prediction of the absolute scores, the pair-wise ranking model is adopted to learn the relative preferences of users. Since NPRFM does not take into account the importance of feature interactions, we propose a new variant of NPRFM, which learns the importance of feature interactions by introducing the attention mechanism . The empirical results on real-world datasets indicate that the proposed neural pairwise ranking factorization machine outperforms the traditional factorization machine models.
Yonghong Yu, Lihong Jiao, Ningning Zhou, Li Zhang 0013, Hongzhi Yin
Pattern Recognit. Lett.4
2019 Distant Pedestrian Detection in the Wild using Single Shot Detector with Deep Convolutional Generative Adversarial Networks
abstract
In this work, we examine the feasibility of applying Deep Convolutional Generative Adversarial Networks (DCGANs) with Single Shot Detector (SSD) as data-processing technique to handle with the challenge of pedestrian detection in the wild. Specifically, we attempted to use in-fill completion to generate random transformations of images with missing pixels to expand existing labelled datasets. In our work, GAN’s been trained intensively on low resolution images, in order to neutralize the challenges of the pedestrian detection in the wild, and considered humans, and few other classes for detection in smart cities. The object detector experiment performed by training GAN model along with SSD provided a substantial improvement in the results. This approach presents a very interesting overview in the current state of art on GAN networks for object detection. We used Canadian Institute for Advanced Research (CIFAR), Caltech, KITTI data set for training and testing the network under different resolutions and the experimental results with comparison been showed between DCGAN cascaded with SSD and SSD itself.
Ranjith Dinakaran, Philip Easom, Li Zhang 0013, Ahmed Bouridane, Richard Jiang 0001, Eran A. Edirisinghe
IJCNN3
2019 Evolving and Ensembling Deep CNN Architectures for Image Classification
abstract
Deep convolutional neural networks (CNNs) have traditionally been hand-designed owing to the complexity of their construction and the computational requirements of their training. Recently however, there has been an increase in research interest towards automatically designing deep CNNs for specific tasks. Ensembling has been shown to effectively increase the performance of deep CNNs, although usually with a duplication of work and therefore a large increase in computational resources required. In this paper we present Swarm Optimised Block Architecture Ensembles (SOBAE), a method for automatically designing and ensembling deep CNN models with a central weight repository to avoid work duplication. The models are trained and optimised together using particle swarm optimisation (PSO), with architecture convergence encouraged. At the conclusion of this combined process a base model nomination method is used to determine the best candidates for the ensemble. Two base model nomination methods are proposed, one using the local best particle positions from the PSO process, and one using the contents of the central weight repository. Once the base model pool has been created, the models inherit their parameters from the central weight repository and are then finetuned and ensembled in order to create a final system. We evaluate our system on the CIFAR-10 classification dataset and demonstrate improved results over the single global best model suggested by the optimisation process, with a minor increase in resources required by the finetuning process. Our system achieves an error rate of 4.27% on the CIFAR-10 image classification task with only 36 hours of combined optimisation and training on a single NVIDIA GTX 1080Ti GPU.
Ben Fielding, Tom Lawrence, Li Zhang 0013
IJCNN3
2019 Integrating Social Circles and Network Representation Learning for Item Recommendation
abstract
With the ever increasing popularity of social network services, social network platforms provide rich and additional information for recommendation algorithms. More and more researchers utilize the trust relationships of users to improve the performance of recommendation algorithms. However, most of the existing social-network-based recommendation algorithms ignore the following problems: (1) In different domains, users tend to trust different friends. (2) the performance of recommendation algorithms is limited by the coarse-grained trust relationships. In this paper, we propose a novel recommendation algorithm that integrates the social circles and the network representation learning for item recommendation. Specifically, we firstly infer the domain-specific social trust circles based on the original users’ rating information and the social network information. Next, we adopt the network representation technique to embed the domain-specific social trust circle into a low-dimensional space, and then utilize the low-dimensional representations of users to infer the fine-grained trust relationships between users. Finally, we integrate the fine-gained trust relationships with the domain-specific matrix factorization model to learn the latent user and item feature vectors. Experimental results on real-world datasets show that our proposed approach outperforms the traditional social-network-based recommendation algorithms.
Yonghong Yu, Li Zhang 0013, Can Wang 0004, Boyu Qi
IJCNN3
2019 Advancing Ensemble Learning Performance through data transformation and classifiers fusion in granular computing context
Han Liu 0002, Li Zhang 0013
Expert Syst. Appl.2
2019 Guest editorial: Automatic facial and bodily expression perception for human behaviour understanding
Li Zhang 0013, Chee Peng Lim, Jungong Han
Multim. Tools Appl.1
2019 A hierarchical and regional deep learning architecture for image description generation
Philip Kinghorn, Li Zhang 0013, Ling Shao 0001
Pattern Recognit. Lett.2
2018 Multi-Population Differential Evolution for Retinal Blood Vessel Segmentation
abstract
The retinal blood vessel segmentation plays a significant role in the automatic or computer-assisted diagnosis of retinopathy. Manual blood vessel segmentation is very time-consuming and requires a great amount of domain knowledge. In addition, the blood vessels are only a few pixels wide and cover the entire fundus image. This further hinders the recent systems from automating the retinal blood vessel segmentation efficiently. In this paper, we propose a modified differential evolution (DE) algorithm to carry out automatic retinal blood vessel segmentation. The modified DE employs cross-communication among multiple populations to select three types of features i.e. thick blood vessels, thin blood vessels and non-blood vessels. Multiple classifiers such as neural networks (NN), Support vector machines (SVM), NN based and SVM based ensembles are used to further measure the performance of segmentation. The proposed algorithm is evaluated on three publicly available retinal image datasets like DRIVE, STARE and HRF. It outperformed the state-of-the-art with a high average accuracy of 98.5% along with high sensitivity and specificity.
Kamlesh Mistry, Biju Issac 0001, Seibu Mary Jacob, Jyoti Jasekar, Li Zhang 0013
ICARCV5
2018 Extended LBP based Facial Expression Recognition System for Adaptive AI Agent Behaviour
abstract
Automatic facial expression recognition is widely used for various applications such as health care, surveillance and human-robot interaction. In this paper, we present a novel system which employs automatic facial emotion recognition technique for adaptive AI agent behaviour. The proposed system is equipped with kirsch operator based local binary patterns for feature extraction and diverse classifiers for emotion recognition. First, we nominate a novel variant of the local binary pattern (LBP) for feature extraction to deal with illumination changes, scaling and rotation variations. The features extracted are then used as input to the classifier for recognizing seven emotions. The detected emotion is then used to enhance the behaviour selection of the artificial intelligence (AI) agents in a shooter game. The proposed system is evaluated with multiple facial expression datasets and outperformed other state-of-the-art models by a significant margin.
Kamlesh Mistry, Jyoti Jasekar, Biju Issac 0001, Li Zhang 0013
IJCNN4
2018 Topological Evolution of Spiking Neural Networks
abstract
Neuro-evolution is often used to generate the parameters, topology, and rules of artificial neural networks. This technique allows for automatic configuration of a neural network. In this paper we propose a method to generate Spiking Neural Networks (SNNs) automatically called NENG (Neuro-Evolutionary Network Generation). The aim was to help alleviate the manual construction and optimization of neural network implementations. The results show the algorithm is successful at generating and improving the design of SNNs for a Classification task. After 812 generations with a population size of 20 the algorithm converges to model the Xor gate with 100% accuracy. The results show improvements to the algorithm execution time and number of neurons over time.
Sam Slade, Li Zhang 0013
IJCNN2
2018 Geographical Proximity Boosted Recommendation Algorithms for Real Estate
Yonghong Yu, Can Wang 0004, Li Zhang 0013, Rong Gao 0001, Hua Wang 0002
WISE (2)3
2018 Detection of online phishing email using dynamic evolving neural network based on reinforcement learning
Sami Smadi, Nauman Aslam, Li Zhang 0013
Decis. Support Syst.3
2018 Feature selection using firefly optimization for classification and regression models
Li Zhang 0013, Kamlesh Mistry, Chee Peng Lim, Siew Chin Neoh
Decis. Support Syst.1
2018 A quantitative image analysis for the cellular cytoskeleton during in vitro tumor growth
M. Abdulla Al Mamun, Worawut Srisukkham, Dewan Md. Farid 0001, Lorna Ravenhill, Li Zhang 0013, M. Alamgir Hossain, Rosemary Bass
Expert Syst. Appl.5
2018 Classifier ensemble reduction using a modified firefly algorithm: An empirical evaluation
Li Zhang 0013, Worawut Srisukkham, Siew Chin Neoh, Chee Peng Lim, Diptangshu Pandit
Expert Syst. Appl.1
2018 A region-based image caption generator with refined descriptions
Philip Kinghorn, Li Zhang 0013, Ling Shao 0001
Neurocomputing2
2018 A scattering and repulsive swarm intelligence algorithm for solving global optimization problems
Diptangshu Pandit, Li Zhang 0013, Samiran Chattopadhyay, Chee Peng Lim, Chengyu Liu 0001
Knowl. Based Syst.2
2018 Intelligent skin cancer detection using enhanced particle swarm optimization
Teck Yan Tan, Li Zhang 0013, Siew Chin Neoh, Chee Peng Lim
Knowl. Based Syst.2
2018 A P2P Botnet detection scheme based on decision tree and adaptive multilayer neural networks
abstract
In recent years, Botnets have been adopted as a popular method to carry and spread many malicious codes on the Internet. These malicious codes pave the way to execute many fraudulent activities including spam mail, distributed denial-of-service attacks and click fraud. While many Botnets are set up using centralized communication architecture, the peer-to-peer (P2P) Botnets can adopt a decentralized architecture using an overlay network for exchanging command and control data making their detection even more difficult. This work presents a method of P2P Bot detection based on an adaptive multilayer feed-forward neural network in cooperation with decision trees. A classification and regression tree is applied as a feature selection technique to select relevant features. With these features, a multilayer feed-forward neural network training model is created using a resilient back-propagation learning algorithm. A comparison of feature set selection based on the decision tree, principal component analysis and the ReliefF algorithm indicated that the neural network model with features selection based on decision tree has a better identification accuracy along with lower rates of false positives. The usefulness of the proposed approach is demonstrated by conducting experiments on real network traffic datasets. In these experiments, an average detection rate of 99.08 % with false positive rate of 0.75 % was observed.
Mohammad Alauthman, Nauman Aslam, Li Zhang 0013, Rafe Alasem, M. Alamgir Hossain
Neural Comput. Appl.3
2017 Facial expression recongition using firefly-based feature optimization
abstract
Automatic facial expression recognition plays an important role in various application domains such as medical imaging, surveillance and human-robot interaction. This research proposes a novel facial expression recognition system with modified Local Gabor Binary Patterns (LGBP) for feature extraction and a firefly algorithm (FA) variant for feature optimization. First of all, in order to deal with illumination changes, scaling differences and rotation variations, we propose an extended overlap LGBP to extract initial discriminative facial features. Then a modified FA is proposed to reduce the dimensionality of the extracted facial features. This FA variant employs Gaussian, Cauchy and Levy distributions to further mutate the best solution identified by the FA to increase exploration in the search space to avoid premature convergence. The overall system is evaluated using three facial expression databases (i.e. CK+, MMI, and JAFFE). The proposed system outperforms other heuristic search algorithms such as Genetic Algorithm and Particle Swarm Optimization and other existing state-of-the-art facial expression recognition research, significantly.
Kamlesh Mistry, Li Zhang 0013, Graham Sexton, Yifeng Zeng, Mengda He
CEC2
2017 Semi-supervised vision-language mapping via variational learning
abstract
Understanding the semantic relations between vision and language data has become a research trend in artificial intelligence and robotic systems. The lack of training data is an essential issue for vision-language understanding. We address the problem of image and sentence cross-modal retrieval when paired training samples are not sufficient. Inspired by recent works in variational inference, in this paper, the autoencoding variational Bayes framework is novelly extended to a semi-supervised model for image-sentence mapping task. Our method does not require all training images and sentences to be paired. The proposed model is an end-to-end system, and consists of a two-level variational embedding structure where unpaired data are involved in the first level embedding to give support to intra-modality statistics so that the lower bound of the joint marginal likelihood of paired data embeddings can be better approximated. The proposed retrieval model is evaluated on two popular datasets, i.e. Flickr30K and Flickr8K, producing superior performances compared with related state-of-the-art methods.
Yuming Shen, Li Zhang 0013, Ling Shao 0001
ICRA2
2017 Deep learning based image description generation
abstract
Describing the contents of images is a challenging task for machines to achieve. It requires not only accurate recognition of objects and humans, but also their attributes and relationships as well as scene information. It would be even more challenging to extend this process to identify falls and hazardous objects to aid elderly or users in need of care. This research makes initial attempts to deal with the above challenges to produce multi-sentence natural language description of image contents. It employs a local region based approach to extract regional image details and combines multiple techniques including deep learning and attribute learning through the use of machine learned features to create high level labels that can generate detailed description of real-world images. The system contains the core functions of scene classification, object detection and classification, attribute learning, relationship detection and sentence generation. We have also further extended this process to deal with open-ended fall detection and hazard identification. In comparison to state-of-the-art related research, our system shows superior robustness and flexibility in dealing with test images from new, unrelated domains, which poses great challenges to many existing methods. Our system is evaluated on a subset from Flickr8k and Pascal VOC 2012 and achieves an impressive average BLEU score of 46 and outperforms related research by a significant margin of 10 BLEU score when evaluated with a small dataset of images containing falls and hazardous objects. It also shows impressive performance when evaluated using a subset of IAPR TC-12 dataset.
Philip Kinghorn, Li Zhang 0013, Ling Shao 0001
IJCNN2
2017 Differential evolution algorithm as a tool for optimal feature subset selection in motor imagery EEG
Muhammad Zeeshan Baig, Nauman Aslam, Hubert P. H. Shum, Li Zhang 0013
Expert Syst. Appl.4
2017 A Micro-GA Embedded PSO Feature Selection Approach to Intelligent Facial Emotion Recognition
abstract
This paper proposes a facial expression recognition system using evolutionary particle swarm optimization (PSO)-based feature optimization. The system first employs modified local binary patterns, which conduct horizontal and vertical neighborhood pixel comparison, to generate a discriminative initial facial representation. Then, a PSO variant embedded with the concept of a micro genetic algorithm (mGA), called mGA-embedded PSO, is proposed to perform feature optimization. It incorporates a nonreplaceable memory, a small-population secondary swarm, a new velocity updating strategy, a subdimension-based in-depth local facial feature search, and a cooperation of local exploitation and global exploration search mechanism to mitigate the premature convergence problem of conventional PSO. Multiple classifiers are used for recognizing seven facial expressions. Based on a comprehensive study using within- and cross-domain images from the extended Cohn Kanade and MMI benchmark databases, respectively, the empirical results indicate that our proposed system outperforms other state-of-the-art PSO variants, conventional PSO, classical GA, and other related facial expression recognition models reported in the literature by a significant margin.
Kamlesh Mistry, Li Zhang 0013, Siew Chin Neoh, Chee Peng Lim, Ben Fielding
IEEE Trans. Cybern.2
2016 An Enhanced Intelligent Agent with Image Description Generation
Ben Fielding, Philip Kinghorn, Kamlesh Mistry, Li Zhang 0013
IVA4
2016 Intelligent facial emotion recognition using moth-firefly optimization
abstract
In this research, we propose a facial expression recognition system with a variant of evolutionary firefly algorithm for feature optimization. First of all, a modified Local Binary Pattern descriptor is proposed to produce an initial discriminative face representation. A variant of the firefly algorithm is proposed to perform feature optimization. The proposed evolutionary firefly algorithm exploits the spiral search behaviour of moths and attractiveness search actions of fireflies to mitigate premature convergence of the Levy-flight firefly algorithm (LFA) and the moth-flame optimization (MFO) algorithm. Specifically, it employs the logarithmic spiral search capability of the moths to increase local exploitation of the fireflies, whereas in comparison with the flames in MFO, the fireflies not only represent the best solutions identified by the moths but also act as the search agents guided by the attractiveness function to increase global exploration. Simulated Annealing embedded with Levy flights is also used to increase exploitation of the most promising solution. Diverse single and ensemble classifiers are implemented for the recognition of seven expressions. Evaluated with frontal-view images extracted from CK+, JAFFE, and MMI, and 45-degree multi-view and 90-degree side-view images from BU-3DFE and MMI, respectively, our system achieves a superior performance, and outperforms other state-of-the-art feature optimization methods and related facial expression recognition models by a significant margin.
Li Zhang 0013, Kamlesh Mistry, Siew Chin Neoh, Chee Peng Lim
Knowl. Based Syst.1
2015 Intelligent facial expression recognition with adaptive feature extraction for a humanoid robot
abstract
Automatic facial expression recognition plays an important role in agent-based interface development and datadriven animation. This paper presents an intelligent facial action and emotion recognition system for a humanoid robot. Motivated by the Facial Action Coding System, this research focuses on the recognition of seven basic emotions and 18 Action Units (AU). Since effective facial representations of original face images are vital for automatic facial emotion recognition, this research implements a novel shape and appearance feature extraction method, which integrates an Independent Active Appearance Model (AAM) with a rotation-invariant feature point detector, BRISK (Binary Robust Invariant Scalable Keypoints). In comparison to AAM with a traditional inverse compositional fitting, our model with BRISK fitting is with less computational cost and is capable of dealing with feature extraction from images of faces with rotations and scaling differences without prior training required. Subsequently shape and appearancebased neural network AU analyzers are used to respectively detect 18 AUs. Emotions are then decoded from the derived AUs using a neural network emotion recognizer. The system is integrated with a modern humanoid robot platform. Evaluation results indicate its high accuracy for AU and emotion recognition. It is also among the top performers on the extended Cohn-Kanade (CK+) database in comparison to other existing state-of-the-art applications.
Kamlesh Mistry, Li Zhang 0013, John A. Barnden
IJCNN2
2015 Adaptive facial point detection and emotion recognition for a humanoid robot
Li Zhang 0013, Kamlesh Mistry, Ming Jiang 0020, Siew Chin Neoh, M. Alamgir Hossain
Comput. Vis. Image Underst.1
2015 Adaptive 3D facial action intensity estimation and emotion recognition
Yang Zhang 0002, Li Zhang 0013, M. Alamgir Hossain
Expert Syst. Appl.2
2015 Intelligent affect regression for bodily expressions using hybrid particle swarm optimization and adaptive ensembles
Yang Zhang 0002, Li Zhang 0013, Siew Chin Neoh, Kamlesh Mistry, M. Alamgir Hossain
Expert Syst. Appl.2
2014 Intelligent Facial Action and emotion recognition for humanoid robots
abstract
This research focuses on the development of a realtime intelligent facial emotion recognition system for a humanoid robot. In our system, Facial Action Coding System is used to guide the automatic analysis of emotional facial behaviours. The work includes both an upper and a lower facial Action Units (AU) analyser. The upper facial analyser is able to recognise six AUs including Inner and Outer Brow Raiser, Upper Lid Raiser etc, while the lower facial analyser is able to detect eleven AUs including Upper Lip Raiser, Lip Corner Puller, Chin Raiser, etc. Both of the upper and lower analysers are implemented using feedforward Neural Networks (NN). The work also further decodes six basic emotions from the recognised AUs. Two types of facial emotion recognisers are implemented, NN-based and multi-class Support Vector Machine (SVM) based. The NN-based facial emotion recogniser with the above recognised AUs as inputs performs robustly and efficiently. The Multi-class SVM with the radial basis function kernel enables the robot to outperform the NN-based emotion recogniser in real-time posed facial emotion detection tasks for diverse testing subjects.
Li Zhang 0013, M. Alamgir Hossain, Ming Jiang 0020
IJCNN1
2014 An intelligent mobile based decision support system for retinal disease diagnosis
Abderrahim Bourouis, Mohammed Feham, M. Alamgir Hossain, Li Zhang 0013
Decis. Support Syst.4
2014 Hybrid decision tree and naïve Bayes classifiers for multi-class classification tasks
Dewan Md. Farid 0001, Li Zhang 0013, Chowdhury Mofizur Rahman, M. Alamgir Hossain, Rebecca Strachan
Expert Syst. Appl.2
2013 A real-time dynamic optimal guidance scheme using a general regression neural network
M. Alamgir Hossain, A. A. Madkour, Keshav P. Dahal, Li Zhang 0013
Eng. Appl. Artif. Intell.4
2013 An adaptive ensemble classifier for mining concept drifting data streams
Dewan Md. Farid 0001, Li Zhang 0013, M. Alamgir Hossain, Chowdhury Mofizur Rahman, Rebecca Strachan, Graham Sexton, Keshav P. Dahal
Expert Syst. Appl.2
2013 Fuzzy association rule mining approaches for enhancing prediction performance
Bilal Sowan, Keshav P. Dahal, M. Alamgir Hossain, Li Zhang 0013, Linda Spencer
Expert Syst. Appl.4
2013 Intelligent facial emotion recognition and semantic-based topic detection for a humanoid robot
Li Zhang 0013, Ming Jiang 0020, Dewan Md. Farid 0001, M. Alamgir Hossain
Expert Syst. Appl.1
2012 Exploration of Affect Detection Using Semantic Cues in Virtual Improvisation
Li Zhang 0013
ITS1
2011 Context-Sensitive Affect Sensing and Metaphor Identification in Virtual Drama
Li Zhang 0013, John A. Barnden
ACII (2)1
2008 EMMA: an automated intelligent actor in e-drama
abstract
We report work on adding an improvisational AI actor and 3D emotional animation to an existing edrama program, a system for dramatic improvisation in simple virtual scenarios. The improvisational AI actor has an affect-detection component, aimed at detecting affective aspects of human-controlled characters' textual input. It also makes an appropriate response to stimulate the improvisation based on this affective understanding. A distinctive feature of our work is a focus on the metaphorical ways in which affect is conveyed. Moreover, we have also introduced how the detected affective states activate the animation engine to produce emotional gestures for human-controlled characters. Finally, we report user testing conducted for the AI actor. Our work contributes to the conference themes on affective user interfaces, natural language processing and emotionally believable gesture generation.
Li Zhang 0013, Marco Gillies, John A. Barnden
IUI1
2007 Don't worry about metaphor: affect detection for conversational agents
Catherine Smith, Timothy H. Rumbell, John A. Barnden, Robert J. Hendley, Mark Lee 0001, Alan M. Wallington, Li Zhang 0013
ACL7
2006 Developments in Affect Detection in E-drama
Li Zhang 0013, John A. Barnden, Robert J. Hendley, Alan M. Wallington
EACL1
2006 Exploitation in Affect Detection in Improvisational E-Drama
Li Zhang 0013, John A. Barnden, Robert J. Hendley, Alan M. Wallington
IVA1
2003 Speech recognition based on syllable recovery
Li Zhang 0013, William H. Edmondson
INTERSPEECH1
2002 Speech recognition using syllable patterns
abstract
This paper presents an account of the use of syllable structure as the basis for a novel approach to speech recognition. This contrasts with the serial organization of more conventional phonetic segments, and their use in speech recognition systems. It is demonstrated that working with syllables provides the basis for linguistically motivated speech recognition using the previously reported notion of the Pseudo-Articulatory Representation (PAR). The results are very promising taking into account the preliminary nature of the work and the novelty of the approach. A related paper [1] deals with theoretical issues in greater depth. 1.
Li Zhang 0013, William H. Edmondson
INTERSPEECH1
2001 Pseudo-articulatory representations and the recognition of syllable patterns in speech
abstract
This paper presents an account of syllable structure as the basis for organizing articulatory activity. This contrasts with the serial organization of more conventional phonetic segments. It is demonstrated that working with syllables in this way can provide the basis for linguistically motivated speech recognition using the previously reported notion of the Pseudo-Articulatory Representation (PAR).
William H. Edmondson, Li Zhang 0013
INTERSPEECH2