VLDB 2026 Research / reviewers in the wild / expert
Mingliang Gao 0001
dblp:169/1302 · also Ming-Liang Gao 0001
· DBLP profile ↗
54ranked-venue papers
5as first author
47since 2021 · last 2026
0000-0001-7273-7499ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 8 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cloud prediction via spatiotemporal-frequency differential and attentional network
Jiabing Liu, Jianhao Sun, Haiwen Wei, Qilei Li, Jun-Zhi Shi, Mingliang Gao 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | FusionFormer : A Multi-Source Data Fusion Transformer for Precipitation NowcastingabstractABSTRACT Precipitation nowcasting plays a critical role in agriculture, transportation, urban planning, and emergency response. With the growing emphasis on sustainable climate resilience, AI‐driven approaches have become vital for enhancing the accuracy and timeliness of meteorological forecasting. However, most existing deep learning models rely on a single data source, employ fixed‐scale feature extraction, and lack effective spatio‐temporal attention mechanisms. To address these issues, we propose a multi‐source data fusion Transformer, termed FusionFormer, for precipitation nowcasting. The FusionFormer integrates a Multi‐Source Fused Embedding (MSFE) module to effectively capture both fine‐grained local features and large‐scale patterns. It utilises parallel multi‐size convolutional kernels and multi‐resolution branches. Additionally, a Multi‐scale Spatio‐Temporal Attention (MSTA) module dynamically identifies the movement trajectory and intensity variations of precipitation systems. We also constructed a high‐resolution, spatiotemporally aligned multi‐source meteorological dataset for Shandong Province. Experiments on this dataset demonstrate that FusionFormer outperforms state‐of‐the‐art methods in both objective and subjective evaluations. Jianhao Sun, Haiwen Wei, Jiabing Liu, Junzhi Shi, Mingliang Gao 0001 |
Expert Syst. J. Knowl. Eng. | 6 |
| 2026 | Trustworthy Precipitation Prediction via Communication-Efficient Lightweight Trend Perception NetworkabstractResponsible and Trustworthy AI is essential for accurate precipitation prediction, which impacts agriculture, transportation, and disaster prevention, where errors can lead to severe economic losses and casualties. In the 6G era, AI models must be reliable in addressing small sample biases and extreme weather data scarcity, while remaining lightweight for efficient real-time deployment and ultra-fast communication. Most deep learning-based methods struggle to capture the temporal evolution of spatial features effectively. The limitations include insufficient awareness of temporal trends and a lack of precise control over the spatial representation of temporal features. To address these issues, we propose a Lightweight Trend Perception Network (LTPNet) for precipitation prediction. It consists of an encoder, a decoder, and two stacked layers of Trend Perception (TP) units. Specifically, in the TP unit, the spatial matching module is proposed to extract the spatio-temporal dependency information of the history in the current situation. At the same time, a time channel fusion module is proposed to capture the spatial variation trend of precipitation. Experimental evaluation on the WeatherBench dataset demonstrates that LTPNet outperforms other state-of-the-art methods objectively and subjectively. Jianhao Sun, Xiangui Meng, Haiwen Wei, Mingliang Gao 0001, Gwanggil Jeon |
IEEE Internet Things J. | 5 |
| 2026 | Towards generalizable and robust deepfake video detection via feature reconstruction and sequential temporal network
Weicheng Song, Qilei Li, Siyou Guo, Mingliang Gao 0001 |
Image Vis. Comput. | 4 |
| 2026 | Federated Learning for Edge Computing Enabled Artificial Intelligence of Things: A comprehensive surveyabstractAmong contemporary AI computing paradigms, Federated Learning (FL) stands out as an innovative method and has shown great potential in conjunction with edge computing. The two techniques combined serve as a building block forthe development of the Artificial Intelligence of Things (AIoT). This paper sheds light on the synergistic integration of FL with edge computing to propel AIoT’s capabilities in decentralized environments. By executing computing tasks closer to the data, FL at the edge not only alleviates latency and bandwidth limitations inherent in cloud-centric architectures, but also presents a robust solution to privacy concerns—a crucial obstacle in traditional centralized training setups. This paper delves into how FL tackles these privacy issues, providing an intricate explanation of its operational principles, applications, and the resultant benefits for AIoT systems. Through this scrutiny, we highlight FL’s potential in bolstering the efficiency and privacy of AIoT deployments while also delineating future research directions and the expected impact across various domains. This study aims to comprehensively comprehend FL for Edge Computing-enabled AIoT and foster developments in intelligent technologies and applications in an interconnected world. Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Wentai Wu, Chen Wang 0011, Ahmed M. Abdelmoniem |
Knowl. Based Syst. | 2 |
| 2026 | Progressive text-semantic-aware generative adversarial network for image fusion
Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, David Camacho |
Pattern Recognit. | 2 |
| 2026 | Infrared and visible image fusion via spatial-frequency edge-aware network
Shuohui Li, Qilei Li, Mingliang Gao 0001, Lucia Cascone |
Signal Process. | 3 |
| 2025 | SCANSleepNet: A spatial-channel attention network for sleep stage classification
Yuyun Liu, Qilei Li, Mingliang Gao 0001, Wenzhe Zhai |
Appl. Intell. | 3 |
| 2025 | Privacy-Preserving Crowd Counting via Quantum-Enhanced Federated LearningabstractABSTRACT Crowd counting plays a crucial role in analyzing group behavior in smart cities. Traditional crowd‐counting models rely on large datasets gathered from diverse individuals for training while ignoring the privacy protection for each training client. Meanwhile, the scale variation has long been a difficult problem in crowd counting and has greatly reduced model accuracy. Therefore, it is essential to achieve privacy‐aware crowd counting and to solve the problem of scale variation in dense scenes. To this end, we propose a Privacy‐preserving Quantum‐enhanced Network (PQNet). The PQNet uses federated learning to share parameters rather than data, which ensures the privacy of each client. Subsequently, a multi‐scale quantum‐driven calibration module is designed to capture multi‐scale information via quantum states. It enhances counting accuracy in dense crowd environments where scale varies. Experiments on four crowd counting and two vehicle counting benchmarks demonstrate that PQNet outperforms state‐of‐the‐art methods subjectively and objectively. The code will be available at: https://github.com/sdutzhangchen/PQNet . Jingan Cheng, Wenzhe Zhai, Mingliang Gao 0001 |
Expert Syst. J. Knowl. Eng. | 5 |
| 2025 | SFINet: A semantic feature interactive learning network for full-time infrared and visible image fusion
Qilei Li, Mingliang Gao 0001, Abdellah Chehri, Gwanggil Jeon |
Expert Syst. Appl. | 3 |
| 2025 | Traffic Density Estimation by Distributed Proxy Model Learning for Internet of VehicleabstractAutonomous driving has been significantly advanced in today’s society, which revolutionized daily routines and facilitated the development of the Internet of Vehicles (IoV). A crucial aspect of this system is understanding traffic density to enable intelligent traffic management. With the rapid improvement in deep neural networks (DNNs), the accuracy of density estimation has markedly improved. However, there are two main issues that remain unsolved. First, current DNN-based models are excessively heavy, characterized by an overwhelming number of training parameters (millions or even billions) and substantial computational complexity, indicated by a high number of FLOPs. These requirements for storage and computation severely limit the practical application of these models, especially on edge devices with limited capacity and computational power. Second, despite the superior performance of DNN models, their effectiveness largely depends on the availability of large-scale data for training. Growing privacy concerns have made individuals increasingly hesitant to allow their data to be publicly used for model training, particularly in vehicle-related applications that might reveal personal movements, which leads to data isolation issues. In this article, we address these two problems at once with a systematic framework. Specifically, we introduce the proxy model distributed learning (PMDL) model for traffic density estimation. PMDL model is composed of two main components. First, we introduce a proxy model learning strategy that transfers fine-grained knowledge from a larger master model to a lightweight proxy model, i.e., a proxy model. Second, we design a distributed learning strategy that trains multiple proxy models with privacy-aware local data and seamlessly aggregates these models via a global parameter server. This ensures privacy protection while significantly improving estimation performance compared to training models with limited, isolated data. We tested the proposed model on four major vehicle density analysis benchmarks and demonstrated its efficiency by outperforming other state-of-the-art competitors. The code is available athttps://github.com/jinyongch/DPML. Qilei Li, Jingan Cheng, Mingliang Gao 0001, Gwanggil Jeon |
IEEE Internet Things J. | 3 |
| 2025 | Cross-Camera Discriminative Person Association by Unsupervised Frame Clustering and SelectionabstractThe objective of cross-camera persson association is to identify individuals captured across disjoint cameras. It is achieved by Re-Identification (ReID) models, which extract unique identity representations from the visual input, dominated by video sequence. Most ReID methods primarily focus on modifying the backbone network architecture to learn more representative features of people. However, these methods often overlook the impact of low-quality frames on the training process. Several studies have confirmed that low-quality data not only hinders the model from learning meaningful content but also diminishes its performance. One possible solution is to manually label the quality of each frames, but this is time-consuming and inefficient. In this paper, we propose a Unsupervised Frame Clustering and Selection framework called UFCS to address this problem by applying the unsupervised clustering to automatically select high-quality frames. Specifically, we applied three unsupervised clustering solutions for high-quality frame selection, namely K-means, Deep K-means, and DBSCAN. These clustering techniques integrate both appearance and, indirectly, temporal consistency by operating within tracklets. These schemes perform clustering in the image or deep feature space to select highquality frames for network training. This straightforward yet effective approach enables the ReID network to generate a more discriminative representations, thereby improving recognition performance. Experimental results obtained on the challenging video-based person ReID datasets MARS indicate that our proposed scheme can outperform related state-of-the-art methods by a large margin. Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai, Gwanggil Jeon |
IEEE Internet Things J. | 2 |
| 2025 | Towards trustworthy image super-resolution via symmetrical and recursive artificial neural network
Mingliang Gao 0001, Jianhao Sun, Qilei Li, Muhammad Attique Khan, Jianrun Shang, Xianxun Zhu, Gwanggil Jeon |
Image Vis. Comput. | 1 |
| 2025 | Generalizable deepfake detection via Spatial Kernel Selection and Halo Attention Network
Siyou Guo, Qilei Li, Mingliang Gao 0001, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 3 |
| 2025 | Deepfake detection via Feature Refinement and Enhancement Network
Weicheng Song, Siyou Guo, Mingliang Gao 0001, Qilei Li, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 3 |
| 2025 | A Dual-branch Progressive Network with spatial-frequency constraint for image fusion
Zenghui Wang 0011, Xuening Xing, Lina Liu 0009, Xianxun Zhu, Mingliang Gao 0001 |
Image Vis. Comput. | 6 |
| 2025 | Composed image retrieval by Multimodal Mixture-of-Expert Synergy
Wenzhe Zhai, Mingliang Gao 0001, Gwanggil Jeon, David Camacho |
Image Vis. Comput. | 2 |
| 2025 | Towards text-refereed multi-modal image fusion by cross-modality interaction
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai |
Signal Process. | 3 |
| 2025 | Context-Aware Deepfake Detection for Securing AI-Driven Financial TransactionsabstractThe rapid advancement of deepfake technology has threatened the community’s sense of security, particularly in the context of face-based payment systems. Thus, deepfake detection has emerged as a critical issue demanding immediate attention. However, the generalization performance of existing detection models is limited as they are overly reliant on specific forged features while ignoring the common forged features. To address this problem, we introduce the context-aware decoupling network (CADNet) for deepfake detection. Specifically, a context self-calibration (CSC) module is constructed to guide the network to focus on local forged regions. It enlarges possible regions to increase the likelihood of forgery cues. Meanwhile, a frequency domain decoupling (FDD) module is introduced to extract and fuse different frequency components. It realizes the collaborative representation optimization of global semantics and local details. The experimental results prove that the proposed model exhibits strong generalization capability across multiple standard datasets. It achieves average area under the curve (AUC) values of 98.64% for in-domain evaluation and 75.52% for cross-dataset generalization. Changcun Liu, Guisheng Zhang, Siyou Guo, Qilei Li, Gwanggil Jeon, Mingliang Gao 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Space-Frequency and Global-Local Attentive Networks for Sequential Deepfake DetectionabstractThe widespread misinformation generated by deepfake systems has emerged as a significant challenge in the dynamic realm of digital media. It poses threats to credibility, privacy, and security of information in daily life. Moreover, the increasing accessibility to facial editing tools further enables users to alter facial characteristics subtly through a series of intricate steps. To address the issue, we introduce a space–frequency and global–local attentive network (SFGLA-Net) for sequential deepfake detection. This method is designed to identify and analyze the sophisticated manipulated attributes of deepfake images. Specifically, we introduce a space–frequency fusion module to leverage the deep feature extracted in spatial and frequency domains, so as to exploit subtle inconsistencies and artifacts that are not perceptible in the spatial domain alone. Additionally, we design a global–local attention module to pinpoint the manipulated areas more accurately. Extensive experiments demonstrate the superior performance of the proposed method by significantly outperforming existing techniques in sequential deepfake detection. The code is available athttps://github.com/guishengzhanga/SFGLA. Guisheng Zhang, Qilei Li, Mingliang Gao 0001, Siyou Guo, Gwanggil Jeon, Ahmed M. Abdelmoniem |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Zero-Shot Object Counting With Vision-Language Prior Guidance NetworkabstractThe majority of existing counting models are designed to operate on a singular object category, such as crowds or vehicles. The emergence of multi-modal foundational models, e.g., Contrastive Language-Image Pre-training (CLIP), has paved the way for class-agnostic counting. This approach facilitates the counting of objects across diverse classes within a single image based on textual indications. However, class-agnostic counting models based on CLIP confront two primary challenges. Firstly, the CLIP model exhibits limited sensitivity towards location information, which prioritizes global content over the precise localization of objects. Therefore, directly employing the CLIP model is regarded as suboptimal. Secondly, these models commonly employ frozen pre-trained vision and language encoders while disregarding potential misalignment within the constructed hypothesis space. In this paper, we propose a unified framework, named the Vision-Language Prior Guidance (VLPG) Network, to tackle these two challenges. The VLPG consists of three key components, namely the Grounding DINO module, Spatial Prior Calibration (SPC) module, and Object-Centric Alignment (OCA) module. The Grounding DINO module utilizes the spatial-awareness capability of extensive pre-trained object grounding models to incorporate the spatial position as an additional prior for a particular query class. This adaptation enables the network to concentrate more precisely on the exact location of the objects. Meanwhile, the SPC module is built to extract the long-range dependencies and local regions of the spatial position. Additionally, to align the feature space across different modalities, we design an OCA module that condenses textual information into an object query which serves as an instruction for cross-modality matching. Through the collaborative efforts of these three modules, multimodal representations are aligned while maintaining their discriminative nature. Comprehensive experiments conducted on various benchmarks validate the effectiveness of the proposed model. Wenzhe Zhai, Xianglei Xing, Mingliang Gao 0001, Qilei Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Generic Representation Learning for Vehicle Association Guided by Foundational ModelsabstractVehicle association is a vital yet complex task to retrieve specific vehicles across various camera angles, time frames, and geographical locations. In environments supported by autonomous driving and 6G networks, this task plays a vital role in urban surveillance and traffic management by enabling the real-time sharing of vehicle location and status information through ultra-high-speed, low-latency 6G communication. The success of a retrieval model largely depends on the quality of the extracted representations, which can be influenced by factors such as background diversity and occlusions. This study proposes a method to extract representations that remain consistent across different domains while retaining the discriminative power necessary to determine a vehicle’s spatial location, regardless of background or environmental variations. To achieve this, we introduce a framework called Generic Representation Learning (GRL). Within GRL, we leverage large-scale pre-trained foundational models to provide spatial priors of vehicles, specifically the Grounding DINO model for object detection and the SAM model for object segmentation. These modules collaborate to help the network understand the spatial context of the object, enabling the feature extractor to focus on discriminative areas while minimizing interference. Additionally, we introduce a complementary feature alignment mechanism based on a memory bank to explore globally applicable knowledge within the learned representation of the object. These constituent elements collectively form SRP, to enhance its capability for outstanding performance in vehicle retrieval. Extensive experimentation demonstrates that SRP significantly outperforms existing models on widely recognized benchmarks. Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Gwanggil Jeon, Ahmed M. Abdelmoniem |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Deepfake Detection via a Progressive Attention NetworkabstractThe rapid advancement of deepfake technology has enabled the creation of highly realistic forged face images or videos. While deepfake technology adds entertainment to people’s lives, it also poses a potential threat to social security. Deepfake detection is a crucial technology for identifying forged images. However, existing deep learning-based models for deepfake detection often overlook subtle forged traces. To solve this problem, we propose a Progressive Attention Network (PANet). The PANet incorporates two attention modules, namely the Efficient Multi-Scale Attention Module (EMAM) and the Spatial and Channel Attention Module (SCAM), in a progressive manner. The EMAM focuses on crucial facial regions, such as the eyes, nose, and mouth, rather than the entire face. The SCAM facilitates fine-grained feature extraction. Experimental results demonstrate that the proposed method achieves state-of-the-art results on deepfake detection datasets. Siyou Guo, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, David Camacho |
IJCNN | 2 |
| 2024 | Dual-branch and triple-attention network for pan-sharpening
Mingliang Gao 0001, Abdellah Chehri, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Appl. Intell. | 2 |
| 2024 | Multiscale aggregation and illumination-aware attention network for infrared and visible image fusionabstractAbstract Image fusion plays a significant role in computer vision since numerous applications benefit from the fusion results. The existing image fusion methods are incapable of perceiving the most discriminative regions under varying illumination circumstances and thus fail to emphasize the salient targets and ignore the abundant texture details of the infrared and visible images. To address this problem, a multiscale aggregation and illumination‐aware attention network (MAIANet) is proposed for infrared and visible image fusion. Specifically, the MAIANet consists of four modules, namely multiscale feature extraction module, lightweight channel attention module, image reconstruction module, and illumination‐aware module. The multiscale feature extraction module attempts to extract multiscale features in the images. The role of the lightweight channel attention module is to assign different weights to each channel so as to focus on the essential regions in the infrared and visible images. An illumination‐aware module is employed to assess the probability distribution regarding the illumination factor. Meanwhile, an illumination perception loss is formulated by the illumination probabilities to enable the proposed MAIANet to better adjust to the changes in illumination. Experimental results on three datasets, that is, MSRS, TNO, and RoadSence, verify the effectiveness of the MAIANet in both qualitative and quantitative evaluations. Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Abdellah Chehri, Gwanggil Jeon |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Multi-scale large kernel convolution and hybrid attention network for remote sensing image dehazing
Lina Liu 0009, Zenghui Wang 0011, Mingliang Gao 0001 |
Image Vis. Comput. | 4 |
| 2024 | SaReGAN: a salient regional generative adversarial network for visible and infrared image fusion
Mingliang Gao 0001, Yi'nan Zhou, Wenzhe Zhai, Qilei Li |
Multim. Tools Appl. | 1 |
| 2024 | Multiscale aggregation network via smooth inverse map for crowd counting
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou |
Multim. Tools Appl. | 2 |
| 2024 | A structure and texture revealing retinex model for low-light image enhancement
Qilei Li, Marco Anisetti, Gwanggil Jeon, Mingliang Gao 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Object counting in remote sensing via selective spatial-frequency pyramid networkabstractAbstract The integration of remote sensing object counting in the Mobile Edge Computing (MEC) environment is of crucial significance and practical value. However, the presence of significant background interference in remote sensing images poses a challenge to accurate object counting, as the results are easily affected by background noise. Additionally, scale variation within remote sensing images presents a further difficulty, as traditional counting methods face challenges in adapting to objects of different scales. To address these challenges, we propose a selective spatial‐frequency pyramid network (SSFPNet). Specifically, the SSFPNet consists of two core modules, namely the pyramid attention (PA) module and the hybrid feature pyramid (HFP) module. The PA module accurately extracts target regions and eliminates background interference by operating on four parallel branches. This enables more precise object counting. The HFP module is introduced to fuse spatial and frequency domain information, leveraging scale information from different domains for object counting, so as to improve the accuracy and robustness of counting. Experimental results on RSOC, CARPK, and PUCPR+ benchmark datasets demonstrate that the SSFPNet achieves state‐of‐the‐art performance in terms of accuracy and robustness. Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Softw. Pract. Exp. | 2 |
| 2024 | Defending Deepfakes by Saliency-Aware AttackabstractWith the rapid development of deep learning, especially the generative adversarial network (GAN), face modification has been substantially advanced and enables the generated images to look more realistic. Given an image or a video frame of a person, such a system can create fake images, which manipulates the movement, expression, and even appearance, e.g., hair color, eye color, and age. Such a system is termed Deepfake, which has raised significant ethical issues, especially for celebrities. With the pretrained Deepfake models being widely available on the Internet, its negative applications, such as face manipulation and pornographic generation, have exposed the dark side of the Deepfake technology to the sociocyber world. In this article, we aim to defend a well-trained Deepfake model by manipulating the raw image with unperceived perturbation. To minimize the alterations to the original image while effectively fooling the Deepfake model, we propose to selectively perturb only the foreground person region and maintain the irrelevant background. This is based on the observation that the salient object in a person’s image is always the foreground face region. Such a strategy introduces negligible alterations to the original image, which makes the attack remain effective. We experimentally demonstrate the superiority of the proposed attacking framework over the existing models and show our approach is ready to be applied for out-of-the-box development. Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | A Trust-Aware and Authentication-Based Collaborative Method for Resource Management of Cloud-Edge Computing in Social Internet of ThingsabstractThe Social Internet of Things (S-IoT) paradigm is focused on topic of the Internet of Things (IoT), which accelerates the object issues by working with the concept of social networks. Searching and finding a new object in the community are considered to manage the number of friends and complex relationships between them and affect the ability to navigate at the cloud-edge layer, and resources, such as battery lifetime of S-IoT devices and energy resources, are important challenges in this field. In the processing of social messages of remote devices, increasing the battery life of devices that require such requirements plays the most important role. In this research, a collaboration scenario is presented to consider object attributes, friend’s functions and intelligent friend selection among objects for group messaging. First, a general reference model is designed and presented to select a friend to access group message remote processing services and minimize cloud-edge resources. The simulation results show that, for the correct communication of friends at the edge of the network and in each service discovery, according to the length of the path in the network, it is possible to establish stable communication and make better service with the least possible. The results show that if we want to develop a method for friendship between objects in communication in cloud computing, the proposed method can greatly improve the effectiveness of providing reliable message processing types. Alireza Souri, Yanlei Zhao, Mingliang Gao 0001, Asghar Mohammadian, Jin Shen, Eyhab Al-Masri |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Object Counting via Group and Graph Attention NetworkabstractObject counting, defined as the task of accurately predicting the number of objects in static images or videos, has recently attracted considerable interest. However, the unavoidable presence of background noise prevents counting performance from advancing further. To address this issue, we created a group and graph attention network (GGANet) for dense object counting. GGANet is an encoder-decoder architecture incorporating a group channel attention (GCA) module and a learnable graph attention (LGA) module. The GCA module groups the feature map into several subfeatures, each of which is assigned an attention factor through the identical channel attention. The LGA module views the feature map as a graph structure in which the different channels represent diverse feature vertices, and the responses between channels represent edges. The GCA and LGA modules jointly avoid the interference of irrelevant pixels and suppress the background noise. Experiments are conducted on four crowd-counting datasets, two vehicle-counting datasets, one remote-sensing counting dataset, and one few-shot object-counting dataset. Comparative results prove that the proposed GGANet achieves superior counting performance. Mingliang Gao 0001, Guofeng Zou, Alessandro Bruno, Abdellah Chehri, Gwanggil Jeon |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | A framework to connect IoT edge networks through 3D Massive MIMO
Talha Younas, Yanlei Zhao, Gwanggil Jeon, Ghulam Farid 0001, Sohaib Tahir, Muluneh Mekonnen, Jin Shen, Mingliang Gao 0001 |
Wirel. Networks | 8 |
| 2023 | Unsupervised person re-identification based on distribution regularization constrained asymmetric metric learning
Guofeng Zou, Guizhen Chen, Mingliang Gao 0001, Liju Yin |
Appl. Intell. | 4 |
| 2023 | FPANet: feature pyramid attention network for crowd counting
Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, Marco Anisetti |
Appl. Intell. | 2 |
| 2023 | Crowd counting in smart city via lightweight Ghost Attention Pyramid Network
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Future Gener. Comput. Syst. | 3 |
| 2023 | Scale-Context Perceptive Network for Crowd Counting and Localization in Smart City SystemabstractThe task of crowd counting and localization is to predict the count and position of people in a crowd, which is a practical and essential sub-task in crowd analysis and smart city systems. However, the inherent problems of scale variation and background disturbance restrain their performance. While recent researches focus on studying counting and localization independently, a few works are capable of executing both tasks simultaneously. To this end, we propose a Scale-Context Perceptive Network (SCPNet) to jointly tackle the crowd counting and localization tasks in a unified framework. Specifically, a scale perceptive (SP) module with a local-global branch schema is designed to capture multiscale information. Meanwhile, a context perceptive (CP) module, by the channel-spatial self-attention mechanism, is derived to suppress the background disturbance. Furthermore, a novel hierarchical scale loss function that combines the Euclidean loss function and structural similarity loss function is designed to prompt the proposed model to fulfill the counting and localization simultaneously. Extensive experiments on challenging crowd datasets prove the superiority of the proposed SCPNet compared with the state-of-the-art competitors in both objective and subjective evaluations. Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon |
IEEE Internet Things J. | 2 |
| 2023 | $\hbox {DA}^2$Net: a dual attention-aware network for robust crowd counting
Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou, Mingliang Gao 0001 |
Multim. Syst. | 7 |
| 2023 | Dense Attention Fusion Network for Object Counting in IoT System
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Kyu Hyung Kim, Gwanggil Jeon |
Mob. Networks Appl. | 2 |
| 2023 | Scale Region Recognition Network for Object Counting in Intelligent Transportation SystemabstractSelf-driving technology and safety monitoring devices in intelligent transportation systems require superb capacity for context awareness. Accurately inferring the counts of crowds and vehicles are the two practical and fundamental tasks in the transportation system. However, the scale variation and background interference in the traffic image hinder the counting performance. To solve the aforementioned problems, a scale region recognition network (SRRNet) is proposed in this paper. It has two key components, termed scale level awareness (SLA) module and object region recognition (ORR) module. The SLA module aims to encode the representations at multiple scales, which are beneficial to address the scale variation. The ORR module is designed to suppress background interference through the visual attention mechanism. Extensive experimental results on four crowd counting datasets and five vehicle counting datasets have demonstrated the superiority of the proposed SRRNet in both counting accuracy and robustness compared with the mainstream competitors. Meanwhile, substantial ablation studies have proved the effectiveness of the proposed SLA and ORS modules. Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Abnormal behavior detection using streak flow acceleration
Mingliang Gao 0001, Jinfeng Pan, Chengyuan Zhao |
Appl. Intell. | 3 |
| 2022 | Visual tracking for UAV using adaptive spatio-temporal regularized correlation filters
Libin Xu, Mingliang Gao 0001, Qilei Li, Guofeng Zou, Jinfeng Pan |
Appl. Intell. | 2 |
| 2022 | Accelerated duality-aware correlation filters for visual tracking
Libin Xu, Mingliang Gao 0001, Zheng Liu 0002, Qilei Li, Gwanggil Jeon |
Neural Comput. Appl. | 2 |
| 2021 | Bayesian regularization restoration algorithm for photon counting images
Ying Li 0096, Liju Yin, Jinfeng Pan, Mingliang Gao 0001, Guofeng Zou, Jiansi Liu |
Appl. Intell. | 5 |
| 2021 | Person re-identification based on metric learning: a survey
Guofeng Zou, Guixia Fu, Mingliang Gao 0001, Zheng Liu 0002 |
Multim. Tools Appl. | 5 |
| 2021 | Deep Learning Thermal Image Translation for Night Vision PerceptionabstractContext enhancement is critical for the environmental perception in night vision applications, especially for the dark night situation without sufficient illumination. In this article, we propose a thermal image translation method, which can translate thermal/infrared (IR) images into color visible (VI) images, called IR2VI. The IR2VI consists of two cascaded steps: translation from nighttime thermal IR images to gray-scale visible images (GVI), which is called IR-GVI; and the translation from GVI to color visible images (CVI), which is known as GVI-CVI in this article. For the first step, we develop the Texture-Net, a novel unsupervised image translation neural network based on generative adversarial networks. Texture-Net can learn the intrinsic characteristics from the GVI and integrate them into the IR image. In comparison with the state-of-the-art unsupervised image translation methods, the proposed Texture-Net is able to address some common challenges, e.g., incorrect mapping and lack of fine details, with a structure connection module and a region-of-interest focal loss. For the second step, we investigated the state-of-the-art gray-scale image colorization methods and integrate the deep convolutional neural network into the IR2VI framework. The results of the comprehensive evaluation experiments demonstrate the effectiveness of the proposed IR2VI image translation method. This solution will contribute to the environmental perception and understanding in varied night vision applications. Shuo Liu 0009, Mingliang Gao 0001, Vijay John, Zheng Liu 0002, Erik Blasch |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2020 | Adaptive Spatio-Temporal Regularized Correlation Filters for UAV-Based Tracking
Libin Xu, Qilei Li, Guofeng Zou, Zheng Liu 0002, Mingliang Gao 0001 |
ACCV (2) | 6 |
| 2020 | A new approach for small sample face recognition with pose variation by fusing Gabor encoding features and deep features
Guofeng Zou, Guixia Fu, Mingliang Gao 0001, Jinfeng Pan, Zheng Liu 0002 |
Multim. Tools Appl. | 3 |
| 2019 | A novel construction method of convolutional neural network model based on data-driven
Guofeng Zou, Guixia Fu, Mingliang Gao 0001, Jin Shen, Liju Yin, Xianye Ben |
Multim. Tools Appl. | 3 |
| 2018 | Orthogonal gradient measurement matrix optimisation methodabstractThe optimisation of measurement matrix that is within the compressive sensing framework is considered in this study. Based on the fact that an information factor with smaller mutual coherence performs better, the gradient measurement matrix optimisation method is improved by an orthogonal search direction revision factor. This algorithm updates the approximation of ideal Gram matrix of information operator and the measurement matrix alternatingly. Using measurement matrix and sparse basis to represent the Gram matrix, the measurement matrix is optimised by the gradient algorithm, in which an orthogonal gradient search direction revision factor is proposed and utilised to further improve the performance of measurement matrix. This orthogonal factor is computed by the Cayley transform of a real skew symmetric matrix that is related to the gradient and the measurement matrix. Results of several experiments show that compared with the initial random matrix, the optimised measurement matrix can lead to better signal reconstruction quality. Jinfeng Pan, Jin Shen, Mingliang Gao 0001, Liju Yin, Faying Liu, Guofeng Zou |
IET Image Process. | 3 |
| 2016 | A novel visual tracking method using bat algorithm
Mingliang Gao 0001, Jin Shen, Liju Yin, Guofeng Zou, Guixia Fu |
Neurocomputing | 1 |
| 2015 | Face tracking based on differential harmony searchabstractOwing to its significant roles in computer vision applications, human face tracking has drawn extensive attention in recent years. Most researchers solve face tracking using particle filter, meanshift and their derivatives. Unlike the traditional methods, in this study, face tracking is treated as an optimisation problem and a new meta‐heuristic optimisation algorithm, differential harmony search (DHS), is introduced to solve face tracking problems. We compare the speed and accuracy of the proposed method with particle filter, meanshift and improved harmony search. Experimental results show that DHS‐based tracker is faster and more accurate and it is easy to handle the parameters tuning. Furthermore, to improve the reliability of tracking, multiple visual cues are applied to DHS‐based tracking system and experimental results demonstrate the increased robustness achieved by fusing multiple cues. Mingliang Gao 0001, Xianming Sun, Dai-Sheng Luo |
IET Comput. Vis. | 1 |
| 2013 | Object tracking using firefly algorithmabstractFirefly algorithm (FA) is a new meta‐heuristic optimisation algorithm that mimics the social behaviour of fireflies flying in the tropical and temperate summer sky. In this study, a novel application of FA is presented as it is applied to solve tracking problem. A general optimisation‐based tracking architecture is proposed and the parameters’ sensitivity and adjustment of the FA in tracking system are studied. Experimental results show that the FA‐based tracker can robustly track an arbitrary target in various challenging conditions. The authors compare the speed and accuracy of the FA with three typical tracking algorithms including the particle filter, meanshift and particle swarm optimisation. Comparative results show that the FA‐based tracker outperforms the other three trackers. Mingliang Gao 0001, Xiaohai He, Dai-Sheng Luo, Qizhi Teng |
IET Comput. Vis. | 1 |