EDBT 2026 Demo / reviewers in the wild / expert
Qilei Li
dblp:85/6368
· DBLP profile ↗
61ranked-venue papers
15as first author
53since 2021 · last 2026
0000-0002-9675-9016ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cloud prediction via spatiotemporal-frequency differential and attentional network
Jiabing Liu, Jianhao Sun, Haiwen Wei, Qilei Li, Jun-Zhi Shi, Mingliang Gao 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | A Comprehensive Review for Agricultural Product Prices Forecasting: Architectural Diversity and Open ChallengesabstractABSTRACT Accurate forecasting of agricultural prices is essential for informed production planning, market stabilisation and effective policy design. This review examines 773 studies published between 2006 and 2025 to synthesise recent advances. We begin by analysing the factor systems and structural characteristics of agri‐price data, and organise forecasting tasks by input–output design, temporal resolution and prediction objectives—ranging from point estimates to trend detection and probabilistic forecasting. Evaluation practices are reviewed across multiple dimensions, including error metrics, trend alignment, model selection and uncertainty estimation. We then trace the evolution of forecasting 15 methods from traditional statistical models to machine learning and deep neural 16 architectures (RNN, CNN, GNN, Transformer), as well as decomposition‐based and 17 ensemble strategies. These developments are contextualised within bibliometric trends, highlighting shifts in research focus and global collaboration. Empirical evidence shows that hybrid pipelines combining decomposition, feature learning and ensemble techniques tend to outperform standalone models, while simple linear models remain competitive for long‐horizon or low‐frequency forecasts. Common challenges include data leakage, inconsistent testing horizons and insufficient treatment of uncertainty. Looking ahead, future research should emphasise integrating diverse data sources—such as weather, trade and policy signals—and building models that can adapt to unexpected market changes. It is equally important to understand how price dynamics respond to policy actions, improve model transferability across regions and commodities and provide well‐calibrated forecasts with interpretable uncertainty estimates to enhance the practical value of agricultural price prediction. Binrong Wu, Qilei Li, Deqian Fu, Lin Wang 0001 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2026 | A Review of Federated Learning Under Data HeterogeneityabstractABSTRACT Federated learning (FL) has emerged as an impactful paradigm for privacy‐preserving machine learning, and allows model training without the need to share raw data. However, data heterogeneity across clients challenges practical FL deployment. Data space heterogeneity and statistical heterogeneity create significant training difficulties. System heterogeneity imposes additional external constraints. These combined factors impair convergence and reduce model performance. They also raise concerns regarding fairness, scalability and robustness. Focused on data heterogeneity, this review provides a structured analysis of FL. It encompasses three key areas: core categorizations of data heterogeneity, algorithmic advances (e.g., personalized FL, mixture‐of‐experts architectures, transfer learning‐based solutions) and system‐level techniques spanning communication optimization, resource adaptation and secure collaboration. We further synthesize benchmark efforts and real‐world applications in healthcare, finance, nuclear power and the Internet of Things (IoT)/edge computing to highlight the practical implications of heterogeneity‐aware FL. Finally, we identify key challenges and outline promising research directions towards scalable, fair and adaptive FL systems capable of operating in complex real‐world settings. This survey aims to serve as a reference point and conceptual roadmap for future research in heterogeneous FL. Wentao Yue, Tianyou Lai, Qingyu Mao, Qilei Li, David Camacho |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | TransHAR: Toward Intent-Aware Transformer-Based Human Activity Recognition in Intelligent IoT Communication SystemsabstractHuman Activity Recognition (HAR) has emerged as a critical component in intent-aware, AI-driven Internet of Things (IoT) Communication systems, enabling context-aware responses in smart environments. Recently, WiFi-based HAR has gained significant attention due to its non-intrusive nature, low deployment cost, and ability to preserve privacy. However, they face a major challenge across domains. To address this limitation, we propose a novel cross-domain HAR framework (called TransHAR) by introducing a lightweight and efficient transformer model. On one hand, we design a feature representation block that processes the Wi-Fi channel frequency response (CFR) phase data to estimate Doppler shifts, capturing motion-related dynamics while remaining invariant to static, environment-specific structures, enhancing generalization across domains. On the other hand, we propose a lightweight Transformer architecture, termed ResDyTFormer, which minimizes reliance on normalization layers by incorporating a novel Residual Dynamic Tanh function. This function dynamically learns to balance between traditional normalization and the Dynamic Tanh operation, thereby maintaining training stability and avoiding gradient vanishing issues often encountered when using Dynamic Tanh alone. Extensive experiments on two benchmark datasets demonstrate that the proposed TransHAR framework achieves state-of-the-art performance in both in-domain and cross-domain HAR tasks with only 0.17M parameters. On the SHARP dataset, it attains an impressive 99.04% F1 score and 98.94% accuracy. On the 3DO dataset, it achieves 86.10% accuracy and 84.95% F1 score. These results highlight the potential of TransHAR as an efficient and scalable framework for real-world WiFi-based human activity sensing. Meng Xu 0022, Qilei Li, Fei Luo 0003, Jiguang Li, Yifeng Zeng, Gwanggil Jeon |
IEEE Internet Things J. | 3 |
| 2026 | Towards generalizable and robust deepfake video detection via feature reconstruction and sequential temporal network
Weicheng Song, Qilei Li, Siyou Guo, Mingliang Gao 0001 |
Image Vis. Comput. | 2 |
| 2026 | Federated Learning for Edge Computing Enabled Artificial Intelligence of Things: A comprehensive surveyabstractAmong contemporary AI computing paradigms, Federated Learning (FL) stands out as an innovative method and has shown great potential in conjunction with edge computing. The two techniques combined serve as a building block forthe development of the Artificial Intelligence of Things (AIoT). This paper sheds light on the synergistic integration of FL with edge computing to propel AIoT’s capabilities in decentralized environments. By executing computing tasks closer to the data, FL at the edge not only alleviates latency and bandwidth limitations inherent in cloud-centric architectures, but also presents a robust solution to privacy concerns—a crucial obstacle in traditional centralized training setups. This paper delves into how FL tackles these privacy issues, providing an intricate explanation of its operational principles, applications, and the resultant benefits for AIoT systems. Through this scrutiny, we highlight FL’s potential in bolstering the efficiency and privacy of AIoT deployments while also delineating future research directions and the expected impact across various domains. This study aims to comprehensively comprehend FL for Edge Computing-enabled AIoT and foster developments in intelligent technologies and applications in an interconnected world. Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Wentai Wu, Chen Wang 0011, Ahmed M. Abdelmoniem |
Knowl. Based Syst. | 1 |
| 2026 | Leverage cross-domain variations for generalizable person ReID representation learning
Qilei Li, Shitong Sun, Weitong Cai, Shaogang Gong |
Pattern Recognit. | 1 |
| 2026 | Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts
Qilei Li, Shitong Sun, Da Li 0001, Timothy M. Hospedales, Shaogang Gong |
Pattern Recognit. | 1 |
| 2026 | Progressive text-semantic-aware generative adversarial network for image fusion
Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, David Camacho |
Pattern Recognit. | 3 |
| 2026 | BenchCIR: Benchmarking robustness in composed image retrieval across modalitiesabstractComposed image retrieval aims to retrieve images based on a query that consists of a reference image and text describing desired modifications to that image. It has recently attracted attention for its ability to tailor image retrieval to user intentions by combining information-rich reference images with concise natural language instructions. Despite its current success, the robustness of composed image retrieval methods to either (1) common corruptions or (2) variations of the textual descriptions have never been systematically evaluated. In this paper, we perform the first robustness study of composed image retrieval, establishing three new benchmarks for a systematic evaluation of robustness to common corruption (in both the textual and visual domains) and robustness in text understanding. For analysis of natural image corruption, we introduce two new large-scale benchmark datasets, CIRR-C and FashionIQ-C, for the open domains and fashion domains respectively–both of which feature 75 visual corruptions and 35 textual corruptions. To facilitate robust evaluation of text understanding, we introduce a new diagnostic dataset CIRR-D by expanding the CIRR dataset with synthetic data, specifically probing text understanding across variations in: numerical, attribute, object removal, and background. We introduce BenchCIR, a testbed for evaluating composed image retrieval model robustness with standardized evaluation protocols. Through benchmarking ten published models in the testbed, we reveal insights into how the composition of visual and textual modalities affects model robustness. The code is in https://suntongtongtong.github.io/BenchCIR/ Shitong Sun, Qilei Li, Shaogang Gong, Weitong Cai, Philip Torr 0001, Jindong Gu |
Pattern Recognit. | 2 |
| 2026 | Infrared and visible image fusion via spatial-frequency edge-aware network
Shuohui Li, Qilei Li, Mingliang Gao 0001, Lucia Cascone |
Signal Process. | 2 |
| 2025 | Discovering Latent Knowledge Prototypes for Heterogeneous Federated LearningabstractFederated learning (FL) is crucial for ensuring data privacy, a major concern in many applications. However, FL faces significant challenges due to data and model heterogeneity arising from diverse learning environments and the varying capabilities of participating entities. Most existing methods primarily concentrate on aggregating knowledge that is represented by models, logits, or features. which rely on specific assumptions that may not hold in real-world scenarios and thus fail to address both data and model heterogeneity simultaneously. In this work, we aim to address these challenges by tackling heterogeneity from both model and data perspectives while maintaining efficiency. To this end, we leverage locally encoded latent prototypes produced from the local knowledge memory bank to represent per-client knowledge updates, which are then aggregated on the server and transferred back to the clients for knowledge decoding and integration as global constraints for further local training. Considering the heterogeneity in model architectures, we design the knowledge encoder and decoder to be compatible with different model architectures and ensure robust prototype aggregation by aligning latent spaces to a common prior distribution, to enhance compatibility under diverse data distributions. We evaluate our method on multiple benchmarks and demonstrate its superior performance in terms of accuracy and effectiveness under various heterogeneous settings. Qilei Li, Ahmed M. Abdelmoniem |
ECAI | 1 |
| 2025 | Hierarchical Knowledge Structuring for Effective Federated Learning in Heterogeneous EnvironmentsabstractFederated learning enables collaborative model training across distributed entities while maintaining individual data privacy. A key challenge in federated learning is balancing the personalization of models for local clients with generalization for the global model. Recent efforts leverage logit-based knowledge aggregation and distillation to overcome these issues. However, due to the non-IID nature of data across diverse clients and the imbalance in the client’s data distribution, directly aggregating the logits often produces biased knowledge that fails to apply to individual clients and obstructs the convergence of local training. To solve this issue, we propose a Hierarchical Knowledge Structuring (HKS) framework that formulates sample logits into a multi-granularity codebook to represent logits from personalized per-sample insights to globalized per-class knowledge. The unsupervised bottom-up clustering method is leveraged to enable the global server to provide multi-granularity responses to local clients. These responses allow local training to integrate supervised learning objectives with global generalization constraints, which results in more robust representations and improved knowledge sharing in subsequent training rounds. The proposed framework’s effectiveness is validated across various benchmarks and model architectures. Wai Fong Tam, Qilei Li, Ahmed M. Abdelmoniem |
IJCNN | 2 |
| 2025 | Boosting Neural Video Representation via Online Structural Reparameterization
Qingyu Mao, Shuai Liu 0022, Qilei Li, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (6) | 4 |
| 2025 | SCANSleepNet: A spatial-channel attention network for sleep stage classification
Yuyun Liu, Qilei Li, Mingliang Gao 0001, Wenzhe Zhai |
Appl. Intell. | 2 |
| 2025 | An efficient framework for general long-horizon time series forecasting with Mamba and Diffusion Probabilistic Models
Qilei Li, Ziwu Jiang, Deqian Fu, David Camacho |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | SUE - TS : A Surrogate Model Based Universal Explanation Framework for Time Series ForecastingabstractABSTRACT Deep learning models have achieved significant success in time series forecasting, substantially improving predictive accuracy across several applications. The interpretability of these intricate models continues to pose a significant problem, especially in high‐stakes fields where transparency and trust are essential. Current explanation strategies are often limited to static assessments, lacking consistency and semantic significance into model behaviour. This article introduces SUE‐TS (Surrogate‐based Universal Explanation for Time Series), a versatile interpretability framework that develops a simple and interpretable surrogate to replicate the predictive behaviour of opaque forecasting models, therefore addressing existing restrictions. SUE‐TS incorporates a SHAP‐based closed‐loop feedback mechanism that enhances prediction accuracy and semantic consistency within the explanation domain. The methodology guarantees fidelity via dual‐consistency evaluation: predictive consistency, which evaluates numerical concordance between surrogate and black‐box models, and explanation‐space consistency, which confirms semantic coherence through high‐dimensional representation analysis. Comprehensive studies on seven benchmark time series datasets and six advanced forecasting systems reveal that SUE‐TS consistently attains high approximation fidelity, interpretability, and robustness. The surrogate models reproduce output behaviours and maintain the reasoning logic of the original models, even in complex temporal dynamics. SUE‐TS offers a scalable, model‐agnostic solution for reliable time series forecasting by reconciling performance with explainability. It provides a significant resource for enhancing interpretability in actual AI applications, facilitating downstream activities such as model audits, knowledge distillation, and decision‐support system integration. Qilei Li, Muhammad Rizwan 0006, Deqian Fu, David Camacho |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | No-Reference Image Quality Assessment: Past, Present, and FutureabstractABSTRACT No‐reference image quality assessment (NR‐IQA) has garnered significant attention due to its critical role in various image processing applications. This survey provides a comprehensive and systematic review of NR‐IQA methods, datasets, and challenges, offering new perspectives and insights for the field. Specifically, we propose a novel taxonomy for NR‐IQA methods based on distortion scenarios and design principles, which distinguishes this work from previous surveys. Representative methods within each category are thoroughly examined, with a focus on their strengths, limitations, and performance characteristics. Additionally, we review 20 widely used NR‐IQA datasets that serve as benchmarks for evaluating these methods, providing detailed information on the number of images, distortion types, and distortion levels for each dataset. Furthermore, we identify and discuss key challenges currently faced by NR‐IQA methods, such as handling diverse and complex distortions, ensuring generalisation across datasets and devices, and achieving real‐time performance. We also suggest potential future research directions to address these issues. In summary, this survey offers a comprehensive and systematic examination of NR‐IQA methods, datasets, and challenges, offering valuable insights and guidance for researchers and practitioners working in the NR‐IQA domain. Qingyu Mao, Shuai Liu 0009, Qilei Li, Gwanggil Jeon, Hyunbum Kim, David Camacho |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | SFINet: A semantic feature interactive learning network for full-time infrared and visible image fusion
Qilei Li, Mingliang Gao 0001, Abdellah Chehri, Gwanggil Jeon |
Expert Syst. Appl. | 2 |
| 2025 | Traffic Density Estimation by Distributed Proxy Model Learning for Internet of VehicleabstractAutonomous driving has been significantly advanced in today’s society, which revolutionized daily routines and facilitated the development of the Internet of Vehicles (IoV). A crucial aspect of this system is understanding traffic density to enable intelligent traffic management. With the rapid improvement in deep neural networks (DNNs), the accuracy of density estimation has markedly improved. However, there are two main issues that remain unsolved. First, current DNN-based models are excessively heavy, characterized by an overwhelming number of training parameters (millions or even billions) and substantial computational complexity, indicated by a high number of FLOPs. These requirements for storage and computation severely limit the practical application of these models, especially on edge devices with limited capacity and computational power. Second, despite the superior performance of DNN models, their effectiveness largely depends on the availability of large-scale data for training. Growing privacy concerns have made individuals increasingly hesitant to allow their data to be publicly used for model training, particularly in vehicle-related applications that might reveal personal movements, which leads to data isolation issues. In this article, we address these two problems at once with a systematic framework. Specifically, we introduce the proxy model distributed learning (PMDL) model for traffic density estimation. PMDL model is composed of two main components. First, we introduce a proxy model learning strategy that transfers fine-grained knowledge from a larger master model to a lightweight proxy model, i.e., a proxy model. Second, we design a distributed learning strategy that trains multiple proxy models with privacy-aware local data and seamlessly aggregates these models via a global parameter server. This ensures privacy protection while significantly improving estimation performance compared to training models with limited, isolated data. We tested the proposed model on four major vehicle density analysis benchmarks and demonstrated its efficiency by outperforming other state-of-the-art competitors. The code is available athttps://github.com/jinyongch/DPML. Qilei Li, Jingan Cheng, Mingliang Gao 0001, Gwanggil Jeon |
IEEE Internet Things J. | 1 |
| 2025 | Cross-Camera Discriminative Person Association by Unsupervised Frame Clustering and SelectionabstractThe objective of cross-camera persson association is to identify individuals captured across disjoint cameras. It is achieved by Re-Identification (ReID) models, which extract unique identity representations from the visual input, dominated by video sequence. Most ReID methods primarily focus on modifying the backbone network architecture to learn more representative features of people. However, these methods often overlook the impact of low-quality frames on the training process. Several studies have confirmed that low-quality data not only hinders the model from learning meaningful content but also diminishes its performance. One possible solution is to manually label the quality of each frames, but this is time-consuming and inefficient. In this paper, we propose a Unsupervised Frame Clustering and Selection framework called UFCS to address this problem by applying the unsupervised clustering to automatically select high-quality frames. Specifically, we applied three unsupervised clustering solutions for high-quality frame selection, namely K-means, Deep K-means, and DBSCAN. These clustering techniques integrate both appearance and, indirectly, temporal consistency by operating within tracklets. These schemes perform clustering in the image or deep feature space to select highquality frames for network training. This straightforward yet effective approach enables the ReID network to generate a more discriminative representations, thereby improving recognition performance. Experimental results obtained on the challenging video-based person ReID datasets MARS indicate that our proposed scheme can outperform related state-of-the-art methods by a large margin. Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai, Gwanggil Jeon |
IEEE Internet Things J. | 1 |
| 2025 | Towards trustworthy image super-resolution via symmetrical and recursive artificial neural network
Mingliang Gao 0001, Jianhao Sun, Qilei Li, Muhammad Attique Khan, Jianrun Shang, Xianxun Zhu, Gwanggil Jeon |
Image Vis. Comput. | 3 |
| 2025 | Generalizable deepfake detection via Spatial Kernel Selection and Halo Attention Network
Siyou Guo, Qilei Li, Mingliang Gao 0001, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 2 |
| 2025 | Deepfake detection via Feature Refinement and Enhancement Network
Weicheng Song, Siyou Guo, Mingliang Gao 0001, Qilei Li, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 4 |
| 2025 | WiKAN: Lightweight Kolmogorov-Arnold Networks for accurate indoor WiFi localization
Yunlong Gu, Meng Xu 0022, Jiguang Li, Qilei Li, Mengshan Li, Lixin Guan, Mikko Valkama |
Pervasive Mob. Comput. | 4 |
| 2025 | Towards text-refereed multi-modal image fusion by cross-modality interaction
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai |
Signal Process. | 1 |
| 2025 | Training-Free 3-D Face Avatars Generation by Knowledge Discovering in Foundational ModelsabstractAn informative 3-D avatar, closely mirroring real-world traits, plays a pivotal role in accessing the metaverse. Traditional methods for creating 3-D avatars usually employ one-to-one training, which restricts avatar diversity. To enhance style diversity in generated 3-D avatars, we utilize synthesized images derived with prompts from ChatGPT in a conversational manner, ultimately resulting in a broader range of 3-D variations. Rather than creating models from scratch, we devise a training-free framework that utilizes established large-scale foundation models. Specifically, we employ a real-world image synthesis technique guided by text prompts that are generated by ChatGPT in a conversational manner, to describe the desired characteristics of the synthesized image. As a result, these informative latent representations can accurately reflect the distinct style of the synthesized image, and further lead to the creation of photorealistic and diverse 3D avatars. Our training-free design allows this proposed method to achieve competitive performance compared to existing generation models, while requiring minimal computational resources. Qilei Li, Wenzhe Zhai, David Camacho, Gwanggil Jeon |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Context-Aware Deepfake Detection for Securing AI-Driven Financial TransactionsabstractThe rapid advancement of deepfake technology has threatened the community’s sense of security, particularly in the context of face-based payment systems. Thus, deepfake detection has emerged as a critical issue demanding immediate attention. However, the generalization performance of existing detection models is limited as they are overly reliant on specific forged features while ignoring the common forged features. To address this problem, we introduce the context-aware decoupling network (CADNet) for deepfake detection. Specifically, a context self-calibration (CSC) module is constructed to guide the network to focus on local forged regions. It enlarges possible regions to increase the likelihood of forgery cues. Meanwhile, a frequency domain decoupling (FDD) module is introduced to extract and fuse different frequency components. It realizes the collaborative representation optimization of global semantics and local details. The experimental results prove that the proposed model exhibits strong generalization capability across multiple standard datasets. It achieves average area under the curve (AUC) values of 98.64% for in-domain evaluation and 75.52% for cross-dataset generalization. Changcun Liu, Guisheng Zhang, Siyou Guo, Qilei Li, Gwanggil Jeon, Mingliang Gao 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Space-Frequency and Global-Local Attentive Networks for Sequential Deepfake DetectionabstractThe widespread misinformation generated by deepfake systems has emerged as a significant challenge in the dynamic realm of digital media. It poses threats to credibility, privacy, and security of information in daily life. Moreover, the increasing accessibility to facial editing tools further enables users to alter facial characteristics subtly through a series of intricate steps. To address the issue, we introduce a space–frequency and global–local attentive network (SFGLA-Net) for sequential deepfake detection. This method is designed to identify and analyze the sophisticated manipulated attributes of deepfake images. Specifically, we introduce a space–frequency fusion module to leverage the deep feature extracted in spatial and frequency domains, so as to exploit subtle inconsistencies and artifacts that are not perceptible in the spatial domain alone. Additionally, we design a global–local attention module to pinpoint the manipulated areas more accurately. Extensive experiments demonstrate the superior performance of the proposed method by significantly outperforming existing techniques in sequential deepfake detection. The code is available athttps://github.com/guishengzhanga/SFGLA. Guisheng Zhang, Qilei Li, Mingliang Gao 0001, Siyou Guo, Gwanggil Jeon, Ahmed M. Abdelmoniem |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Zero-Shot Object Counting With Vision-Language Prior Guidance NetworkabstractThe majority of existing counting models are designed to operate on a singular object category, such as crowds or vehicles. The emergence of multi-modal foundational models, e.g., Contrastive Language-Image Pre-training (CLIP), has paved the way for class-agnostic counting. This approach facilitates the counting of objects across diverse classes within a single image based on textual indications. However, class-agnostic counting models based on CLIP confront two primary challenges. Firstly, the CLIP model exhibits limited sensitivity towards location information, which prioritizes global content over the precise localization of objects. Therefore, directly employing the CLIP model is regarded as suboptimal. Secondly, these models commonly employ frozen pre-trained vision and language encoders while disregarding potential misalignment within the constructed hypothesis space. In this paper, we propose a unified framework, named the Vision-Language Prior Guidance (VLPG) Network, to tackle these two challenges. The VLPG consists of three key components, namely the Grounding DINO module, Spatial Prior Calibration (SPC) module, and Object-Centric Alignment (OCA) module. The Grounding DINO module utilizes the spatial-awareness capability of extensive pre-trained object grounding models to incorporate the spatial position as an additional prior for a particular query class. This adaptation enables the network to concentrate more precisely on the exact location of the objects. Meanwhile, the SPC module is built to extract the long-range dependencies and local regions of the spatial position. Additionally, to align the feature space across different modalities, we design an OCA module that condenses textual information into an object query which serves as an instruction for cross-modality matching. Through the collaborative efforts of these three modules, multimodal representations are aligned while maintaining their discriminative nature. Comprehensive experiments conducted on various benchmarks validate the effectiveness of the proposed model. Wenzhe Zhai, Xianglei Xing, Mingliang Gao 0001, Qilei Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Generic Representation Learning for Vehicle Association Guided by Foundational ModelsabstractVehicle association is a vital yet complex task to retrieve specific vehicles across various camera angles, time frames, and geographical locations. In environments supported by autonomous driving and 6G networks, this task plays a vital role in urban surveillance and traffic management by enabling the real-time sharing of vehicle location and status information through ultra-high-speed, low-latency 6G communication. The success of a retrieval model largely depends on the quality of the extracted representations, which can be influenced by factors such as background diversity and occlusions. This study proposes a method to extract representations that remain consistent across different domains while retaining the discriminative power necessary to determine a vehicle’s spatial location, regardless of background or environmental variations. To achieve this, we introduce a framework called Generic Representation Learning (GRL). Within GRL, we leverage large-scale pre-trained foundational models to provide spatial priors of vehicles, specifically the Grounding DINO model for object detection and the SAM model for object segmentation. These modules collaborate to help the network understand the spatial context of the object, enabling the feature extractor to focus on discriminative areas while minimizing interference. Additionally, we introduce a complementary feature alignment mechanism based on a memory bank to explore globally applicable knowledge within the learned representation of the object. These constituent elements collectively form SRP, to enhance its capability for outstanding performance in vehicle retrieval. Extensive experimentation demonstrates that SRP significantly outperforms existing models on widely recognized benchmarks. Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Gwanggil Jeon, Ahmed M. Abdelmoniem |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Feature-Distribution Perturbation and Calibration for Generalized ReidabstractPerson Re-identification (ReID) has been advanced remarkably over the last 10 years. However, the i.i.d. (independent and identically distributed) assumption is somewhat non-applicable to ReID considering its objective to identify images of the same pedestrian across cameras at different locations. In this work, we propose a Feature-Distribution Perturbation and Calibration (PECA) method to derive generic feature representations for person ReID. Specifically, we perform per-domain feature-distribution perturbation to refrain the model from overfitting to the domain-biased distribution of each source (seen) domain by enforcing feature invariance to distribution shifts caused by perturbation. Furthermore, we design a global calibration mechanism to align feature distributions across all the source domains to improve the model’s generalization capacity by eliminating domain bias. These local perturbation and global calibration are conducted simultaneously, which share the same principle to avoid models overfitting by regularization respectively on the perturbed and the original distributions. Extensive experiments were conducted and the proposed PECA model outperformed the state-of-the-art competitors by significant margins. Qilei Li, Jiabo Huang, Jian Hu 0002, Shaogang Gong |
ICASSP | 1 |
| 2024 | Deepfake Detection via a Progressive Attention NetworkabstractThe rapid advancement of deepfake technology has enabled the creation of highly realistic forged face images or videos. While deepfake technology adds entertainment to people’s lives, it also poses a potential threat to social security. Deepfake detection is a crucial technology for identifying forged images. However, existing deep learning-based models for deepfake detection often overlook subtle forged traces. To solve this problem, we propose a Progressive Attention Network (PANet). The PANet incorporates two attention modules, namely the Efficient Multi-Scale Attention Module (EMAM) and the Spatial and Channel Attention Module (SCAM), in a progressive manner. The EMAM focuses on crucial facial regions, such as the eyes, nose, and mouth, rather than the entire face. The SCAM facilitates fine-grained feature extraction. Experimental results demonstrate that the proposed method achieves state-of-the-art results on deepfake detection datasets. Siyou Guo, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, David Camacho |
IJCNN | 3 |
| 2024 | Mitigate Domain Shift by Primary-Auxiliary Objectives Association for Generalizing Person ReIDabstractWhile deep learning has significantly improved ReID model accuracy under the independent and identical distribution (IID) assumption, it has also become clear that such models degrade notably when applied to an unseen novel domain due to unpredictable/unknown domain shift. Contemporary domain generalization (DG) ReID models struggle in learning domain-invariant representation solely through training on an instance classification objective. We consider that a deep learning model is heavily influenced and therefore biased towards domain-specific characteristics, e.g., background clutter, scale and viewpoint variations, limiting the generalizability of the learned model, and hypothesize that the pedestrians are domain invariant owning they share the same structural characteristics. To enable the ReID model to be less domain-specific from these pure pedestrians, we introduce a method that guides model learning of the primary ReID instance classification objective by a concurrent auxiliary learning objective on weakly la-beled pedestrian saliency detection. To solve the problem of conflicting optimization criteria in the model parameter space between the two learning objectives, we introduce a Primary-Auxiliary Objectives Association (PAOA) mechanism to calibrate the loss gradients of the auxiliary task towards the primary learning task gradients. Benefiting from the harmonious multitask learning design, our model can be extended with the recent test-time diagram to form the PAOA+, which performs on-the-fly optimization against the auxiliary objective in order to maximize the model’s generative capacity in the test target domain. Experiments demonstrate the superiority of the proposed PAOA model. Qilei Li, Shaogang Gong |
WACV | 1 |
| 2024 | Dual-branch and triple-attention network for pan-sharpening
Mingliang Gao 0001, Abdellah Chehri, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Appl. Intell. | 5 |
| 2024 | Multiscale aggregation and illumination-aware attention network for infrared and visible image fusionabstractAbstract Image fusion plays a significant role in computer vision since numerous applications benefit from the fusion results. The existing image fusion methods are incapable of perceiving the most discriminative regions under varying illumination circumstances and thus fail to emphasize the salient targets and ignore the abundant texture details of the infrared and visible images. To address this problem, a multiscale aggregation and illumination‐aware attention network (MAIANet) is proposed for infrared and visible image fusion. Specifically, the MAIANet consists of four modules, namely multiscale feature extraction module, lightweight channel attention module, image reconstruction module, and illumination‐aware module. The multiscale feature extraction module attempts to extract multiscale features in the images. The role of the lightweight channel attention module is to assign different weights to each channel so as to focus on the essential regions in the infrared and visible images. An illumination‐aware module is employed to assess the probability distribution regarding the illumination factor. Meanwhile, an illumination perception loss is formulated by the illumination probabilities to enable the proposed MAIANet to better adjust to the changes in illumination. Experimental results on three datasets, that is, MSRS, TNO, and RoadSence, verify the effectiveness of the MAIANet in both qualitative and quantitative evaluations. Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Abdellah Chehri, Gwanggil Jeon |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Efficient blind super-resolution imaging via adaptive degradation-aware estimation
Haoran Yang 0008, Qilei Li, Bin Meng 0001, Gwanggil Jeon, Kai Liu 0012, Xiaomin Yang |
Knowl. Based Syst. | 2 |
| 2024 | SaReGAN: a salient regional generative adversarial network for visible and infrared image fusion
Mingliang Gao 0001, Yi'nan Zhou, Wenzhe Zhai, Qilei Li |
Multim. Tools Appl. | 5 |
| 2024 | Multiscale aggregation network via smooth inverse map for crowd counting
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou |
Multim. Tools Appl. | 4 |
| 2024 | A structure and texture revealing retinex model for low-light image enhancement
Qilei Li, Marco Anisetti, Gwanggil Jeon, Mingliang Gao 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Object counting in remote sensing via selective spatial-frequency pyramid networkabstractAbstract The integration of remote sensing object counting in the Mobile Edge Computing (MEC) environment is of crucial significance and practical value. However, the presence of significant background interference in remote sensing images poses a challenge to accurate object counting, as the results are easily affected by background noise. Additionally, scale variation within remote sensing images presents a further difficulty, as traditional counting methods face challenges in adapting to objects of different scales. To address these challenges, we propose a selective spatial‐frequency pyramid network (SSFPNet). Specifically, the SSFPNet consists of two core modules, namely the pyramid attention (PA) module and the hybrid feature pyramid (HFP) module. The PA module accurately extracts target regions and eliminates background interference by operating on four parallel branches. This enables more precise object counting. The HFP module is introduced to fuse spatial and frequency domain information, leveraging scale information from different domains for object counting, so as to improve the accuracy and robustness of counting. Experimental results on RSOC, CARPK, and PUCPR+ benchmark datasets demonstrate that the SSFPNet achieves state‐of‐the‐art performance in terms of accuracy and robustness. Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Softw. Pract. Exp. | 5 |
| 2024 | Defending Deepfakes by Saliency-Aware AttackabstractWith the rapid development of deep learning, especially the generative adversarial network (GAN), face modification has been substantially advanced and enables the generated images to look more realistic. Given an image or a video frame of a person, such a system can create fake images, which manipulates the movement, expression, and even appearance, e.g., hair color, eye color, and age. Such a system is termed Deepfake, which has raised significant ethical issues, especially for celebrities. With the pretrained Deepfake models being widely available on the Internet, its negative applications, such as face manipulation and pornographic generation, have exposed the dark side of the Deepfake technology to the sociocyber world. In this article, we aim to defend a well-trained Deepfake model by manipulating the raw image with unperceived perturbation. To minimize the alterations to the original image while effectively fooling the Deepfake model, we propose to selectively perturb only the foreground person region and maintain the irrelevant background. This is based on the observation that the salient object in a person’s image is always the foreground face region. Such a strategy introduces negligible alterations to the original image, which makes the attack remain effective. We experimentally demonstrate the superiority of the proposed attacking framework over the existing models and show our approach is ready to be applied for out-of-the-box development. Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | FPANet: feature pyramid attention network for crowd counting
Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, Marco Anisetti |
Appl. Intell. | 3 |
| 2023 | Crowd counting in smart city via lightweight Ghost Attention Pyramid Network
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
Future Gener. Comput. Syst. | 5 |
| 2023 | Scale-Context Perceptive Network for Crowd Counting and Localization in Smart City SystemabstractThe task of crowd counting and localization is to predict the count and position of people in a crowd, which is a practical and essential sub-task in crowd analysis and smart city systems. However, the inherent problems of scale variation and background disturbance restrain their performance. While recent researches focus on studying counting and localization independently, a few works are capable of executing both tasks simultaneously. To this end, we propose a Scale-Context Perceptive Network (SCPNet) to jointly tackle the crowd counting and localization tasks in a unified framework. Specifically, a scale perceptive (SP) module with a local-global branch schema is designed to capture multiscale information. Meanwhile, a context perceptive (CP) module, by the channel-spatial self-attention mechanism, is derived to suppress the background disturbance. Furthermore, a novel hierarchical scale loss function that combines the Euclidean loss function and structural similarity loss function is designed to prompt the proposed model to fulfill the counting and localization simultaneously. Extensive experiments on challenging crowd datasets prove the superiority of the proposed SCPNet compared with the state-of-the-art competitors in both objective and subjective evaluations. Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon |
IEEE Internet Things J. | 4 |
| 2023 | $\hbox {DA}^2$Net: a dual attention-aware network for robust crowd counting
Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou, Mingliang Gao 0001 |
Multim. Syst. | 2 |
| 2023 | Dense Attention Fusion Network for Object Counting in IoT System
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Kyu Hyung Kim, Gwanggil Jeon |
Mob. Networks Appl. | 4 |
| 2023 | Rapid Person Re-Identification via Sub-space Consistency Regularization
Qingze Yin, Guan'an Wang, Guodong Ding, Qilei Li, Shaogang Gong, Zhenmin Tang |
Neural Process. Lett. | 4 |
| 2023 | Scale Region Recognition Network for Object Counting in Intelligent Transportation SystemabstractSelf-driving technology and safety monitoring devices in intelligent transportation systems require superb capacity for context awareness. Accurately inferring the counts of crowds and vehicles are the two practical and fundamental tasks in the transportation system. However, the scale variation and background interference in the traffic image hinder the counting performance. To solve the aforementioned problems, a scale region recognition network (SRRNet) is proposed in this paper. It has two key components, termed scale level awareness (SLA) module and object region recognition (ORR) module. The SLA module aims to encode the representations at multiple scales, which are beneficial to address the scale variation. The ORR module is designed to suppress background interference through the visual attention mechanism. Extensive experimental results on four crowd counting datasets and five vehicle counting datasets have demonstrated the superiority of the proposed SRRNet in both counting accuracy and robustness compared with the mainstream competitors. Meanwhile, substantial ablation studies have proved the effectiveness of the proposed SLA and ORS modules. Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Visual tracking for UAV using adaptive spatio-temporal regularized correlation filters
Libin Xu, Mingliang Gao 0001, Qilei Li, Guofeng Zou, Jinfeng Pan |
Appl. Intell. | 3 |
| 2022 | Accelerated duality-aware correlation filters for visual tracking
Libin Xu, Mingliang Gao 0001, Zheng Liu 0002, Qilei Li, Gwanggil Jeon |
Neural Comput. Appl. | 4 |
| 2021 | Local-Global Associative Frame Assemble in Video Re-ID
Qilei Li, Jiabo Huang, Shaogang Gong |
BMVC | 1 |
| 2021 | Pansharpening multispectral remote-sensing images with guided filter for monitoring impact of human behavior on environmentabstractSummary Human behavior would lead to a significant impact on the environment. By monitoring the environment, we can indirectly monitor human behavior. Remote sensing (RS) technology provides a large number of multispectral (MS) images. When combining the Internet of things (IoT) technology, those images can be used for human behavioral monitoring. However, due to the limitation of the optical sensors embedded in satellites, the spatial resolution of MS image is relatively low, which poses a huge problem for further understanding these images. Pansharpening, also known as multisensor image fusion, aims to sharp an MS image to a high‐resolution multisensor image (HMS) by integrating a corresponding high‐resolution panchromatic (PAN) image. By doing so, the redundancy among big data can be effectively reduced. Traditional Intensity‐Hue‐Saturation (IHS)–based methods often suffer from spectral distortion. To address this problem, a novel pansharpening method is proposed in this paper. Different from those traditional IHS methods, the proposed method first decomposes MS and PAN into high‐frequency‐component (HFC) and low‐frequency‐component (LFC), respectively. Then, the guided filter (GF) is utilized to enhance the spectral information on the detail map. Furthermore, the detail map is refined according to the adaptive coefficients for each band of MS. By performing experiments, we demonstrate the proposed method can obtain satisfying results in both visual quality and object assessment among existing methods. Qilei Li, Xiaomin Yang, Wei Wu 0002, Kai Liu 0012, Gwanggil Jeon |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Adaptive Spatio-Temporal Regularized Correlation Filters for UAV-Based Tracking
Libin Xu, Qilei Li, Guofeng Zou, Zheng Liu 0002, Mingliang Gao 0001 |
ACCV (2) | 2 |
| 2020 | Deep recursive up-down sampling networks for single image super-resolution
Zhen Li 0031, Qilei Li, Wei Wu 0002, Jinglei Yang, Xiaomin Yang |
Neurocomputing | 2 |
| 2020 | Clustering based multiple branches deep networks for single image super-resolution
Zhen Li 0031, Qilei Li, Wei Wu 0002, Zongjun Wu, Lu Lu 0005, Xiaomin Yang |
Multim. Tools Appl. | 2 |
| 2019 | Gated Multiple Feedback Network for Image Super-Resolution
Qilei Li, Zhen Li 0031, Lu Lu 0005, Gwanggil Jeon, Kai Liu 0012, Xiaomin Yang |
BMVC | 1 |
| 2009 | Performance-driven motion choreographing with accelerometersabstractAbstract Live performance is an intuitive way to naturally draft the desired motion in the choreographer's mind. In this paper we present a novel approach to choreographing motions by live performance captured with degree of freedom (3‐DOF) accelerometers. The process begins by placing the accelerometers on the user's limbs according to the pre‐specified positions. The computer then recognizes the performed actions using Hidden Markov Model (HMM), which is pre‐trained by the acceleration data samples automatically generated from a pre‐segmented motion capture database. At last, the captured actions are further synthesized with motion retiming and exaggeration based on the acceleration signals from the accelerometers. This method can intuitively rapid‐prototype the choreographed motions for pre‐production of animation, the avatar control in virtual reality and game‐like scenarios, etc. The experimental results show that it can effectively recognize actions with spatial‐time variance, and is easy‐to‐use especially for a novice with little experience. Copyright © 2009 John Wiley & Sons, Ltd. Xiubo Liang, Qilei Li, Weidong Geng |
Comput. Animat. Virtual Worlds | 2 |
| 2008 | Interactive Animation of Virtual Characters: Application to Virtual Kung-Fu FightingabstractThis paper aims at proposing a framework for animating virtual humans that can efficiently interact with real users in virtual reality (VR). If the user's order can be modeled as targets and commands, the system searches a database for the most convenient behavior. In order to avoid using a huge database that can deal with any kind of situation, we propose to associate this searching process to an adaptation module. Hence, even if the selected motion is not perfectly suited with the situation, it can be adapted in order to reach accurately the target specified by the user. This framework is illustrated with a kung-fu fighter example. Two people are involved in this example: the user and the supervisor. The user is displacing in the real environment while the position of his head is tracked in real-time thanks to reflective markers. The virtual opponent follows the displacement of the user to stay close to him. At any time, the supervisor can ask the virtual character to kick or punch the user. Our system automatically searches for the convenient motion in an average-size database (75 motions compared to hundreds of motions required for motion graphs) and adapts it to the current situation. Nicolas Pronost, Franck Multon, Qilei Li, Weidong Geng, Richard Kulpa, Georges Dumont |
CW | 3 |
| 2005 | Mocap data editing via movement notationsabstractIn most motion editing approaches, the users often make changes on mocap data by specifying numerical parameters. However, it is not intuitive for a novice to edit motion sequences in motion planning tasks. In this paper, we present a method of notation-based motion editing, in which Labanotation, a well-developed notation language for human movements, is employed as the editing interface. It allows the user to specify the editing requirements via notation score, and the system will semi-automatically generate the desired motion data. The algorithmic steps of its pipeline and the core issues of converting Labanotation into motion sequences are discussed in detail. We show the results of our systems using martial arts motion as its testing data. XiaoJie Shen, Qilei Li, Weidong Geng, Newman Lau |
CAD/Graphics | 2 |
| 2005 | Motion retrieval based on movement notation languageabstractAbstract With the increased availability of motion capture data, the volume of motion library grows so large that it is difficult for animators to manually browse the dataset to search desired motions for reuse. To address this issue, we implement a framework, which allows the user to retrieve motions via Labanotation. For each motion clip in the library, we generate a corresponding Labanotation sequence as additional motion property. A similarity metric for Labanotation sequences is proposed and used to search the motions that have similar Laban descriptions. Our search algorithm is able to retrieve motion segments that only match part of the query Laban sequence. Then based on dynamic programming, these segments are stitched together to form a smooth output motion that is in an optimal sense of matching query Laban sequence. Experimental results demonstrate our method could effectively improve the utilization of motion data. Copyright © 2005 John Wiley & Sons, Ltd. XiaoJie Shen, Qilei Li, Weidong Geng |
Comput. Animat. Virtual Worlds | 3 |