EDBT 2026 Demo / reviewers in the wild / expert
Weilong Chen
dblp:19/6157
· DBLP profile ↗
23ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Socially Aware Load Forecasting Utilizing Large Language Models
Weilong Chen, Xinran Zhang 0006, Zheng Chang 0001, Zhu Han 0001, Yanru Zhang |
IEEE Trans. Ind. Informatics | 1 |
| 2026 | Semantic Communication Based on Large Language Model for Underwater Image TransmissionabstractUnderwater communication is essential for environmental monitoring, marine biology research, and underwater exploration. Traditional underwater communication faces limitations like low bandwidth, high latency, and susceptibility to noise, while semantic communication (SC) offers a promising solution by focusing on the exchange of semantics rather than symbols or bits. However, SC encounters challenges in underwater environments, including semantic information mismatch and difficulties in accurately identifying and transmitting critical information that aligns with the diverse requirements of underwater applications. To address these challenges, we propose a novel SC framework based on Large Language Models (LLMs). Our framework leverages visual LLMs to perform semantic compression and prioritization of underwater image data according to the query from users. By identifying and encoding key semantic elements within the images, the system selectively transmits high-priority information while applying higher compression rates to less critical regions. On the receiver side, an LLM-based recovery mechanism, along with Global Vision ControlNet and Key Region ControlNet networks, aids in reconstructing the images, thereby enhancing communication efficiency and robustness. Our framework reduces the overall data size to 0.8% of the original. Experimental results demonstrate that our method significantly outperforms existing approaches, ensuring high-quality, semantically accurate image reconstruction. Weilong Chen, Xinran Zhang 0006, Zhijin Qin, Yanru Zhang, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose EstimationabstractCategory-agnostic pose estimation aims to locate keypoints on query images according to a few annotated support images for arbitrary novel classes. Existing methods generally extract support features via heatmap pooling, and obtain interacted features from support and query via cross-attention. Hence, these works neglect to mine fine-grained and structure-aware (FGSA) features from both support and query images, which are crucial for pixel-level keypoint localization. To this end, we propose a novel yet concise framework, which recurrently mines FGSA features from both support and query images. Specifically, we design a FGSA mining module based on deformable attention mechanism. On the one hand, we mine fine-grained features by applying deformable attention head over multi-scale feature maps. On the other hand, we mine structure-aware features by offsetting the reference points of keypoints to their linked keypoints. By means of above module, we recurrently mine FGSA features from support and query images, and thus obtain better support features and query estimations. In addition, we propose to use mixup keypoints to pad various classes to a unified keypoint number, which could provide richer supervision than the zero padding used in existing works. We conduct extensive experiments and in-depth studies on large-scale MP-100 dataset, and outperform SOTA method dramatically (+3.2%[email protected]). Weilong Chen |
CVPR | 2 |
| 2025 | Large Language Model-Enhanced Deep Reinforcement Learning for Interpretable Microgrid ManagementabstractWith the intensification of energy crises and climate change, intelligent microgrid management has become critical for achieving energy resilience and emission reduction. Traditional deep reinforcement learning (DRL) methods excel in optimizing power dispatch and storage utilization but suffer from insufficient interpretability due to their black-box nature. This paper proposes a hybrid framework integrating large language models (LLMs) with DRL to enhance both decision-making performance and interpretability in microgrid management. Leveraging a multi-agent interactive scenario and a modular reward function, the framework balances supply-demand matching, grid interaction, and energy storage coordination. Through carefully engineered prompts, LLMs refine DRL actions and provide real-time natural language explanations for control adjustments. Experimental results demonstrate that the LLM+DRL composite model improves the overall performance score by 0.5%–1.6% compared to standalone DRL while enabling transparent reasoning. The study also explores LLMs scheduling capabilities under varying prior knowledge, highlighting the importance of contextual information for decision quality. This work offers a novel pathway for intelligent and interpretable microgrid management systems, bridging the gap between data-driven optimization and operational transparency. Zhuo Lan, Yuanqing Cai, Weilong Chen, Yanru Zhang |
GLOBECOM | 4 |
| 2025 | Feature Disentangling Dual-stream Network for User Bias Alleviation in Social Media PredictionabstractSocial media popularity prediction is increasingly crucial for optimizing user engagement and guiding content recommendation systems. However, existing methods suffer from an excessive reliance on user information, which disproportionately influences predictions and leads to the neglect of content diversity. This oversight results in user bias, which adversely impacts the accuracy of predictions. In this paper, an approach named Feature Disentangling Dual-Stream Network (FDDN) is introduced to address this gap. In FDDN, we introduce the Multimodal Extraction Module to extract content features from different modalities. Additionally, the User Popularity Extraction Module helps to analyze the impact of user features on popularity. The Disentangled Adaptation Module distinguishes the impact of user features from content features, thus alleviating user bias and ensuring a more comprehensive prediction. Extensive experiments on a large public dataset demonstrate the robustness and effectiveness of our approach, indicating superior performance compared to existing methods. Weilong Chen, Weimin Yuan, Xiaolu Chen, Yanru Zhang, Zhu Han 0001 |
ICASSP | 2 |
| 2025 | Privacy-Preserving Socio-Aware Short-Term Residential Load ForecastingabstractThis paper introduces a novel approach named the Privacy-Preserving Socio-Aware Model (PSocLF) for Short-Term Residential Load Forecasting, which addresses the need for forecasting models tailored to district-specific socio-demographic characteristics. By leveraging sociodemographic characteristics and personalized socio-aware knowledge sharing, PSocLF develops district-level forecasting models that enhance load forecasting precision while safeguarding individual privacy. Within PSocLF, we propose a new model structure named the Self-Gating TSMixer (SGTSMixer), which integrates self-gating mixing procedures with stacked multi-layer perceptrons (MLPs). It can efficiently extract temporal patterns and incorporate socioaware information to improve prediction accuracy. Simulation results based on real-world data demonstrate the effectiveness of the proposed PSocLF framework, outperforming alternative training paradigms and benchmarks in model structure design, particularly in scenarios with varying sociodemographic characteristics among districts. This paper contributes to advancing federated residential load forecasting and highlights the practical benefits of integrating sociodemographic information for improved forecasting accuracy and effectiveness. Weilong Chen, Yixin Liang, Zheng Chang 0001, Yanru Zhang, Zhu Han 0001 |
ICC | 1 |
| 2025 | Tri-Modal Transformers With Mixture-of-Modality-Experts for Social Media PredictionabstractWith billions of users worldwide, accurately predicting social media popularity is crucial for assessing user behavior, forecasting trends, and enhancing social interactions and business strategies. However, this task presents significant challenges. Firstly, the extraction of valuable insights is complicated by the presence of tri-modal data (visual, text, structured) and pervasive noise. Secondly, the applicability of knowledge acquired during the pre-training phase is often limited due to discrepancies with downstream prediction tasks during the fine-tuning phase. Existing methods for Social Media Popularity Prediction (SMPP), including traditional models and Visual-and-language Models (VLMs), struggle to overcome these challenges, thereby failing to achieve satisfactory accuracy. To tackle these challenges, we propose a novel approach named Tri-Modal Transformers with Mixture-of-Modality-Experts (TTME) for SMPP. TTME integrates Artificial Intelligence Generated Content to mitigate data noise and incorporate a mix of Modality Experts in pre-training phases to effectively utilize tri-modal data. Moreover, to address training disparity, we explore strategies for downstream task adaptation including the integration of diverse pre-training experts and the implementation of DistillSoftmax. Through empirical evaluation, we demonstrate that the TTME significantly improves the accuracy of social media popularity predictions, effectively utilizes tri-modal data with noise, and enhances transferring knowledge from pre-training to downstream tasks. Weilong Chen, Xiaolu Chen, Weimin Yuan, Yan Wang 0083, Yanru Zhang, Zhu Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Zero-Trust Based Robust Federated Learning Against Betrayal BehaviorsabstractDue to its advantage of protecting data privacy and reducing communication overhead, Federated Learning (FL) is becoming a promising machine learning paradigm. However, resource limitations and unstable communication connections on the participating client end can lead to unintentional failures that degrade FL performance. Moreover, as FL systems scale and interconnect increasingly, they face growing exposure to intentional network risks. Furthermore, the assumption of continued trust in historically benign clients introduces vulnerabilities to potential internal betrayal within FL systems. In this paper, we enhance the robustness of FL by incorporating the zero-trust principle, which eliminates implicit trust in clients and mitigates unintentional failures, intentional attacks, and strategic betrayal risks. The framework incorporates dynamic client selection and aggregation weight allocation through trustworthiness evaluation and sustained skepticism toward each potential betrayal behavior. Specifically, a Dirichlet-based trust evaluation technique is presented to update clients' trustworthiness with evolving observations. Then, to reduce potential betrayal loss, we formulate a min-max optimization problem that minimizes the worst-case betrayal loss. Next, we transform the formulation into a convex programming problem for solution. Extensive simulations are conducted to demonstrate the efficacy of the zero-trust based FL in the accurate trust assessment and the system's betrayal-aware robustness enhancement. Xinran Zhang 0006, Dan Wang 0002, Yifei Zhu 0001, Weilong Chen, Zheng Chang 0001, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | When Zero-Trust Meets Federated LearningabstractNowadays, Federated Learning (FL) has emerged as a promising and critical machine learning scheme to protect data privacy and reduce communication overhead. As the scale and connectivity expand in the FL system, enhancing the model’s robustness against security threats from malicious clients grows ever more critical. An effective defensive solution involves selecting benign clients appropriately, thereby mitigating the vulnerability of the FL system to malicious attacks. However, clients exhibit varying behaviors over time, which complicates the task of accurately modeling their future trustworthiness. Moreover, blindly trusting clients with high trust values poses risks, given the potential for severe losses from betrayal. To tackle these problems, we propose a zero-trust policy in FL aimed at establishing continuous trust in each client while maintaining skepticism towards potential betrayal attacks. Specifically, we develop a Dirichlet-based trust evaluation technique to enable a comprehensive selection of trustworthy participants. This technique leverages the posterior distribution to estimate clients’ trust values from their evolving behavior records over time. Then, we anticipate potential betrayal from a selected client and formulate a min-max optimization problem to minimize the worst-case betrayal loss, thereby boosting the system’s betrayalaware robustness. Next, we convert this problem into a convex optimization problem and utilize the interior point method for resolution. We conduct extensive simulations to validate the efficacy of our proposed zero-trust policy in accurately assessing trust and enhancing the model’s robustness to betrayal. Xinran Zhang 0006, Dan Wang 0002, Yifei Zhu 0001, Weilong Chen, Zheng Chang 0001, Zhu Han 0001 |
GLOBECOM | 4 |
| 2024 | Multi-dimensional Resource Allocation in HAP-assisted UAV Wireless Networks for IoRT Data CollectionabstractIn this paper, we propose a multi-dimensional resource allocation scheme for Internet of Remote Things (IoRT) data collection in a high altitude platform (HAP)-assisted unmanned aerial vehicle (UAV) network. Considering the quality of service (QoS) requirements of delay-sensitive IoRT data, we propose a UAV-HAP double relay data transmission mode to reduce the transmission delay for delay-sensitive data. Since the resources of the UAV are limited, we jointly optimize communications, computing and storage resources to maximize the utility of the considered system. Due to the high dimensionality of the solution space, we design a Twin Delayed Deep Deterministic policy gradient-based multi-dimensional resource allocation (TD3-MDRA) algorithm to find the optimal resource allocation strategy. Extensive simulation results are presented to demonstrate the superior performance of TD3-MDRA for IoRT data collection with delay and resource constraints. Xinran Zhang 0006, Weilong Chen, Xiaobin Xu 0004, Li Wang 0039, Zheng Chang 0001 |
GLOBECOM | 3 |
| 2024 | Dual-Stream Pre-Training Transformer to Enhance Multimodal Learning for Social Media PredictionabstractSocial media has emerged as a vital platform for communication, information sharing, and acquisition. Predictive analysis of social media data has wide applications, such as sentiment examination and social network analysis. However, existing work often directly utilizes social media data for training, neglecting the issue of mismatched text and images. This neglect can lead to confusion about the contents, thereby affecting the identification of trending topics and the accuracy of social media predictions. In this paper, an approach named Dual-Stream Pre-training Transformer (DSPT) is introduced to address this gap. In DSPT, we use a Visual-Language Model (VLM) and a Language Model (LM) to separately learn from image and text data, mitigating the impact of text-image mismatches. Moreover, to enhance the understanding of the model to social media data, we conduct incremental pre-training for both models. To achieve better feature interaction, we construct an integrated regression module combining LightGBM and CatBoost, jointly predicting the extracted feature embeddings. This dual-stream multimodal feature extraction method improves the performance of predictive tasks. Experimental results validate the effectiveness of our approach, demonstrating its potential and providing deeper insights into multimodal data mining in social media. Weilong Chen, Weimin Yuan, Yan Wang 0083, Shimin Cai, Yanru Zhang |
ACM Multimedia | 2 |
| 2024 | Joint Accuracy and Latency Optimization for Quantized Federated Learning in Vehicular NetworksabstractNowadays, vehicular networks have emerged as a boosting technology to enhance traffic efficiency and safety within transportation systems. As the amount of onboard data increases and data privacy concerns grow, federated learning (FL) has gained popularity for harnessing the data for intelligent transportation operations. To satisfy the strict latency criteria in vehicular networks, a quantization scheme is employed within FL to reduce the size of local models before uplink transmission. In this paper, considering the nature of vehicles’ high mobility, we aim to optimize both the learning performance and latency simultaneously by jointly considering the communication resource budget and quantization strategies. Specifically, we first analyze the convergence performance of the quantized FL, which demonstrates the effects of both quantization error and the number of clients on the convergence rate. Then, we formulate a multi-objective optimization problem (MOP) to maximize the number of participating clients and minimize the overall latency, by jointly optimizing the quantization level, wireless resource allocation and client selection. To deal with the MOP, we decompose the MOP into a set of scalar optimization subproblems, each formulated as a Markov Decision Process (MDP). To solve the MDP in high-mobile vehicular networks, we propose a novel deep reinforcement learning-based vehicle heterogeneous quantization FL (DRL-VQFL) method, which leverages a DRL framework built upon the proximal policy optimization algorithm. Then, a parameter transfer strategy is employed to solve the neighboring subproblems efficiently. Our extensive simulations demonstrate the effectiveness and efficiency of the DRL-VQFL approach, showcasing its superiority over other benchmark methods. Xinran Zhang 0006, Weilong Chen, Zheng Chang 0001, Zhu Han 0001 |
IEEE Internet Things J. | 2 |
| 2024 | HQ-DCGAN: Hybrid quantum deep convolutional generative adversarial network approach for ECG generationabstractThe class imbalance of electrocardiogram (ECG) data is a serious impediment to the development of diagnostic systems for heart disease. To address this issue, this paper proposes HQ-DCGAN, a hybrid quantum deep convolutional generative adversarial network, specifically designed for the generation of ECGs. The proposed algorithm employs different quantum convolutional layers for the generator and discriminator as feature extractors and utilizes parameterized quantum circuits (PQCs) to enhance computational capabilities, along with the model-feature mapping process. Moreover, this algorithm preserves the nonlinearity and scalability inherent to classical convolutional neural networks (CNNs), thereby optimizing the utilization of quantum resources, and ensuring compatibility with contemporary quantum devices. In addition, this paper proposes a novel evaluation metric, 1D Fréchet Inception Distance (1DFID), to assess the quality of the generated ECG signals. Simulation experiments show that HQ-DCGAN exhibits strong performance in ECG signal generation. Furthermore, the generated signals achieve an average classification accuracy of 82.2%, outperforming the baseline algorithms. It has been experimentally proven that HQ-DCGAN is friendly to currently noisy intermediate-scale quantum (NISQ) computers, in terms of both number of qubits and circuit depths, while improving the stability. Zhiguo Qu, Weilong Chen, Prayag Tiwari |
Knowl. Based Syst. | 2 |
| 2024 | CIPPO: Contrastive Imitation Proximal Policy Optimization for Recommendation Based on Reinforcement LearningabstractRecommendation systems, widely adopted in social networks, personalize user experiences through advanced technologies such as Reinforcement Learning (RL), known for producing high-performance, list- wise recommendations. However, RL-based recommendation methods exhibit biases, specifically: 1) Online bias, which stems from a complex real-worldonline policycomposed of various rules and models rather than a single policy; 2) Training bias, a distributional shift resulting from differences between thetarget policyand thebehavior policy. To address these issues, we introduce a novel framework named Contrastive Imitation Proximal Policy Optimization (CIPPO) for recommendation based on RL. This approach leverages extensively labeled feedback data and incorporates a Masked Imitation Network (MIN) that closely emulates the online policy, thus reducing discrepancies between online and offline environments. Additionally, the clipping function in Proximal Policy Optimization, combined with a specially designed contrastive module, effectively reduces the distributional shift between the behavior and target policies. We conduct offline and online experiments to show the improvements of CIPPO, providing details including ablation tests and parameter analysis to validate the effectiveness and robustness. CIPPO gains 12.79% on ACN and in WeChat Top Stories, a large media platform with over 50 million users. Weilong Chen, Ruobing Xie, Feng Xia 0006, Leyu Lin, Xinran Zhang 0006, Yan Wang 0083, Yanru Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Vehicle Selection and Resource Allocation for Federated Learning-Assisted Vehicular NetworkabstractTo exploit the massive amounts of onboard data in vehicular networks while protecting data privacy and security, federated learning (FL) is regarded as a promising technology to support enormous vehicular applications. Despite that FL has great potential to improve the architecture of intelligent vehicular networks, the mobility of the vehicles and the dynamic nature of wireless channels make the integration of FL and vehicular networks more challenging. In this paper, we propose a vehicle mobility- and channel dynamic-aware FL (MADCA-FL) scheme to fit vehicular networks and enhance learning performances. This novel scheme enables the RSU to select appropriate vehicles and weightedly average the local models. Afterward, MADCA-FL formulates a problem to maximize the model accuracy while assuring the latency and energy restrictions, by jointly optimizing the computation and communication resources. With a mixed- integer non-linear programming structure, the problem is NP-hard. Firstly, we utilize the successive convex approximation algorithm to handle the non-convexity, and then apply the Lagrange multiplier method and the block coordinate descent method to obtain the optimal solution. Extensive experiments are conducted to confirm the effectiveness of our proposed scheme. Xinran Zhang 0006, Zheng Chang 0001, Tao Hu 0012, Weilong Chen, Xin Zhang 0122, Geyong Min |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer for Social Media Popularity PredictionabstractSocial media popularity prediction aims to predict future interaction or attractiveness of new posts. However, in most existing works, there is a notable deficiency in the effective treatment of numerical features. Despite their significant potential to provide ample information, these features are often inadequately processed, leading to insufficiency of information acquirement. In this paper, we introduce a method, named Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer (DFT-MOVLT). To supplement the information in vision-and-language pre-training (VLP), we propose compound text, which is concatenated by numerical data and text. Furthermore, during VLP, a transformer is trained using 3 objectives to ensure thorough feature extraction. Finally, for more generalized prediction, we fine-tune 2 models using different training ways and ensemble them. To evaluate the effectiveness of each mechanism adopted in the proposed method, we conduct an array of ablation experiments. Our team achieve the 3rd place in Social Media Prediction (SMP) Challenge 2023. Xiaolu Chen, Weilong Chen, Zhongjian Zhang, Lixin Duan, Yanru Zhang |
ACM Multimedia | 2 |
| 2022 | DearFSAC: A DRL-based Robust Design for Power Demand Forecasting in Federated Smart GridabstractPower demand forecasting plays a significant role in the operation of power plants and utility companies. For data privacy, federated learning (FL) is widely adopted to aggregate local models of utility companies to a global model with very few data leaks. However, defects such as malicious updates, poisoning attacks, and low-quality data, may exist in multiple FL processes. As the general resistance to various defects is not considered by most FL approaches, a design with strong generalization is strongly needed. In this paper, we adopt DEfect-AwaRe federated soft actor-critic (DearFSAC), which dynamically assigns weights to FL's local models according to their quality. For fast and stable convergence, a deep neural network based on auto-encoder is designed for model quality evaluation and dimension reduction. Then, a deep reinforcement learning (DRL) algorithm soft actor-critic (SAC) is adopted to achieve the optimal weights assignment, considering SAC's near-optimum and sufficient exploration. We conduct simulations on power consumption data in real world. The results show that our approach performs well no matter if there exist defects or not. Weilong Chen, Feng Hong 0005, Shunji Yang, Shengrong Bu, Changkun Jiang, Yingjie Zhou 0001, Yanru Zhang |
GLOBECOM | 2 |
| 2022 | Title-and-Tag Contrastive Vision-and-Language Transformer for Social Media Popularity PredictionabstractSocial media is an indispensable part of modern life, and social media popularity prediction (SMPP) plays a vital role in practice. In current work, the inconsistency of words in labels and titles, user feature transformation, etc have not been well noticed. In this paper, we propose a novel approach named Title-and-Tag Contrastive Vision-and-Language Transformer (TTC-VLT), combining two pre-trained vision and language transformers and other two dense feature parts for this prediction task. On one hand, in order to learn the differences between titles and tags, we design title-tag contrastive learning for title-visual and tag-visual, which separately extracts multimodal information from two types of text. On the other hand, user identification features are transformed to embedding vectors to capture user attribute details. From the extensive experiments, our approach outperforms the other methods on the social media prediction dataset. Our team achieve the 2nd place on the leader board of the Social Media Prediction Challenge 2022. Weilong Chen, Weimin Yuan, Xiaolu Chen, Xinran Zhang 0006, Yanru Zhang |
ACM Multimedia | 1 |
| 2020 | Curriculum Learning for Wide Multimedia-Based Transformer with Graph Target DetectionabstractThe social media prediction task is aiming at predicting content popularity which includes social multimedia data such as photos, videos, and news. The task can not only help make better decisions for recommendation, but also reveals the public attention from evolutionary social systems. In this paper, we propose a novel approach named curriculum learning for wide multimedia-based transformer with graph target detection(CL-WMTG). The curriculum learning is designed for the transformer to improve the efficiency of model convergence. The mechanism of wide multimedia-based transformer is to make the model capable of learning cross information from text, pictures and other features(e.g. categories, location). Moreover, the graph target detection part can extract different features in the picture by pretrained model and reconstruct the features with a homogeneous graph network. We achieved third place in the SMP Challenge 2020. Weilong Chen, Feng Hong 0005, Rui Wang 0068, Ruobing Xie, Feng Xia 0006, Leyu Lin, Yanru Zhang, Yan Wang 0083 |
ACM Multimedia | 1 |
| 2019 | Time-aware Session Embedding for Click-Through-Rate PredictionabstractTV series correlation computing is one of the most important tasks of personalized online streaming services. With the relevance of TV series and viewer feedback, we can calculate the TV series correlation table based on the viewer's implicit feedback which does not perform well for the newly added "cold start" TV series. In this paper, we aim to improve correlation computing within the cold-start phase. We propose a framework named Time-aware Session Embedding (TSE), with Item Embedding in Session and Time Decay Factor for a multimodal recommendation. We apply an lower- dimensional vector as item embedding and calculate their factor considering the time decay. The framework performed well in the Content-based Video Relevance Prediction Challenge and we get the first place in this competition. Qidi Xu, Haocheng Xu, Weilong Chen, Chaojun Han, Haoyang Li 0002, Wenxin Tan, Fumin Shen, Heng Tao Shen |
ACM Multimedia | 3 |
| 2005 | PCA and LDA in DCT domain
Weilong Chen, Meng Joo Er, Shiqian Wu |
Pattern Recognit. Lett. | 1 |
| 2005 | High-speed face recognition based on discrete cosine transform and RBF neural networksabstractIn this paper, an efficient method for high-speed face recognition based on the discrete cosine transform (DCT), the Fisher's linear discriminant (FLD) and radial basis function (RBF) neural networks is presented. First, the dimensionality of the original face image is reduced by using the DCT and the large area illumination variations are alleviated by discarding the first few low-frequency DCT coefficients. Next, the truncated DCT coefficient vectors are clustered using the proposed clustering algorithm. This process makes the subsequent FLD more efficient. After implementing the FLD, the most discriminating and invariant facial features are maintained and the training samples are clustered well. As a consequence, further parameter estimation for the RBF neural networks is fulfilled easily which facilitates fast training in the RBF neural networks. Simulation results show that the proposed system achieves excellent performance with high training and recognition speed, high recognition rate as well as very good illumination robustness. Meng Joo Er, Weilong Chen, Shiqian Wu |
IEEE Trans. Neural Networks | 2 |
| 2004 | Illumination compensation and normalization using logarithm and discrete cosine transformabstractThis paper presents a novel illumination normalization approach for face recognition under varying lighting conditions. First, we demonstrate that illumination compensation can be efficiently implemented in the logarithm domain. In the proposed approach, discrete cosine transform (DCT) is employed to compensate for illumination variations in the logarithm domain. Since illumination variations mainly lie in the low-frequency band, an appropriate number of DCT coefficients are truncated to reduce the variations under different lighting conditions. The salient feature of our approach is that it does not need any training or modelling step and can be easily implemented with high speed. Weilong Chen, Meng Joo Er, Shiqian Wu |
ICARCV | 1 |