Wei-Bin Kou

dblp:350/5553 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-9817-791XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MGFN++: Magnitude-contrastive glance-and-focus network for weakly-supervised video anomaly detection
Yingxian Chen, Wei-Bin Kou, Wilton W. T. Fok, Zhengzhe Liu, Xiaojuan Qi 0001, Yik-Chung Wu
Pattern Recognit.2
2026 Exploring Performance-Fairness Trade-Offs in Federated Learning
abstract
Fairness in federated learning (FL) has emerged as a critical concern, aiming to develop an unbiased model among groups (e.g., male or female) of diverse sensitive features. However, there is a trade-off between model performance and fairness, i.e., improving fairness will decrease performance. Existing approaches have characterized such a trade-off by introducing hyperparameters to quantify client’s preferences over fairness and performance. Nevertheless, these approaches are limited to scenarios where each client has only a single pre-defined preference, and fail to address in cases where each client has multiple preferences. In this work, we aim to design algorithms that allow the trained model to adapt to each client’s diverse preferences in real time. The key challenges lie in (I) associating preferences with the trained model; (II) mitigating data heterogeneity; (III) preventing preference leakage; and (IV) handling data noise. To address these, we propose two preference-aware schemes, PraFFL and PraFFL-R, designed to generate models tailored to specific client preferences. PraFFL tackles challenges (I)–(III) by incorporating a hypernetwork that adaptively adjusts the model according to each client’s preferences, thereby better satisfying individual needs. To further address data noise (challenge (IV)), we introduce PraFFL-R, which enhances the robustness of the learned Pareto front by optimizing worst-case scenarios among preference-specific models. We provide theoretical guarantees demonstrating that PraFFL and PraFFL-R can produce an optimal model customized for any client preference, with proofs of linear convergence in the strongly convex setting and sublinear convergence in the non-convex setting. Experimental results show that our proposed PraFFL and PraFFL-R outperform five fair FL algorithms in terms of the model’s capability of adapting to clients’ different preferences.
Rongguang Ye, Wei-Bin Kou, Ming Tang 0006
IEEE Trans. Netw.2
2025 Label Anything: An Interpretable, High-Fidelity and Prompt-Free Annotator
abstract
Learning-based street scene semantic understanding in autonomous driving (AD) has advanced significantly recently, but the performance of the AD model is heavily dependent on the quantity and quality of the annotated training data. However, traditional manual labeling involves high cost to annotate the vast amount of required data for training robust model. To mitigate this cost of manual labeling, we propose a Label Anything Model (denoted as LAM), serving as an interpretable, high-fidelity, and prompt-free data annotator. Specifically, we firstly incorporate a pretrained Vision Transformer (ViT) to extract the latent features. On top of ViT, we propose a semantic class adapter (SCA) and an optimization-oriented unrolling algorithm (OptOU), both with a quite small number of trainable parameters. SCA is proposed to fuse ViT-extracted features to consolidate the basis of the subsequent automatic annotation. OptOU consists of multiple cascading layers and each layer contains an optimization formulation to align its output with the ground truth as closely as possible, though which OptOU acts as being interpretable rather than learning-based blackbox nature. In addition, training SCA and OptOU requires only a single pre-annotated RGB seed image, owing to their small volume of learnable parameters. Extensive experiments clearly demonstrate that the proposed LAM can generate high-fidelity annotations (almost 100% in mIoU) for multiple real-world datasets (i.e., Camvid, Cityscapes, and Apolloscapes) and CARLA simulation dataset.
Wei-Bin Kou, Guangxu Zhu, Rongguang Ye, Shuai Wang 0004, Ming Tang 0006, Yik-Chung Wu
ICRA1
2025 Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
abstract
To improve the generalization of the autonomous driving (AD) perception model, vehicles need to update the model over time based on the continuously collected data. As time progresses, the amount of data fitted by the AD model expands, which helps to improve the AD model generalization substantially. However, such ever-expanding data is a double-edged sword for the AD model. Specifically, as the fitted data volume grows to exceed the AD model’s fitting capacities, the AD model is prone to under-fitting. To address this issue, we propose to use a pretrained Large Vision Models (LVMs) as backbone coupled with downstream perception head to understand AD semantic information. This design can not only surmount the aforementioned under-fitting problem due to LVMs’ powerful fitting capabilities, but also enhance the perception generalization thanks to LVMs’ vast and diverse training data. On the other hand, to mitigate vehicles’ computational burden of training the perception head while running LVM backbone, we introduce a Posterior Optimization Trajectory (POT)-Guided optimization scheme (POTGui) to accelerate the convergence. Concretely, we propose a POT Generator (POTGen) to generate posterior (future) optimization direction in advance to guide the current optimization iteration, through which the model can generally converge within 10 epochs. Extensive experiments demonstrate that the proposed method improves the performance by over 66.48% and converges faster over 6 times, compared to the existing state-of-the-art approaches.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Jingreng Lei, Shuai Wang 0004, Rongguang Ye, Guangxu Zhu, Yik-Chung Wu
IROS1
2025 FedEMA: Federated Exponential Moving Averaging with Negative Entropy Regularizer in Autonomous Driving
abstract
Street Scene Semantic Understanding (denoted as S3U) is a crucial but complex task for autonomous driving (AD) vehicles. Their inference models typically face poor generalization due to domain-shift. Federated Learning (FL) has emerged as a promising paradigm for enhancing the generalization of AD models through privacy-preserving distributed learning. However, these FL AD models face significant temporal catastrophic forgetting when deployed in dynamically evolving environments, where continuous adaptation causes abrupt erosion of historical knowledge. This paper proposes Federated Exponential Moving Average (FedEMA), a novel framework that addresses this challenge through two integral innovations: (I) Server-side model’s historical fitting capability preservation via fusing current FL round’s aggregation model and a proposed previous FL round’s exponential moving average (EMA) model; (II) Vehicle-side negative entropy regularization to prevent FL models’ possible overfitting to EMA-introduced temporal patterns. Above two strategies empower FedEMA a dual-objective optimization that balances model generalization and adaptability. In addition, we conduct theoretical convergence analysis for the proposed FedEMA. Extensive experiments both on Cityscapes dataset and Camvid dataset demonstrate FedEMA’s superiority over existing approaches, showing 7.12% higher mean Intersectionover-Union (mIoU).
Wei-Bin Kou, Guangxu Zhu, Bingyang Cheng, Shuai Wang 0004, Ming Tang 0006, Yik-Chung Wu
IROS1
2025 PraFFL: A Preference-Aware Scheme in Fair Federated Learning
abstract
Fairness in federated learning has emerged as a critical concern, aiming to develop an unbiased model among groups (e.g., male or female) of diverse sensitive features. However, there is a trade-off between model performance and fairness, i.e., improving model fairness will decrease model performance. Existing approaches have characterized such a trade-off by introducing hyperparameters to quantify client's preferences for model fairness and model performance. Nevertheless, these approaches are limited to scenarios where each client has only a single pre-defined preference, and fail to work in practical systems where each client generally has multiple preferences. To this end, we propose a Preference-aware scheme in Fair Federated Learning (called PraFFL) to generate preference-specific models in real time. PraFFL can adaptively adjust the model based on each client's preferences to meet their needs. We theoretically prove that PraFFL can offer the optimal model tailored to an arbitrary preference of each client, and show its linear convergence. Experimental results show that our proposed PraFFL outperforms six fair federated learning algorithms in terms of the model's capability of adapting to clients' different preferences. Our implementation is available at https://github.com/rG223/PraFFL.
Rongguang Ye, Wei-Bin Kou, Ming Tang 0006
KDD (1)2
2025 pFedLVM: A Large Vision Model (LVM)-Driven and Latent Feature-Based Personalized Federated Learning Framework in Autonomous Driving
abstract
Deep learning-based Autonomous Driving (AD) perception models often exhibit poor generalization due to data heterogeneity in an ever domain-shifting environment. While Federated Learning (FL) could improve the generalization of an AD model (known as FedAD system), conventional models often struggle with under-fitting as the amount of accumulated training data progressively increases. To address this issue, instead of conventional small models, employing Large Vision Models (LVMs) in FedAD is a viable option for better learning of representations from a vast volume of data. However, implementing LVMs in FedAD introduces three challenges:(I)the extremely high communication overheads associated with transmitting LVMs between participating vehicles and a central server;(II)lack of computing resource to deploy LVMs on each vehicle;(III)the performance drop due to LVM focusing on shared features but overlooking local vehicle characteristics. To overcome these challenges, we propose pFedLVM, a LVM-Driven, Latent Feature-Based Personalized Federated Learning framework. In this approach, the LVM is deployed only on central server, which effectively alleviates the computational burden on individual vehicles. Furthermore, the exchange between central server and vehicles are the learned features rather than the LVM parameters, which significantly reduces communication overhead. In addition, we utilize both shared features from all participating vehicles and individual characteristics from each vehicle to establish a personalized learning mechanism. This enables each vehicle’s model to learn features from others while preserving its personalized characteristics, thereby outperforming globally shared models trained in general FL. As a demonstration of the proposed pFedLVM, this paper focuses on the semantic segmentation (SSeg) task. Extensive experiments demonstrate that pFedLVM outperforms the existing state-of-the-art approach by 18.47%, 25.60%, 51.03% and 14.19% in terms of mIoU, mF1, mPrecision and mRecall, respectively.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Rongguang Ye, Yang Leng, Shuai Wang 0004, Guofa Li, Zhenyu Chen 0001, Guangxu Zhu, Yik-Chung Wu
IEEE Trans. Intell. Transp. Syst.1
2025 Fast-Convergent and Communication-Alleviated Heterogeneous Hierarchical Federated Learning in Autonomous Driving
abstract
Street Scene Semantic Understanding (denoted as TriSU) is a complex task for autonomous driving (AD). However, inference model trained from data in a particular geographical region faces poor generalization when applied in other regions due to inter-city data domain-shift. Hierarchical Federated Learning (HFL) offers a potential solution for improving TriSU model generalization by collaborative privacy-preserving training over distributed datasets from different cities. Unfortunately, it suffers from slow convergence because the data from different cities are with disparate statistical properties. Going beyond existing HFL methods, we propose a Gaussian heterogeneous HFL algorithm (FedGau) to address inter-city data heterogeneity so that convergence can be accelerated. In the proposed FedGau algorithm, both single RGB image and RGB dataset are modelled as Gaussian distributions for aggregation weight design. This approach not only differentiates each RGB image by respective statistical distribution, but also exploits the statistics of dataset from each city in addition to the conventionally considered data volume. With the proposed approach, the convergence is accelerated by 35.5%-40.6% compared to existing state-of-the-art (SOTA) HFL methods. On the other hand, to reduce the involved communication resource, we further introduce a novel performance-aware adaptive resource scheduling (AdapRS) policy. Unlike the traditional static resource scheduling policy that exchanges a fixed number of models between two adjacent aggregations, AdapRS adjusts the number of model aggregation at different levels of HFL so that unnecessary communications are minimized. Extensive experiments demonstrate that AdapRS saves 29.65% communication overhead compared to conventional static resource scheduling policy while maintaining almost the same performance.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Rongguang Ye, Shuai Wang 0004, Guangxu Zhu, Yik-Chung Wu
IEEE Trans. Intell. Transp. Syst.1
2024 FedRC: A Rapid-Converged Hierarchical Federated Learning Framework in Street Scene Semantic Understanding
abstract
Street Scene Semantic Understanding (denoted as TriSU) is a crucial but complex task for world-wide distributed autonomous driving (AD) vehicles (e.g., Tesla). Its inference model faces poor generalization issue due to inter-city domain-shift. Hierarchical Federated Learning (HFL) offers a potential solution for improving TriSU model generalization, but suffers from slow convergence rate because of vehicles’ surrounding heterogeneity across cities. Going beyond existing HFL works that have deficient capabilities in complex tasks, we propose a rapid-converged heterogeneous HFL framework (FedRC) to address the inter-city data heterogeneity and accelerate HFL model convergence rate. In our proposed FedRC framework, both single RGB image and RGB dataset are modelled as Gaussian distributions in HFL aggregation weight design. This approach not only differentiates each RGB sample instead of typically equalizing them, but also considers both data volume and statistical properties rather than simply taking data quantity into consideration. Extensive experiments on the TriSU task using across-city datasets demonstrate that FedRC converges faster than the state-of-the-art benchmark by 38.7%, 37.5%, 35.5%, and 40.6% in terms of mIoU, mPrecision, mRecall, and mF1, respectively. Furthermore, qualitative evaluations in the CARLA simulation environment confirm that the proposed FedRC framework delivers top-tier performance.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Shuai Wang 0004, Guangxu Zhu, Yik-Chung Wu
IROS1
2024 Pareto Front Shape-Agnostic Pareto Set Learning in Multi-Objective Optimization
abstract
Pareto set learning (PSL) is an emerging approach for acquiring the complete Pareto set of a multi-objective optimization problem. Existing methods primarily rely on the mapping of preference vectors in the objective space to Pareto optimal solutions in the decision space. However, the sampling of preference vectors theoretically requires prior knowledge of the Pareto front shape to ensure high performance of the PSL methods. Designing a sampling strategy of preference vectors is difficult since the Pareto front shape cannot be known in advance. To make Pareto set learning work effectively in any Pareto front shape, we propose a Pareto front shape-agnostic Pareto _Set Learning (GPSL) that does not require the prior information about the Pareto front. The fundamental concept behind GPSL is to treat the learning of the Pareto set as a distribution transformation problem. Specifically, GPSL can transform an arbitrary distribution into the Pareto set distribution. We demonstrate that training a neural network by maximizing hypervolume enables the process of distribution transformation. Our proposed method can handle any shape of the Pareto front and learn the Pareto set without requiring prior knowledge. Experimental results show the high performance of our proposed method on diverse test problems compared with recent Pareto set learning algorithms.
Rongguang Ye, Longcan Chen, Wei-Bin Kou, Hisao Ishibuchi
SMC3
2024 A hierarchical federated learning framework for collaborative quality defect inspection in construction
Heng Li 0001, Hung-Lin Chi, Wei-Bin Kou, Yik-Chung Wu, Shuai Wang 0004
Eng. Appl. Artif. Intell.4
2024 Communication-Efficient Activity Detection for Cell-Free Massive MIMO: An Augmented Model-Driven End-to-End Learning Framework
abstract
A great amount of endeavour has recently been devoted to activity detection for cell-free massive multiple-input multiple-output (MIMO) systems, where multiple access points (APs) jointly identify the active devices from a large number of potential devices. In practice, the APs and the central processing unit (CPU) are connected by capacity-limited fronthauls and the signals at the APs need to be compressed/quantized before they are forwarded to the CPU. However, existing approaches treat the compression/quantization and activity detection as separate tasks, which makes it difficult to achieve global system optimality. To tackle the above problem, this paper proposes an augmented model-driven end-to-end learning framework which jointly optimizes the compression modules, quantization modules at the APs, and the decompression module and detection module at the CPU. Specifically, deep unfolding is leveraged for designing the detection module in order to inherit the domain knowledge derived from the optimization algorithm, and other modules are constructed by judiciously designed neural network architectures for improving the learning capability. Furthermore, we design an enhanced scheme so that the proposed framework is adaptable to different compression rates. We demonstrate numerically that the proposed framework significantly reduces the computational complexity and achieves better detection performance than the conventional approaches. Moreover, it costs a much smaller number of bits on the fronthauls while still maintaining the detection performance.
Qingfeng Lin, Yang Li 0035, Wei-Bin Kou, Tsung-Hui Chang, Yik-Chung Wu
IEEE Trans. Wirel. Commun.3
2023 Communication-Efficient Joint Signal Compression and Activity Detection in Cell-Free Massive MIMO
abstract
A great amount of endeavour has recently been devoted to device activity detection in massive machine-type communications. This paper targets at a practical issue: communication-efficient joint signal compression and activity detection in cell-free massive MIMO with capacity-limited fronthauls. To this end, we propose a novel deep learning framework which jointly optimizes the compression modules, quantization modules at the access points, and the decompression module and detection module at the central processing unit. Specifically, deep unfolding is leveraged for designing the detection module in order to inherit the domain knowledge derived from the optimization algorithm, and the other modules are constructed by generic layers for increasing the learning capability. A joint training strategy is proposed to optimize all the modules in an end-to-end manner. Numerical results demonstrate the superiority of the proposed end-to-end learning framework compared with classical optimization methods.
Qingfeng Lin, Yang Li 0035, Wei-Bin Kou, Tsung-Hui Chang, Yik-Chung Wu
ICC3
2023 Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
abstract
While federated learning (FL) improves the generalization of end-to-end autonomous driving by model aggregation, the conventional single-hop FL (SFL) suffers from slow convergence rate due to long-range communications among vehicles and cloud server. Hierarchical federated learning (HFL) overcomes such drawbacks via introduction of mid-point edge servers. However, the orchestration between constrained communication resources and HFL performance becomes an urgent problem. This paper proposes an optimization-based Communication Resource Constrained Hierarchical Federated Learning (CRCHFL) framework to minimize the generalization error of the autonomous driving model using hybrid data and model aggregation. The effectiveness of the proposed CRCHFL is evaluated in the Car Learning to Act (CARLA) simulation platform. Results show that the proposed CRCHFL both accelerates the convergence rate and enhances the generalization of federated learning autonomous driving model. Moreover, under the same communication resource budget, it outperforms the HFL by 10.33% and the SFL by 12.44%.
Wei-Bin Kou, Shuai Wang 0004, Guangxu Zhu, Bin Luo 0004, Yingxian Chen, Derrick Wing Kwan Ng, Yik-Chung Wu
IROS1