EDBT 2026 Demo / reviewers in the wild / expert
Minsu Kim 0003
dblp:25/6052-3
· DBLP profile ↗
9ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-7556-8958ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Generalization for Multi-Modal Wireless Networks Under Data Scarcity
Minsu Kim 0003, Walid Saad 0001, Doru Calin |
ICC | 1 |
| 2026 | Transformer Architecture With Minimal Inference Latency for Multimodal Wireless NetworksabstractNext-generation wireless networks are expected to leverage multi-modal data sources in order to execute various wireless communication tasks such as beamforming and blockage prediction with situational-awareness. To do so, multi-modal transformers emerged as an effective tool, however, existing transformer-based approaches suffer from high inference latency and large memory footprints when processing multi-modal data. Hence, such existing solutions cannot handle wireless communication tasks that require fast inference to track a dynamically changing environment with moving vehicles and blockages. One major bottleneck is the reliance on attention mechanisms whose complexity grows quadratically with respect to the number of tokens. Hence, in this paper, a novel, fast multi-modal transformer inference framework is designed to practically support wireless communication tasks by processing only important tokens. To this end, an optimization problem is formulated to find the optimal number of tokens under a target floating point operations (FLOPs) for a given wireless communication task while maintaining the task accuracy. To solve this problem, modality-specific tokenizers are first designed to project each modality into the same embedding dimension. Then, a token router is introduced to learn the importance of each token and process only important tokens. Subsequently, a trainable keep ratio is introduced to learn how many tokens should be processed for each layer under the target FLOPs. Simulation results show that, on DeepSense 6G beamforming tasks, the proposed framework can reduce the inference latency, GPU memory, and FLOPs by 86.2% 35%, and 80%, respectively, with negligible accuracy loss compared to a baseline that processes all tokens. To further validate the feasibility of the proposed framework for real-world deployments, a multi-modal handover dataset is developed using a real-world testbed. Emulation results on the developed dataset show that the proposed framework can proactively initiate handover before blockage, while the baseline experiences significant received signal strength degradation due to higher inference latency. Minsu Kim 0003, Walid Saad 0001, Kui Wang 0004, Zongdian Li, Tao Yu 0011, Kei Sakaguchi |
IEEE Internet Things J. | 1 |
| 2025 | Edge vs Cloud: How Do We Balance Cost, Latency, and Quality for Large Language Models Over 5G Networks?abstractLarge language models (LLMs) can perform a plethora of tasks, however, they often require cloud servers for deployment due to their computing cost and size. Meanwhile, small and cost-effective LLMs can be deployed on edge devices (e.g., mobile devices), but they often exhibit lower response quality than larger models. In this paper, a measurement-driven training framework is proposed for a device-side artificial intelligence (AI)-enabled router that selects the best LLM in terms of cost, latency, and performance. In the considered framework, a mobile device uses its local LLM (sub-billion LLM), while having access to a bigger server LLM (GPT-4) hosted on a 5G core network. The mobile device has a router that sends an input prompt to the local or server LLM to optimize the cost, latency, and performance of output LLM responses. To train the router, the dataset is constructed by measuring the cost, latency, and performance of the server LLM and the local LLM on a mobile device. The structural causal model (SCM) of the measured dataset is identified. To further improve performance, a causality-driven data augmentation method is also proposed based on the discovered SCM. Real-world experimental results show that the proposed framework can improve the cost and latency by 50% and 31%, respectively, with only a 2.13% performance loss compared to a baseline that only uses the server LLM. Minsu Kim 0003, Pinyarash Pinyoanuntapong, Bong-Ho Kim, Walid Saad 0001, Doru Calin |
WCNC | 1 |
| 2024 | Analysis of the Memorization and Generalization Capabilities of AI Agents: are Continual Learners Robust?abstractIn continual learning (CL), an AI agent (e.g., autonomous vehicles or robotics) learns from non-stationary data streams under dynamic environments. For the practical deployment of such applications, it is important to guarantee robustness to unseen environments while maintaining past experiences. In this paper, a novel CL framework is proposed to achieve robust generalization to dynamic environments while retaining past knowledge. The considered CL agent uses a capacity-limited memory to save previously observed environmental information to mitigate forgetting issues. Then, data points are sampled from the memory to estimate the distribution of risks over environmental change so as to obtain predictors that are robust with unseen changes. The generalization and memorization performance of the proposed framework are theoretically analyzed. This analysis showcases the tradeoff between memorization and generalization with the memory size. Experiments show that the proposed algorithm outperforms memory-based CL baselines across all environments while significantly improving the generalization performance on unseen target environments. Minsu Kim 0003, Walid Saad 0001 |
ICASSP | 1 |
| 2024 | SpaFL: Communication-Efficient Federated Learning With Sparse Models And Low Computational OverheadabstractThe large communication and computation overhead of federated learning (FL) is one of the main challenges facing its practical deployment over resource-constrained clients and systems. In this work, SpaFL: a communication-efficient FL framework is proposed to optimize sparse model structures with low computational overhead. In SpaFL, a trainable threshold is defined for each filter/neuron to prune its all connected
parameters, thereby leading to structured sparsity. To optimize the pruning process itself, only thresholds are communicated between a server and clients instead of parameters, thereby learning how to prune. Further, global thresholds are used to update model parameters by extracting aggregated parameter importance. The generalization bound of SpaFL is also derived, thereby proving key insights on the relation between sparsity and performance. Experimental results show that SpaFL improves accuracy while requiring much less communication and computing resources compared to sparse baselines. The code is available at https://github.com/news-vt/SpaFL_NeruIPS_2024 Minsu Kim 0003, Walid Saad 0001, Mérouane Debbah, Choong Seon Hong |
NeurIPS | 1 |
| 2024 | Green, Quantized Federated Learning Over Wireless Networks: An Energy-Efficient DesignabstractThe practical deployment of federated learning (FL) over wireless networks requires balancing energy efficiency, convergence rate, and a target accuracy due to the limited available resources of devices. Prior art on FL often trains deep neural networks (DNNs) to achieve high accuracy and fast convergence using 32 bits of precision level. However, such scenarios will be impractical for resource-constrained devices since DNNs typically have high computational complexity and memory requirements. Thus, there is a need to reduce the precision level in DNNs to reduce the energy expenditure. In this paper, a green-quantized FL framework, which represents data with a finite precision level in both local training and uplink transmission, is proposed. Here, the finite precision level is captured through the use of quantized neural networks (QNNs) that quantize weights and activations in fixed-precision format. In the considered FL model, each device trains its QNN and transmits a quantized training result to the base station. Energy models for the local training and the transmission with quantization are rigorously derived. To minimize the energy consumption and the number of communication rounds simultaneously, a multi-objective optimization problem is formulated with respect to the number of local iterations, the number of selected devices, and the precision levels for both local training and transmission while ensuring convergence under a target accuracy constraint. To solve this problem, the convergence rate of the proposed FL system is analytically derived with respect to the system control variables. Then, the Pareto boundary of the problem is characterized to provide efficient solutions using the normal boundary inspection method. Design insights on balancing the tradeoff between the two objectives while achieving a target accuracy are drawn from using the Nash bargaining solution and analyzing the derived convergence rate. Simulation results show that the proposed FL framework can reduce energy consumption until convergence by up to 70% compared to a baseline FL algorithm that represents data with full precision without damaging the convergence rate. Minsu Kim 0003, Walid Saad 0001, Mohammad Mozaffari, Mérouane Debbah |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | A Bargaining Game for Personalized, Energy Efficient Split Learning over Wireless NetworksabstractSplit learning (SL) is an emergent distributed learning framework which can mitigate the computation and wireless communication overhead of federated learning. It splits a machine learning model into a device-side model and a server-side model at a cut layer. Devices only train their allocated model and transmit the activations of the cut layer to the server. However, SL can lead to data leakage as the server can reconstruct the input data using the correlation between the input and intermediate activations. Although allocating more layers to a device-side model can reduce the possibility of data leakage, this will lead to more energy consumption for resource-constrained devices and more training time for the server. Moreover, non-iid datasets across devices will reduce the convergence rate leading to increased training time. In this paper, a new personalized SL framework is proposed. For this framework, a novel approach for choosing the cut layer that can optimize the tradeoff between the energy consumption for computation and wireless transmission, training time, and data privacy is developed. In the considered framework, each device personalizes its device-side model to mitigate non-iid datasets while sharing the same server-side model for generalization. To balance the energy consumption for computation and wireless transmission, training time, and data privacy, a multiplayer bargaining problem is formulated to find the optimal cut layer between devices and the server. To solve the problem, the Kalai-Smorodinsky bargaining solution (KSBS) is obtained using the bisection method with the feasibility test. Simulation results show that the proposed personalized SL framework with the cut layer from the KSBS can achieve the optimal sum utilities by balancing the energy consumption, training time, and data privacy, and it is also robust to non-iid datasets. Minsu Kim 0003, Alexander C. DeRieux, Walid Saad 0001 |
WCNC | 1 |
| 2022 | On the Tradeoff between Energy, Precision, and Accuracy in Federated Quantized Neural NetworksabstractDeploying federated learning (FL) over wireless networks with resource-constrained devices requires balancing between accuracy, energy efficiency, and precision. Prior art on FL often requires devices to train deep neural networks (DNNs) using a 32-bit precision level for data representation to improve accuracy. However, such algorithms are impractical for resource-constrained devices since DNNs could require execution of millions of operations. Thus, training DNNs with a high precision level incurs a high energy cost for FL. In this paper, a quantized FL framework, that represents data with a finite level of precision in both local training and uplink transmission, is proposed. Here, the finite level of precision is captured through the use of quantized neural networks (QNNs) that quantize weights and activations in fixed-precision format. In the considered FL model, each device trains its QNN and transmits a quantized training result to the base station. Energy models for the local training and the transmission with the quantization are rigorously derived. An energy minimization problem is formulated with respect to the level of precision while ensuring convergence. To solve the problem, we first analytically derive the FL convergence rate and use a line search method. Simulation results show that our FL framework can reduce energy consumption by up to 53% compared to a standard FL model. The results also shed light on the tradeoff between precision, energy, and accuracy in FL over wireless networks. Minsu Kim 0003, Walid Saad 0001, Mohammad Mozaffari, Mérouane Debbah |
ICC | 1 |
| 2022 | Ensuring Data Freshness for Blockchain-Enabled Monitoring NetworksabstractThe Age of Information (AoI) is a recently proposed metric for quantifying data freshness in real-time status monitoring systems, where timeliness is of importance. In this article, the problem of characterizing and controlling the AoI is studied in the context of blockchain-enabled monitoring networks (BeMNs). In BeMN, status updates from sources are transmitted and recorded in a blockchain. To investigate the statistical characteristics of the AoI in BeMN, the transmission latency and the consensus latency are first rigorously modeled. Then, the average AoI, the AoI violation probability, and the peak AoI violation probability are derived in a closed form so as to quantify the performance of BeMN. Furthermore, a simplified form is derived for the AoI violation probability, and it is shown that this quantity can capture the upper or lower bounds of the actual AoI violation probability. Simulation results show that each BeMN parameters (i.e., target successful transmission probability, block size, and timeout) can have conflicting effects on the AoI-related performance. Subsequently, design insights are provided to maintain the freshness of the status data in BeMN. Then, experimental results with a real Hyperledger Fabric platform further validate the accuracy of our modeling and analysis. Minsu Kim 0003, Chanwon Park, Jemin Lee 0002, Walid Saad 0001 |
IEEE Internet Things J. | 1 |