VLDB 2026 Research / reviewers in the wild / expert
Kahou Tam
dblp:349/7780 · also KaHou Tam
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5816-6837ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Efficient and distributed learning · 73% Autonomous driving · 17% Video understanding and tracking · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 50% Memory systems · 50% | |
| Computer networks
4 papers |
Edge and fog computing · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
3.5 | 4 | 2026 | Floe: Federated Specialization for Real-Time LLM-SLM Inference · IEEE Trans. Parallel Distributed Syst. 2026 Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge · ACL (1) 2026 Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting · IEEE Trans. Parallel Distributed Syst. 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
2.0 | 2 | 2026 | Floe: Federated Specialization for Real-Time LLM-SLM Inference · IEEE Trans. Parallel Distributed Syst. 2026 Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge · ACL (1) 2026 |
Robotics › Autonomous driving › risk assessment
accident anticipation |
1.5 | 2 | 2024 | CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus Attentions · ACM Multimedia 2024 When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language Models · ACM Multimedia 2024 |
Edge and fog computing
edge inference |
1.5 | 3 | 2026 | CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge · USENIX ATC 2025 Floe: Federated Specialization for Real-Time LLM-SLM Inference · IEEE Trans. Parallel Distributed Syst. 2026 Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.2 | 2 | 2026 | Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge · ACL (1) 2026 FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024 |
Machine learning › Efficient and distributed learning › federated learning
federated fine-tuning |
1.0 | 1 | 2026 | Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
1.0 | 1 | 2026 | Floe: Federated Specialization for Real-Time LLM-SLM Inference · IEEE Trans. Parallel Distributed Syst. 2026 |
Memory systems
data layout optimization |
1.0 | 1 | 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs · MobiSys 2026 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
1.0 | 1 | 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs · MobiSys 2026 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge · USENIX ATC 2025 |
Machine learning › Efficient and distributed learning › federated learning
client selection |
0.8 | 1 | 2024 | FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024 |
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning |
0.8 | 1 | 2024 | Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting · IEEE Trans. Parallel Distributed Syst. 2024 |
Machine learning › Efficient and distributed learning › federated learning › resource-efficient federated learning
memory-efficient federated learning |
0.8 | 1 | 2024 | FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024 |
Machine learning › Efficient and distributed learning › distributed training › parallelization
model partitioning |
0.8 | 1 | 2024 | Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting · IEEE Trans. Parallel Distributed Syst. 2024 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal interaction modeling |
0.8 | 1 | 2024 | CDSTraj: Characterized Diffusion and Spatial-Temporal Interaction Network for Trajectory Prediction in Autonomous Driving · IJCAI 2024 |
Robotics › Autonomous driving
trajectory prediction |
0.8 | 1 | 2024 | CDSTraj: Characterized Diffusion and Spatial-Temporal Interaction Network for Trajectory Prediction in Autonomous Driving · IJCAI 2024 |
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning |
0.3 | 1 | 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs · MobiSys 2026 |
Machine learning › Efficient and distributed learning › model compression › neural network compression
activation compression |
0.2 | 1 | 2024 | FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language Models · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
tile-based layout transformation · 2.0logit-level fusion · 2.0layer-wise training · 2.0heterogeneity-aware LoRA · 2.0dynamic layer co-tuning · 2.0adaptive tuning · 2.0activation-guided layout selection · 2.0LLM customization · 1.7spatial-temporal interaction network · 0.8hierarchical training management · 0.8diffusion model · 0.8cost-aware checkpointing · 0.8adaptive cutting layer selection · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the EdgeabstractFederated fine-tuning enables privacypreserving LLM adaptation but faces a critical bottleneck: the disparity between LLMs' high memory demands and edge devices' limited capacity.To break the memory barrier, we propose Chain Federated Fine-Tuning (CHAINFED), an innovative paradigm that forgoes end-to-end updates in favor of a sequential, layer-by-layer manner.It first trains the initial adapter to convergence, freezes its weights, and then proceeds to the next.This iterative train-and-freeze process forms an optimization chain, gradually enhancing the model's task-specific proficiency.CHAINFED further integrates three core techniques: 1) Dynamic Layer Co-Tuning to bridge semantic gaps between sequentially tuned layers and facilitate information flow; 2) Globally Perceptive Optimization to endow each adapter with foresight beyond its local objective; 3) Function-Oriented Adaptive Tuning to automatically identify the optimal fine-tuning starting point.Extensive experiments on multiple benchmarks demonstrate the superiority of CHAINFED over existing methods, boosting average accuracy by up to 46.46%. Yebo Wu, Jingguang Li, Chunlin Tian, Kahou Tam, Zhijiang Guo, Li Li 0064 |
ACL (1) | 4 |
| 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUsabstractTransformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs due to severe memory constraints and frequent layout transformations in attention mechanism during training. Existing mobile training frameworks either use unified layouts for forward and backward passes - leading to fragmented memory access and poor GPU utilization during backpropagation - or rely on explicit layout conversions, which introduce significant transformation overhead.To overcome this, we propose FBLayout, a layout-aware framework that co-designs tensor organization with mobile GPU platforms. FBLayout introduces: (1) a unified R-Tile layout for multidimensional reductions across forward/backward passes; (2) tile-based index transformation to eliminate physical data movement; and (3) activation-guided layout selection to propagate efficient layouts globally. Evaluations on seven transformer models across different mobile phones (including ARM Mali and Qualcomm Adreno GPUs) show that FBLayout achieves 2.2-5.7× speedup over MNN, TFLite, and TVM, while significantly improving cache efficiency and reducing memory footprint, enabling practical on-device large model fine-tuning. Kahou Tam, Wei Niu 0002, Xiaomin Ouyang, Cheng-Zhong Xu 0001, Li Li 0064 |
MobiSys | 1 |
| 2026 | Floe: Federated Specialization for Real-Time LLM-SLM InferenceabstractDeploying large language models (LLMs) in realtime systems is challenging due to their high resource demands and privacy concerns. We propose Floe a hybrid federated learning framework designed for latency-sensitive, resourceconstrained environments. Floeombines a cloud-based blackbox LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy ensures efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Experiments demonstrate that Floenhances user privacy and personalization, while significantly improving model performance and reducing inference latency on edge devices under real-time constraints, compared to baseline approaches. Chunlin Tian, Kahou Tam, Yebo Wu, Shuaihang Zhong, Li Li 0064, Nicholas D. Lane, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Chunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 0064, Yuanzhe Zhao, Minglei Zhang, Cheng-Zhong Xu 0001 |
USENIX ATC | 3 |
| 2025 | Federated Noisy Client LearningabstractFederated learning (FL) collaboratively trains a shared global model depending on multiple local clients, while keeping the training data decentralized to preserve data privacy. However, standard FL methods ignore the noisy client issue, which may harm the overall performance of the shared model. We first investigate the critical issue caused by noisy clients in FL and quantify the negative impact of the noisy clients in terms of the representations learned by different layers. We have the following two key observations: 1) the noisy clients can severely impact the convergence and performance of the global model in FL and 2) the noisy clients can induce greater bias in the deeper layers than the former layers of the global model. Based on the above observations, we propose federated noisy client learning (Fed-NCL), a framework that conducts robust FL with noisy clients. Specifically, Fed-NCL first identifies the noisy clients through well estimating the data quality and model divergence. Then robust layerwise aggregation is proposed to adaptively aggregate the local models of each client to deal with the data heterogeneity caused by the noisy clients. We further perform label correction on the noisy clients to improve the generalization of the global model. Experimental results on various datasets demonstrate that our algorithm boosts the performances of different state-of-the-art systems with noisy clients. Our code is available at https://github.com/TKH666/Fed-NCL. Kahou Tam, Li Li 0064, Bo Han 0003, Cheng-Zhong Xu 0001, Huazhu Fu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | CDSTraj: Characterized Diffusion and Spatial-Temporal Interaction Network for Trajectory Prediction in Autonomous Driving
Haicheng Liao, Xuelin Li, Yongkang Li 0003, Hanlin Kong, Chengyue Wang 0001, Bonan Wang, Yanchen Guan, Kahou Tam, Zhenning Li 0001 |
IJCAI | 8 |
| 2024 | When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language ModelsabstractAs autonomous driving systems increasingly become part of daily transportation, the ability to accurately anticipate and mitigate potential traffic accidents is paramount. Traditional accident anticipation models primarily utilizing dashcam videos are adept at predicting when an accident may occur but fall short in localizing the incident and identifying involved entities. Addressing this gap, this study introduces a novel framework that integrates Large Language Models (LLMs) to enhance predictive capabilities across multiple dimensions-what, when, and where accidents might occur. We develop an innovative chain-based attention mechanism that dynamically adjusts to prioritize high-risk elements within complex driving scenes. This mechanism is complemented by a three-stage model that processes outputs from smaller models into detailed multimodal inputs for LLMs, thus enabling a more nuanced understanding of traffic dynamics. Empirical validation on the DAD, CCD, and A3D datasets demonstrates superior performance in Average Precision (AP) and Mean Time-To-Accident (mTTA), establishing new benchmarks for accident prediction technology. Our approach not only advances the technological framework for autonomous driving safety but also enhances human-AI interaction, making predictive insights generated by autonomous systems more intuitive and actionable. Haicheng Liao, Yongkang Li 0003, Chengyue Wang 0001, Yanchen Guan, Kahou Tam, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
ACM Multimedia | 5 |
| 2024 | CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus AttentionsabstractAccurately and promptly predicting accidents among surrounding traffic agents from camera footage is crucial for the safety of autonomous vehicles (AVs). This task presents substantial challenges stemming from the unpredictable nature of traffic accidents, their long-tail distribution, the intricacies of traffic scene dynamics, and the inherently constrained field of vision of onboard cameras. To address these challenges, this study introduces a novel accident anticipation framework for AVs, termed CRASH. It seamlessly integrates five components: object detector, feature extractor, object-aware module, context-aware module, and multi-layer fusion. Specifically, we develop the object-aware module to prioritize high-risk objects in complex and ambiguous environments by calculating the spatial-temporal relationships between traffic agents. In parallel, the context-aware is also devised to extend global visual information from the temporal to the frequency domain using the Fast Fourier Transform (FFT) and capture fine-grained visual features of potential objects and broader context cues within traffic scenes. To capture a wider range of visual cues, we further propose a multi-layer fusion that dynamically computes the temporal dependencies between different scenes and iteratively updates the correlations between different visual features for accurate and timely accident prediction. Evaluated on real-world datasets-Dashcam Accident Dataset (DAD), Car Crash Dataset (CCD), and AnAn Accident Detection (A3D) datasets-our model surpasses existing top baselines in critical evaluation metrics like Average Precision (AP) and mean Time-To-Accident (mTTA). Importantly, its robustness and adaptability are particularly evident in challenging driving scenarios with missing or limited training data, demonstrating significant potential for application in real-world autonomous driving systems. Haicheng Liao, Huanming Shen, Chengyue Wang 0001, Chunlin Tian, Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
ACM Multimedia | 6 |
| 2024 | FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor ManagementabstractFederated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, one fundamental and prevailing challenge that hinders the deployment of FL on mobile devices is the memory limitation. This paper proposes FedHybrid, a novel framework that effectively reduces the memory footprint during the training process while guaranteeing the model accuracy and the overall training progress. Specifically, FedHybrid first selects the participating devices for each training round by jointly evaluating their memory budget, computing capability, and data diversity. After that, it judiciously analyzes the computational graph and generates an execution plan for each selected client in order to meet the corresponding memory budget while minimizing the training delay through employing a hybrid of recomputation and compression techniques according to the characteristic of each tensor. During the local training process, FedHybrid carries out the execution plan with a well-designed activation compression technique to effectively achieve memory reduction with minimum accuracy loss. We conduct extensive experiments to evaluate FedHybrid on both simulation and off-the-shelf mobile devices. The experiment results demonstrate that FedHybrid achieves up to a 39.1% increase in model accuracy and a 15.5X reduction in wall clock time under various memory budgets compared with the baselines. Kahou Tam, Chunlin Tian, Li Li 0064, Haikai Zhao, Cheng-Zhong Xu 0001 |
SenSys | 1 |
| 2024 | Breaking the Memory Wall for Heterogeneous Federated Learning via Model SplittingabstractFederated Learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. Ever-increasing model complexity coupled with limited memory resources on the participating devices severely bottlenecks the deployment of FL in real-world scenarios. Thus, a framework that can effectively break the memory wall while jointly taking into account the hardware and statistical heterogeneity in FL is urgently required. In this article, we proposeSmartSplita framework that effectively reduces the memory footprint on the device side while guaranteeing the training progress and model accuracy for heterogeneous FL through model splitting. Towards this end,SmartSplitemploys a hierarchical structure to adaptively guide the overall training process. In each training round, the central manager, hosted on the server, dynamically selects the participating devices and sets the cutting layer by jointly considering the memory budget, training capacity, and data distribution of each device. The MEC manager, deployed within the edge server, proceeds to split the local model and perform training of the server-side portion. Meanwhile, it fine-tunes the splitting points based on the time-evolving statistical importance. The on-device manager, embedded inside each mobile device, continuously monitors the local training status while employing cost-aware checkpointing to match the runtime dynamic memory budget. Extensive experiments on representative datasets are conducted on both commercial off-the-shelf mobile device testbeds. The experimental results show thatSmartSplitexcels in FL training on highly memory-constrained mobile SoCs, offering up to a 94% peak latency reduction and 100-fold memory savings. It enhances accuracy performance by 1.49%-57.18% and adaptively adjusts to dynamic memory budgets through cost-aware recomputation Chunlin Tian, Li Li 0064, Kahou Tam, Yebo Wu, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | FedCoop: Cooperative Federated Learning for Noisy LabelsabstractFederated Learning coordinates multiple clients to collaboratively train a shared model while preserving data privacy. However, the training data with noisy labels located on the participating clients severely harm the model performance. In this paper, we propose FedCoop, a cooperative Federated Learning framework for noisy labels. FedCoop mainly contains three components and conducts robust training in two phases, data selection and model training. In the data selection phase, in order to mitigate the confirmation bias caused by a single client, the Loss Transformer intelligently estimates the probability of each sample’s label to be clean through cooperating with the helper clients, which have high data trustability and similarity. After that, the Feature Comparator evaluates the label quality for each sample in terms of latent feature space in order to further improve the robustness of noisy label detection. In the model training phase, the Feature Matcher trains the model on both the noisy and clean data in a semi-supervised manner to fully utilize the training data and exploits the feature of global class to increase the consistency of pseudo labeling across the clients. The experimental results show FedCoop outperforms the baselines on various datasets with different noise settings. It effectively improves the model accuracy up to 62% and 27% on average compared with the baselines. Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001 |
ECAI | 1 |