Kaijie Gong

dblp:346/2601 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-2872-7327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EdgeGen: Efficient LLM-Empowered Model Generation with Quantization-Aware NAS
abstract
The rapid evolution of the Web of Things (WoT) has created new opportunities for connectivity and standardization across heterogeneous devices, enabling the development of increasingly complex systems. However, edge devices deployed in resource-constrained scenarios face significant challenges. These devices require lightweight and efficient models to achieve high accuracy while operating within strict memory constraints. Typical approaches to model generation, Neural Architecture Search (NAS), have proven effective in automating the search for optimal architectures. However, existing NAS methods suffer from two critical limitations: (1) they fail to incorporate quantization into the search space, which can result in overlooking larger models that might perform better after quantization; and (2) current model evaluation methods struggle to provide accurate assessments within a short time period. To address these challenges, we propose EdgeGen, a novel NAS framework that integrates multiple quantization methods into the search process, enabling the discovery of larger models that need quantization to satisfy the constraint. EdgeGen employs a multi-beam Monte Carlo Tree Search (MCTS) algorithm and a constraint validator to explore the expanded search space efficiently, searching vast original and quantized models. Furthermore, EdgeGen evaluates the model performance following a GNN-based performance predictor, which provides a rapid and precise prediction. Across multiple benchmarks, EdgeGen consistently outperforms state-of-the-art NAS methods. Code available at: https://doi.org/10.5281/zenodo.18323232.
Yingqi Peng, Kaijie Gong, Yi Gao 0001, Wei Dong 0001
WWW3
2026 A Dual-Encoder Convolutional Neural Networks-Frequency Transformer framework with bidirectional attention for precise brain tumor segmentation
abstract
Brain tumors present significant neurological challenges with high mortality rates, where early diagnosis based on magnetic resonance imaging (MRI) is clinically vital but is hindered by labor-intensive manual segmentation. To overcome limitations in existing approaches, where convolutional neural networks (CNNs) struggle with global context dependencies and Transformers face computational inefficiency, we propose a dual-branched framework called bifurcated fusion dual-encoder convolutional neural networks –frequency Transformer (BiF-DTNet). The BiF-DTNet employs a dual-encoder architecture, where the CNN encoder captures local features. In contrast, the frequency-domain vision Transformer (FViT) encoder efficiently models global context through fast Fourier transform and self-attention. The model incorporates bidirectional spatial-channel attention (BISC) blocks for effective multi-scale feature fusion, enhancing segmentation accuracy. While BiF-DTNet performs feature encoding on two-dimensional (2D) axial slices for computational efficiency, the model reconstructs the final segmentation in a volumetric manner, enabling precise and coherent three-dimensional (3D) tumor prediction. Evaluated on the brain tumor segmentation (BraTS) 2020 and BraTS 2021 benchmarks, BiF-DTNet achieves Dice scores of 80.42%/85.15%/91.34% (enhancing tumor (ET)/tumor core (TC)/whole tumor (WT)) on BraTS 2020 and 85.98%/91.25%/93.64% on BraTS 2021, outperforming state-of-the-art baselines. These results conclusively demonstrate BiF-DTNet’s superiority in precise 3D tumor segmentation, particularly for irregular boundaries and small lesions, through its synergistic integration of local and global features.
Kaijie Gong, Dichao Pan, Mohammed A. A. Al-qaness, Jianguo Shen
Eng. Appl. Artif. Intell.1
2026 Exploiting Partial JPEG Decoding to Mitigate On-Device Image Processing
abstract
Device-cloud collaborative inference is often necessary for resource-constrained IoT devices that cannot support full on-device models. To minimize bandwidth and support concurrency, existing methods typically compress images before transmission. However, these approaches often ignore the significant overhead of decoding native JPEG camera output, especially for high-resolution frames. Our measurements show that the on-device (Raspberry Pi 4B) decoding overhead for 700KB JPEG format is$\sim$14.4x the latency of on-cloud (GeForce RTX 3090) meter recognition inference. To reduce on-device decoding overhead, we design DC Camera, which is built upon a JPEG camera and leverages partial JPEG decoding to efficiently extract DC features from high-resolution images, significantly mitigating on-device image processing overhead. These DC features can preserve structural information better than conventional downsampled images. We utilize DC Camera to implement fast meter recognition system and deploy the system in material science laboratory to monitor multiple meters. Our evaluation demonstrates that compared to state-of-the-art (SOTA) methods, DC Camera can reduce on-device computation overhead by$\sim$5.8x and decrease transmission volume by$\sim$90.9x, without inference accuracy degradation.
Kaijie Gong, Hao Wang 0238, Yi Gao 0001, Weijie Fang, Wei Dong 0001
IEEE Trans. Mob. Comput.1
2025 Programming Embedded IoT Applications in Natural Language with IoTPilot
abstract
In recent years, the swift expansion of Internet of Things (IoT) applications has been notable. However, developing a comprehensive IoT application is highly challenging for non-expert developers due to the highly diverse characteristics of embedded operating systems. LLM-based methods provide a paradigm for code generation through natural language, which can greatly simplify and accelerate the development of IoT applications. While promising, existing works have failed to account for the specific characteristics of embedded operating system, resulting in the lower quality of generated IoT code. In this paper, we present IoTPilot, an LLM-driven embedded IoT programming tool. We have observed that the conflicts between LLM internal APIs/headers and external OS-specific APIs/headers are key factors leading to the low quality of generated embedded IoT applications. Thus, we introduce two effective self-thinking chains to integrate internal LLM knowledge with external documentation, addressing conflicts in APIs and headers. We provide embedded IoT benchmarks (IoTEval), which are built on RIOT, Zephyr, Contiki and FreeRTOS. Results show that IoTPilot can improve the performance of IoT code generation on all the three embedded OSes compared with existing state-of-the-art (SOTA) methods.
Kaijie Gong, Wei Dong 0001, Hao Wang 0238, Yingqi Peng, Yi Gao 0001
MobiSys1
2025 Optimizing WebAssembly Bytecode for IoT Devices Using Deep Reinforcement Learning
abstract
WebAssembly has shown promising potential on various IoT devices to achieve the desired features such as multi-language support and seamless device-cloud integration. The execution performance of WebAssembly bytecode is directly influenced by compilation sequences. While existing research has explored the optimization of compilation sequences for native code, these approaches are not suitable to WebAssembly bytecode due to its unique instruction format and control flow graph structure. In this work, we propose WasmRL, a novel efficient deep reinforcement learning (DRL)-based compiler optimization framework tailored for WebAssembly bytecode. We conduct a fine-grained analysis of the characteristics of WebAssembly instructions and associated compilation flags. We observe that the same compilation sequence may yield contrasting performance outcomes in WebAssembly and native code. Motivated by our observation, we introduce a WebAssembly-specific DRL state representation that simultaneously captures the impact of various compilation sequences on the WebAssembly bytecode and its runtime performance. To enhance the training efficiency of the DRL model, we propose a tree-based action space refinement method. Furthermore, we develop a pluggable cross-platform training strategy to optimize WebAssembly bytecode across different IoT devices. We evaluate the performance of WasmRL extensively on PolybenchC, MiBench, Shootout public datasets and real-world IoT applications. Experimental results show: (1) The DRL model trained on a specific device achieves 1.4x/1.1x speedups over -O3 for seen/unseen programs; (2) The DRL model trained on different devices simultaneously achieves 1.21x/1.06x improvements respectively. The code has been available at https://github.com/CarrollAdmin/WasmRL .
Kaijie Gong, Yi Gao 0001, Wei Dong 0001
ACM Trans. Internet Techn.1
2024 Poster: Enabling IoT Application Programming in Natural Language with IoTPilot
abstract
In recent years, the swift expansion of Internet of Things (IoT) applications has been notable. However, developing a comprehensive IoT application is highly challenging for non-expert developers due to the highly diverse characteristics of embedded operating systems. The LLM-based approach shows promise in generating code from natural language, but its performance in IoT code generation is poor. This stems from the LLM's insufficient understanding of the embedded IoT code context, leading to missed and conflicting OS-specific APIs. In this paper, we present IoTPilot, a LLM-driven multi-agent IoT programming framework. We develop a clustering-based progressive RAG strategy and auto-calibrating self-debug mechanism to enhance the quality of generated IoT applications.
Kaijie Gong, Wei Dong 0001, Yingqi Peng, Hao Wang 0238, Yi Gao 0001
SenSys1
2023 LinkLab 2.0: A Multi-tenant Programmable IoT Testbed for Experimentation with Edge-Cloud Integration
Wei Dong 0001, Borui Li 0001, Kaijie Gong, Wenzhao Zhang, Yi Gao 0001
NSDI5