EDBT 2026 Demo / reviewers in the wild / expert
Kaiyu Chen
dblp:08/2232
· DBLP profile ↗
13ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structured Time-Frequency Feature-Driven Meta-Learning for Fall Detection With mmWave RadarabstractMillimeter-wave radar has emerged as a preferred modality for privacy-preserving human activity recognition. However, the effectiveness of current deep learning approaches is often severely compromised by data scarcity and incomplete feature representation in complex, real-world environments. To address these challenges, this paper presents a novel meta-learning framework tailored for few-shot radar recognition. Structured time-frequency feature strategy is proposed, representing the systematic extraction of multi-spectral joint temporal-spectral features from raw IQ data. Unlike traditional passive mapping, this strategy utilizes structured temporal alignment to construct distinctive two-dimensional maps that actively compensate for information loss. Synergizing with this, we propose the SAMSCNN architecture, which departs from conventional layer stacking to employ a customized hierarchical multi-scale topology. This architecture innovatively integrates spatially-aware self-attention with meta-learning optimization, enabling the network to rapidly adapt to new tasks while preserving fine-grained signal structures. Experimental results demonstrate state-of-the-art performance, achieving 99.21% accuracy specifically in fall detection. Notably, under severe few-shot conditions with only five training samples, the method sustains an average accuracy of 90.24% across six distinct human activities, significantly outperforming existing methods. Extensive comparative analyses further confirm the framework’s superior robustness across varying scene configurations and environmental interference. Kaiyu Chen, Shaoxi Wang, Yongxin Guo 0002, Hao Zhang 0076 |
IEEE Internet Things J. | 1 |
| 2026 | mmWave Radar-Based Continuous Sign Language Recognition: Lightweight Modeling, Contextual Optimization, and Embedded ImplementationabstractCommunication barriers for the hearing-impaired represent a significant societal concern. Existing sign language recognition solutions are constrained by lighting sensitivity, wearable device burdens, and limited vocabulary coverage. This work presents the first FMCW mmWave radar-based continuous sign language recognition framework, with three key innovations: (1) the Radar Continuous Chinese Sign Language-108 (RCCSL-108) dataset with 108 isolated signs and 66 natural sentences to address the critical absence of sentence-level radar data; (2) a novel lightweight lightweight Continuous Sign Language Recognition Temporal Convolutional Network (CSLR-TCN) that incorporates dual-dilated convolutions for variable-length inputs, achieving 92.66% frame accuracy and robust cross-user generalization; and (3) a novel contextual reordering module that enhances semantic understanding by 6.33% on unseen data by mitigating emotion loss in translation. The implemented embedded system achieves 87.49% recognition accuracy in untrained user experiments with ≥15 FPS throughput, marking a pivotal advancement toward practical radar-based sign language interfaces. Zhiyan Lin, Minming Gu, Kaiyu Chen, Keyu Pan |
IEEE Internet Things J. | 3 |
| 2025 | Accelerating Shortest Path Counting on Road NetworksabstractCounting the number of shortest paths between two query vertices on road networks has a wide range of applications and has recently drawn significant research attention. The state-of-the-art solution builds a tree-based index using the concept of tree decomposition. However, its performance deteriorates when the tree decomposition results in an unbalanced tree and may not perform well when the query vertices are close to each other. This paper aims to improve the efficiency of shortest path counting. We propose a novel indexing scheme that combines hub labeling with a balanced tree hierarchy. This approach significantly reduces the number of visited labels compared to the state-of-the-art solution. Furthermore, we introduce several optimizations to enhance the efficiency of index construction and minimize its size. Extensive experiments conducted on real-world road networks demonstrate that our method achieves up to 4.1 times higher query efficiency and reduces the index size by a factor of 2.35 compared to the state-of-the-art solution. Kaiyu Chen, Dong Wen 0001, Zhengyi Yang 0001, Wentao Li 0001, Ying Zhang 0001 |
ICDE | 2 |
| 2025 | Covering K-Cliques in Billion-Scale GraphsabstractThe k-clique structure in graphs has been investigated in various real-world applications, such as community detection in complex networks, functional module discovery in biological networks, and link spam detection in web graphs. Despite extensive research on k-clique enumeration, the large number of k-cliques in many graphs poses a challenge for practical application and computation. To address this, we explore the k-clique τ-cover problem, a generalization of the vertex cover problem. The problem aims to find a small set of vertices that can effectively represent all k-cliques in the graph. We prove the NP-hardness of finding the minimum k-clique cover. We propose a hierarchical solution that computes a small cover without enumerating k-cliques. Extensive experiments on real-world graphs verify the efficiency and effectiveness of our solution. Kaiyu Chen, Dong Wen 0001, Hanchen Wang 0001, Zhengyi Yang 0001, Wenjie Zhang 0001, Xuemin Lin 0001 |
WWW | 1 |
| 2025 | Liquid metal microfluidic cooling system for high-efficiency thermal management via learning-based genetic algorithmabstractHigh heat flux density is a critical factor that limits the performance and reliability of miniaturized, high-power microelectronic systems. This study proposes a liquid metal (LM)-based microfluidic cooling system optimized through a data-driven computational framework based on an enhanced Genetic Algorithm (LC-GA), aiming to deliver an efficient thermal management solution for high-density integrated systems. By integrating LM near-junction cooling with microchannel heat dissipation in a silicon substrate, we developed a heterogeneous three-dimensional interconnect cooling architecture capable of optimizing thermal performance through algorithm-guided parameter tuning. To validate the proposed method, four distinct microchannel configurations were designed, fabricated, and experimentally tested. LM was introduced into the channels to conduct both experimental cooling tests and thermal performance simulations on a simulated heat source. The results demonstrate that this LM-based microfluidic cooling system, optimized through computational parameter determination, can effectively dissipate heat from chips with power consumption up to 800 W while maintaining stable thermal performance. Additionally, a response surface methodology combined with enhanced LC-GA was utilized for multi-factor sensitivity analysis and multi-objective optimization, enabling automatic determination of optimal design and operating parameters to balance thermal resistance and pressure drop. The optimized configuration reduced the maximum chip temperature to approximately 357.54 K, lowered the system pressure requirement, and improved the Performance Evaluation Criterion (PEC) to 2.327. This work provides a data-driven optimization approach that supports the development of high-performance integrated microsystems through algorithm-assisted thermal design. Yucheng Wang 0007, Antong Bi, Kaiyu Chen, Shenxin Yu, Wanping Gao, Yuwan Wu, Shaoxi Wang |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | On Compressing Historical Cliques in Temporal Graphs
Kaiyu Chen, Dong Wen 0001, Wentao Li 0001, Zhengyi Yang 0001, Wenjie Zhang 0001 |
DASFAA (1) | 1 |
| 2024 | Querying Structural Diversity in Streaming GraphsabstractStructural diversity of a vertex refers to the diversity of connections within its neighborhood and has been applied in various fields such as viral marketing and user engagement. The paper studies querying the structural diversity of a vertex for any query time windows in streaming graphs. Existing studies are limited to static graphs which fail to capture vertices' structural diversities in snapshots evolving over time. We design an elegant index structure to significantly reduce the index size compared to the basic approach. We propose an optimized incremental algorithm to update the index for continuous edge arrivals. Extensive experiments on real-world streaming graphs demonstrate the effectiveness of our framework. Kaiyu Chen, Dong Wen 0001, Wenjie Zhang 0001, Ying Zhang 0001, Xiaoyang Wang 0002, Xuemin Lin 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Perimeter Control and Route Guidance of Multi-Region MFD Systems With Boundary Queues Using Colored Petri NetsabstractPerimeter control based on Macroscopic Fundamental Diagram (MFD) aims to meter the number of transferring vehicles at the periphery of the protected urban region in order to obtain the desired number of vehicles in that region. The advantage of perimeter control is less computational effort, while its drawback is that it may create long queues and delays at the perimeter of the controlled area. For capturing boundary queue dynamics, an enhanced accumulation-based MFD model is proposed using colored Petri Nets by considering transfer flows, boundary queues and travel delays simultaneously. The gated intersections and related road segments on the border of a protected region are modeled as so-called boundary buffers. Based on the enhanced MFD model, anintegrated perimeter control framework is proposed with consideration of travel time and queuing time in buffers. In this framework, the controllers between peripheral and protected region are optimized using model predictive control theory. Then, internal flow controllers are adopted to homogenize traffic density among subregions, and route guidance is also used to balance the number of queuing vehicles among boundary buffers. Simulation results verify the effectiveness of the proposed integrated perimeter control. Furthermore, the impacts of buffer storage capacity on region heterogeneity and trip completion rates are also investigated in this paper.Note to Practitioners—It is challenging to manage traffic congestion in large-scale urban network. Perimeter control provides an accumulation-based methodology with consideration of the existing correlation between traffic density and flow, which is known as MFD. For practical application, the efficiency as well asweakness of potential perimeter control strategies need to be evaluated and improved using customized traffic simulations. Accumulation-based traffic model using Petri Nets is introduced to serve for perimeter control, in which the intersections and road segments on the boundary of each pair of adjacent subregions are modeled as a boundary buffer. Both perimeter control and route guidance are integrated in the proposed control framework considering the queuing vehicles in the boundary buffers. Moreover, the effect of buffer storage capacity on network performances is tested, which is the essential for traffic engineers to design and implement management measures in practice. Saifei Chen, Kaiyu Chen, Anastasios Kouvelas, Nikolas Geroliminis |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | C-for-Metal: High Performance Simd Programming on Intel GPUsabstractThe SIMT execution model is commonly used for general GPU development. CUDA and OpenCL developers write scalar code that is implicitly parallelized by compiler and hardware. On Intel GPUs, however, this abstraction has profound performance implications as the underlying ISA is SIMD and important hardware capabilities cannot be fully utilized. To close this performance gap we introduce C- For- Metal (CM), an explicit SIMD programming framework designed to deliver close-to-the-metal performance on Intel GPUs. The CM programming language and its vector/matrix types provide an intuitive interface to exploit the underlying hardware features, allowing fine-grained register management, SIMD size control and cross-lane data sharing. Experimental results show that CM applications from different domains outperform the best-known SIMT-based OpenCL implementations, achieving up to 2.7x speedup on the latest Intel GPU. Guei-Yuan Lueh, Kaiyu Chen, Joel Fuentes, Wei-Yu Chen, Fangwen Fu, Hongzheng Li, Daniel Rhee |
CGO | 2 |
| 2019 | Image Block Augmentation for One-Shot LearningabstractGiven one or a few training instances of novel classes, oneshot learning task requires that the classifier generalizes to these novel classes. Directly training one-shot classifier may suffer from insufficient training instances in one-shot learning. Previous one-shot learning works investigate the metalearning or metric-based algorithms; in contrast, this paper proposes a Self-Training Jigsaw Augmentation (Self-Jig) method for one-shot learning. Particularly, we solve one-shot learning by directly augmenting the training images through leveraging the vast unlabeled instances. Precisely our proposed Self-Jig algorithm can synthesize new images from the labeled probe and unlabeled gallery images. The labels of gallery images are predicted to help the augmentation process, which can be taken as a self-training scheme. Intrinsically, we argue that we provide a very useful way of directly generating massive amounts of training images for novel classes. Extensive experiments and ablation study not only evaluate the efficacy but also reveal the insights, of the proposed Self-Jig method. Zitian Chen, Yanwei Fu 0001, Kaiyu Chen, Yu-Gang Jiang 0001 |
AAAI | 3 |
| 2018 | Register allocation for Intel processor graphicsabstractRegister allocation is a well-studied problem, but surprisingly little work has been published on assigning registers for GPU architectures. In this paper we present the register allocator in the production compiler for Intel HD and Iris Graphics. Intel GPUs feature a large byte-addressable register file organized into banks, an expressive instruction set that supports variable SIMD-sizes and divergent control flow, and high spill overhead due to relatively long memory latencies. These distinctive characteristics impose challenges for register allocation, as input programs may have arbitrarily-sized variables, partial updates, and complex control flow. Not only should the allocator make a program spill-free, but it must also reduce the number of register bank conflicts and anti-dependencies. Since compilation occurs in a JIT environment, the allocator also needs to incur little overhead. Wei-Yu Chen, Guei-Yuan Lueh, Pratik Ashar, Kaiyu Chen, Buqi Cheng |
CGO | 4 |
| 2008 | Runtime validation of memory ordering using constraint graph checkingabstractAn important correctness issue for emerging multi/many-core shared memory systems is to ensure that the inter-processor communication through shared memory conforms to the memory ordering rules, as specified by the architecturepsilas memory consistency model. This presents a significant validation challenge. Growing system complexity makes it increasingly hard to identify all deep-state logic bugs in pre-silicon verification. Further, aggressive technology scaling makes hardware more vulnerable to dynamic errors that can only be detected at runtime. In this paper, we propose an approach for runtime validation of memory ordering. This allows us to survive bugs that escape pre-silicon verification, as well as deal with emerging dynamic errors. Our solution consists of two parts: 1) at the microarchitecture level, we add efficient hardware support to capture the observed ordering among shared-memory operations; 2) we perform online verification of the observed memory ordering by checking for cycles in the constraint graph. We combine these to achieve end-to-end correctness validation of the system execution with respect to the memory ordering specification. There are several challenges that need to be addressed to make this approach practical. We describe these, as well as optimization techniques for reducing the hardware overhead. Estimates obtained from preliminary chip multiprocessor simulation experiments show that the proposed techniques are very effective in achieving acceptable hardware overhead and minimal performance impact. Kaiyu Chen, Sharad Malik, Priyadarsan Patra |
HPCA | 1 |
| 2006 | Dependable Multithreaded Processing Using Runtime ValidationabstractModern processors face growing verification and reliability challenges posed by increasing micro-architecture complexity and aggressive technology scaling. While viable approaches have been proposed to address these challenges in the context of uniprocessors, little work has been done for emerging multithreaded processors. Multithreading raises new issues for validation due to inter-thread interactions and inherent complexity of the underlying hardware. We propose an extension of the DIVA approach, which employs a simple checker processor to effectively validate the complex superscalar processor, to perform instruction-level runtime validation for both intra-thread and inter-thread correctness properties for multithreaded execution. We present the validation methodology using a representative simultaneous-multithreaded (SMT) architecture, and briefly discuss its general applicability to other forms of multithreading. Detailed timing simulation shows this solution has low performance penalty, while providing general robustness against both operational and functional errors with relatively small hardware overhead Kaiyu Chen, Sharad Malik |
PRDC | 1 |