VLDB 2026 Research / reviewers in the wild / expert
Hongyu Cai
dblp:207/5088
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0003-0048-5626ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoding inner speech via frequency-specific cortical EEG representationsabstractNeural decoding of inner speech provides a crucial pathway for silent communication and cognitive state recognition. However, existing circular brain electrical activity mapping (BEAM)-based decoding methods suffer from the problem that edge distortion and information loss easily occur, and they are mismatched with the input structure of convolutional networks, ultimately limiting decoding accuracy. To overcome this problem, this paper first proposed a frequency-specific rectangular BEAM sequence construction method that employed an adaptive variation function to generate high-fidelity EEG sequences across different frequency bands, preserving cortical spatial correlations while perfectly matching the rectangular input structure of convolutional networks. Then, based on the continuous temporal and multi-frequency features of rectangular BEAM maps, 3D continuous frequency-spatial and temporal-spatial dynamic representation modules, respectively, were constructed. Finally, a dual-branch 3D convolutional network was designed to achieve joint decoding of inner speech brain activity. The manuscript reorganized the publicly available InnerSpeech dataset, involving 10 subjects, 3 task modalities, and 4 categories, with the first 10 trials selected for each condition. In total, 1200 samples were used to evaluate the proposed model’s ability to accurately classify inner speech. Experimental results showed that the proposed method outperformed the baseline model EEGNet, achieving accuracies of 61.79%, 62.67%, and 53.25% in the Inner, Pron, and Vis classification tasks, respectively, while also providing new insights into the mechanisms of frequency–task coupling and the development of high-precision brain–computer interfaces. Hongyu Cai, Zejia Yang, Jian Zhao 0011, Zhejun Kuang |
Inf. Process. Manag. | 2 |
| 2026 | U-RWKV: Accurate and Efficient Volumetric Medical Image Segmentation via RWKVabstractAccurate and efficient volumetric medical image segmentation is vital for clinical diagnosis, pre-operative planning, and disease-progression monitoring. Conventional convolutional neural networks (CNNs) struggle to capture long-range contextual information, whereas Transformer-based methods suffer from quadratic computational complexity, making it challenging to couple global modeling with high efficiency. To address these limitations, we explore an effective yet accurate segmentation model for volumetric data. Specifically, we introduce a novel linear-complexity sequence modeling technique, RWKV, and leverage it to design a Tri-directional Spatial Enhancement RWKV (TSE-R) block; this module performs global modeling via RWKV and incorporates two optimizations tailored to three-dimensional data: 1) a spatial-shift strategy that enlarges the local receptive field and facilitates inter-block interaction, thereby alleviating the structural information loss caused by sequence serialization; and 2) a tri-directional scanning mechanism that constructs sequences along three distinct directions, applies global modeling via WKV, and fuses them with learnable weights to preserve the inherent 3D spatial structure. Building upon the TSE-R block, we develop an end-to-end 3D segmentation network, termed U-RWKV, and extensive experiments on three public 3D medical segmentation benchmarks demonstrate that U-RWKV outperforms state-of-the-art CNN-, Transformer-, and Mamba-based counterparts, achieving a Dice score of 87.21% on the Synapse multi-organ abdominal dataset while reducing parameter count by a factor of 16.08 compared with leading methods. Hongyu Cai, Zhejun Kuang |
IEEE Trans. Image Process. | 1 |
| 2025 | A Progressive Transformer for Unifying Binary Code Embedding and Knowledge TransferabstractLanguage models have recently been applied to binary analysis tasks, such as function similarity detection and function signature recovery. These models typically employ a two-stage training process: pre-training via Masked Language Modeling (MLM) on machine code and fine-tuning for specific tasks. While MLM helps to understand binary code structures, it ignores essential code characteristics, including control and data flow, which negatively affect model generalization. Recent work leverages domain-specific features (e.g., control flow graphs and dynamic execution traces) in transformer-based approaches to improve binary code semantic understanding. This approach, however, involves complex feature engineering, a cumbersome and time-consuming process that can introduce predictive uncertainty when dealing with stripped or obfuscated code, which leads to a performance drop. We introduce PROTST, a novel transformer-based methodology for binary code embedding. PROTST employs a hierarchical training process based on a unique tree-like structure, where knowledge progressively flows from fundamental tasks at the root to more specialized tasks at the leaves. This progressive teacher-student paradigm allows the model to build upon previously learned knowledge, resulting in high-quality embeddings that can be effectively leveraged for diverse downstream binary analysis tasks. The effectiveness of PROTST is evaluated in seven binary analysis tasks, demonstrating an average of 14.8% improvement in F1 and MRR compared to traditional two-stage training, and a 16.6 % improvement when analyzing obfuscated code. Hanxiao Lu, Hongyu Cai, Yiming Liang, Antonio Bianchi, Z. Berkay Celik |
SANER | 2 |
| 2021 | NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip AcceleratorsabstractThe revolution of machine learning poses an unprecedented demand for computation resources, urging more transistors on a single monolithic chip, which is not sustainable in the Post-Moore era. The multichip integration with small functional dies, called chiplets, can reduce the manufacturing cost, improve the fabrication yield, and achieve die-level reuse for different system scales. DNN workload mapping and hardware design space exploration on such multichip systems are critical, but missing in the current stage.This work provides a hierarchical and analytical framework to describe the DNN mapping on a multichip accelerator and analyze the communication overhead. Based on this framework, we propose an automatic tool called NN-Baton with a pre-design flow and a post-design flow. The pre-design flow aims to guide the chiplet granularity exploration with given area and performance budgets for the target workload. The post-design flow focuses on the workload orchestration on different computation levels -package, chiplet, and core - in the hierarchy. Compared to Simba, NN-Baton generates mapping strategies that save 22.5%∼44% energy under the same computation and memory configurations.The architecture exploration demonstrates that area is a decisive factor for the chiplet granularity. For a 2048-MAC system under a 2 mm2chiplet area constraint, the 4-chiplet implementation with 4 cores and 16 lanes of 8-size vector-MAC is always the top-pick computation allocation across several benchmarks. In contrast, the optimal memory allocation policy in the hierarchy typically depends on the neural network models. Zhanhong Tan, Hongyu Cai, Runpei Dong, Kaisheng Ma |
ISCA | 2 |