EDBT 2026 Demo / reviewers in the wild / expert
Fan Cui
dblp:190/7297
· DBLP profile ↗
17ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cayman: Custom Accelerator Generation with Control Flow and Data Access OptimizationabstractCustom accelerators enhance System-on-Chips’ performance through hardware specialization. High-level synthesis (HLS) can automatically synthesize accelerators for given kernels but requires manual selection and extraction of kernels from applications. This paper proposes Cayman, the first end-to-end framework to synthesize high-performance custom accelerators with both control flow and data access optimization. Cayman automatically selects kernels for hardware acceleration based on a hierarchical program representation, which captures kernel candidates with general control flows. Besides, Cayman optimizes accelerators with specialized processor-accelerator interfaces for data access acceleration. Cayman further introduces a novel accelerator merging mechanism to synthesize reusable accelerators. Experiments on various benchmarks demonstrate that Cayman outperforms two state-of-the-art frameworks by $8.0 \times$ and $14.4 \times$. Youwei Xiao, Fan Cui, Zizhang Luo, Weijie Peng, Yun Liang 0001 |
DAC | 2 |
| 2025 | An Empirical Comparision of LLM-based Hardware Design and High-level SynthesisabstractField-Programmable Gate Arrays (FPGAs) are increasingly used for accelerating diverse applications due to their reconfigurability and ability to implement custom hardware architectures. However, programming FPGAs remains challenging, traditionally relying on low-level Hardware Description Languages (HDLs) like Verilog, which are intricate and time-consuming. High-Level Synthesis (HLS) tools, such as Vitis HLS, have emerged to address these issues by allowing hardware functionality description in high-level languages like C/C++, but they come with their own limitations, including less efficient hardware implementations, delay overhead caused by conservative scheduling strategies, and unpredictable solutions due to semantic differences between software and hardware. Fan Cui, Youwei Xiao, Kexing Zhou, Yun Liang 0001 |
FPGA | 1 |
| 2025 | GLCLAP: A Novel Contrastive Learning Pre-trained Model for Contextual Biasing in ASR
Yuxiang Kong, Fan Cui, Liyong Guo, Heinrich Dinkel, Lichun Fan, Jian Luan 0001 |
INTERSPEECH | 2 |
| 2025 | PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language ModelsabstractCurrent benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9\% accuracy compared to human experts' 61.9\%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204\% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/. Shi Qiu 0016, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun, Jiashen Wei, Tianyu Luo, Yixuan Yin 0003, Haoxu Zhang, Chencheng Tang, Haoling Chang, Jingtian Zhang, Zhangyi Liu, Yuku Zhang, Boxuan Jing, Xianqi Yin, Yutong Ren, Zizhuo Fu, Jiaming Ji, Anqi Lv, Laifu Man, Jianxiang Li, Feiyu Tao, Qihua Sun, Zhou Liang, Yushu Mu, Zhongxuan Li, Jing-Jun Zhang, Xingqi Xia, Zheyu Shen, Jiahang Chen, Qiuhao Xiong, Binran Wang, Fengyuan Wang, Ziyang Ni, Fan Cui, Changkun Shao, Qing-Hong Cao, Ming-xing Luo, Muhan Zhang, Hua Xing Zhu |
NeurIPS | 48 |
| 2025 | Depiction of Subsurface Leak Areas Based on Adaptive Sensitive Frequency Attribute AnalysisabstractGround penetrating radar (GPR) frequency attributes are commonly used to describe subsurface structures and characterize anomalous media. Single-frequency slices face challenges in capturing the broadband characteristics of GPR data, so the fusion of extracted multi-frequency components using a fusion algorithm can be effective. However, selecting appropriate attributes and mapping them to media characterization remain unresolved challenges. In this study, we propose a workflow based on adaptive sensitive frequency attribute analysis (ASFAA) to address these issues. First, the generalized S-transform (GST) is used to calculate the multi-frequency attributes of GPR data. Then, a sensitive feature analysis method combining hierarchical clustering and correlation analysis is employed to reduce redundancy in frequency attributes. Multi-frequency data are fused using the potential of heat-diffusion for affinity-based transition embedding (PHATE), which performs affinity-based diffusion embedding. The workflow is tested with synthetic and field data, yielding characterization results consistent with both the forward model results and actual leak extents. Therefore, the proposed workflow effectively integrates multi-frequency components, demonstrating its capability to delineate leak extents. Fan Cui, Guoqi Dong, Guixin Zhang, Mengli Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Bonspiel: Low Tail Latency Transactions in Geo-Distributed DatabasesabstractTail latency is crucial as it impacts user satisfaction and service-level objectives (SLOs). However, geo-distributed databases have long struggled with this issue due to wide-area network access, resulting in tail latencies of several or even exceeding ten seconds. In this paper, we highlight that further optimizing atomic commit protocols does not help but hit a tail latency wall. Instead, making concurrency control and access method selection geo-aware can mitigate this issue. To this end, we present Bonspiel, a new geo-distributed database equipped with geo-aware concurrency control and access method selection. In our experiments, Bonspiel successfully caps the tail latency of TPC-C at 1.8 seconds. Remarkably, it achieves this while maintaining full generality - it is fully SQL-compliant and strongly consistent, with both average latency and system throughput remaining at the top of the field. Fan Cui, Eric Lo 0001, Srijan Srivastava, Ziliang Lai |
Proc. VLDB Endow. | 1 |
| 2024 | Occam's Razor for Distributed ProtocolsabstractOptimizing distributed protocols has traditionally been a real pain, requiring experts to figure out where improvements can be made, along with rigorous correctness proofs and meticulous implementation. This paper presents a theory to systematize this process. The proposed theory can optimize any existing distributed protocols while preserving all their original quality attributes (e.g., generality, correctness). Crucially, applying the optimizations derived from this theory does not necessitate touching the original implementation --- all you need is just to bolt on a few new message adapters. Case studies demonstrate the effectiveness of this approach. For instance, applying the theory to optimize Spanner's atomic commit protocol results in a mere 4 new message adapters, but improving its latency by up to 51% and peak throughput by up to 56%. Ziliang Lai, Fan Cui, Hua Fan 0002, Eric Lo 0001, Wenchao Zhou, Feifei Li 0001 |
SoCC | 2 |
| 2024 | OriGen: Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-ReflectionabstractRecent studies have demonstrated the significant potential of Large Language Models (LLMs) in generating Register Transfer Level (RTL) code, with notable advancements showcased by commercial models such as GPT-4 and Claude3-Opus. However, these proprietary LLMs often raise concerns regarding privacy and security. While open-source LLMs offer solutions to these concerns, they typically underperform commercial models in RTL code generation tasks, primarily due to the scarcity of high-quality open-source RTL datasets. To address this challenge, we introduce OriGen, a fully open-source framework that incorporates self-reflection capabilities and a novel dataset augmentation methodology for generating high-quality, large-scale RTL code. Our approach employs a code-to-code augmentation technique to enhance the quality of open-source RTL code datasets. Furthermore, OriGen can rectify syntactic errors through a self-reflection process that leverages compiler feedback. Fan Cui, Chenyang Yin, Kexing Zhou, Youwei Xiao, Guangyu Sun 0003, Qiang Xu 0001, Qipeng Guo, Yun Liang 0001, Xingcheng Zhang, Demin Song, Dahua Lin |
ICCAD | 1 |
| 2024 | A Method of Reconstructing Ground Penetrating Radar Bscan for Advanced Detection Based on High-Order Synchrosqueezing TransformabstractAccurately predicting geological structures (GS) ahead of tunnel excavation is of paramount importance for ensuring safety in coal mines. Standard time or frequency analysis methods face significant challenges due to limitations in downhole environments and ground penetrating radar (GPR) instrument performance, hindering the acquisition of precise geological information. Time-frequency (TF) analysis, capable of simultaneously examining non-stationary signals in both time and frequency domains, has become an important tool for interpreting GPR data. High-order synchrosqueezing transform (HSST) has a solid mathematical foundation and higher TF resolution than linear TF transforms and synchrosqueezing transforms. In this study, we propose a method utilizing HSST to extract TF features of geological structure responses and reconstruct them into novel Bscan for advanced detection of geological structures. Experimental examples using GPR forward simulation data containing GSs demonstrate the method’s precise characterization and intuitive visualization capabilities for GS interfaces, even in noise-polluted environments. Additionally, we present a real GPR data example that demonstrates the method’s ability to intuitively and accurately characterize the front GS interface. Field revelations provide robust verification of the method’s accuracy and practical applicability. Fan Cui, Guoqi Dong, Shuai Li 0021 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based NetworkabstractElectroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the EEG signal. It adopts the Attention-based Correlation Module (ACM) to discover the connection between auditory speech and EEG from global aspect, and the Shallow-Deep Similarity Classification Module (SDSCM) to decide the classification result via the embeddings learned from the shallow and deep layers. Moreover, various training strategies and data augmentation are used to boost the model robustness. Experiments are conducted on the dataset provided by Auditory EEG challenge (ICASSP Signal Processing Grand Challenge 2023). Results show that the proposed model has a significant gain over the baseline on the match-mismatch track. Fan Cui, Liyong Guo, Jiyao Liu, Ercheng Pei, Dongmei Jiang |
ICASSP | 1 |
| 2023 | Predicting Multi-Codebook Vector Quantization Indexes for Knowledge DistillationabstractKnowledge distillation (KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour of a teacher model. However, traditional KD methods suffer from teacher label storage issue, especially when the training corpora are large. Although on-the-fly teacher label generation tackles this issue, the training speed is significantly slower as the teacher model has to be evaluated every batch. In this paper, we reformulate the generation of teacher label as a codec problem. We propose a novel Multi-codebook Vector Quantization (MVQ) approach that compresses teacher embeddings to codebook indexes (CI). Based on this, a KD training framework (MVQ-KD) is proposed where a student model predicts the CI generated from the embeddings of a self-supervised pre-trained teacher model. Experiments on the LibriSpeech clean-100 hour show that MVQ-KD framework achieves comparable performance as traditional KD methods (11, 12), while requiring 256 times less storage. When the full LibriSpeech dataset is used, MVQ-KD framework results in 13.8% and 8.2% relative word error rate reductions (WERRs) for non -streaming transducer on test-clean and test-other and 4.0% and 4.9% for streaming transducer. The implementation of this work is already released as a part of the open-source project icefall1. Liyong Guo, Xiaoyu Yang 0005, Quandong Wang, Yuxiang Kong, Zengwei Yao, Fan Cui, Wei Kang 0006, Long Lin, Mingshuang Luo, Piotr Zelasko, Daniel Povey |
ICASSP | 6 |
| 2023 | Improving Weakly Supervised Sound Event Detection with Causal InterventionabstractExisting weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds, so they would be inevitably entangled, causing misclassification and biased localization results with only clip-level supervision. To tackle this issue, we first establish a structural causal model (SCM) to reveal that the context is the main cause of co-occurrence confounders that mislead the model to learn spurious correlations between frames and clip-level labels. Based on the causal analysis, we propose a causal intervention (CI) method for WSSED to remove the negative impact of co-occurrence confounders by iteratively accumulating every possible context of each class and then re-projecting the contexts to the frame-level features for making the event boundary clearer. Experiments show that our method effectively improves the performance on multiple datasets and can generalize to various baseline models. Yifei Xin, Dongchao Yang, Fan Cui, Yuexian Zou |
ICASSP | 3 |
| 2022 | Multi-Scale Refinement Network Based Acoustic Echo CancellationabstractRecently, deep encoder-decoder networks have shown outstanding performance in acoustic echo cancellation (AEC). However, the subsampling operations like convolution striding in the encoder layers significantly decrease the feature resolution lead to fine-grained information loss. This paper proposes an encoder-decoder network for acoustic echo cancellation with mutli-scale refinement paths to exploit the information at different feature scales. In the encoder stage, high-level features are obtained to get a coarse result. Then, the decoder layers with multiple refinement paths can directly refine the result with fine-grained features. Refinement paths with different feature scales are combined by learnable weights. The experimental results show that using the proposed multi-scale refinement structure can significantly improve the objective criteria. In the ICASSP 2022 Acoustic echo cancellation Challenge, our submitted system achieves an overall MOS score of 4.439 with 4.37 million parameters at a system latency of 40ms. Fan Cui, Liyong Guo, Peng Gao 0013 |
ICASSP | 1 |
| 2022 | Exploring representation learning for small-footprint keyword spottingabstractIn this paper, we investigate representation learning for low-resource keyword spotting (KWS). The main challenges of KWS are limited labeled data and limited available device resources. To address those challenges, we explore representation learning for KWS by self-supervised contrastive learning and self-training with pretrained model. First, local-global contrastive siamese networks (LGCSiam) are designed to learn similar utterance-level representations for similar audio samplers by proposed local-global contrastive loss without requiring ground-truth. Second, a self-supervised pretrained Wav2Vec 2.0 model is applied as a constraint module (WVC) to force the KWS model to learn frame-level acoustic representations. By the LGCSiam and WVC modules, the proposed small-footprint KWS model can be pretrained with unlabeled data. Experiments on speech commands dataset show that the self-training WVC module and the self-supervised LGCSiam module significantly improve accuracy, especially in the case of training on a small labeled dataset. Fan Cui, Liyong Guo, Quandong Wang, Peng Gao 0013 |
INTERSPEECH | 1 |
| 2019 | Learning image convolutional representations and complete tags jointly
Yanbin Wu, Hongbin Zhai, Mengna Li, Fan Cui, Li Wang 0055, Nitin Patil |
Neural Comput. Appl. | 4 |
| 2018 | Cross-model convolutional neural network for multiple modality data representation
Yanbin Wu, Li Wang 0055, Fan Cui, Hongbin Zhai, Baoming Dong |
Neural Comput. Appl. | 3 |
| 2018 | Coarse-to-fine salient object detection based on deep convolutional neural networks
Ying Li 0017, Fan Cui, Xizhe Xue, Jonathan Cheung-Wai Chan |
Signal Process. Image Commun. | 2 |