EDBT 2026 Demo / reviewers in the wild / expert
Jiaxuan Cai
dblp:277/6694
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HQC Post-Quantum Cryptography Decryption with Soft-Decision Reed-Solomon Decoder
Jiaxuan Cai, Xinmiao Zhang 0001 |
ISCAS | 1 |
| 2026 | Toward WAN-Aware LLM Training Across Heterogeneous, Geo-Distributed SitesabstractLarge Language Model (LLM) training is increasingly concentrated in homogeneous datacenters, while private data and underutilized GPUs across universities, laboratories, and edge sites remain difficult to use. This extended abstract presents preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns. The prototype connects three heterogeneous GPU sites via cloud-hosted parameter servers, outbound-only gRPC streams, two-stage delta compression (INT8 quantization + Huffman coding, achieving up to 4× payload reduction), and fault-tolerant rejoin. In real deployments, GPT-2 Medium pretraining achieves stable loss reduction and reaches the target loss 15.2% faster in wall-clock time than the best tested baseline; Llama3-1B pretraining remains stable under larger communication pressure; and cross-site latency traces reveal site-dependent WAN spikes of up to 200s. These results motivate adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training. Ziyue Luo, Jiaxuan Cai, Cedric Le Denmat, Srijith Nair, Fatemeh Nourzad, Rohith Krishnan Sudha, Qinhang Wu, Jifan Zhang, Zhe Li 0083, Peiwen Qiu, Siddharth Shah, Yinglun Xia, Xue Zheng, Bicheng Ying, Kaushik R. Chowdhury, Gauri Joshi, Yingbin Liang, Robert D. Nowak, Srinivasan Parthasarathy 0001, Saurav Prakash, Balaraman Ravindran, Sanjay Shakkottai, Ness Shroff, Sundararajan Srinivasan, Haibo Yang 0001, Aylin Yener, Jia Liu 0002 |
SIGCOMM | 2 |
| 2025 | Efficient Layered New Bit-Flipping QC-MDPC Decoder for BIKE Post-Quantum CryptographyabstractThe medium-density parity-check (MDPC) code-based Bit Flipping Key Encapsulation (BIKE) mechanism remains a candidate for post-quantum cryptography standardization. The latest version utilizes a new bit-flipping (BF) decoding algorithm, which decides the BF threshold by an affine function with high-precision coefficients. Previous BF decoder implementations can be extended to the new algorithm. However, they suffer from large memories that dominate the overall complexity. This paper proposes a column-layered decoder for the new BIKE BF decoding algorithm to substantially reduce the memory requirement, and optimizes the affine BF threshold function coefficients to reduce the code length needed for the same security level. For the first time, our work also investigates the impact of finite precision representation of the threshold coefficients on the decoding performance. For an example MDPC code considered for the standard, the proposed layered BF decoder achieves 20% complexity reduction compared to the best prior effort with a very small latency overhead. Jiaxuan Cai, Xinmiao Zhang 0001 |
ISCAS | 1 |
| 2025 | Low-Complexity Linear Feedback Shift Register Architecture For CRC En/DecodingabstractCyclic redundancy check (CRC) is utilized in digital communication and storage systems for error detection. CRC en/decoding is implemented using linear feedback shift registers (LFSRs). To achieve high throughput, a parallel LFSR can be implemented by registers with a feedback matrix and a pre-processing matrix multiplication. In previous designs, the feedback matrix is decided by look-ahead computations of the LFSR, and its multiplication contributes to a significant portion of the overall complexity. This paper proposes to search over a wide range of powers of the companion matrix describing the LFSR to minimize the gate count of the feedback matrix multiplication. This is enabled by an alternative interpretation of data inputs. Although the achievable data length protected by CRC is reduced by the proposed scheme, it still meets the requirement of IEEE standards. Besides, our scheme does not affect the pre-processing matrix and the input tap of the LFSR can still be shifted to reduce the complexity. For an example case that the parallelism equals the generator polynomial degree, the proposed design can reduce the gate count by 18%-53% and achieve shorter critical path for various CRCs. Yok Jye Tang, Jiaxuan Cai, Xinmiao Zhang 0001 |
ISCAS | 2 |
| 2023 | Low-Complexity Parallel Min-Sum Medium-Density Parity-Check Decoder for McEliece CryptosystemabstractThe McEliece cryptosystem based on medium-density parity-check (MDPC) codes remains a candidate in the fourth round submission of post-quantum cryptography standard. The low-density parity-check (LDPC) decoders used in digital communications have been extensively studied. However, the MDPC codes for the McEliece cryptosystem have much higher column weight and different structure in their parity-check matrices. As a result, simplification techniques for LDPC decoders are not applicable to MDPC decoders. Besides, existing MDPC decoder designs have been focusing on the simplest bit-flipping algorithm, whose performance is inferior compared to that of the Min-sum algorithm. This paper first optimizes the scaled Min-sum algorithm for codes with high column weight to improve the performance with simple scalar multiplications. The overall decoder architecture is re-designed to take into account the sparsity of the parity-check matrix and nontrivial min-sum check node processing. Besides, a flexible message storage scheme is proposed to reduce the worst-case decoding latency of the randomly constructed codes utilized in the McEliece cryptosystem. Then a 2-stage scaling scheme is developed to reduce the long critical path caused by the high column weight and a group size re-balancing scheme is introduced to mitigate the precision loss caused by the 2-stage scaling in parallel decoders. For an example MDPC decoder, the proposed optimized 2-stage scaled Min-sum algorithm leads to orders of magnitude error-correcting performance improvement and 16% higher clock frequency with negligible silicon area overhead compared to unoptimized Min-sum decoders. Jiaxuan Cai, Xinmiao Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | PBCStereo: A Compressed Stereo Network with Pure Binary Convolutional Operations
Jiaxuan Cai, Keqi Fu, Xulong Shi, Zan Li 0005 |
ACCV (3) | 1 |
| 2022 | Towards High Performance and Accurate BNN Inference on FPGA with Structured Fine-Grained PruningabstractAs the extreme case of quantization networks, Binary Neural Networks (BNNs) have received tremendous attention due to many hardware-friendly properties in terms of storage and computation. To reach the limit of compact models, we attempt to combine binarization with pruning techniques, further exploring the redundancy of BNNs. However, coarse-grained pruning methods may cause server accuracy drops, while traditional fine-grained ones induce irregular sparsity hard to be utilized by hardware. In this paper, we propose two advanced fine-grained BNN pruning modules, i.e., structured channel-wise kernel pruning and dynamic spatial pruning, from a joint perspective of algorithm and hardware. The pruned BNN models are trained from scratch and present not only a higher precision but also a high degree of parallelism. Then, we develop an accelerator architecture that can effectively exploit the sparsity caused by our algorithm. Finally, we implement the pruned BNN models on an embedded FPGA (Ultra96v2). The results show that our software and hardware codesign achieves 5.4x inference-speedup than the baseline BNN, with higher resource and energy efficiency compared with prior FPGA implemented BNN works. Keqi Fu, Jiaxuan Cai, Xulong Shi |
ICCAD | 3 |
| 2020 | MABNet: A Lightweight Stereo Network Based on Multibranch Adjustable Bottleneck Module
Jiabin Xing, Jiying Dong, Jiaxuan Cai, Hao Liu 0013 |
ECCV (28) | 4 |