VLDB 2026 Research / reviewers in the wild / expert
Jiaao Ma
dblp:350/2153
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashTFHE: A Scalable Architecture for Efficient Multi-Bit Fully Homomorphic Encryption
Jiaao Ma, Ceyu Xu, Lisa Wu Wills |
ISCA | 1 |
| 2025 | COCOSSim: A Cycle-Accurate Simulator for Heterogeneous Systolic Array ArchitecturesabstractPerformance modeling is an essential tool for enabling cost-effective and efficient exploration of architectural and microarchitectural design decisions for hardware development. However, existing simulators face limitations in architectural flexibility, memory hierarchy modeling, and support for modern accelerators that cater to the computational demands of contemporary machine learning models, where non-linear operations like softmax and other vectorized activations have become increasingly important. To address these gaps, we present COCOSSim, a cycleaccurate performance simulator designed for evaluating architectural and microarchitectural modifications in heterogeneous systolic array-based accelerators. COCOSSim supports a wide range of neural network inference workloads, integrating systolic arrays with vector units for non-linear operations and DRAMSim3 for memory modeling. By offering a PyTorch frontend for seamless model integration, flexibility in scheduling strategies, and the ability to make architectural and microarchitectural modifications, COCOSSim enables detailed performance analysis and design exploration. When validated against the Google TPU v3, COCOSSim achieves an average error rate of 13 %. Additionally, COCOSSim is highly scalable and offers simulation speeds an order of magnitude faster than prior work while providing significantly greater flexibility and enhanced modeling capabilities. We present two case studies demonstrating how COCOSSim can be used to evaluate architectural modifications to modern heterogeneous accelerators: model parallelism and operation fusion. Mansi Choudhary, Chris Kjellqvist, Jiaao Ma, Lisa Wu Wills |
ISPASS | 3 |
| 2023 | PyTFHE: An End-to-End Compilation and Execution Framework for Fully Homomorphic Encryption ApplicationsabstractFully Homomorphic Encryption (FHE) is a powerful cryptographic scheme that enables computation on encrypted data, which allows clients to offload computation to an untrusted third party without compromising data privacy. However, FHE has not yet been widely adopted due to its enormous computational overhead. Further, cryptographic software development requires specialized expertise and presents a significant challenge when applying FHE to a broad range of applications. We present PyTFHE, a framework that tackles these difficulties by enabling highly productive FHE application development and orders of magnitude more efficient FHE application execution. PyTFHE is built on top of the TFHE (Fast Fully Homomorphic Encryption over the Torus) scheme, which is an FHE scheme that supports gate-level evaluation and arbitrary depth of boolean circuits. PyTFHE is designed in a TFHE-specific approach, allowing state-of-the-art optimizations for TFHE applications. Specifically, PyTFHE features ChiselTorch, the first compiler that allows easy generations of privacy-preserving deep neural network models with PyTorch-compatible APIs. PyTFHE is also the first FHE framework that employs a powerful backend enabling efficient execution of TFHE applications on distributed CPU systems and high-performance GPUs. We demonstrate the effectiveness of PyTFHE by benchmarking the framework using VIP-Bench and implementing privacy-preserving deep neural networks and evaluating their performance on various systems. We compare the performance of our generated TFHE program execution with three existing frameworks, Google Transpiler, Cingulata, and E3. We show that PyTFHE achieves one to two orders of magnitude performance advantage. Jiaao Ma, Ceyu Xu, Lisa Wu Wills |
ISPASS | 1 |