VLDB 2026 Research / reviewers in the wild / expert
Junyao Zhang 0003
dblp:14/10083-3
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0008-3913-7879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix MultiplicationabstractThe rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantization, there are abundant opportunities for results reuse, and thus it can be boosted with lookup tables (LUTs) based acceleration. However, existing LUT-based methods suffer from computation and hardware overheads for LUT construction, and rely solely on bit-serial computation, which is suboptimal for ternary-weight networks. We propose Platinum, a lightweight ASIC accelerator for integer weight mixed-precision matrix multiplication (mpGEMM) using LUTs. Platinum reduces LUT construction overhead via offline-generated construction paths and supports both general bit-serial and optimized ternaryweight execution through adaptive path switching. On BitNet b1.58-3B, Platinum achieves up to $73.6 \times, 4.09 \times$, and $2.15 \times$ speedups over SpikingEyeriss, Prosperity, and 16-thread T-MAC (CPU), respectively, along with energy reductions of $32.4 \times, 3.23 \times$, and $20.9 \times$, all within a $0.96 \mathrm{~mm}^{2}$ chip area. This demonstrates the potential of LUT-based ASICs as efficient, scalable solutions for ultra-low-bit neural networks on edge platforms. Haoxuan Shan, Cong Guo 0003, Chiyue Wei, Junyao Zhang 0003, Hai Li 0001, Yiran Chen 0001 |
ASP-DAC | 5 |
| 2026 | Focus: A Streaming Concentration Architecture for Efficient Vision-Language ModelsabstractVision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inputs lead to significant computational and memory overhead, posing challenges for real-time deployment on hardware accelerators. While prior work attempts to reduce redundancy via token pruning or merging, these methods typically operate at coarse granularity and incur high runtime overhead due to global token-level operations. In this study, we propose Focus, a Streaming Concentration Architecture that efficiently accelerates VLM inference through progressive, fine-grained redundancy elimination. Focus introduces a multilevel concentration paradigm that hierarchically compresses vision-language inputs at three levels: (1) semantic-guided token pruning based on textual prompts, (2) spatial-temporal blocklevel concentration using localized comparisons, and (3) vectorlevel redundancy removal via motion-aware matching. All concentration steps are tightly co-designed with the architecture to support streaming-friendly, on-chip execution. Focus leverages GEMM tiling, convolution-style layout, and cross-modal attention to minimize off-chip access while enabling high throughput. Implemented as a modular unit within a systolic-array accelerator, Focus achieves$2.4 \times$speedup and$3.3 \times$reduction in energy, significantly outperforming state-of-the-art accelerator in both performance and energy efficiency. Full-stack implementation of Focus is open-sourced at https://github.com/dubcyfor3/Focus. Chiyue Wei, Cong Guo 0003, Junyao Zhang 0003, Haoxuan Shan, Qinsi Wang, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001 |
HPCA | 3 |
| 2025 | qGDP: Quantum Legalization and Detailed Placement for Superconducting Quantum ComputersabstractQuantum computers (QCs) are currently limited by qubit numbers. A major challenge in scaling these systems is crosstalk, which arises from unwanted interactions among neighboring components such as qubits and resonators. An inno-vative placement strategy tailored for superconducting QCs can systematically address crosstalk within limited substrate areas. Legalization is a crucial stage in placement process, refining post-global-placement configurations to satisfy design constraints and enhance layout quality. However, existing legalizers are not supported to legalize quantum placements. We aim to address this gap with qGDP, developed to meticulously legalize quantum components by adhering to quantum spatial constraints and reducing resonator crossing to alleviate various crosstalk effects. Our results indicate that qGDP effectively legalizes and fine-tunes the layout, addressing the quantum-specific spatial constraints inherent in various device topologies. By evaluating diverse benchmarks. qGDP consistently outperforms state-of-the-art legalization engines, delivering substantial improvements in fidelity and reducing spatial violation, with average gains of 34.4 x and 16.9 x, respectively. Junyao Zhang 0003, Guanglei Zhou, Jonathan Hao-Cheng Ku, Jiaqi Gu 0002, Hanrui Wang 0002, Hai Li 0001, Yiran Chen 0001 |
DATE | 1 |
| 2025 | AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Tunhou Zhang, Junyao Zhang 0003, Jonathan Hao-Cheng Ku, Yitu Wang, Xiaoxuan Yang 0001, Hai Li 0001, Yiran Chen 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Diffusion-Model-Enhanced Layout Pattern Generation for Sub-3nm DFMabstractModern VLSI layout pattern generation for design for manufacturability (DFM) at sub-3 nm nodes faces two challenges: 1) the rapid evolution of intricate design rules; 2) the scarcity of high-quality, rule-compliant layout data during the development of new process technologies. To address these challenges, we introduce a diffusion-based framework that re-frames complex layout synthesis as a sequence of template-guided inpainting tasks, which significantly reduces training sample requirements for legal pattern generation. This approach leverages the knowledge of a pre-trained image foundation model to generate layout variations that satisfy complex 2D metal interconnect design rule constraints, and introduces a novel template-based denoising scheme to eliminate residual noisy pixels. Through few-shot fine-tuning, our approach uniquely produces legal layouts conforming to a full sign-off rule deck at sub-3nm nodes while delivering superior pattern diversity, offering a production-ready, data-efficient solution for next-generation technology node development. Guanglei Zhou, Chen-Chia Chang, Junyao Zhang 0003, Jingyu Pan, Yiran Chen 0001 |
ICCAD | 3 |
| 2025 | Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-Aware Cache CompressionabstractLarge language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memory and computational demands, especially in resource-constrained environments.Quantization techniques have emerged as a critical solution, reducing data precision to enhance memory and computational efficiency.However, existing methods often suffer from high runtime overheads and potential accuracy degradation.To address these challenges, we propose Ecco, an entropy-based cache compression technique tailored for LLMs.Ecco combines group-wise and nonuniform quantization with pre-defined shared k-means patterns and Huffman coding to exploit the inherent entropy characteristics of LLM cache data.Recognizing the inefficiencies of traditional Huffman coding in terms of parallelism and latency, we introduce a novel parallel Huffman-based decoding process with a multi-stage pipeline design, reducing latency by two orders of magnitude and achieving throughput comparable to GPU L2 caches.Comprehensive evaluations demonstrate that Ecco achieves an up to 2.9× and 1.9× speedup over the state-of-the-art AWQ and SmoothQuant framework, 2.4× over the Olive accelerator, all while increasing memory capacity by nearly 4× and maintaining state-of-the-art LLM accuracy.These results underscore the effectiveness of our Cong Guo 0003, Chiyue Wei, Junyao Zhang 0003, Changchun Zhou 0001, Edward Hanson, Jiaqi Zhang 0002, Xiaoxiao Liu 0001, Hai Li 0001, Yiran Chen 0001 |
ISCA | 4 |
| 2025 | QPlacer: Frequency-Aware Component Placement for Superconducting Quantum ComputersabstractQuantum Computers face a critical limitation in qubit numbers, hindering their progression towards large-scale and fault-tolerant quantum computing.A significant challenge impeding scaling is crosstalk, characterized by unwanted interactions among neighboring components on quantum chips, including qubits, resonators, and substrates.We motivate a general approach to systematically resolving multifaceted crosstalks in a limited substrate area.We propose QPlacer, a frequency-aware electrostatic-based placement framework tailored for superconducting quantum computers, to alleviate crosstalk by isolating these components in spatial and frequency domains alongside compact substrate design.QPlacer commences with a frequency assigner that ensures frequency domain isolation for qubits and resonators.It then incorporates a padding strategy and resonator partitioning for layout flexibility.Central to our approach is the conceptualization of quantum components as charged particles, enabling strategic spatial isolation through a 'frequency repulsive force' concept.Our results demonstrate that QPlacer carefully crafts the physical component layout in mitigating various crosstalk impacts while maintaining a compact substrate size.On various device topologies and NISQ benchmarks, QPlacer improves fidelity by an average of 37.5× and reduces spatial violations (susceptible to crosstalk) by an average of 12.76×, compared to classical placement engines.Regarding area Junyao Zhang 0003, Hanrui Wang 0002, Jiaqi Gu 0002, Reouven Assouly, William D. Oliver, Song Han 0003, Kenneth R. Brown, Hai Li 0001, Yiran Chen 0001 |
ISCA | 1 |
| 2024 | ModSRAM: Algorithm-Hardware Co-Design for Large Number Modular Multiplication in SRAMabstractElliptic curve cryptography (ECC) is widely used in security applications such as public key cryptography (PKC) and zero-knowledge proofs (ZKP). ECC is composed of modular arithmetic, where modular multiplication takes most of the processing time. Computational complexity and memory constraints of ECC limit the performance. Therefore, hardware acceleration on ECC is an active field of research. Processing-in-memory (PIM) is a promising approach to tackle this problem. In this work, we design ModSRAM, the first 8T SRAM PIM architecture to compute large-number modular multiplication efficiently. In addition, we propose R4CSA-LUT, a new algorithm that reduces the cycles for an interleaved algorithm and eliminates carry propagation for addition based on look-up tables (LUT). ModSRAM is co-designed with R4CSA-LUT to support modular multiplication and data reuse in memory with 52% cycle reduction compared to prior works with only 32% area overhead. Jonathan Hao-Cheng Ku, Junyao Zhang 0003, Haoxuan Shan, Saichand Samudrala, Jiawen Wu 0006, Qilin Zheng, Ziru Li, Jeyavijayan Rajendran, Yiran Chen 0001 |
DAC | 2 |
| 2021 | Trust-aware Control for Intelligent Transportation SystemsabstractMany intelligent transportation systems are multiagent systems, i.e., both the traffic participants and the subsystems within the transportation infrastructure can be modeled as interacting agents. The use of AI-based methods to achieve coordination among the different agents systems can provide greater safety over transportation systems containing only human-operated vehicles, and also improve the system efficiency in terms of traffic throughput, sensing range, and enabling collaborative tasks. However, increased autonomy makes the transportation infrastructure vulnerable to compromised vehicular agents or infrastructure. This paper proposes a new framework by embedding the trust authority into transportation infrastructure to systematically quantify the trustworthiness of agents using an epistemic logic known as subjective logic. In this paper, we make the following novel contributions: (i) We propose a framework for using the quantified trustworthiness of agents to enable trust-aware coordination and control. (ii) We demonstrate how to synthesize trust-aware controllers using an approach based on reinforcement learning. (iii) We comprehensively analyze an autonomous intersection management (AIM) case study and develop a trust-aware version called AIM-Trust that leads to lower accident rates in scenarios consisting of a mixture of trusted and untrusted agents. Mingxi Cheng, Junyao Zhang 0003, Shahin Nazarian, Jyotirmoy V. Deshmukh, Paul Bogdan |
IV | 2 |