Jingyun Gu

dblp:231/3371 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0006-9175-8598ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 442.42 TOPS/W RRAM-based Digital Computing-in-Memory Accelerator for BF16×1-bit Vision Transformer
Jingyun Gu, Jiaqi Yang 0009, Jingyao Dong, Qilong Chen, Dingbang Liu, Wei Mao 0002, Hao Yu 0001
ISCAS1
2025 A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM Accelerator
abstract
Transformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2.
Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001
ISLPED4
2025 Abuttable Analog Cell Library and Automatic AMS Layout
abstract
The state of the art analog circuit design applies mainly a full-custom layout methodology. This demands high expertise and heavy manual workload. Additionally, neither can the resulting layout be re-used easily across different designs or different PDKs. Learning from digital standard cells, existing work has proposed stem cells that are abuttable. But stem cells have a fixed area ratio of 2 over same-sized Pcells, limiting its wide application. In this paper we develop a new type of abuttable analog cells (called Acells) for transistors and passive elements. Acells are compatible with digital standard cells and can be abutted in all directions, enabling the use of automatic digital place and route (PnR) engines. We automate Acell generation and show that the average area ratio over same-sized Pcell is 1.49 for 65nm technology and 1.3 for 28nm technology, and is expected to decrease for more advanced technologies. We then use digital PnR to automatically layout a number of analog and mixed-signal (AMS) circuits mainly in 28nm, and show that compared to Pcell-based manual layout, Acell-based layout obtains similar performance and its circuit level layout area is about 2% higher for large scale AMS circuits in our experiments.
Tianjia Zhou, Jingyun Gu, Zexin Ji, Hailang Liang, Zhanfei Chen, Ting-Jung Lin, Na Bai, Zhengping Li, Lei He 0001
ISPD4
2023 Distributed processing of spatiotemporal ocean data: a survey
Xiaoyong Li 0002, Jingyun Gu, Guolong Tan, Wenjing Jiang, Ao Cui, Leiming Shu, Kaijun Ren, Haoyang Zhu, Jedi S. Shang, Zichen Xu 0001
World Wide Web (WWW)2
2021 Cost risk analysis for instance recommendation in a sustainable Cloud-cyber-physical system framework
abstract
Abstract Cloud markets advocate powerful instances to take computation over from the cyber‐physical system (CPS). Combining the Cloud and CPS layer, the whole Cloud–CPS framework is designed to achieve both accurate data sensing and fast data analysis. While most researchers trust the computation side, and focus on the actuator in the physical space to ensure the service‐level objectives, SLO, that is, deadline misses, cloud can be a threat to the service sustainability as instance may fail, especially when one tries to make a cost‐effective design. Specifically, users must bear the risk of instance failure. These risks can cause the entire cyber‐physical system to collapse. Our work tackles the cloud aspect of the sustainability challenge from the cloud side in a cloud–CPS framework. We have studied the instance selection problem for the CPS systems, and propose a Cost‐Risk Analysis for Instance Recommendation, or CRAIR, to support a sustainable Cloud–CPS framework. We have adopted the classic risk analysis process from the portfolio management in hedge financial market, combining with the system modeling for the CPS instance selection, as an optimization problem. To solve this problem, we formulate it as a multi‐armed bandit problem and solve it with our upper confidence bound bandit algorithm together, our CRAIR can provide an online risk analysis to maximize the profit with a comparative ratio of O(1+ ). We have evaluated CRAIR based on simulations using real‐world Google and Alibaba workloads and cloud market numbers. The results show that, compared to traditional approaches, our approach provides the best tradeoff between SLOs and costs. All users achieve their SLOs goals while minimizing their average expenses by 34.6%. By using CRAIR for instance selection, the CPS service can maximize its benefit under a controlled risk.
Wenjing Jiang, Zichen Xu 0001, Cuiying Gao, Jingyun Gu, Yuhao Wang 0001
Softw. Pract. Exp.4
2018 RISC: Risk Assessment of Instance Selection in Cloud Markets
Jingyun Gu, Zichen Xu 0001, Cuiying Gao
ICA3PP (1)1