Jinpeng Ye

dblp:375/4955 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0008-7451-3200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CUTE-XS: Bottleneck Analysis and Collaborative Optimization for Integrating CUTE Into XiangShan
Junyu Yue, Chongxi Wang, Jinpeng Ye, Jianan Xie, Longbing Zhang, Fuxin Zhang
APPT3
2026 SPSRL: Open-vocabulary semantic segmentation with spatial prior and semantic relation learning
Mingzhu Ping, Jinpeng Ye, Sijia Cui, Pengpeng Xu
Neurocomputing2
2026 FU-Mamba: A frequency-enhanced dynamic scanning framework for oralscan image segmentation
Xinxin Zhao, Jinpeng Ye, Liqin Wu, Mahmoud Hassaballah, Karen Egiazarian, Aura Conci, Victor Hugo C. de Albuquerque, Abdulkadir Sengür, Leszek Rutkowski
Neurocomputing2
2025 LitTLS: Lightweight Thread-Level Speculation on Little Cores
abstract
Thread-Level Speculation (TLS) utilizes speculative parallelization to accelerate hard-to-parallelize serial codes on multi-cores. As the heterogeneous multi-core architecture is becoming ubiquitous, it presents an opportunity for TLS to reorganize little cores for the acceleration of these serial codes instead of a big core with similar or more area and power. However, previous TLS designs significantly suffer from extended hardware overhead and costly speculative forwarding. We present LitTLS, a lightweight TLS design with versioning caches to eliminate significant extended hardware overhead by storing versions in caches without speculative write buffers and memory undo-logs. Additionally, LitTLS introduces the Speculative Address Table, a novel component to accelerate speculative forwarding with a central structure to trace memory dependencies. Evaluations on four little cores show that LitTLS achieves an average performance speedup of 2.87× compared to a little core, outperforming a big core by 94% with similar area and less power. The extended area size is only 0.07 mm 2 , and the maximum increase in dynamic power consumption is limited to 0.3%, compared to four little cores.
Xin Cheng 0021, Jinpeng Ye, Haoyu Deng 0001
ACM Trans. Archit. Code Optim.2
2025 ETBench: Characterizing Hybrid Vision Transformer Workloads Across Edge Devices
abstract
Lightweight Convolution and Vision Transformer hybrid models have increasingly dominated the frontiers of deep learning (DL) on edge devices; however, to the best of our knowledge, no prior work has provided comprehensive evaluation on hybrid models’ performance and analyzed their characteristics by diving deep into the edge ecosystem with diversified modern DL inference engines and heterogeneous hardware. This paper proposes a comprehensive open-source benchmark suite,ETBench, to allow power-efficiency, performance and accuracy assessment for state-of-the-art (SOTA) hybrid models across 11 most widely-used DL engines deployed on diverse edge devices. After building ETBench that satisfies 6 design requirements proposed in our work, we conduct extensive experiments on 14 devices including 19 CPUs, 11 GPUs and 5 NPUs, and obtain benchmark results from all deployment scenarios (combinations of models, quantization formats, software engines, and hardware platforms). Valuable observations and insightful implications are finally summarized. For example, within current DL engines, the INT8 quantization is significantly underperformed in terms of accuracy and speed against FP16 for hybrid models. Overall, ETBench serves as a collaborative platform that assists model architects in better evaluating their models and makes it possible for future co-optimizations of DL engines and hardware accelerators.
Yingkun Zhou, Zhengshuyuan Tian, Jinpeng Ye, Chenji Han, Fuxin Zhang
IEEE Trans. Computers5
2024 CUTE: A scalable CPU-centric and Ultra-utilized Tensor Engine for convolutions
Jinpeng Ye, Fuxin Zhang
J. Syst. Archit.2