EDBT 2026 Demo / reviewers in the wild / expert
Jieming Yin
dblp:82/9927
· DBLP profile ↗
37ranked-venue papers
6as first author
25since 2021 · last 2026
0009-0008-2878-1853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 6 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-Based Islandization
Jieming Yin, ZhiLei Chai, Jiliang Zhang 0011, Herman Lam |
ISCA | 2 |
| 2026 | From pseudo- to non-correspondences: Robust point cloud registration via thickness-guided self-correction
Yifei Tian, Jieming Yin |
Comput. Graph. | 3 |
| 2026 | A review of cross-source point cloud fusion: Registration and repair
Yifei Tian, Jieming Yin |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | MOA: Efficient Scene-Aware Multi-Object Arrangement in VRabstract3D multi-object arrangement is a fundamental task in VR that relies on accurate and natural initial selection alongside rapid and convenient subsequent manipulation to ensure high efficiency. However, existing methods fail to support efficient multi-object arrangement in highly occluded scenes with densely packed candidate objects through controller-free natural interactions. In this article, we propose an efficient, scene-aware multi-object arrangement method (MOA) designed for fast, precise, and convenient object arrangement. First, MOA introduces an importance-driven multi-object initial selection algorithm that assigns higher spatiotemporally correlated object importance (IMP) to target objects, establishing a natural multi-object initial selection mode that enables quick and accurate selection of high-IMP objects. Subsequently, it presents an auxiliary-structure-guided multi-object manipulation algorithm that constructs an auxiliary manipulation structure to assist subsequent multi-object manipulation, alongside a multi-modal interaction mode that facilitates swift and natural manipulation. Compared to state-of-the-art controller-free and controller-based methods, MOA significantly improves task performance, reduces task load, and enhances convenience in complex multi-object arrangement scenes involving hundreds of highly occluded objects need to be arranged. Xuehuai Shi, Yuhan Duan, Ziteng Wang 0002, Jian Wu 0033, Zhiwen Shao, Jieming Yin, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | FineRR-ZNS: Enabling Fine-Granularity Read Refreshing for ZNS SSDsabstractZoned namespace (ZNS) SSDs are emerging storage devices offering low cost, high performance, and software definability. By adopting host-managed zone-based sequential programming, ZNS SSDs effectively eliminate the space overhead associated with on-board DRAM memory and garbage collection. However, while background read refreshing serves as a data protection mechanism in conventional block-interface SSDs, state-of-the-art ZNS SSDs lack read refreshing functionality to guarantee data reliability. Moreover, implementing zonelevel read refreshing in ZNS SSDs incurs significant overhead due to the large volume of valid data movements in a zone, leading to degraded I/O performance. To efficiently enable read refreshing for ZNS SSDs, this paper proposes FineRR-ZNS, a fine-granularity read refreshing mechanism for ZNS SSDs. FineRR-ZNS employs a host-controlled fine-granularity read refreshing scheme that selectively determines block-level read refreshing via metadata remapping. A zone reconstruction method is also designed to retrieve remapped data forming complete data during zone-level RR. Specifically, the remapped data after zone reconstruction are still available and prioritized for read access until their respective blocks need the next RR. Evaluation results show that FineRR-ZNS significantly enhances read refreshing efficiency and I/O throughput compared to zone-level read refreshing implemented in the state-of-the-art ZenFS file system. Jun Li 0062, Zhibing Sha, Fan Yang 0110, Xiaofei Xu 0002, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DAC | 6 |
| 2025 | CoupledCB: Eliminating Wasted Pages in Copyback-based Garbage Collection for SSDsabstractThe management of garbage collection poses significant challenges in high-density NAND flash-based SSDs. The introduction of the copyback command aims to expedite the migration of valid data. However, its odd/even constraint causes wasted pages during migrations, limiting the efficiency of garbage collection. Additionally, while full-sequence programming en-hances write performance in high-density SSDs, it increases write granularity and exacerbates the issue of wasted pages. To address the problem of wasted pages, we propose a novel method called CoupledCB, which utilizes coupled blocks to fill up the wasted space in copyback-based garbage collection. By taking into account the access characteristics of the candidate coupled blocks and workloads, we develop a coupled block selection model assisted by logistic regression. Experimental results show that our proposal significantly enhances garbage collection efficiency and 1/O performance compared to state-of-the-art schemes. Jun Li 0062, Xiaofei Xu 0002, Zhibing Sha, Xiaobai Chen, Jieming Yin, Jianwei Liao 0001 |
DATE | 5 |
| 2025 | CAMC: A Multi-Chiplet Accelerator With Heterogeneous Memory-Based Computing Architecture For DNN TrainingabstractDeep Neural Networks (DNNs) are extensively utilized in various fields due to their remarkable performance. However, as DNN models increase in complexity and size, the training process incurs substantial data transfer costs between computation and storage. The slowdown of Moore’s Law further challenges the integration of additional resources on a single chip, making it difficult to improve storage capacity and reduce off-chip data transfers. To address these challenges, we propose CAMC, a multi-chiplet DNN accelerator with a heterogeneous memory computing architecture. CAMC integrates SRAM-based in-memory computing with TSV-stacked DRAM-based near-memory computing. In addition, an efficient mapping strategy was developed to optimize resource utilization and performance. The experimental results demonstrate that CAMC enhances energy efficiency by 11.84 times and reduces data transfer costs by 8.78 times compared to the baseline design. Xiaobai Chen, Jiacheng Mei, Yifei Tian, Jieming Yin, Fu Xiao 0001 |
ISCAS | 6 |
| 2025 | DPAcc: An FPGA-based Differential Privacy Acceleration FrameworkabstractIn the data-driven era, privacy protection has become a critical concern. Differential privacy is an effective technique that incorporates random noise during data processing to ensure that alterations to individual data points do not significantly affect overall outputs. However, the additional operations required by differential privacy can result in prolonged training times and degraded model performance. This work proposes DPAcc, an FPGA-based acceleration framework for differential privacy that utilizes hardware implementation to decrease training time. The designed FPGA module efficiently executes clipping and noise addition operations, significantly reducing the overhead compared to standard training. Experimental results demonstrate that DPAcc improves training efficiency across multiple models, achieving up to 2× speedup compared to standard differential privacy training methods. Ao Dong, Pengyang Li, Yifei Tian, Xiaobai Chen, Jieming Yin |
ISCAS | 6 |
| 2025 | FlexAcc: Accelerating Batch Normalization through GPU-FPGA IntegrationabstractConvolutional neural networks are fundamental to deep learning, especially in computer vision. However, their computational demands, particularly during batch normalization, create significant inefficiencies due to excessive data movement between memory and processing units. To address this, we propose FlexAcc, a novel architecture that integrates GPUs and FPGAs to offload BN computations to FPGAs, reducing data movement and improving hardware utilization. FlexAcc accelerates end-to-end training performance by up to 1.1× across various models. This approach bridges the performance gap between convolutional and non-convolutional layers, advancing deep learning model deployment. Haishuai Zhang, Pengyang Li, Xuehuai Shi, Xiaobai Chen, Jieming Yin |
ISCAS | 6 |
| 2025 | MMG: Manipulation-Aware Holistic Human Motion Generation from Sparse Tracking SignalsabstractGenerating realistic avatar motion via sparse tracking signals through VR devices is essential for enhancing the immersive user experience. Human-object manipulation behaviors not only affect hand motion but also significantly impact body motion. However, existing motion generation methods for human-object interactions overlook the coordinated coupling between body and hand motions during manipulations. Due to the diversity and complexity of holistic motion (body and hand motions simultaneously) in the latent motion space, generating physically plausible and temporally consistent holistic motion in real time, via the joint constraints imposed by sparse tracking signals and manipulation content, is a major challenge in the human motion generation task. We propose the manipulation-aware holistic human motion generation method (MMG) to help resolve this issue. In MMG, first, we construct a manipulation-aware holistic human motion generation framework that serially compresses the latent motion space distribution of the body and hand to generate realistic holistic human motion with object manipulation enabled. Second, to enhance the impact of object manipulation on holistic motion generation, MMG designs a novel object manipulation representation to extract effective manipulation features. Third, MMG is trained by an elaborate progressive manipulation-guided training algorithm to improve motion generation robustness and inference performance. Compared to state-of-the-art methods, MMG achieves up to a 39% improvement in the generated holistic motion quality with a 3.55 × speedup in generation performance. In manipulation-enabled scenes, MMG generates holistic motion in real time ($\geq 24 f p s$). Compared to the state-of-the-art methods, its perceived quality is significantly improved, and the task performance of holistic motion-required VR manipulation is high-significantly improved. This paper's code is at https://github.com/XRZ-BUAA/MMG. Xuehuai Shi, Renzhi Xiao, Yilun Sheng, Xiaobai Chen, Jieming Yin, Qingshan Liu 0001 |
ISMAR | 7 |
| 2025 | LEGOSim: A Unified Parallel Simulation Framework for Multi-chiplet Heterogeneous Integration
Tiantian Lin, Xiaohang Wang 0001, Ling Wang 0005, Zhulin Zheng, Yingtao Jiang, Amit Kumar Singh 0002, Jieming Yin, Sihai Qiu, Mingzhe Zhang 0005, Kui Ren 0001 |
MICRO | 8 |
| 2025 | Audio-Visual Aware Foveated RenderingabstractWith the increasing complexity of geometry and rendering effects in virtual reality (VR) scenes, existing foveated rendering methods for VR head-mounted displays (HMDs) struggle to meet users' demands for VR scene rendering with high frame rates ($\geq 60fps$≥60fps for rendering binocular foveated images in VR scenes containing over 50 m triangles). Current research validates that auditory content affects the perception of the human visual system (HVS). However, existing foveated rendering methods primarily model the HVS's eccentricity-dependent visual perception ability on the visual content in VR while ignoring the impact of auditory content on the HVS's visual perception. In this article, we introduce an auditory-content-based perceived rendering quality analysis to quantify the impact of visual perception under different auditory conditions in foveated rendering. Based on the analysis results, we propose an audio-visual aware foveated rendering method (AvFR). AvFR first constructs an audio-visual feature-driven perception model that predicts the eccentricity-based visual perception in real time by combining the scene's audio-visual content, and then proposes a foveated rendering cost optimization algorithm to adaptively control the shading rate of different regions with the guidance of the perception model. In complex scenes with visual and auditory content containing over 1.17 m triangles, AvFR renders high-quality binocular foveated images at an average frame rate of 116$fps$fps. The results of the main user study and performance evaluation validate that AvFR achieves significant performance improvement (up to 1.4× speedup) without lowering the perceived visual quality compared with the state-of-the-art VR-HMD foveated rendering method. Xuehuai Shi, Jian Wu 0033, Jieming Yin, Xiaobai Chen, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Rubick: A Unified Infrastructure for Analyzing, Exploring, and Implementing Spatial Architectures via Dataflow DecompositionabstractThe fast-growing tensor applications expose tremendous dataflow alternatives when implemented on spatial architectures that feature large PE arrays and abundant interconnection resources. Prior works develop various notations and performance models for dataflows. Though these notations are very useful for understanding the reuse, bandwidth, and performance of dataflows, they do not define the underlying hardware implementation. Due to the semantic gap, analysis based on these notations cannot capture the detailed architectural features between different dataflows, leading to inefficient design space exploration and suboptimal designs. To address these issues, we propose Rubick, a unified infrastructure for analyzing, exploring, and implementing spatial architectures. The main innovation of Rubick is it decomposes the dataflow into two low-level intermediate representations: access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. These two representations allow us to infer the hardware implementation details such as PE interconnection and memory structure, which are amenable for structural analysis and systematic exploration. Based on this decomposition analysis, Rubick provides opportunities for micro-architecture optimization and efficient design space exploration. Our experiments demonstrate that Rubick can reduce 82.4% of wire resources with only a 2.7% latency increase by optimizing access entry IR, and achieve 70.8% memory overhead reduction by optimizing data layout IR. Rubick also accelerates the DSE time of dataflows by up to 1.1×105X, saving the time from several days to minutes. The source code of Rubick is publically available on (https://link-omitted-for-blind-review). Liqiang Lu, Zizhang Luo, Size Zheng 0001, Jieming Yin, Jason Cong, Yun Liang 0001, Jianwei Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Rubick: A Synthesis Framework for Spatial Architectures via Dataflow DecompositionabstractDataflows are critical for spatial architectures designed for tensor applications. Prior works develop various notations and hardware generation frameworks for dataflows. However, due to the semantic gap between notations and low-level details, analysis based on these notations cannot capture the detailed architectural features between different dataflows, so these works failed to provide architectural optimization and efficient design space exploration (DSE) at the same time.We propose Rubick, a synthesis framework for spatial architecture. Rubick decomposes the dataflow into two low-level intermediate representations including access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. Based on this decomposition, Rubick provides efficient DSE and generates optimized hardware. Experiments show that the DSE time is accelerated by up to 1.1×105X and performance on FPGA is improved by 13%. Zizhang Luo, Liqiang Lu, Size Zheng 0001, Jieming Yin, Jason Cong, Jianwei Yin, Yun Liang 0001 |
DAC | 4 |
| 2023 | Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingabstractMulti-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average. Bingyao Li 0001, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang 0002, Xulong Tang |
HPCA | 2 |
| 2023 | Monad: Towards Cost-Effective Specialization for Chiplet-Based Spatial AcceleratorsabstractAdvanced packaging offers a new design paradigm in the post-Moore era, where many small chiplets can be assembled into a large system. Based on heterogeneous integration, a chiplet-based accelerator can be highly specialized for a specific workload, demonstrating extreme efficiency and cost reduction. To fully leverage this potential, it is critical to explore both the architectural design space for individual chiplets and different integration options to assemble these chiplets, which have yet to be fully exploited by existing proposals. This paper proposes Monad, a cost-aware specialization approach for chiplet-based spatial accelerators that explores the tradeoffs between PPA and fabrication costs. To evaluate a specialized system, we introduce a modeling framework considering the non-uniformity in dataflow, pipelining, and communications when executing multiple tensor workloads on different chiplets. We propose to combine the architecture and integration design space by uniformly encoding the design aspects for both spaces and exploring them with a systematic ML-based approach. The experiments demonstrate that Monad can achieve an average of 16% and 30% EDP reduction compared with the state-of-the-art chiplet-based accelerators, Simba and NN-Baton, respectively. Xiaochen Hao, Zijian Ding, Jieming Yin, Yuan Wang 0001, Yun Liang 0001 |
ICCAD | 3 |
| 2023 | COLA: Orchestrating Error Coding and Learning for Robust Neural Network Inference Against Hardware DefectsabstractError correcting output codes (ECOCs) have been proposed to improve the robustness of deep neural networks (DNNs) against hardware defects of DNN hardware accelerators. Unfortunately, existing efforts suffer from drawbacks that would greatly impact their practicality: 1) robust accuracy (with defects) improvement at the cost of degraded clean accuracy (without defects); 2) no guarantee on better robust or clean accuracy using stronger ECOCs. In this paper, we first shed light on the connection between these drawbacks and error correlation, and then propose a novel comprehensive error decorrelation framework, namely COLA. Specifically, we propose to reduce inner layer feature error correlation by 1) adopting a separated architecture, where the last portions of the paths to all output nodes are separated, and 2) orthogonalizing weights in common DNN layers so that the intermediate features are orthogonal with each other. We also propose a regularization technique based on total correlation to mitigate overall error correlation at the outputs. The effectiveness of COLA is first analyzed theoretically, and then evaluated experimentally, e.g. up to 6.7% clean accuracy improvement compared with the original DNNs and up to 40% robust accuracy improvement compared to the state-of-the-art ECOC-enhanced DNNs. Anlan Yu, Ning Lyu, Jieming Yin, Zhiyuan Yan 0001, Wujie Wen |
ICML | 3 |
| 2023 | CompoundEye: A 0.24-4.17 TOPS Scalable Multi-Node DNN Processor for Image RecognitionabstractThis paper proposes a scalable DNN processor that can be flexibly reconfigured to maximize inference efficiency on a wide range of DNN models. The processor consists of 18 computing nodes with various precision modes support. To improve the computation throughput, we propose a sub-image parallelization strategy, where the original input image is divided into multiple sub-images and computed on multiple nodes in parallel. In addition, the cross-layer pipeline is implemented to improve resource utilization. The proposed processor is implemented in 28nm CMOS technology and achieves a peak performance of 4.17 TOPS and an energy efficiency of 2.08 TOPS/W. Xiaobai Chen, Qiurun Hu, Fu Xiao 0001, Jieming Yin |
ISCAS | 4 |
| 2023 | QuCT: A Framework for Analyzing Quantum Circuit by Extracting Contextual and Topological FeaturesabstractIn the current Noisy Intermediate-Scale Quantum era, quantum circuit analysis is an essential technique for designing high-performance quantum programs. Current analysis methods exhibit either accuracy limitations or high computational complexity for obtaining precise results. To reduce this tradeoff, we propose QuCT, a unified framework for extracting, analyzing, and optimizing quantum circuits. The main innovation of QuCT is to vectorize each gate with each element, quantitatively describing the degree of the interaction with neighboring gates. Extending from the vectorization model, we propose two representative downstream models for fidelity prediction and unitary decomposition. The fidelity prediction model performs a linear transformation on all gate vectors and aggregates the results to estimate the overall circuit fidelity. By identifying critical weights in the transformation matrix, we propose two optimizations to improve the circuit fidelity. In the unitary decomposition model, we significantly reduce the search space by bridging the gap between unitary and circuit via gate vectors. Experiments show that QuCT improves the accuracy of fidelity prediction by 4.2 × on 5-qubit and 18-qubit quantum devices and achieves 2.5 × fidelity improvement compared to existing quantum compilers [19, 55]. In unitary decomposition, QuCT achieves 46.3 × speedup for 5-qubit unitary and more than hundreds of speedup for 8-qubit unitary, compared to the state-of-the-art method [87]. Siwei Tan, Congliang Lang, Shudi Wang, Xinghui Jia, Tingting Li 0004, Jieming Yin, Yongheng Shang, Andre Python, Liqiang Lu, Jianwei Yin |
MICRO | 8 |
| 2023 | NeuroPots: Realtime Proactive Defense against Bit-Flip Attacks in Neural Networks
Qi Liu 0017, Jieming Yin, Wujie Wen, Chengmo Yang, Shi Sha |
USENIX Security Symposium | 2 |
| 2022 | CryptoGCN: Fast and Scalable Homomorphically Encrypted Graph Convolutional Network InferenceabstractRecently cloud-based graph convolutional network (GCN) has demonstrated great success and potential in many privacy-sensitive applications such as personal healthcare and financial systems. Despite its high inference accuracy and performance on the cloud, maintaining data privacy in GCN inference, which is of paramount importance to these practical applications, remains largely unexplored. In this paper, we take an initial attempt towards this and develop CryptoGCN--a homomorphic encryption (HE) based GCN inference framework. A key to the success of our approach is to reduce the tremendous computational overhead for HE operations, which can be orders of magnitude higher than its counterparts in the plaintext space. To this end, we develop a solution that can effectively take advantage of the sparsity of matrix operations in GCN inference to significantly reduce the encrypted computational overhead. Specifically, we propose a novel Adjacency Matrix-Aware (AMA) data formatting method along with the AMA assisted patterned sparse matrix partitioning, to exploit the complex graph structure and perform efficient matrix-matrix multiplication in HE computation. In this way, the number of HE operations can be significantly reduced. We also develop a co-optimization framework that can explore the trade-offs among the accuracy, security level, and computational overhead by judicious pruning and polynomial approximation of activation modules in GCNs. Based on the NTU-XVIEW skeleton joint dataset, i.e., the largest dataset evaluated homomorphically by far as we are aware of, our experimental results demonstrate that CryptoGCN outperforms state-of-the-art solutions in terms of the latency and number of homomorphic operations, i.e., achieving as much as a 3.10$\times$ speedup on latency and reduces the total Homomorphic Operation Count (HOC) by 77.4\% with a small accuracy loss of 1-1.5$\%$. Our code is publicly available at https://github.com/ranran0523/CryptoGCN. Quan Gang, Jieming Yin, Wujie Wen |
NeurIPS | 4 |
| 2021 | Distilling Arbitration Logic from Traces using Machine Learning: A Case Study on NoCabstractArbitration logic is extensively used in modern computer architectures to dynamically determine how shared hardware resources are allocated or accessed. Recent work has shown that machine learning techniques can learn non-obvious yet effective arbitration policies, which in simulation demonstrate superior performance over human-designed heuristics. However, existing methods based on deep learning are too expensive to be directly implemented as an arbitration unit in hardware. While some prior efforts managed to manually analyze and reduce a deep learning model into relatively small circuits in certain cases, such ad hoc and labor-intensive approaches cannot easily generalize. In this work, we propose a new methodology to automatically “distill” the arbitration logic from simulation traces. Starting by training a deep learning model, we leverage tree-based models as a bridge to convert the more complex model to a compact logic implementation. This paper presents a case study of the proposed methodology on a network-on-chip port arbitration task. Compared with an array of combinational multipliers that exactly computes the neural network output, our arbitration logic achieves up to 282x area reduction without significant performance degradation. Under the training traffic, our arbitration logic achieves up to 64x reduction in average packet latency and up to 5% increase in network throughput over the FIFO arbitration policy. The distilled arbitration policy is also able to generalize to different injection rates and traffic patterns. Hanyu Wang 0005, Jieming Yin, Zhiru Zhang |
DAC | 3 |
| 2021 | Designing a Cost-Effective Cache Replacement Policy using Machine LearningabstractExtensive research has been carried out to improve cache replacement policies, yet designing an efficient cache replacement policy that incurs low hardware overhead remains a challenging and time-consuming task. Given the surging interest in applying machine learning (ML) to challenging computer architecture design problems, we use ML as an offline tool to design a cost-effective cache replacement policy. We demonstrate that ML is capable of guiding and expediting the generation of a cache replacement policy that is competitive with state-of-the-art hand-crafted policies. In this work, we use Reinforcement Learning (RL) to learn a cache replacement policy. After analyzing the learned model, we are able to focus on a few critical features that might impact system performance. Using the insights provided by RL, we successfully derive a new cache replacement policy – Reinforcement Learned Replacement (RLR). Compared to the state-of-the-art policies, RLR has low hardware overhead, and it can be implemented without needing to modify the processor’s control and data path to propagate information such as program counter. On average, RLR improves single-core and four-core system performance by 3.25% and 4.86% over LRU, with an overhead of 16.75KB for 2MB last-level cache (LLC) and 67KB for 8MB LLC. Subhash Sethumurugan, Jieming Yin, John Sartori |
HPCA | 2 |
| 2021 | TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric NotationabstractAccelerating tensor applications on spatial architectures provides high performance and energy-efficiency, but requires accurate performance models for evaluating various dataflow alternatives. Such modeling relies on the notation of tensor dataflow and the formulation of performance metrics. Recent proposed compute-centric and data-centric notations describe the dataflow using imperative directives. However, these two notations are less expressive and thus lead to limited optimization opportunities and inaccurate performance models.In this paper, we propose a framework TENET that models hardware dataflow of tensor applications. We start by introducing a relation-centric notation, which formally describes the hardware dataflow for tensor computation. The relation-centric notation specifies the hardware dataflow, PE interconnection, and data assignment in a uniform manner using relations. The relation-centric notation is more expressive than the compute-centric and data-centric notations by using more sophisticated affine transformations. Another advantage of relation-centric notation is that it inherently supports accurate metrics estimation, including data reuse, bandwidth, latency, and energy. TENET computes each performance metric by counting the relations using integer set structures and operators. Overall, TENET achieves 37.4% and 51.4% latency reduction for CONV and GEMM kernels compared with the state-of-the-art data-centric notation by identifying more sophisticated hardware dataflows. Liqiang Lu, Naiqing Guan, Yuyue Wang 0001, Liancheng Jia, Zizhang Luo, Jieming Yin, Jason Cong, Yun Liang 0001 |
ISCA | 6 |
| 2021 | Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignabstractIn recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively. Bingyao Li 0001, Jieming Yin, Youtao Zhang, Xulong Tang |
MICRO | 2 |
| 2020 | Kite: A Family of Heterogeneous Interposer Topologies Enabled via Accurate Interconnect ModelingabstractRecent advances in die-stacking and 2.5D chip integration technologies introduce in-package network heterogeneities that can complicate the interconnect design. Integrating chiplets over a silicon interposer offers new opportunities of optimizing interposer topologies. However, limited by the capability of existing network-on-chip (NoC) simulators, the full potential of the interposer-based NoCs has not been exploited. In this paper, we address the shortfalls of prior NoC designs and present a new family of chiplet topologies called Kite. Kite topologies better utilize the diverse networking and frequency domains existing in new interposer systems and outperform the prior chiplet topology proposals. Kite decreased synthetic traffic latency by 7% and improved the maximum throughput by 17% on average versus Double Butterfly and Butter Donut, two previous proposals developed using less accurate modeling. Srikant Bharadwaj, Jieming Yin, Bradford M. Beckmann, Tushar Krishna |
DAC | 2 |
| 2020 | Experiences with ML-Driven Design: A NoC Case StudyabstractThere has been a lot of recent interest in applying machine learning (ML) to the design of systems, which purports to aid human experts in extracting new insights leading to better systems. In this work, we share our experiences with applying ML to improve one aspect of networks-on-chips (NoC) to uncover new ideas and approaches, which eventually led us to a new arbitration scheme that is effective for NoCs under heavy contention. However, a significant amount of human effort and creativity was still needed to optimize just one aspect (arbitration) of what is only one component (the NoC) of the overall processor. This leads us to conclude that much work (and opportunity!) remains to be done in the area of ML-driven architecture design. Jieming Yin, Subhash Sethumurugan, Yasuko Eckert, Chintan Patel, Alan Smith 0003, Eric Morton, Mark Oskin, Natalie D. Enright Jerger, Gabriel H. Loh |
HPCA | 1 |
| 2020 | In-Network Memory Access Ordering for Heterogeneous Multicore SystemsabstractIn heterogeneous multicore systems, implementing a programmer-friendly memory consistency model while maximizing memory-level parallelism is challenging. Ideally, memory accesses can be performed out of order as long as program order is not violated. But enforcing memory access order at the end-point (e.g., a core) prohibits a number of architecture optimizations and limits memory-level parallelism. In this work, we explore the opportunity of preserving memory access order inside the on-chip interconnection network. We propose a hybrid switching networks-on-chip (NoC) attached with a light-weight token ring network to guarantee global memory access order. The hybrid switching NoC that supports both packet and circuit switching serves as the underlying communication infrastructure, while the token ring network is used to preserve memory order among multiple ordering points. Our proposed design enables strong memory consistency models and deterministic program execution, with negligible performance overhead compared to an un-ordered packet switching network. Jieming Yin, Antonia Zhai |
NOCS | 1 |
| 2019 | Northup: Divide-and-Conquer Programming in Systems with Heterogeneous Memories and ProcessorsabstractIn recent years we have seen rapid development in both frontiers of emerging memory technologies and accelerator architectures. Future memory systems are becoming deeper and more heterogeneous. Adopting NVM and die-stacked DRAM on each HPC node is a new trend of development. On the other hand, GPUs and many-core processors have been widely deployed in today's supercomputers. However, software for programming and managing a system that consists of heterogeneous memories and processors is still in its very early stage of development. How to exploit such deep memory hierarchy and heterogeneous processors with minimal programming effort is an important issue to address. In this paper, we propose Northup, a programming and runtime framework, using a divide-and-conquer approach to map an application efficiently to heterogeneous systems. The proposed solution presents a portable layer that abstracts the system architecture, providing flexibility to support easy integration of new memories and processor nodes. We show that Northup outof-core execution with SSD is only an average of 17% slower than in-memory processing for the evaluated applications. Shuai Che, Jieming Yin |
IPDPS | 2 |
| 2019 | An orchestrated NoC prioritization mechanism for heterogeneous CPU-GPU systems
Xiangwei Cai, Jieming Yin, Pingqiang Zhou |
Integr. | 2 |
| 2018 | Lost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediate Language LevelabstractModern GPU frameworks use a two-phase compilation approach. Kernels written in a high-level language are initially compiled to an implementation agnostic intermediate language (IL), then finalized to the machine ISA only when the target GPU hardware is known. Most GPU microarchitecture simulators available to academics execute IL instructions because there is substantially less functional state associated with the instructions, and in some situations, the machine ISA's intellectual property may not be publicly disclosed. In this paper, we demonstrate the pitfalls of evaluating GPUs using this higher-level abstraction, and make the case that several important microarchitecture interactions are only visible when executing lower-level instructions. Our analysis shows that given identical application source code and GPU microarchitecture models, execution behavior will differ significantly depending on the instruction set abstraction. For example, our analysis shows the dynamic instruction count of the machine ISA is nearly 2× that of the IL on average, but contention for vector registers is reduced by 3× due to the optimized resource utilization. In addition, our analysis highlights the deficiencies of using IL to model instruction fetching, control divergence, and value similarity. Finally, we show that simulating IL instructions adds 33% error as compared to the machine ISA when comparing absolute runtimes to real hardware. Anthony Gutierrez, Bradford M. Beckmann, Alexandru Dutu, Joseph Gross, Michael LeBeane, John Kalamatianos, Onur Kayiran, Matthew Poremba, Brandon Potter, Sooraj Puthoor, Matthew D. Sinclair, Mark Wyse, Jieming Yin, Xianwei Zhang 0001, Akshay Jain 0005, Timothy G. Rogers |
HPCA | 13 |
| 2018 | Modular Routing Design for Chiplet-Based SystemsabstractSystem-on-Chip (SoC) complexity and the increasing costs of silicon motivate the breaking of an SoC into smaller "chiplets." A chiplet-based SoC design process has the promise to enable fast SoC construction by using advanced packaging technologies to tightly integrate multiple disparate chips (e.g., CPU, GPU, memory, FPGA). However, when assembling chiplets into a single SoC, correctness validation becomes a significant challenge. In particular, the network-on-chip (NoC) used within the individual chiplets and across chiplets to tie them together can easily have deadlocks, especially if each chip is designed in isolation. We introduce a simple, modular, yet elegant methodology for ensuring deadlock-free routing in multi-chiplet systems. As an example, we focus on future systems combining chiplets on an active silicon interposer. To maximize modularity, each individual chiplet is free to implement its own NoC topology and local routing algorithm, and the interposer can implement its own independent topology and routing. Our methodology imposes a few simple turn restrictions applied only to traffic as it flows into or out of the chiplets from the interposer, and we provide a way to determine these restrictions. The end result is an overall approach that enables highly-modular, chiplet-based SoC construction while eliminating deadlocks with high performance. Jieming Yin, Zhifeng Lin, Onur Kayiran, Matthew Poremba, Muhammad Shoaib Bin Altaf, Natalie D. Enright Jerger, Gabriel H. Loh |
ISCA | 1 |
| 2017 | There and Back Again: Optimizing the Interconnect in Networks of Memory Cubes
Matthew Poremba, Itir Akgun, Jieming Yin, Onur Kayiran, Yuan Xie 0001, Gabriel H. Loh |
ISCA | 3 |
| 2016 | Efficient synthetic traffic models for large, complex SoCsabstractThe interconnect or network on chip (NoC) is an increasingly important component in processors. As systems scale up in size and functionality, the ability to efficiently model larger and more complex NoCs becomes increasingly important to the design and evaluation of such systems. Recent work proposed the "SynFull" methodology that performs statistical analysis of a workload's NoC traffic to create compact traffic generators based on Markov models. While the models generate synthetic traffic, the traffic is statistically similar to the original trace and can be used for fast NoC simulation. However, the original SynFull work only evaluated multi-core CPU scenarios with a very simple cache coherence protocol (MESI). We find the original SynFull methodology to be insufficient when modeling the NoC of a more complex system on a chip (SoC). We identify and analyze the shortcomings of SynFull in the context of a SoC consisting of a heterogeneous architecture (CPU and GPU), a more complex cache hierarchy including support for full coherence between CPU, GPU, and shared caches, and heterogeneous workloads. We introduce new techniques to address these shortcomings. Furthermore, the original SynFull methodology can only model a NoC with N nodes when the original application analysis is performed on an identically-sized N-node system, but one typically wants to model larger future systems. Therefore, we introduce new techniques to enable SynFull-like analysis to be extrapolated to model such larger systems. Finally, we present a novel synthetic memory reference model to replace SynFull's fixed latency model; this allows more realistic evaluation of the memory subsystem and its interaction with the NoC. The result is a robust NoC simulation methodology that works for large, heterogeneous SoC architectures. Jieming Yin, Onur Kayiran, Matthew Poremba, Natalie D. Enright Jerger, Gabriel H. Loh |
HPCA | 1 |
| 2014 | Energy-Efficient Time-Division Multiplexed Hybrid-Switched NoC for Heterogeneous Multicore SystemsabstractNoCs are an integral part of modern multicore processors, they must continuously support high-throughput low-latency on-chip data communication under a stringent energy budget when system size scales up. Heterogeneous multicore systems further push the limit of NoC design by integrating cores with diverse performance requirements onto the same die. Traditional packet-switched NoCs, which have the flexibility of connecting diverse computation and storage devices, are facing great challenges to meet the performance requirements within the energy budget due to latency and energy consumption associated with buffering and routing at each router. In this paper, we take advantage of the diversity in performance requirements of on-chip heterogeneous computing devices by designing, implementing, and evaluating a hybrid-switched network that allows the packet-switched and circuit-switched messages to share the same communication fabric by partitioning the network through time-division multiplexing (TDM). In the proposed hybrid-switched network, circuit-switched paths are established along frequently communicating nodes. Our experiments show that utilizing these paths can improve system performance by reducing communication latency and alleviating network congestion. Furthermore, better energy efficiency is achieved by reducing buffering in routers and in turn enabling aggressive power gating. Jieming Yin, Pingqiang Zhou, Sachin S. Sapatnekar, Antonia Zhai |
IPDPS | 1 |
| 2012 | Energy-efficient non-minimal path on-chip interconnection network for heterogeneous systemsabstractNetwork-on-Chips (NoCs) in heterogeneous systems containing both CPU and GPU cores must be designed to satisfy the performance requirements of both latency-sensitive CPU traffic and throughput-intensive GPU traffic. DVFS and adaptive routing can potentially improve NoC energy and performance efficiency. We further notice that GPU traffic can sometimes tolerate a slack defined as the number of cycles a packet can be delayed without causing performance penalty. In this work, we take advantage of the slack in GPU packets to route packets through non-minimal path, so that routers can operate at a lower frequency without suffering performance penalty. Jieming Yin, Pingqiang Zhou, Anup Holey, Sachin S. Sapatnekar, Antonia Zhai |
ISLPED | 1 |
| 2011 | NoC frequency scaling with flexible-pipeline routers
Pingqiang Zhou, Jieming Yin, Antonia Zhai, Sachin S. Sapatnekar |
ISLPED | 2 |