EDBT 2026 Demo / reviewers in the wild / expert
Li Shen 0007
dblp:91/3680-7
· DBLP profile ↗
67ranked-venue papers
2as first author
33since 2021 · last 2026
0000-0001-9043-2998ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 1 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoCT: Hybrid Compressor Tree Optimization via Reinforcement Learning with Graph ModelingabstractCompressor tree optimization is a critical step in the design of high-performance arithmetic circuits, such as multipliers or Multiply-Accumulate (MAC) units, where the goal is to efficiently accumulate partial products. Traditional methods often rely on heuristic rules or manual design, which struggle to achieve optimal performance across diverse multiplier scales. In this paper, we propose AutoCT, a novel framework for hybrid compressor tree optimization that leverages reinforcement learning (RL) with graph neural networks (GNNs) to model and optimize compressor trees. By representing the compressor tree as a graph and employing a deep Q-network enhanced with graph attention networks, AutoCT dynamically selects compressor types and configurations to minimize both area and delay. Experimental results highlight that AutoCT reduces the area-delay product by up to $\mathbf{1 5. 0 3} \boldsymbol{\%}$ compared with commercial multiplier IPs for 32-bit designs. Our source code is publicly available at https://github.com/shangshnagyao/AutoCT. Shangshang Yao, Kunlong Li, Li Shen 0007 |
ASP-DAC | 3 |
| 2026 | Row-wise Inter-Phase Pipelining for Hardware-Efficient GCN AccelerationabstractGraph Convolutional Networks (GCNs) are widely deployed for learning on graph-structured data. Yet, their canonical two-phase execution, consisting of feature transformation followed by neighborhood aggregation, suffers from severe inefficiencies on existing accelerators. The prevailing phase-decoupled model explicitly stores dense intermediate results to off-chip memory after transformation and reloads them for aggregation, incurring substantial inter-phase data movement. Additionally, irregular access patterns cause intra-phase redundancy through repeated fetches of sparse inputs. To overcome both forms of redundancy, we propose IPRS-GCN, a streaming accelerator built on Inter-Phase Row-Streaming (IPRS). IPRS fuses transformation and aggregation at row granularity into a single pipeline, forwarding each node’s transformed features directly to its neighbors via the Row Queue as soon as they are computed. This eliminates off-chip storage of intermediates and enforces single-pass streaming access to inputs while keeping weights resident on-chip. The architecture realizes this dataflow through lightweight mechanisms including conflict-aware non-zero packing and degree-adaptive buffering, achieving near-ideal compute utilization with minimal on-chip footprint. Evaluated across five real-world datasets, IPRS-GCN achieves a geometric mean speedup of 107.74 × over an NVIDIA V100 GPU, outperforms HyGCN and AWB-GCN by 7.74 × and 2.12 ×, respectively, and improves energy efficiency by 1.37 ×. Junsheng Chang, Yang Guo 0003, Li Shen 0007 |
CF | 4 |
| 2026 | Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei |
ISCA | 7 |
| 2026 | Research on the development and challenges of PIM
Qingjie Lang, Ruo-Xi Wang, DongHuan Xie, Zhen-Yu Gao, Li Shen 0007 |
Frontiers Comput. Sci. | 6 |
| 2026 | Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM ComputingabstractThe growing demand for high-performance, energy-efficient execution of data-parallel workloads has driven the resurgence of vector processors, yet their expanding instruction sets exacerbate the area-performance tradeoff of vector processing units (VPUs). Processing-in-memory (PIM) technique offers a promising path to mitigate this tradeoff by offloading vector operations near data. However, integrating the computing-capable SRAM (C-SRAM) with conventional VPUs introduces significant architectural challenges, including inefficient coordination between heterogeneous devices, the lack of a unified hardware/software interface, and underutilized parallelism within the C-SRAM arrays due to unoptimized data handling. To address these challenges, this article proposes a heterogeneous VPU (HVPU) that seamlessly integrates a standard VPU inside a vector processor with a C-SRAM for more efficient vector processing. HVPU introduces a standardized interface between the processor frontend and the C-SRAM, enabling instruction dispatch, dynamic hazard resolution, and concurrent execution. Furthermore, it employs multiple independent PIM blocks coupled with dual controlling pipelines inside the C-SRAM for further performance improvement. This design facilitates pipelined data loading and computing, effectively hiding memory latency and fully unlocking the parallel potential of C-SRAM. The system is supported by a user-friendly and generic programming model featuring a two-layer extended ISA system and a vector batch pipelining mechanism. Experimental results show that our HVPU-enhanced processor achieves significant speedups of 5.11× to 39.0× over the Xuantie-910 baseline on vector benchmarks, respectively, while reducing energy consumption by 75% on average. Meanwhile, it outperforms state-of-the-art PIM accelerators by 1.33× to 1.97× on the same benchmarks, with minimal area overhead. This demonstrates that our architectural co-design effectively alleviates the area-performance tradeoff in vector processors and offers a scalable heterogeneous architecture template for efficient vector processing. Dunbo Zhang, Shangshang Yao, Qingjie Lang, Junyi Zhu 0016, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 6 |
| 2026 | Communication-Partition Co-Optimization for Quantum Circuit Simulation on CPU+GPU ClustersabstractSimulating the behavior of quantum circuits on classical computers is one of the widely used approaches for current quantum computing device and quantum algorithm research. State- vector simulators keep all the quantum states in main memory, thus consuming a large amount of memory and resulting in significant memory access and communication overhead far longer than calculation time. Consequently, data communication time often becomes the primary bottleneck in quantum circuit simulation. Many classical simulators, such as QuEST, are designed to execute quantum gates serially without exploiting data locality, and all amplitudes need to be exchanged. Such a gate-unaware full-data communication scheme forces powerful computation such as GPUs to spend substantial time waiting for data transfers. In this paper, we propose a gate-aware on-demand communication quantum simulation framework to optimize communication overhead. We analyze the on-demand communication process and develop a communication cost model to elucidate the relationship between communication costs and circuit partition. Since greedy circuit partition method can only yield locally optimal solutions, we propose a choice-start-node quantum circuit partition method guided by the communication cost model to further reduce communication overhead in quantum circuit simulation. We evaluate on a 4-node cluster with 2 NVIDIA A100 GPUs on each node and our design achieves 59.32× speedup compared to the gate-unaware full-data communication scheme. Furthermore, compared to the greedy circuit partition method, our partition method can achieve simulation acceleration of 1.79×. Chenyang Jiao, Li Shen 0007 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Accelerating DFS-based Subgraph Matching on GPU via Reusing IntersectionabstractSubgraph matching is a well-known NP-hard problem widely applied in fields such as bioinformatics, cheminformatics, and social network analysis. It aims to enumerate all embeddings in a data graph that are isomorphic to a query graph. Subgraph matching algorithms can be roughly classified into BFS-based and DFS-based algorithms. The intersection operation is the core operation in both types of algorithms and consumes a significant amount of time. There are numerous repeated intersection operations in subgraph matching, and their results can be reused. Recent studies have focused on implementing the DFS-based algorithm on GPUs with an explicit stack. We categorize the reuse in the DFS-based algorithm into two types: within-stack reuse and across-stack reuse. Previous works only considered within-stack reuse, which has a narrow scope of application and lacks generality for some query graphs. In this paper, we are the first to propose a method for across-stack reuse in DFS-based algorithms on GPUs. We introduce a tree-structured copy stack to reuse duplicate intersections and design a secondary checking mechanism to ensure the correctness of the reuse. We use the non-blocking lock mechanism to avoid read-write races and conduct a rigorous theoretical analysis to prove that the impact of the non-blocking lock on reuse is negligible. Moreover, we propose a GPU-specific optimization method that uses key nodes to reduce memory transactions during intersection operations. Compared with the state-of-the-art DFS-based matching algorithms, our work achieves $1.12 \times$ to $1.95 \times$ speedup, and the reuse rate reaches 78.68%.ACM Reference Format:Chen Chen, Shanzhi Gu, Junsheng Chang, and Li Shen. 2025. Accelerating DFS-based Subgraph Matching on GPU via Reusing Intersection. In. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnnn Chen Chen 0016, Shanzhi Gu, Junsheng Chang, Li Shen 0007 |
PACT | 4 |
| 2025 | Combination of Storage and Accumulation for Synchronous SpMV Acceleration on FPGAs with HBMabstractSparse matrix-vector multiplication (SpMV) is a crucial computational operation in various fields, such as graph computation, machine learning, and molecular dynamics.However, due to the irregular data distribution and low density of non-zero elements, the performance of SpMV is typically inferior to that of dense matrix computations.To tackle this issue, numerous optimization efforts have been made on FPGAs equipped with high-bandwidth memory (HBM), addressing problems like excessive transmission latency and load imbalance between channels.Nevertheless, several key challenges remain that hinder the performance of FPGA-based SpMV accelerators, including (1) the tendency for cache blocks to frequently miss and be replaced due to the irregular distribution of non-zero elements in sparse matrices; (2) control divergence issues within the single instruction multiple data (SIMD) SpMV accelerator architecture.To overcome these challenges, this study introduces CoSpMV, an accelerator design for FPGAs with HBM.It incorporates (1) a matrix block synchronous processing technique under ping-pong buffering, (2) a more efficient data compression format known as row-column compressed coordinate format (R3Coo), and (3) a storage accumulation module.R3Coo enhances data transmission efficiency by compressing bit width and improving the efficiency of flag bits; the matrix block synchronous processing technique under ping-pong buffering conceals replacement overhead and prevents irregular cache access by dividing and synchronously processing matrix data into rows and columns; the storage accumulation module eliminates inconsistent control flow by decoupling the addition DongHuan Xie, Qingjie Lang, Dunbo Zhang, Junsheng Chang, Li Shen 0007 |
CF | 6 |
| 2025 | ILP-Driven FPGA Multiplier Synthesis: A Scalable Framework for Area-Latency Co-OptimizationabstractModern computing paradigms impose diverging requirements on arithmetic circuits, with cloud applications prioritizing throughput and edge devices demanding area efficiency under power constraints. While Field-Programmable Gate Arrays (FPGAs) leverage heterogeneous DSP-LUT fabrics for flexibility, their rigid DSP layouts and LUT-centric architectural constraints hinder scalable multiplier designs. Existing FPGA-based approaches face intrinsic scalability limitations from primitive cascading techniques and inflexible performance-resource tradeoffs. This paper proposes an Integer Linear Programming (ILP)-driven framework for Pareto-optimal multiplier synthesis, enabling arbitrary bit-widths via LUT-compressor modeling, application-aware configurations (performance-focused vs. area-minimized modes), and an automated toolchain translating ILP solutions to synthesizable Verilog. Evaluations on Xilinx UltraScale+ series FPGAs demonstrate 27.8% critical path delay reduction in high-performance mode and 21.4% LUT resource savings in areaefficient mode for 16-bit multipliers versus Xilinx LogiCORE™IP. The framework’s adaptive optimization bridges cloud-edge computational divergence, achieving a 15.5-34.2% area-delay product improvement across 8-16b designs. Shangshang Yao, Kunlong Li, Li Shen 0007 |
ICCAD | 3 |
| 2025 | Cross-Dimension Feature Fusion for Real-Time Object Detection
En Zhu, Li Shen 0007, Tie Hong |
PRCV (16) | 3 |
| 2025 | A comprehensive survey on graph neural network accelerators
Li Shen 0007 |
Frontiers Comput. Sci. | 3 |
| 2025 | Decoupled Vector Processing Unit: Past, Present, and Future
Ruo-Xi Wang, Dun-Bo Zhang, Qingjie Lang, DongHuan Xie, Zhen-Yu Gao, Li Shen 0007 |
J. Comput. Sci. Technol. | 7 |
| 2025 | TSN Cache: Exploiting Data Localities in Graph Computing ApplicationsabstractThis article finds that the reusability of vertices in the same graph in graph processing differs, and the high-reuse and low-reuse vertices are stored together. These phenomena lead to the inability of existing GPU architectures to capture the reusability of graph processing. The most advanced cache optimization strategies cannot implement different management strategies for data with different reusability, which is an essential reason for graph processing’s poor performance. Therefore, we propose a TSN cache scheme for the GPU platform. This scheme employs distinct management strategies for data with varying reusability in the cache, effectively leveraging the locality of these different data types. In addition, the TSN cache scheme can also reduce the probability of cache thrashing caused by low-reuse data. This article evaluates multiple graph algorithms and datasets and shows that the TSN cache scheme achieves an average speedup of 1.38 compared with the baseline scheme. Chaoyang Jia, Kai Lu 0001, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | In-SRAM Parallel Data ShuffleabstractWhile Single Instruction Multiple Data (SIMD) units are widely employed in processors for neural networks, signal processing, and high-performance computing, they suffer from expensive shuffle operations dedicated to data alignment. In fact, shuffle operations only change the layout of data and ideally should be done entirely within memory. To this end, we propose Shuffle SRAM in this article, which can shuffle multiple data elements simultaneously across SRAM banks. The key idea is exploiting inter-bank word line wise data movement to shuffle data in parallel, where all data elements on the same word line of SRAM can be shuffled simultaneously, achieving a high level of parallelism. Through suitable data layout preparation and proper control, Shuffle SRAM efficiently supports a wide range of commonly used shuffle operations. Our evaluation results show that the Shuffle SRAM can reap performance benefits of 14.3× for data reorganization only applications and 1.97× for data reorganization + computation applications over conventional shuffle architecture on general-purpose processors. With Shuffle SRAM, the state-of-the-art vector processor can obtain 2.58× energy efficiency. Compared with traditional SRAM, Shuffle SRAM only increases 3.5% additional area overhead. Chaoyang Jia, Dunbo Zhang, Qingjie Lang, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | Eliminate Data Divergence in SpMV via Processor and Memory Co-Computing FrameworkabstractSparse matrix-vector multiplication (SpMV) is a performance-critical kernel in various application domains, including high-performance computing, artificial intelligence, and big data. However, the performance of SpMV on SIMD devices is greatly affected by data divergences. To address this issue, we propose an In-SRAM Computing-based Processor Memory Co-Compute SpMV optimization framework that divides the SpMV kernel into two stages: a compute-intensive stage and a control-intensive stage. For optimizing the first stage, we leverage the parallel random access feature of multi-bank SRAM to eliminate overheads caused by memory divergences and use the Aggregate Table (AT) to reduce bank conflicts. For optimizing the second stage, we convert control divergences into memory divergences and utilize the Accumulate ScratchPad Memory (AccSPM) for executing reduction operations while eliminating overheads caused by memory divergences. Experimental results demonstrate that our solution achieves significant throughput increase over highly optimized vector SpMV kernels under CSR, CSR5, and CVR compression formats with performance speedups up to 4.74x, 5.58x, and 4.83x (3.11x, 3.04x, and 3.07x on average), respectively. Dunbo Zhang, Li Shen 0007, Kai Lu 0001 |
IEEE Trans. Computers | 2 |
| 2024 | Extending the RISC-V Instruction Set for High Performance Data Compression Hardware AccelerationabstractWith the advent of the big data era, the exponen-tially growing data processing requirements pose a huge challenge to data compression. Existing FPGA hardware acceleration schemes have many problems and a new hardware acceleration scheme needs to be explored. There are a large number of parallelizable loops in the data compression algorithm, so they can be accelerated by vectorization. In this paper, we improve RISC- V Vector Extension (RVV) for the data compression. We analyze five common compression algorithms and design a class of vector adjacency instructions for vectorization acceleration for hotspot loops in compression algorithms that cannot use RVV vectorization. We design a decoupled vector arithmetic unit for these instructions that is able to complete computations with data-dependent loops in a non-blocking way. The open source vector processor Ara is used to implement our ideas and is synthesized and implemented on the Alveo U50. The results show that our work brings a maximum speedup of 13.24x in cycle count. Junzhe Huang, Qiang Dou, Li Shen 0007 |
ASAP | 3 |
| 2024 | A Distributed Framework for Subgraph Isomorphism Leveraging CPU and GPU Heterogeneous ComputingabstractSubgraph isomorphism enumerates all embeddings in a data graph that are identical to a query graph. It is a well-known NP-hard problem widely used in various domains, such as bioinformatics, chem-informatics, and social network analysis. Recent works are focused on using GPUs for subgraph isomorphism. Due to the massive scale of intermediate results, current GPU implementations face challenges in scaling across multiple nodes due to high communication costs. The computational power of CPUs is not fully utilized in this process. We present a distributed framework for subgraph isomorphism that leverages CPU and GPU heterogeneous computing. It eliminates the intermediate results on GPU and significantly reduces communication overhead during the load-balancing process. The experiments indicate that our algorithm can be extended to multiple nodes with an almost linear efficiency improvement. Furthermore, our method also significantly outperforms other existing works on GPUs. It can reach an improvement of up to 21 × compared to the state-of-the-art implementation CuTS in the distributed environment. Chen Chen 0016, Li Shen 0007, Yingwen Chen 0001 |
ICPP | 2 |
| 2024 | ScalaQC: a scalability optimization framework for full-state quantum simulation on CPU+GPU heterogeneous clusters
Chenyang Jiao, Zhikai Qin, Li Shen 0007 |
CCF Trans. High Perform. Comput. | 3 |
| 2024 | Extension VM: Interleaved Data Layout in Vector MemoryabstractWhile vector architecture is widely employed in processors for neural networks, signal processing, and high-performance computing; however, its performance is limited by inefficient column-major memory access. The column-major access limitation originates from the unsuitable mapping of multidimensional data structures to two-dimensional vector memory spaces. In addition, the traditional data layout mapping method creates an irreconcilable conflict between row- and column-major accesses. Ideally, both row- and column-major accesses can take advantage of the bank parallelism of vector memory. To this end, we propose the Interleaved Data Layout (IDL) method in vector memory, which can distribute vector elements into different banks regardless of whether they are in the row- or column-major category, so that any vector memory access can benefit from bank parallelism. Additionally, we propose an Extension Vector Memory (EVM) architecture to achieve IDL in vector memory. EVM can support two data layout methods and vector memory access modes simultaneously. The key idea is to continuously distribute the data that needs to be accessed from the main memory to different banks during the loading period. Thus, EVM can provide a larger spatial locality level through careful programming and the extension ISA support. The experimental results showed a 1.43-fold improvement of state-of-the-art vector processors by the proposed architecture, with an area cost of only 1.73%. Furthermore, the energy consumption was reduced by 50.1%. Dunbo Zhang, Qingjie Lang, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AIabstractThe dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs. Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Local Sample-Weighted Multiple Kernel Clustering With Consensus Discriminative GraphabstractMultiple kernel clustering (MKC) is committed to achieving optimal information fusion from a set of base kernels. Constructing precise and local kernel matrices is proven to be of vital significance in applications since the unreliable distant-distance similarity estimation would degrade clustering performance. Although existing localized MKC algorithms exhibit improved performance compared with globally designed competitors, most of them widely adopt the KNN mechanism to localize kernel matrix by accounting for τ -nearest neighbors. However, such a coarse manner follows an unreasonable strategy that the ranking importance of different neighbors is equal, which is impractical in applications. To alleviate such problems, this article proposes a novel local sample-weighted MKC (LSWMKC) model. We first construct a consensus discriminative affinity graph in kernel space, revealing the latent local structures. Furthermore, an optimal neighborhood kernel for the learned affinity graph is output with naturally sparse property and clear block diagonal structure. Moreover, LSWMKC implicitly optimizes adaptive weights on different neighbors with corresponding samples. Experimental results demonstrate that our LSWMKC possesses better local manifold representation and outperforms existing kernel or graph-based clustering algorithms. The source code of LSWMKC can be publicly accessed from https://github.com/liliangnudt/LSWMKC. Liang Li 0041, Siwei Wang 0001, Xinwang Liu 0002, En Zhu, Li Shen 0007, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MPRTA: An Efficient Multilevel Parallel Mobile Accelerator for High-Performance Ray TracingabstractRay tracing has been regarded as the future of graphics rendering technology for a long time. However, interactive ray tracing still faces challenges, especially in mobile devices, such as high computational intensity and multiple branches. In this brief, we aim to maximize overall efficiency by leveraging all forms of potential parallelism, including task, basic block, loop, and pipeline levels. We present multilevel parallel ray tracing accelerator (MPRTA), an innovative mobile accelerator that offers high performance and optimal efficiency for ray tracing. Experimental results indicate that MPRTA is$1.67\times $more efficient than the currently best-reported mobile accelerator. Run Yan, Yin Su, Hui Guo 0004, Yashuai Lü, Nong Xiao 0001, Li Shen 0007, Yongwen Wang, Libo Huang 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | ImprLM: An Improved Logarithmic Multiplier Design Approach via Iterative Linear-Compensation and Modified Dynamic SegmentabstractIn this paper, we present ImprLM, an improved logarithmic multiplier design approach via iterative linear-compensation and dynamic segment. Firstly, we have optimized the computational flow of the Mitchell-based logarithmic algorithm to devise a more concise and compact logarithmic multiplier. Subsequently, we introduce an innovative iterative linear error compensation technique, which significantly improves the precision of the logarithmic multiplier compared with the cutting-edge Look-Up Table (LUT) based error compensation strategies. For the purpose of achieving additional logic savings, we introduce modified dynamic segment methods for intermediate variables during multiplication. The synthesis results demonstrate that ImprLM can achieve up to a 22.5% reduction in the power-delay product, while simultaneously enhancing accuracy by 26.8% compared to the state-of-the-art Mitchell-based logarithmic multipliers. Furthermore, we incorporated ImprLM into image processing and the results reveal that ImprLM has a negligible effect on the output quality, while maintaining a low resource consumption. Shangshang Yao, Li Shen 0007 |
ICCD | 2 |
| 2023 | Communication Optimizations for State-vector Quantum Simulator on CPU+GPU ClustersabstractSimulating the behavior of quantum circuits on classical computers are one of the widely used approaches for current quantum computing device and quantum algorithm research. State-vector simulators keep all the quantum states in main memory, thus consuming a large amount of memory and resulting in great memory access and communication overhead far longer than calculation time. Many classical simulators, such as QuEST, are designed by serial execution of quantum gates without using the data locality, and all amplitudes need to be exchanged. Such a gate-unaware full data communication scheme forces powerful computation such as GPUs to spend a lot of time on waiting for data transmission. In this paper, we propose a gate-aware on-demand communication quantum simulation framework to optimize communication overhead. A quantum circuit partition method through gate fusion is first proposed to avoid unnecessary communications. Moreover, a on-demand data communication scheme is proposed to optimize data transfer among computation nodes. Based on these designs, a prototype has been implemented on QuEST. We evaluated on an 8-node cluster with four NVIDIA Tesla V100 GPUs on each node and our designs can achieve 3.0x-15.6x speedup compared to QuEST for 32-34-qubit quantum systems. Moverover, it can scale to simulate 37-qubit quantum system instead of 34-qubit quantum system for QuEST. Chenyang Jiao, Li Shen 0007 |
ICPP | 3 |
| 2023 | Dense Crosstalk Feature Aggregation for Classification and Localization in Object DetectionabstractThe misalignment between classification and localization is a significant performance improvement point for object detection. To cope with the misalignment problem, more attempts have been made to separate different tasks (e.g., Classification, Bounding Box Regression) by introducing extra heads, which emphasizes the separation of multiple tasks to cope with their variability. In this paper, we consider that both separation and crosstalk are important between classification and localization. Considering that the two types of tasks are different and have different regions and features of interest, they are in conflict with each other and therefore need to be separated. However, they also need to be fused, because classification and localization are, after all, about understanding the same object. To realize this idea, we introduce bidirectional crosstalk detection head in a systematic manner to provide a full deep cross-fusion between classification and localization. To our best knowledge, it is the first time that full bidirectional crosstalk is introduced between classification and localization for one-stage detector. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method. With a ResNet-50 backbone, our method can significantly improve the GFLV1 baseline by 2.0 AP with similar inference speed (18.5 fps vs. 18.3 fps) and further boost GFLV1 with a big margin (4.3 AP) by increasing our model size. Fair comparisons also show that the proposed head outperforms state-of-the-art heads (T-Head, DyHead) with comparable or faster inference speed under the same ATSS baseline model. With a Res2Net-DCN backbone, our model achieves 51.7 AP at single-model single-scale testing. The code and pretrained models will be made publicly available. En Zhu, Jiyong Tan, Li Shen 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Late Fusion Multiple Kernel Clustering With Local Kernel Alignment MaximizationabstractMulti-view clustering, which appropriately integrates information from multiple sources to reveal data’s inherent structure, is gaining traction in clustering. Though existing procedures have yielded satisfactory results, we observe that they have neglected the inherent local structure in the base kernels. This may cause adverse effects on clustering. To solve the problem, we introduce LF-MKC-LKA, a simple yet effective late fusion multiple kernel clustering with local kernel alignment maximisation approach. In particular, we first determine the nearest$k$neighbours in the average kernel space for each sample and record the information in the nearest neighbor indicator matrix. Then, the nearest neighbor indicator matrix can be used to generate local structure matrix of each sample. The local kernels of each view may then be generated using the local structure matrix, retaining just the highly confident local similarities for learning the intrinsic global manifold of data. They can also be utilised to keep the block diagonal structure and improve the robustness of the underlying kernels against noise.We input the local kernels of each view into the kernel$k$-means (KKM) algorithm and get the local base partitions. Finally, we use a three-step iterative optimization approach to maximize the alignment of the consensus partition using base partitions and a regularisation term. As demonstrated, a significant number of trials on 11 multi-kernel benchmark datasets have shown that the proposed LF-MKC-LKA is effective and efficient. A number of experiments are also designed to demonstrate the fast convergence, excellent performance, robustness and low parameter sensitivity of the algorithm. Our code can be find athttps://github.com/TiejianZhang/TMM21-LF-MKC-LKA. Tiejian Zhang, Xinwang Liu 0002, Lei Gong 0008, Siwei Wang 0001, Xin Niu 0002, Li Shen 0007 |
IEEE Trans. Multim. | 6 |
| 2022 | DGEMM Optimization Oriented to ARM SVE Instruction Set ArchitectureabstractIn current computer performance evaluation, linear mathematics library is an important test program and matrix multiplication is its main calculation part. Matrix multiplication algorithm contains multi-layer loops and can be parallelized flexibly. It is very suitable to run on multi-core processor with vector registers. With the popularity and application of ARMV8 vector processors in the field of high performance computing, the peak performance of processors is required to be higher. In this paper, we propose a double precision general matrix multiplication DGEMM vectorization method based on OpenBLAS and implemented on Phytium processor. The main work includes: the data block size is optimized for the storage structure of the processor and the data rearrangement is redesigned; based on the SVE vector instructions, a mathematical model is established to find the optimal kernel function size; an efficient assembly kernel is realized by using computing instructions to hide the memory delay. The experimental results show that the optimized code improves the measured performance of OpenBLAS original DGEMM algorithm from 45.07% of the theoretical peak performance of single core to 80.18%. Sizheng Sun, Li Shen 0007 |
ICPADS | 5 |
| 2022 | DLC: An Optimization Framework for Full-State Quantum Simulation
Zhikai Qin, Li Shen 0007 |
NPC | 3 |
| 2022 | Compressed page walk cache
Dunbo Zhang, Chaoyang Jia, Li Shen 0007 |
Frontiers Comput. Sci. | 3 |
| 2021 | Multi-level PWB and PWC for Reducing TLB Miss Overheads on GPUs
Dunbo Zhang, Chaoyang Jia, Qiong Wang 0001, Li Shen 0007 |
ICA3PP (2) | 5 |
| 2021 | A Multi-precision Quantized Super-Resolution Model Framework
Dunbo Zhang, Qiong Wang 0001, Li Shen 0007 |
ICA3PP (1) | 4 |
| 2021 | An Efficient Hybrid Parallel Compression Approximate MultiplierabstractApproximate computing has been widely used in many fault-tolerant applications. Multiplication as a key kernel in such applications, it is significant to improve the efficiency of approximate multiplier to achieve high computational performance. This paper proposes a novel approximate multiplier design based on using different compressors for different regions of partial products. We designed two Preprocessing Units (PUs) to explore the best efficiency via increasing the number of sparse partial products. Multiple 8-bit multipliers are designed using Verilog and synthesized under the 45-nm CMOS technology. Compared with the conventional Wallace Tree multiplier, experimental results indicate that one of our proposed multipliers reduce Power-Delay Product (PDP) by 58.5% at most with 0.42% normalized mean error distance. Moreover, a case study of image processing applications is also investigated. Our proposed multipliers can achieve a high peak signal-to-noise ratio of 51.87dB. Compared to the state-of-the-art, the proposed multiplier has a better comprehensive performance in accuracy, area and power consumption. Shangshang Yao, Qiong Wang 0001, Li Shen 0007 |
ICCD | 4 |
| 2021 | GraphPEG: Accelerating Graph Processing on GPUsabstractDue to massive thread-level parallelism, GPUs have become an attractive platform for accelerating large-scale data parallel computations, such as graph processing. However, achieving high performance for graph processing with GPUs is non-trivial. Processing graphs on GPUs introduces several problems, such as load imbalance, low utilization of hardware unit, and memory divergence. Although previous work has proposed several software strategies to optimize graph processing on GPUs, there are several issues beyond the capability of software techniques to address. In this article, we present GraphPEG, a graph processing engine for efficient graph processing on GPUs. Inspired by the observation that many graph algorithms have a common pattern on graph traversal, GraphPEG improves the performance of graph processing by coupling automatic edge gathering with fine-grain work distribution. GraphPEG can also adapt to various input graph datasets and simplify the software design of graph processing with hardware-assisted graph traversal. Simulation results show that, in comparison with two representative highly efficient GPU graph processing software framework Gunrock and SEP-Graph, GraphPEG improves graph processing throughput by 2.8× and 2.5× on average, and up to 7.3× and 7.0× for six graph algorithm benchmarks on six graph datasets, with marginal hardware cost. Ya-Shuai Lü, Hui Guo 0004, Libo Huang 0002, Qi Yu 0003, Li Shen 0007, Nong Xiao 0001, Zhiying Wang 0003 |
ACM Trans. Archit. Code Optim. | 5 |
| 2020 | High-Performance Computing and Engineering Educational Development and PracticeabstractThis Innovate Practice Full Paper presents HPC and engineering educational development and practice. Educators and researchers are witnessing the emergence of parallel and high-performance computing (HPC) in many computing and engineering environments worldwide. Single-core processing is slowly becoming obsolete while parallel solutions to application development are now replacing serial solutions to achieve higher performance, particularly in many scientific, engineering, and computing fields. Some universities or research institutes worldwide have developed and have become centers of supercomputing. Most of these centers, particularly the world acclaimed Tianhe-2 supercomputer developed by China's National University of Defense Technology (NUDT), often have their own HPC educational approaches as well as their own curricula at both undergraduate and graduate levels. This paper presents some high-level HPC educational concepts, educational framework designs, and strategies for cultivating HPC talented graduates. In this work, the authors show unique perspectives and practical experiences on student capability, oriented toward HPC curricular development. They also show a step-by-step practical system in which real platform, real application, and research-teaching integration modes can immerse students into an HPC world for better understanding. The authors show how researchers with rich computing and engineering experience can become deeply involved in course design, in-class teaching, and in laboratory instruction. They also discuss the effectiveness of HPC courses at the NUDT and they show how HPC can become a stable element in computing and engineering education. Juan Chen 0001, John Impagliazzo, Li Shen 0007 |
FIE | 3 |
| 2020 | A Multi-model Super-Resolution Training and Reconstruction Framework
Ninghui Yuan, Dunbo Zhang, Qiong Wang 0001, Li Shen 0007 |
NPC | 4 |
| 2020 | Transparent partial page migration between CPU and GPU
Shiqing Zhang, Zheng Qin 0002, YaoHua Yang, Li Shen 0007, Zhiying Wang 0003 |
Frontiers Comput. Sci. | 4 |
| 2019 | A Lightweight Method for Handling Control Divergence in GPGPUsabstractAt present, graphics processing units (GPUs) has been widely used for scientific and high performance acceleration in the general purpose computing area, which is inseparable from the SIMT (Single-Instruction, Multiple-Thread) execution model. With SIMT, GPUs can fully utilize the advantages of SIMD parallel computing. However, when threads in a warp do not follow the same execution path, control divergence generates and affects the hardware utilization. In response to this problem, warp regrouping method has been proposed to combine threads executing the same branch path, which can significantly improve thread-level parallelism. But it is found that not all warps can be regrouped effectively because that may introduce a lot of unnecessary overheads, limiting further performance improvement. In this paper, we analyze the source of overheads and propose a lightweight warp regrouping method --- Partial Warp Regrouping (PWR) that controls the scope of reorganization and avoids most of the unnecessary warp regrouping by setting thresholds. In this method, it also can reduce the complexity of hardware design. Our experimental results show that this mechanism can improve the performance by 12% on average and up to 27% compared with immediate post-dominator. YaoHua Yang, Shiqing Zhang, Li Shen 0007 |
HPC Asia | 3 |
| 2019 | MMSR: A Multi-model Super Resolution Framework
Ninghui Yuan, Xinzhou Wu, Li Shen 0007 |
NPC | 4 |
| 2019 | A statistic approach for power analysis of integrated GPU
Qiong Wang 0001, Li Shen 0007, Zhiying Wang 0003 |
Soft Comput. | 3 |
| 2018 | Adaptive VC Partitioning for NoCs in GPGPUsabstractThe design of efficient Networks-on-Chip (NoCs) is essential for GPGPUs. The asymmetry of GPGPU traffic has a significant effect on the overall system performance. An existing VC partitioning design statically assigns more VCs to the heavier reply traffic. Yet, its static partitioning cannot adapt to the dynamic variation of NoC traffic. Thus, we propose an adaptive VC partitioning (A-VCP) mechanism , which dynamically chooses the optimal VC partitioning by sampling the traffic status. Compared with the static configuration, A-VCP averagely improves the system performance by 10.5%, and reduces the energy-delay product by 9.1%. Sheng Ma, Hongyi Lu, Libo Huang 0002, Li Shen 0007, Yang Guo 0003, Zhiying Wang 0003, Wenliang Xue |
ISCAS | 4 |
| 2018 | GPU Memory Management Solution Supporting Incomplete Pages
Li Shen 0007, Shiqing Zhang, YaoHua Yang, Zhiying Wang 0003 |
NPC | 1 |
| 2018 | Design of Practical Experiences to Improve Student Understanding of Efficiency and Scalability Issues in High Performance Computing: (Abstract Only)abstractWith the increasing demand of big data technology, there has been a growing interest of introducing high performance computing in computer science curriculum. One challenge in helping students understand the nature of efficiency and scalability issues in high performance computing is the lack of opportunities for them to be engaged in large-scale applications that run on supercomputer system architecture. This poster presents a collection of example projects that have been used in a parallel computing course in multiple universities in China, including National University of Defense Technology, Sun Yat-sen University and Hunan University. These projects were adopted from a wide range of scientific computing applications such as CFD, text mining of biomedical literature and so on. The large-scale computing resource for courses is supported by two National Supercomputing Centers, one in Guangzhou and the other in Changsha. The poster describes the background, objective, structure, task, practice process and outcome for each project. It also discusses the impact on student understanding all kinds of key topics and major challenges related to computational efficiency and scalability. Such projects build a positive practical environment to make students indulge in doing all kinds of interesting and helpful trials to validate their assumptions, especially when they have different perspectives or results for one problem. The poster presents our design evaluation rubric to reflect the effectiveness of our practice, as well as the statistics about the students" achievements for the last three semesters. Juan Chen 0001, Li Shen 0007, Jianping Yin, Chunyuan Zhang |
SIGCSE | 2 |
| 2018 | Resolving the GPU responsiveness dilemma through program transformations
Bo Wu 0002, Xipeng Shen, Li Shen 0007, Zhiying Wang 0003 |
Frontiers Comput. Sci. | 5 |
| 2017 | POSTER: DaQueue: A Data-Aware Work-Queue Design for GPGPUsabstractWork-queue is an effective approach for mapping irregular-parallel workloads to GPGPUs. It can improve the utilization of SIMD units by only processing useful works which are dynamically generated during execution. As current GPGPUs lack necessary supports for work-queues, a software-based work-queue implementation often suffers from memory contention and load balancing issues. We present a novel hardware work-queue design named DaQueue, which incorporates data-aware features to improve the efficiency of work-queues on GPGPUs. We evaluate our proposal on irregular-parallel workloads with a cycle-level simulator. Experimental results show that the DaQueue significantly improves the performance over software-based implementation for these workloads. Compared with an idealized hardware worklist approach which is the state-of-the-art prior work, the DaQueue can achieve an average of 29.54% extra speedup. Ya-Shuai Lü, Libo Huang 0002, Li Shen 0007 |
PACT | 3 |
| 2017 | Co-Run Scheduling with Power Cap on Integrated CPU-GPU SystemsabstractThis paper presents the first systematic study on co-scheduling independent jobs on integrated CPU-GPU systems with power caps considered. It reveals the performance degradations caused by the co-run contentions at the levels of both memory and power. It then examines the problem of using job co-scheduling to alleviate the degradations in this less understood scenario. It offers several algorithms and a lightweight co-run performance and power predictive model for computing the performance bounds of the optimal co-schedules and finding appropriate schedules. Results show that the method can efficiently find co-schedules that significantly improve the system throughput (9-46% on average over the default schedules). Bo Wu 0002, Xipeng Shen, Li Shen 0007, Zhiying Wang 0003 |
IPDPS | 4 |
| 2017 | Unleashing the power of GPU for physically-based rendering via dynamic ray shufflingabstractComputer graphics is generally divided into two branches: real-time rendering and physically-based rendering. Conventional graphics processing units (GPUs) were designed to accelerate the former which is based on the standard Z-buffer algorithm. However, many applications in entertainment, science, and industry require high quality visual effects such as soft-shadows, reflections, and diffuse lighting interactions which are difficult to achieve with the Z-buffer algorithm, but are straightforward to implement using physically-based rendering methods. Physically-based rendering can already be implemented on present programmable GPUs. However, for physically-based rendering on GPUs, a large portion of the processing power is wasted due to low utilization of SIMD units. This is because the core algorithm of physically-based rendering, ray tracing, suffers from Single Instruction, Multiple Thread (SIMT) control flow divergences. In this paper, we propose the Dynamic Ray Shuffling (DRS) architecture for GPUs to address this problem. Our key insight is that the primary control flow divergences are caused by inconsistent ray traversal states of a warp, and can be eliminated by dynamically shuffling rays. Experimental results show that, for an estimated 0.11% area cost, DRS significantly improves the SIMD efficiency for the tested benchmarks from 41.06% to 81.04% on average. With this, the performance of a physically-based rendering method such as path tracing can be improved by 1.67X--1.92X, and 1.79X on average. Ya-Shuai Lü, Libo Huang 0002, Li Shen 0007, Zhiying Wang 0003 |
MICRO | 3 |
| 2017 | Understanding co-run performance on CPU-GPU integrated processors: observations, insights, directions
Bo Wu 0002, Xipeng Shen, Li Shen 0007, Zhiying Wang 0003 |
Frontiers Comput. Sci. | 5 |
| 2017 | Improving the Efficiency of GPGPU Work-Queue Through Data AwarenessabstractThe architecture and programming model of current GPGPUs are best suited for applications that are dominated by structured control and data flows across large regular datasets. Parallel workloads with irregular control and data structures cannot easily harness the processing power of the GPGPU. One approach for mapping these irregular-parallel workloads to GPGPUs is using work-queues. The work-queue approach improves the utilization of SIMD units by only processing useful works that are dynamically generated during execution. As current GPGPUs lack necessary supports for work-queues, a software-based work-queue implementation often suffers from memory contention and load balancing issues. In this article, we present a novel hardware work-queue design named DaQueue , which incorporates three data-aware features to improve the efficiency of work-queues on GPGPUs. We evaluate our proposal on the irregular-parallel workloads and carry out a case study on a path tracing pipeline with a cycle-level simulator. Experimental results show that for the tested workloads, DaQueue improves performance by 1.53× on average and up to 1.91×. Compared to a hardware worklist approach that is the state-of-the-art prior work, DaQueue can achieve an average of 33.92% extra speedup with less hardware area cost. Libo Huang 0002, Ya-Shuai Lü, Li Shen 0007, Zhiying Wang 0003 |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Dynamic Power-Performance Adjustment on Clustered Multi-Threading ProcessorsabstractDynamic optimization techniques such as dynamic voltage and frequency scaling (DVFS), power gating (PG) and thread migration are widely used in current multi-core processor platforms to boost performance, lower power and improve energy efficiency. To obtain the best optimization results, an important issue is to predict the performance of each thread quantitatively. In this paper, a separated and clustered performance predictor, called SCP, is proposed oriented to clustered multi- threading (CMT) processors. SCP model is implemented in AMD's FX-8320 processor and its accuracy is about 6% for SPEC CPU 2006 benchmark suite. To illustrate the application and effectiveness of SCP model furthermore, we propose PPEP-SCP model for power capping, which is a combination of SCP model and PPEP model. Compared to power capping strategy based on PPEP model, strategy based on PPEP-SCP performs better in terms of the performance and the difference between the actual and the target power consumption, because PPEP-SCP based strategy can perform optimization through both DVFS and PG techniques. Li Shen 0007, Zhiying Wang 0003, Yemao Xu |
NAS | 2 |
| 2016 | Optimization Strategies Oriented to Loop Characteristics in Software Thread Level Speculation Systems
Li Shen 0007, Zhiying Wang 0003 |
J. Comput. Sci. Technol. | 1 |
| 2014 | Improving Speculation Accuracy with Inter-thread Fetching Value Prediction
Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009 |
ICA3PP (2) | 2 |
| 2014 | PPEP: Online Performance, Power, and Energy Prediction Framework and DVFS Space ExplorationabstractPerformance, power, and energy (PPE) are critical aspects of modern computing. It is challenging to accurately predict, in real time, the effect of dynamic voltage and frequency scaling (DVFS) on PPE across a wide range of voltages and frequencies. This results in the use of reactive, iterative, and inefficient algorithms for dynamically finding good DVFS states. We propose PPEP, an online PPE prediction framework that proactively and rapidly searches the DVFS space. PPEP uses hardware events to implement both a cycles-per-instruction (CPI) model as well as a per-core power model in order to predict PPE across all DVFS states. We verify on modern AMD CPUs that the PPEP power model achieves an average error of 4.6% (2.8% standard deviation) on 152 benchmark combinations at 5 distinct voltage-frequency states. Predicting average chip power across different DVFS states achieves an average error of 4.2% with a 3.6% standard deviation. Further, we demonstrate the usage of PPEP by creating and evaluating a highly responsive power capping mechanism that can meet power targets in a single step. PPEP also provides insights for future development of DVFS technologies. For example, we find that it is important to carefully consider background workloads for DVFS policies and that enabling north bridge DVFS can offer up to 20% additional energy saving or a 1.4x performance improvement. Junli Gu, Li Shen 0007, Wei Huang 0004, Joseph L. Greathouse, Zhiying Wang 0003 |
MICRO | 3 |
| 2014 | Implementing a Leading Loads Performance Predictor on Commodity Processors
Joseph L. Greathouse, Junli Gu, Michael Boyer, Li Shen 0007, Zhiying Wang 0003 |
USENIX ATC | 5 |
| 2014 | Binary compatibility for embedded systems using greedy subgraph mapping
Xuhao Chen 0001, Li Shen 0007, Zhiying Wang 0003, Wei Chen 0009 |
Sci. China Inf. Sci. | 2 |
| 2014 | Novel Flow Control for Fully Adaptive Routing in Cache-Coherent NoCsabstractRouting algorithms for cache-coherent NoCs only have limited VCs at their disposal, which poses challenges to the design of routing algorithms. Existing fully adaptive routing algorithms apply conservative VC re-allocation: only empty VCs can be re-allocated, which limits performance. We propose two novel flow control designs. First, whole packet forwarding (WPF) re-allocates a nonempty VC if the VC has enough free buffers for an entire packet. WPF does not induce deadlock if the routing algorithm is deadlock-free using conservative VC re-allocation. It is an important extension to several deadlock avoidance theories. Second, we extend Duato's theory to apply aggressive VC re-allocation on escape VCs without deadlock. Finally, we propose a design which maintains maximal routing flexibility with low hardware cost. For synthetic traffic, our design performs averagely 88.9 percent better than existing fully adaptive routing. Our design is superior to partially adaptive and deterministic routing. Sheng Ma, Zhiying Wang 0003, Natalie D. Enright Jerger, Li Shen 0007, Nong Xiao 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | HEUSPEC: A Software Speculation Parallel ModelabstractConventional software speculative parallel models are facing challenge due to the increasing number of the processor core and the diversification of the application. The performance of the guest program under the software speculative parallel execution model is closely related to the speculation accuracy, the control overhead and the rollback overhead of the model. In order to improve the speculative accuracy and the load balance, as well as improve the overhead of the conventional model, in this paper, we proposed a novel speculative parallel model named HEUSPEC. The HEUSPEC includes 2 key techniques, the heuristic value prediction(HVP) and the dynamic task granularity resizing(DTGR). We have implemented the runtime system of the model in ANSI C language. The experiment results show that when the speedup of the HEUSPEC model can reach 4.51 on the average (12% higher than conventional model) when speculative depth equals to 7. Besides, it shows good scalability and lower memory cost. Li Shen 0007, Zhiying Wang 0003, Hui Guo 0004, Wei Chen 0009 |
ICPP | 2 |
| 2012 | Dynamic Optimization on Multi-core PlatformabstractApplications involve dynamic linked library or legacy code which is difficult to optimize for static compiler, however, can be applied with optimization techniques by run-time system who focuses on binary code instead of source code. In this paper, we present a run-time system which observes the behaviors of the applications and performs optimizations base on the information it collects, and it works on the multi-core platform to boost performance. First, we establish a framework of traditional dynamic optimization system, base on which we propose the ways to reduce the overhead of the system, including different ways of trace construction and basic block linking. Second, we take advantage of hardware resources that multi-core platform provides and exploit parallelism in our system. Finally, we evaluate the performance of our system and have two major results. First, basic block linking can greatly improve the performance of the system, because without it, the system spends much of its time on interpreting and context switching. Second, the benchmarks that running on the system have a speedup of 1.05 on average, depending on the program's behavior. Jiahui Wen, Li Shen 0007, Zhiying Wang 0003 |
TrustCom | 2 |
| 2012 | Low-Cost Binary128 Floating-Point FMA Unit Design with SIMD SupportabstractBinary64 arithmetic is rapidly becoming inadequate to cope with today's large-scale computations due to an accumulation of errors. Therefore, binary128 arithmetic is now required to increase the accuracy and reliability of these computations. At the same time, an obvious trend emerging in modern processors is to extend their instruction sets by allowing single instruction multiple data (SIMD) execution, which can significantly accelerate the data-parallel applications. To address the combined demands mentioned above, this paper presents the architecture of a low-cost binary128 floating-point fused multiply add (FMA) unit with SIMD support. The proposed FMA design can execute a binary128 FMA every other cycle with a latency of four cycles, or two binary64 FMAs fully pipelined with a latency of three cycles, or four binary32 FMAs fully pipelined with a latency of three cycles. We use two binary64 FMA units to support binary128 FMA which requires much less hardware than a fully pipelined binary128 FMA. The presented binary128 FMA design uses both segmentation and iteration hardware vectorization methods to trade off performance, such as throughput and latency, against area and power. Compared with a standard binary128 FMA implementation, the proposed FMA design has 30 percent less area and 29 percent less dynamic power dissipation. Libo Huang 0002, Sheng Ma, Li Shen 0007, Zhiying Wang 0003, Nong Xiao 0001 |
IEEE Trans. Computers | 3 |
| 2011 | A specialized low-cost vectorized loop buffer for embedded processorsabstractCurrent loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting loop buffer is to obtain its performance benefit. In this paper, we propose an application specific loop buffer organization for vectorized processing kernels, to achieve low-power and high-performance goals. The vectorized loop buffer (VLB) is simplified with single loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices. Libo Huang 0002, Zhiying Wang 0003, Li Shen 0007, Hongyi Lu, Nong Xiao 0001, Cong Liu 0009 |
DATE | 3 |
| 2011 | A novel shared-buffer router for network-on-chip based on Hierarchical Bit-line BufferabstractBuffer resources are key components of the on-chip router, shared-buffer structures are proposed to improve performance and reduce power consumption. This paper presents a novel on-chip network router with a shared-buffer based on Hierarchical Bit-line Buffer (HiBB). HiBB can be configured flexibly according to traffics and its inherent characteristic of low power is also noticeable. Moreover, we propose two schemes to further optimize the router. First, a congestion-aware output-port allocation scheme is used to assign higher priority to packets heading to light-loaded directions, and the congestion situation of the total network will be addressed. Second, an efficient run-time Virtual Channel (VC) regulation scheme is proposed to configure the shared buffer, so that VCs are allocated according to the loads of network. Experimental results show that the proposed HiBB router with about 6.9% area savings outperforms the generic router under different traffic patterns. The power consumption of the HiBB router can also be reduced up to about 70% of the generic router under light traffics, while it may exceed that of the generic one up to about 3.7–5.7% under heavy traffics for the increased flit transmissions. Weixia Xu 0001, Hongguang Ren, Qiang Dou, Zhiying Wang 0003, Li Shen 0007, Cong Liu 0009 |
ICCD | 6 |
| 2011 | Characterizing Fine-Grain Parallelism on Modern Multicore PlatformabstractSince chip multiprocessors have dominated the processor market, developing a parallel programming model with proper trade-off between productivity and efficiency become increasingly important. As a typical fine-grain parallelism model, Intel Threading Building Blocks (TBB) simplifies parallel programming by runtime schedule. Despite its simplicity, it costs non-trivial runtime overhead which may increase as the thread counts increase. In this work, we conduct an experiment on real commodity hardware to evaluate performance scalability of TBB using PARSEC benchmark suite. We first compare TBB with Pthreads to show that TBB applications can achieve comparable performance as Pthreads applications. To find the performance bottleneck of TBB applications, we measure the runtime overhead of TBB focused on 3 basic TBB runtime activities. The result provides valuable implications which can be used to develop scalable runtime libraries and architectural support for alleviating performance bottlenecks. Xuhao Chen 0001, Wei Chen 0009, Li Shen 0007, Zhiying Wang 0003 |
ICPADS | 5 |
| 2010 | SIF: Overcoming the limitations of SIMD devices via implicit permutationabstractSIMD devices have gained widespread acceptance in modern microprocessor designs for their superior performance for multimedia applications. However, there are three remaining limitations to the efficient utilization of SIMD devices in general-purpose computer systems: memory alignment, data reorganization and control flow. This paper presents SIF, an efficient SIMD interface framework that addresses these three shortcomings without modifying existing ISA. It is designed around a permutation vector register file (PVRF) and it adds new extended instructions to set internal permutation state in SIMD datapath rather than putting the permutation state setting bits in every instruction. The implicit permutation capability provided by PVRF results in zero overhead, which frees the handling of three limitations by using permutation instructions. To further reduce the state setting instructions in SIMD datapath, a technique that moves the workloads from SIMD pipeline into scalar pipeline is also introduced. With the help of proposed compilation algorithm, SIF can efficiently transform regular SIMD codes into SIF codes which make it easily integrated in all existing SIMD devices. We implemented these techniques in a vectorizing compiler and experimental results show that most of the permutation overhead instructions can be eliminated and distinct performance speedup can be achieved, which is 37% higher than current SIMD techniques on average. Libo Huang 0002, Li Shen 0007, Zhiying Wang 0003, Nong Xiao 0001, Sheng Ma |
HPCA | 2 |
| 2010 | Permutation optimization for SIMD devicesabstractSingle-instruction-multiple-data (SIMD) devices have been widely incorporated into baseline instruction level parallelism (ILP) processors to enable more efficient data level parallelism (DLP) support. This paper addresses the unsolved problem of the need to permute the SIMD elements packed in registers for maximum parallelism performance. An implicit data permutation (IDP) mechanism is proposed for handling various permutation operations without performance overhead. Various ways can be used to implement IDP mechanism. One way is to modify the baseline processors with permutation vector register file (PVRF) and associated new extended instructions. The PVRF allows accessing the data by using permutation pattern in addition to the existing row pattern. This method is described in detail and experimental results show that distinct performance speedup can be achieved, which is 47% higher than current SIMD techniques on average. Libo Huang 0002, Li Shen 0007, Zhiying Wang 0003 |
ISCAS | 2 |
| 2009 | A Light-weight Code Cache Design for Dynamic Binary TranslationabstractInterpretation and basic block translation (BBT) are two typical strategies for cold code emulation in a dynamic binary translation (DBT) system. More and more DBT systems employ BBT as the generated native code runs more efficient than the interpretation routines. We observe that BBT's high efficiency is based on those special hardware assists. With certain simple hardware techniques, interpretation could outperform BBT. In our pervious work, we proposed a hardware interpreted code cache (Pcache) mechanism to speedup interpretation by saving the decoded instruction information during interpretation. This light-weight code cache design could be extended to assist the hotspots translation, thus further reduce the DBT systems' overhead. We add the translation entry into the Pcache design thus saving most decoding operations during translation. We use eight SPEC 2000 integer benchmarks on our DBT simulator. Results show that the modified Pcache design causes a speedup of 1.94 according to the referenced DBT with basic interpretation and the interpretation based DBT system assisted by the modified Pcache performs more efficiently than the DBT system which employs BBT for the cold code. Wei Chen 0009, Li Shen 0007, Hongyi Lu, Zhiying Wang 0003, Nong Xiao 0001 |
ICPADS | 2 |
| 2009 | Using Pcache to Speedup Interpretation in Dynamic Binary TranslationabstractAbstract— Dynamic binary translation (DBT) converts codes written for a source instruction set architecture (ISA) into optimized code for a target ISA. DBT has emerged as an important tool with real world applications. Interpretation is always adopted to handle the non-hotspot code in a two-stage DBT system. An important consideration in such DBT systems is the interpretation overhead. We investigate that repeated redecoding operations are the bottleneck of interpretation overhead. We propose interpreted code cache (Pcache), a hardware assist to save the information of the decoded instruction for reuse. We analyze and model Pcache performance via simulation on a DBT system simulator. Results from SPEC2000 integer benchmarks show that Pcache could significantly reduce redecoding operations and the overhead of interpretation in a DBT system. The speedup of interpretation is up to 17.12 on average with assist of Pcache. We also analyze the extra overhead caused by Pcache, which is neglectable compared to the performance gains. Wei Chen 0009, Hongyi Lu, Li Shen 0007, Zhiying Wang 0003, Nong Xiao 0001 |
ISPA | 3 |
| 2008 | Customizing computation accelerators for extensible multi-issue processors with effective optimization techniquesabstractCompared with single-issue general purpose processors (GPPs), extensible multi-issue/VLIW processors can exploit instruction-level parallelism, which are more suitable for computation intensive tasks. Moreover, they offer the ability of customizing computation accelerators for an application domain. In this paper, we present an automated methodology that customizes computation accelerators for the multi-issue/VLIW extensible processors, where several techniques are also proposed to optimize the design of an accelerator. Ya-Shuai Lü, Li Shen 0007, Libo Huang 0002, Zhiying Wang 0003, Nong Xiao 0001 |
DAC | 2 |
| 2007 | A New Architecture For Multiple-Precision Floating-Point Multiply-Add Fused Unit DesignabstractThe floating-point multiply-add fused (MAF) unit sets a new trend in the processor design to speed up floatingpoint performance in scientific and multimedia applications. This paper proposes a new architecture for the MAF unit that supports multiple IEEE precisions multiply-add operation (AtimesB+C) with Single Instruction Multiple Data (SIMD) feature. The proposed MAF unit can perform either one double-precision or two parallel single-precision operations using about 18% more hardware than a conventional double-precision MAF unit and with 9% increase in delay. To accommodate the simultaneous computation of two single-precision MAF operations, several basic modules of double-precision MAF unit are redesigned. They are either segmented by precision mode dependent multiplexers or attached by the duplicated hardware. The proposed MAF unit can be fully pipelined and the experimental results show that it is suitable for processors with floatingpoint unit (FPU). Libo Huang 0002, Li Shen 0007, Kui Dai, Zhiying Wang 0003 |
IEEE Symposium on Computer Arithmetic | 2 |