EDBT 2026 Demo / reviewers in the wild / expert
Peng Liu 0016
dblp:21/6121-16
· DBLP profile ↗
41ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0001-9107-6673ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comprehensive RISC- V Floating-Point Verification: Efficient Coverage Models and Constraint-Based Test GenerationabstractThe increasing complexity of processor architectures necessitates more rigorous functional verification. Floating-point operations, in particular, present significant challenges due to their extensive range of computational cases that require verification. This paper proposes a comprehensive approach for generating floating-point instruction sequences to enhance the verification of RISC-V. We introduce a constraint-based method for floating-point test generation and design efficient coverage models as input constraints for this process. The resulting representative floating-point tests are integrated with RISC-V instruction sequence generation through a memory-bound register update method. Experimental results demonstrate that our approach improves the functional coverage of RISC-V floating-point instruction sequences from 93.32% to 98.34%, while simultaneously reducing the number of required instructions by 66.67% compared to the Google RISCV-DV generator. Additionally, our method achieves more comprehensive coverage of floating-point types in instruction write-back data compared to RISCV-DV. Using the proposed approach, we successfully detect representative floating-point-related faults injected into the RISC-V processor CV32E40P, thereby demonstrating its effectiveness. Tianyao Lu, Anlin Liu, Bingjie Xia, Peng Liu 0016 |
DATE | 4 |
| 2025 | Efficient Multiple-Precision Floating-Point Multiply-Add Architecture for Deep Learning ApplicationsabstractTo fully exploit the potential of parallel computing while minimizing hardware costs, floating-point multiply-add (FMA) units that support multiple precisions are widely used in deep learning. However, achieving a balance between accuracy, performance, and hardware overhead remains a significant challenge. This paper presents a unified multiple-precision floating-point FMA architecture that supports four precision formats: single-precision (SP), Bfloat16 (BF16), TensorFloat32+ (TF32+), and INT8. The proposed architecture allows for the parallel execution of nine BF16 FMA operations, two TF32+ FMA operations, one SP FMA operation, or nine INT8 multiply-add (MA) operations. Through careful data format selection and an optimized architectural design, the architecture achieves multiplier utilization rates of 100%, 88.9%, 100%, and 100% for the four precision modes, respectively, with all multipliers operating at full bit width. Compared to state-of-the-art multiple-precision FMA designs, this architecture delivers over nine times the BF16 throughput while increasing the area by only 49%. The flexible data format configuration makes the proposed architecture suitable for a wide range of deep learning applications. Songtai Liang, Bingjie Xia, Wen Wang 0015, Peng Liu 0016 |
ISCAS | 5 |
| 2025 | Modern Hopfield-Heuristic Physical Computing Architecture for Diabetes PredictionabstractThe classical Hopfield network exhibits the ability to solve specific combinatorial optimization problems; however, its inherent storage capacity constrains its accuracy in multi-class classification tasks. In contrast, the modern Hopfield algorithm, which leverages exponential functions, significantly enhances storage capacity and accelerates convergence. This paper presents a novel CMOS-integrable physical computing system that exploits the expanded capacity of modern Hopfield networks. By optimizing network architecture and hidden node circuits, the proposed hardware system effectively stores multiple patterns and retrieves the one that best matches a given input state. Experimental results on a binary diabetes prediction dataset demonstrate that the proposed hardware system achieves 91% accuracy. The system’s CMOS compatibility makes it a promising candidate for embedded health monitoring applications. Guofan Jiang, Peng Liu 0016 |
ISCAS | 3 |
| 2025 | Enhancing Computer Organization Education: A Reformative Teaching ApproachabstractThis paper presents an innovative teaching approach to address key challenges in contemporary computer architecture and organization courses. Traditional courses often overemphasize conceptual definitions, leading to a surface-level understanding while neglecting in-depth principles analysis. As a result, students struggle to grasp design motivations and practical challenges. To overcome these limitations, we implement a comprehensive course reform, using the RISC-V instruction set architecture (ISA) as a central example. Our approach integrates in-depth principle analysis, historical context, active learning strategies, practical applications, and collaborative projects. A key highlight is the introduction of tailored student projects that focus on designing RISC-V extension instructions for cryptographic applications. This hands-on experience enhances students’ understanding of computer architecture principles while fostering innovation and problem-solving skills. Preliminary results show significant improvements in students’ ability theoretical knowledge and design practical systems. To further enhance the curriculum, we propose strengthening industry collaboration, adopting interdisciplinary methods, establishing feedback mechanisms, and implementing capstone projects. This comprehensive strategy aims to cultivate critical thinking skills and better prepare students for careers in the rapidly evolving technological landscape. Cansong Zhou, Guofan Jiang, Peng Liu 0016 |
ISCAS | 3 |
| 2024 | Flow-Aware Scheduling with Graph Neural Network Routing for Resource Efficiency in Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) emerges as a promising technique offering deterministic transmission through effective scheduling mechanisms. However, the coexistence of non-critical flows and time-sensitive flows in TSN presents a predicament as the expansion of solution space renders it arduous to solve. It also faces the dilemma of lacking adaptability to dynamic network environments due to limited information under increasingly complex requirements. To attack this problem, a flow-aware scheduling with graph neural network routing for resource efficiency in TSN, named DTRJ, is proposed. DTRJ utilizes a multi-agent deep reinforcement learning (DRL) method to dynamically allocate queue resources for flows, adjust link weights to enhance network load balancing, and jointly achieve resource allocation optimization. The flow-aware strategy and lightweight neural network of the scheduling agent, combined with the link-level features extraction ability of graph neural network-based routing agent, contribute to the effectiveness of DTRJ in improving network performance. Specifically, DTRJ introduces a parallel design for joint scheduling and routing tasks, enabling real-time decision-making in multi-agent systems to expedite computation. Experimental results show that DTRJ outperforms the state-of-the-art Tabu/DRL methods in resource utilization, while maintaining robust performance across various topologies and flow periods, achieving an average schedulability of 91%. Xinyi Hong, Yuhao Xi, Peng Liu 0016 |
IJCNN | 3 |
| 2024 | Mantissa-Aware Floating-Point Eight-Term Fused Dot Product UnitabstractFloating-point dot product is widely used in various applications. A conventional discrete construction of multipliers and adders often leads to accumulated errors and lower speed. This paper presents a floating-point eight-term fused dot product unit with mantissa-aware hardware design to attack these problems. In the proposed design, multiplication and addition of numbers are operated in a fused manner with exception controller. A pre-shift and post-shift combined scheme for mantissa alignment is utilized to eliminate the latency between significand multiplication and mantissa alignment. A one-path mantissa compress and addition structure is employed that can effectively reduce the footprint. An impact of internal mantissa datapath width on calculation accuracy is analyzed. Compared to the discrete method, the proposed design delivers a significant reduction up to 45.7%, 26.9%, and 33.0%, in terms of latency, area, and power, respectively. Wen Wang 0015, Bingjie Xia, Peng Liu 0016 |
ISCAS | 5 |
| 2024 | Enhancing Functional Verification with Dynamic Instruction Generation by Exploiting Processor Runtime StatesabstractAs the architectural complexity of processors increases dramatically, rigorous functional verification remains essential to ensure performance and immunity from design bugs. A pivotal yet challenging aspect of functional verification is the generation of test instruction streams that are not only highly effective in coverage but also compact enough to significantly reduce verification time. In this paper, we introduce DIG, a novel Dynamic Instruction Generator that leverages processor runtime architectural states through an instruction set simulator. In essence, DIG accesses processor runtime information to produce instruction streams with valid semantics, effectively avoiding illegal memory accesses and infinite loops. The quality of these instruction streams is further enhanced by incorporating both intra-instruction and inter-instruction test knowledge. The effectiveness of DIG is demonstrated through a case study involving Western Digital’s open-source RISC-V core, VeeR EH2. Our experimental results indicate that DIG reduces the number of test instructions by as much as 62.50% and 86.11%, compared to the state-of-the-art random instruction generators, RISC-V DV and RISC-V Torture, respectively, while simultaneously achieving superior functional coverage. Anlin Liu, Tianyao Lu, Yuhao Xi, Yangfan Liu, Peng Liu 0016 |
ITC | 5 |
| 2024 | Resource-Aware Online Traffic Scheduling for Time-Sensitive NetworkingabstractTraditional Ethernet technology struggles to meet the ever-increasing demands of modern cyber-physical systems in terms of transmission bandwidth and network distribution. Time-sensitive networking (TSN) has emerged to address these challenges by providing mechanisms for accurate clock synchronization, intelligent network configuration, and high-precision traffic scheduling. However, resource-limited TSN switches restrict the length of internal queues, necessitating improved resource utilization while ensuring successful scheduling. This article introduces an online traffic scheduling framework using deep reinforcement learning (DRL) to optimize scheduling and resource allocation in TSN. By incorporating the resource allocation quality metric, known as bandwidth satisfaction, our goal is to achieve load balancing in the network's switch queues while ensuring deterministic traffic transmission. By integrating both network protocol constraints and user-defined constraints, the proposed DDTA-CNN algorithm leverages a convolutional neural network (CNN) to extract flow features and allocate sending time slots, maximizing queue resource utilization. Evaluations under various TSN settings demonstrate that DDTA-CNN significantly outperforms state-of-the-art algorithms, SMT and Tabu, by 10× and 31.8% in terms of resource allocation quality, and by 23.5% and 13.6% in terms of schedulable flows. Our framework improves network resource utilization, ensures real-time transmission, and achieves higher scheduling success rates, addressing the critical need for efficient and scalable traffic scheduling in TSN environments. Xinyi Hong, Yuhao Xi, Peng Liu 0016 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Adaptive Caching Policies for Chiplet Systems Based on Reinforcement LearningabstractChiplet packaging becomes a popular solution to integrate more hardware components. However, shared memory access across chiplets suffers from high miss penalty due to long route latency and low bandwidth of inter-chiplet interconnects. We observe that the aggregated last-level cache (LLC) miss penalty takes approximately 35% of time on data access, and that the miss is dominated by coherence miss as a result of shared reads and writes from other LLCs. To address this problem, we propose a caching manager which speculatively enforces (or discards) LLC caching via online reinforcement learning. On every invalidated cacheline, the caching manager receives the cacheline access features, evaluates the current caching policy, and makes the next caching policy adaptively. Experimental evaluation justifies that the caching manager can reduce more than 10% coherence miss and offers a 3% speedup against a state-of-the-art cache coherence protocol. Chongyi Yang, Xiaohang Wang 0001, Peng Liu 0016 |
ISCAS | 4 |
| 2023 | RUPA: A High Performance, Energy Efficient Accelerator for Rule-Based Password Generation in Heterogenous Password Recovery SystemabstractThere has been a growing demand for energy efficient and high performance password recovery systems. As password generation and password validation are two integral components of any password recovery system, the former yet lags behind the latter in performance particularly for the case when the popular rule-based password generation method is applied to the heterogeneous CPU-FPGA system. In this paper, we thus present a high performance, energy efficient accelerator to speed up the rule functions in rule-based password generation. Dubbed RUPA, this proposed accelerator explores previously undiscovered computational features and memory access patterns for processing the rule functions. Specially, we show that the rule functions can be mapped to three distinct groups according to their character dependency graphs. Correspondingly, three kinds of datapath units, referred to as rule logic units, are created, and the rule functions from the same group will be processed in their shared rule logic unit. Compared with the state-of-the-art password recovery system built upon a CPU-GPU platform, the FPGA-based RUPA system achieves 5.3x speed improvement and is 33.1x more energy efficient. If RUPA is integrated into the popular password recovery tool John the Ripper (JtR), JtR's rule-based attack performance can soar by more than 48.7x. Peng Liu 0016, Yingtao Jiang |
IEEE Trans. Computers | 2 |
| 2022 | IMSC: Instruction set architecture monitor and secure cache for protecting processor systems from undocumented instructionsabstractAbstract A secure processor requires that no secret, undocumented instructions be executed. Unfortunately, as today's processor design and supply chain are increasingly complex, undocumented instructions that can execute some specific functions can still be secretly introduced into the processor system as flaws or vulnerabilities. To address this problem that may cause potentially serious security breaches, the instruction set architecture (ISA) monitor and secure cache (IMSC) is proposed. As a lightweight solution, IMSC employs an ISA monitor to discover and correct any potential threats imposed by undocumented instructions, and it relies on a secure cache to ensure the credibility of the system. The authors’ case studies have confirmed that IMSC can effectively protect a processor system from being exploited by undocumented instructions and thus provide a trustworthy computing environment, all at low hardware and run‐time costs. Yuze Wang 0001, Peng Liu 0016, Yingtao Jiang |
IET Inf. Secur. | 2 |
| 2022 | On a Consistency Testing Model and Strategy for Revealing RISC Processor's Dark Instructions and VulnerabilitiesabstractOne major security vulnerability of a microprocessor can be attributed to its underlying instruction set architecture (ISA). Generally, it is required that no secret instructions be included in the ISA or implemented in the processor micro-architecture. Such a requirement is particularly important for the reduced instruction set computing (RISC) processors that are widely used nowadays, and applying the proposed consistency testing approach is poised to ensure this requirement is met. Capable of revealing any possible dark instructions (i.e., executable instructions but without clear definitions of their behavior) in RISC processors, a consistency test comes in three phases. During the generation phase, based on the instruction set encoding rules, all the undefined instructions are generated. Even with a smaller test space, this step guarantees the test coverage needed to reveal all the dark instructions that may exist. In the next phase, all the undefined instructions obtained from the previous phase are executed on the processor under test, following a set of persistence strategies; any instruction exhibiting unusual execution result will be deemed suspicious and recorded so. During the last analysis phase, each of those recorded suspicious instructions will be checked and analyzed to decide whether it truly constitutes a dark instruction. We have applied the proposed testing model and strategy to several RISC processors and found that all of them have a few dark instructions previously unknown. The potential vulnerabilities of these processors introduced by their respective dark instructions have thus been evaluated and exposed. Yuze Wang 0001, Peng Liu 0016, Xiaohang Wang 0001, Yingtao Jiang |
IEEE Trans. Computers | 2 |
| 2019 | Energy-Efficient RAR3 Password Recovery with Dual-Granularity Data Path StrategyabstractPassword recovery tools are used to recover lost passwords and regain access to precious data. Due to the extremely large time and energy consumption of password recovery, efficient hardware accelerators are demanded to accelerate the recovery process. However, simple and regular data interconnect paths between data sources and non-blocking hash pipelines are hard to construct for Roshal ARchive version 3 (RAR3) algorithm based on field programmable gate array (FPGA) devices. The difficulty comes from the fact that the message format of the hash pipeline inputs vary with the password length and the secure hash algorithm 1 (SHA-1) iteration phase. To attack this problem, a dual-granularity data path adjustment strategy is proposed to eliminate the randomness of message block formats caused by the irregularity of password length and to efficiently schedule the data through the regular data interconnect paths. Experimental results show that the proposed hardware accelerator for RAR3 password recovery is 3.3 × more energy-efficient than a state-of-the-art implementation Hashcat on NVIDIA GTX 1060 GPU. Qingyuan Ding, Shunbin Li, Peng Liu 0016 |
ISCAS | 4 |
| 2019 | Hardware Trojans Detection at Register Transfer Level Based on Machine LearningabstractTo accurately detect Hardware Trojans in integrated circuits design process, a machine-learning-based detection method at the register transfer level (RTL) is proposed. In this method, circuit features are extracted from the RTL source codes and a training database is built using circuits in a Hardware Trojans library. The training database is used to train an efficient detection model based on the gradient boosting algorithm. In order to expand the Hardware Trojans library for detecting new types of Hardware Trojans and update the detection model in time, a server-client mechanism is used. The proposed method can achieve 100% true positive rate and 89% true negative rate, on average, based on the benchmark from Trust-Hub. Yuze Wang 0001, Peng Liu 0016 |
ISCAS | 3 |
| 2019 | Ensemble-Learning-Based Hardware Trojans Detection Method by Detecting the Trigger NetsabstractWith the globalization of integrated circuit (IC) design and manufacturing, malicious third-party vendors can easily insert hardware Trojans into their intellect property (IP) cores during IC design phase, threatening the security of IC systems. It is strongly required to develop hardware-Trojan detection methods especially for the IC design phase. As the particularity of Trigger nets in Trojan circuits, in this paper, we propose an ensemble-learning-based hardware-Trojan detection method by detecting the Trigger nets at the gate level. We extract the Trigger-net features for each net from known netlists and use the ensemble learning method to train two detection models according to the Trojan types. The detection models are used to identify suspicious Trigger nets in an unknown detected netlist and give results of suspiciousness values for each detected net. By flagging the top n% suspicious nets of each detection model as the suspicious Trigger nets based on the suspiciousness values, the proposed method can achieve, on average, 88% true positive rate, 90% true negative rate, and 90% Accuracy. Yuze Wang 0001, Peng Liu 0016 |
ISCAS | 4 |
| 2019 | An Energy-Efficient Accelerator Based on Hybrid CPU-FPGA Devices for Password RecoveryabstractPassword recovery tools are needed to recover lost and forgotten passwords so as to regain access to valuable information. As the process of password recovery can be extremely compute-intensive, hardware accelerators are often needed to expedite the recovery process. This paper thus presents a high performance, energy-efficient accelerator built upon modern hybrid CPU-FPGA SoC devices. The proposed password recovery accelerator relies on the development of a set of intellectual property (IP) cores for implementing variety of encryption algorithms with vastly different characteristics and complexities. To keep the resource requirements of each IP core running on a resource-strapped FPGA to the minimum, while achieving the highest throughput possible, the most performance critical computational hash functions are mapped to the FPGA with two specific optimization techniques, namely the fixed message padding for hashing and loop transformation for deep pipelining. The proposed password recovery accelerator implements a non-blocking deep pipeline design that does not incur any data and structural hazards, which is made possible by applying a task scheduling scheme through the use of block RAMs. Synchronization between tasks that are mapped to run separately on CPU and FPGA is achieved through task reordering and a communication protocol for maximum parallelism and low overhead. The proposed design is evaluated on Xilinx XC7Z030-3 device, and it is compared much favorably with other known implementations. The proposed hardware accelerator design is found 12.5 and 3.1 times more resource-efficient than the pure FPGA-based password recovery accelerators for TrueCrypt and WPA-2, respectively. The proposed implementation also shows more than 200 percent improvement in energy efficiency over a state-of-the-art implementation on NVIDIA GTX 750 Ti GPU. Peng Liu 0016, Shunbin Li, Qingyuan Ding |
IEEE Trans. Computers | 1 |
| 2017 | Adaptive Coherence Granularity for Multi-Socket SystemsabstractThe emerging large-scale multi-socket systems make the need for more sophisticated large-scale coherence management of necessity. Directory-based coherence has been an ad hoc solution and a clear candidate for large-scale shared-memory systems. A vanilla directory design, however, suffers from inefficient use of storage to keep coherence metadata, resulting in a high storage overhead for large-scale systems. In this paper, we propose a dynamic multi-grain directory for large multi-socket systems. The idea is to track coherence of regions of different sizes which requires storing much less information in the directory than having a directory entry per each data block. It dynamically refines granularity according to the application phase and therefore tracks coherence information for regions of varying sizes. The results show that the proposal allows to reduce the directory storage by an order of magnitude, while the loss of precision does not cause performance penalty. The paper demonstrates that different applications and different application phases have different requirements for the region size. Performance results are compared against two state-of-the-art multi-grain directories and it is the one that obtains the better results. Peng Liu 0016, Xingcheng Hua |
IEEE Trans. Computers | 1 |
| 2017 | An Adaptive PAM-4 Analog Equalizer With Boosting-State Detection in the Time DomainabstractThis paper introduces an improved adaptive analog equalizer that is required in high speed serial receivers using four-level pulse amplitude modulation signaling. By performing boosting-state detection in the time domain, the proposed adaptive analog equalizer can effectively overcome a serious problem that the received signal's eye-opening tends to be compromised by the convergence accuracy of the adaptive control loop. To suppress the pattern-dependent jitters (PDJs), an inductor-less, cross-stage feedback structure is employed in the proposed analog equalizer to help broaden its effective tuning bandwidth. Multichannel simulations have confirmed that the proposed equalizer is able to achieve a 42% improvement in eye-height opening when compared with the equalizers employing popular spectrum-comparing schemes. Trellis diagram analyses under different date rates have revealed that the bandwidth of the proposed equalizer can be extended by as much as 60%, thus effectively bringing the PDJs from 44% down to 27.5%. Shunbin Li, Yingtao Jiang, Peng Liu 0016 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Thread-Aware Adaptive Prefetcher on Multicore Systems: Improving the Performance for Multithreaded WorkloadsabstractMost processors employ hardware data prefetching techniques to hide memory access latencies. However, the prefetching requests from different threads on a multicore processor can cause severe interference with prefetching and/or demand requests of others. The data prefetching can lead to significant performance degradation due to shared resource contention on shared memory multicore systems. This article proposes a thread-aware data prefetching mechanism based on low-overhead runtime information to tune prefetching modes and aggressiveness, mitigating the resource contention in the memory system. Our solution has three new components: (1) a self-tuning prefetcher that uses runtime feedback to dynamically adjust data prefetching modes and arguments of each thread, (2) a filtering mechanism that informs the hardware about which prefetching request can cause shared data invalidation and should be discarded, and (3) a limiter thread acceleration mechanism to estimate and accelerate the critical thread which has the longest completion time in the parallel region of execution. On a set of multithreaded parallel benchmarks, our thread-aware data prefetching mechanism improves the overall performance of 64-core system by 13% over a multimode prefetch baseline system with two-level cache organization and conventional modified, exclusive, shared, and invalid-based directory coherence protocol. We compare our approach with the feedback directed prefetching technique and find that it provides 9% performance improvement on multicore systems, while saving the memory bandwidth consumption. Peng Liu 0016, Jiyang Yu, Michael C. Huang 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | Building Expressive and Area-Efficient Directories with Hybrid Representation and Adaptive Multi-Granular TrackingabstractMainstream chip multiprocessors already include a significant number of cores that make straightforward snooping-based cache coherence less appropriate. Further increase in core count will almost certainly require more sophisticated tracking of data sharing to minimize unnecessary messages and cache snooping. Directory-based coherence has been the standard solution for large-scale shared-memory multiprocessors and is a clear candidate for on-chip coherence maintenance. A vanilla directory design, however, suffers from inefficient use of storage to keep coherence metadata. The result is a high storage overhead for larger scales. Reducing this overhead leads to saving of resources that can be redeployed for other purposes. In this paper, we exploit familiar characteristics of coherence metadata, but with novel angles and propose two practical techniques to increase the expressiveness of directory entries, particularly for chip-multiprocessors. First, it is well known that the vast majority of cache lines have a small number of sharers. We exploit a related fact with a subtle but important difference: that a significant portion of directory entries only need to track one node. We can thus use a hybrid representation of sharers list for the directory. Second, contiguous memory regions often share the same coherence characteristics and can be tracked by a single entry. We propose an adaptive multi-granular mechanism that does not rely on any profiling, compiler, or operating system support to identify such regions. Moreover, it allows co-existence of line and region entries in the same locations, thus making regions more applicable. We show that both techniques improve the expressiveness of directory entries, and, when combined, can reduce directory storage by more than an order of magnitude with negligible loss of precision. Peng Liu 0016, Michael C. Huang 0001, Guofan Jiang |
IEEE Trans. Computers | 1 |
| 2015 | Exploiting Transmission Lines on Heterogeneous Networks-on-Chip to Improve the Adaptivity and Efficiency of Cache CoherenceabstractEmerging heterogeneous interconnects have shown lower latency and higher throughput, which can improve the efficiency of communication and create new opportunities for memory system designs. In this paper, transmission lines are employed as a latency-optimized network and combined with a packet-switched network to create heterogeneous interconnects improving the efficiencies of on-chip communication and cache coherence. We take advantage of this heterogeneous interconnect design, and keep cache coherence adaptively based on data locality. Different type of messages are adaptively directed through selected medium of the heterogeneous interconnects to enhance cache coherence effectiveness. Compared with a state-of-the-art coherence mechanism, the proposed technique can reduce the coherence overhead by 24%, reduce the network energy consumption by 35%, and improve the system performance by 25% on a 64-core system. Peng Liu 0016, Michael C. Huang 0001, Xianghui Xie 0001 |
NOCS | 2 |
| 2015 | Physical-based modeling and fast simulation of wireline linksabstractTo efficiently exploit the performance of wireline channels and alleviate the computation time due to the large amounts of simulation, this paper proposes a physical-based modeling mechanism for the performance studies of electrical links on multilayered printed circuit boards (PCBs). Our mechanism includes: 1) using a physical-based method to achieve high-speed and high precision simulation; and 2) combining configurable parameters and intermediate data of simulation to simplify evaluation process and support parameters sweeping. Experimental results show a good correlation with full-wave method up to 40 GHz and a significant acceleration in computation time of at least two orders of magnitude. Cooperating with system simulator, our approach also gives a good assistance in PCB channel design for high-speed interconnect systems. Peng Liu 0016 |
VLSI-SoC | 2 |
| 2014 | A Thread-Aware Adaptive Data PrefetcherabstractMost processors employ hardware data prefetching to hide memory access latencies. However the prefetching requests from different threads on a multi-core processor can cause severe interference with prefetching and/or demand requests of others. The data prefetching can lead to significant performance degradation due to shared resource contention on shared memory multi-core systems. This paper proposes a thread-aware data prefetching mechanism based on low-overhead run-time information to tune prefetching modes and aggressiveness, mitigating the resource contention in the memory system. Our solution has two new components: 1) a filtering mechanism that informs the hardware about which prefetching requests can cause shared data invalidation and should be discarded, and 2) a self-tuning prefetcher that uses run-time feedback to adjust each thread's data prefetching mode and arguments. On a set of parallel benchmarks, our thread-aware data prefetching mechanisms improve the overall performance of 64-core system by 11% and reduce the energy-delay product by 13% over a multi-mode prefetch baseline system with a two level cache organization and a conventional MESI-based directory coherence protocol. We compare our approach to the feedback directed prefetching (FDP) technique and find that it provides better performance on multi-core systems, while reducing the energy delay product. Jiyang Yu, Peng Liu 0016 |
ICCD | 2 |
| 2014 | A novel signaling technique for high-speed wireline backplane transceiver: Four phase-shifted sinusoid symbol (PSS-4)abstractThis paper proposes a novel four phase-shifted sinusoid symbol (PSS-4) signaling technique to relief high-speed wireline backplane transceiver to reduce large intersymbol interference. The four-level pulse amplitude modulation (PAM-4) and other existing amplitude modulation techniques have been widely used to simplify transceiver equalization design for highly dispersive channels. Unfortunately, PAM-4 and other methods reduce the symbol rate at the expense of signal-to-noise ratio (SNR) and thus limit the bit-error-rate (BER). The proposed PSS-4 signaling avoids large SNR degradation by using four phase-shifted symbols to transmit two-bit data. The experimental results show that with sufficient equalization, our PSS-4 signaling can achieve over 2 dB larger SNR than the conventional non-return-to-zero (NRZ), Duobinary, and PAM-4 signaling techniques. In addition, for a target BER of 1E-12, the proposed PSS-4 signaling has the largest sampling range comparing against existing backplane signaling techniques. Furthermore, compared with NRZ-based backplane transceiver, PSS-4 signaling reduces the equalizer power consumption by 24%. Kejun Wu, Peng Liu 0016, Qiaoyan Yu |
ISCAS | 2 |
| 2014 | A new fault injection method for evaluation of combining SEU and SET effects on circuit reliabilityabstractWe propose a new dual-level fault injection method for evaluating combination effect of single event upsets (SEUs) and single event transients (SETs). The proposed interaction method allows collaborative simulation on register-transfer level (RTL) and gate level. Conventional fault injection methods or fault model techniques typically aim at SEUs or SETs, rather than the combination of SETs and SEUs. As a logic depth and clock period decrease, SEUs and SET are likely to co-exist, which further challenges circuit reliability. To facilitate the investigation of advanced SEU and SET management methods, our fault injection method considers both SETs and SEUs. We apply the proposed method to two ITC'99 benchmark circuits to analyze the mutual masking effect between SETs and SEUs. Simulations performed on the two circuits show that SET duration time is the dominant factor affecting the mutual masking effect. If SEU duration time changes (but not beyond one cycle), the maximum masked error ratio is up to five times the minimum masked error ratio. We also observed that doubling clock frequency results in the average masked error ratio varying from 3% to 10%. Kejun Wu, Hoda Pahlevanzadeh, Peng Liu 0016, Qiaoyan Yu |
ISCAS | 3 |
| 2014 | DEAM: Decoupled, Expressive, Area-Efficient Metadata Cache
Peng Liu 0016, Michael C. Huang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | On self-tuning networks-on-chip for dynamic network-flow dominance adaptationabstractModern network-on-chip (NoC) systems are required to handle complex runtime traffic patterns and unprecedented applications. Data traffics of these applications are difficult to fully comprehend at design time so as to optimize the network design. However, it has been discovered that the majority of dataflows in a network are dominated by less than 10% of the specific pathways. In this article, we introduce a method that is capable of identifying critical pathways in a network at runtime and can then dynamically reconfigure the network to optimize for network performance subject to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerging dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance dataflows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulations. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016, Masoud Daneshtalab, Maurizio Palesi, Terrence S. T. Mak |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | Building expressive, area-efficient coherence directoriesabstractMainstream chip multiprocessors already include a significant number of cores that make straightforward snooping-based cache coherence less appropriate. Further increase in core count will almost certainly require more sophisticated tracking of data sharing to minimize unnecessary messages and cache snooping. Directory-based coherence has been the standard solution for large-scale shared-memory multiprocessors and is a clear candidate for on-chip coherence maintenance. A vanilla directory design, however, suffers from inefficient use of storage to keep coherence metadata. The result is a high storage overhead for larger scales. Reducing this overhead leads to saving of resources that can be redeployed for other purposes. In this paper, we exploit familiar characteristics of coherence metadata, but with novel angles and propose two practical techniques to increase the expressiveness of directory entries, particularly for chip-multiprocessors. First, it is well known that the vast majority of cache lines have a small number of sharers. We exploit a related fact with a subtle but important difference: that a significant portion of directory entries only need to track one node. We can thus use a hybrid representation of sharers list for the whole set. Second, contiguous memory regions often share the same coherence characteristics and can be tracked by a single entry. We propose a multi-granular mechanism that does not rely on any profiling, compiler, or OS support to identify such regions. Moreover, it allows co-existence of line and region entries in the same locations, thus making regions more applicable. We show that both techniques improve the expressiveness of directory entries, and, when combined, can reduce directory storage by more than an order of magnitude with negligible loss of precision. Peng Liu 0016, Michael C. Huang 0001, Guofan Jiang |
PACT | 2 |
| 2013 | A novel energy-efficient serializer design method for gigascale systemsabstractSerial communication facilitates the high-speed communication in gigascale systems. Serializer designs typically use the current-mode logic to achieve high speed at the cost of large power consumption. For the latches in the serializer, the power-hungry current-mode logic is replaced with differential cascaded pass-gate to reduce the power and delay. For the selectors in the serializer, the conventional differential cascode voltage switch is modified with pass-gate logic by replacing a PMOS load with a resistor load and adding an inductive peaking structure. Simulation results show that the proposed method reduces the power-delay-product by up to 70%, compared to the conventional current-mode-logic-based serializer. Kejun Wu, Peng Liu 0016, Qiaoyan Yu |
ISCAS | 2 |
| 2013 | Scalable-Grain Pipeline Parallelization Method for Multi-core Systems
Peng Liu 0016, Chunming Huang, Yang Geng, Mei Yang 0001 |
NPC | 1 |
| 2013 | Energy Efficient Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Peng Liu 0016, Mei Yang 0001, Maurizio Palesi, Yingtao Jiang, Michael C. Huang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2013 | Efficient multicast schemes for 3-D Networks-on-Chip
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Maurizio Palesi, Peng Liu 0016, Terrence S. T. Mak, Nader Bagherzadeh |
J. Syst. Archit. | 5 |
| 2013 | Avoiding request-request type message-dependent deadlocks in networks-on-chips
Xiaohang Wang 0001, Peng Liu 0016, Mei Yang 0001, Yingtao Jiang |
Parallel Comput. | 2 |
| 2013 | An efficient protocol with synchronization accelerator for multi-processor embedded systems
Jiyang Yu, Peng Liu 0016, Chunming Huang, Yingtao Jiang, Qingdong Yao |
Parallel Comput. | 2 |
| 2011 | A design space exploration of transmission-line links for on-chip interconnect
Aaron Carpenter, Jianyun Hu, Michael C. Huang 0001, Hui Wu 0007, Peng Liu 0016 |
ISLPED | 5 |
| 2011 | An Efficient Architectural Design of Hardware Interface for Heterogeneous Multi-core System
Xiongli Gu, Xiamin Wu, Chunming Huang, Peng Liu 0016 |
NPC | 5 |
| 2011 | Power-Aware Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016 |
NPC | 6 |
| 2011 | Low latency and energy efficient multicasting schemes for 3D NoC-based SoCsabstractIn this paper, two topology oriented multicast routing algorithms, MXYZ and AL+XYZ, are proposed to support multicasting in 3D Networks on Chips (NoCs). In specific, MXYZ is a dimension order multicast routing algorithm that targets 3D NoC systems built upon regular topologies, while AL+XYZ is applicable to NoCs with irregular topologies. If the output channel found by MXYZ is not available (i.e. in the same region), an alternative output channel is used to forward/replicate the packets in AL+XYZ. MXYZ is evaluated against a path based regular topology oriented multicast routing and AL+XYZ against an irregular region oriented multiple unicast routing algorithm. Our experimental results have demonstrated that the proposed MXYZ and AL+XYZ schemes have lower latency and energy consumption than the conventional path based multicast routing and the multiple unicast routing algorithms, meriting them to be more suitable for supporting multicasting in 3D NoC systems. Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016 |
VLSI-SoC | 6 |
| 2010 | An intra-chip free-space optical interconnectabstractContinued device scaling enables microprocessors and other systems-on-chip (SoCs) to increase their performance, functionality, and hence, complexity. Simultaneously, relentless scaling, if uncompensated, degrades the performance and signal integrity of on-chip metal interconnects. These systems have therefore become increasingly communications-limited. The communications-centric nature of future high performance computing devices demands a fundamental change in intra- and inter-chip interconnect technologies. Alok Garg, Berkehan Ciftcioglu, Jianyun Hu, Ioannis Savidis, Rebecca Berman, Peng Liu 0016, Michael C. Huang 0001, Hui Wu 0007, Eby G. Friedman, Gary Wicks, Duncan Moore |
ISCA | 9 |
| 2010 | A power-aware mapping approach to map IP cores onto NoCs under bandwidth and latency constraintsabstractIn this article, we investigate the Intellectual Property (IP) mapping problem that maps a given set of IP cores onto the tiles of a mesh-based Network-on-Chip (NoC) architecture such that the power consumption due to intercore communications is minimized. This IP mapping problem is considered under both bandwidth and latency constraints as imposed by the applications and the on-chip network infrastructure. By examining various applications' communication characteristics extracted from their respective communication trace graphs, two distinguishable connectivity templates are realized: the graphs with tightly coupled vertices and those with distributed vertices. These two templates are formally defined in this article, and different mapping heuristics are subsequently developed to map them. In general, tightly coupled vertices are mapped onto tiles that are physically close to each other while the distributed vertices are mapped following a graph partition scheme. Experimental results on both random and multimedia benchmarks have confirmed that the proposed template-based mapping algorithm achieves an average of 15% power savings as compared with MOCA, a fast greedy-based mapping algorithm. Compared with a branch-and-bound--based mapping algorithm, which produces near optimal results but incurs an extremely high computation cost, the proposed algorithm, due to its polynomial runtime complexity, can generate the results of almost the same quality with much less CPU time. As the on-chip network size increases, the superiority of the proposed algorithm becomes more evident. Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016 |
ACM Trans. Archit. Code Optim. | 4 |
| 2002 | Hardware/software codesign for HDTV source decoder on system level
Peng Liu 0016, Qingdong Yao, Weijian Yang |
VCIP | 1 |