Jianhua Li 0003

dblp:93/3389-3 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-7086-7625ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Fabdb: a low-latency fault-tolerant architecture based on dynamic bypass for network-on-chip
Zhuoxuan Ji, Jianhua Li 0003, Huaguo Liang
J. Supercomput.3
2024 F-Bypass: A Low-Power Network-on-Chip Design Utilizing Bypass to Improve Network Connectivity
abstract
With the development of transistor feature size to nanometer level, static power consumption has gradually become the main factor affecting the overall power consumption of network-on-chip (NoC). Power gating is an effective technology to reduce static power consumption, but it also brings new challenges, such as BET violation, wake-up latency and network connectivity. Therefore, a power gating method is needed to improve NoC performance and reduce static power consumption. This article proposes a low-power bypass method, namely Forwarding bypass (F-Bypass). First, F-Bypass adds bypass paths between all input and output ports and the network interface (NI) and connects the pop-up port and injection port in NI through the bypass path. When the router is powered off, F-Bypass performs wake-up-free packet transmission, which reduces the break-even time (BET) violation and cumulative wake-up latency while ensuring network connectivity. Secondly, this article adds the modified VC state table to NI so that the power-off router can perform normal traffic control. Finally, a new wake-up criterion is proposed, which can effectively avoid the frequent wake-up of power-off routers, and the detailed hardware implementation of F-Bypass is provided. The simulation results under integrated traffic load show that compared with the traditional scheme, the delay of F-Bypass is reduced by 2.2%, the throughput is increased by 13.1%, and the total static power consumption is reduced by 75.2%. Key performance indicators are superior to other solutions, and the increased area cost is moderate.
Shuaijie Yuan, Jianhua Li 0003, Huaguo Liang
ACM J. Emerg. Technol. Comput. Syst.3
2023 A transparent virtual channel power gating method for on-chip network routers
Wu Zhou 0007, Jianhua Li 0003
Integr.3
2023 Transit ring: bubble flow control for eliminating inter-ring communication congestion
Chenglong Sun, Qi Wang 0027, Jianhua Li 0003
J. Supercomput.5
2022 WiGRUNT: WiFi-Enabled Gesture Recognition Using Dual-Attention Network
abstract
Gestures constitute an important form of nonverbal communication where bodily actions are used for delivering messages alone or in parallel with spoken words. Recently, there exists an emerging trend of WiFi sensing-enabled gesture recognition due to its inherent merits like remote sensing, non-line-of-sight covering, and privacy-friendly. However, current WiFi-based approaches mainly reply on domain-specific training since they don’t know “where to look” and “when to look.” To this end, we propose WiGRUNT, a WiFi-enabled gesture recognition system using dual-attention network, to mimic how a keen human being intercepting a gesture regardless of the environment variations. The key insight is to train the network to dynamically focus on the domain-independent features of a gesture on the WiFi channel state information via a spatial-temporal dual-attention mechanism. WiGRUNT roots in a deep residual network (ResNet) backbone to evaluate the importance of spatial-temporal clues and exploit their inbuilt sequential correlations for fine-grained gesture recognition. We evaluate WiGRUNT on the open Widar3 dataset and show that it significantly outperforms its state-of-the-art rivals by achieving the best-ever performance in-domain or cross-domain.
Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Meng Wang 0001, Huan Yan 0005, Yusheng Ji, Zhi Liu 0002, Jianhua Li 0003, Mianxiong Dong
IEEE Trans. Hum. Mach. Syst.8
2019 CPCA: An efficient wireless routing algorithm in WiNoC for cross path congestion awareness
Jianhua Li 0003, Chenglong Sun, Huaguo Liang, Gaoming Du
Integr.3
2017 Thread Criticality Assisted Replication and Migration for Chip Multiprocessor Caches
abstract
Non-Uniform Cache Architecture (NUCA) is a viable solution to mitigate the problem of large on-chip wire delay due to the rapid increase in the cache capacity of chip multiprocessors (CMPs). Through partitioning the last-level cache (LLC) into smaller banks connected by on-chip network, the access latency will exhibit non-uniform distribution. Various works have well explored the NUCA design, including block migration, block replication and block searching. However, all of the previous mechanisms designed for NUCA are thread-oblivious when multi-threaded applications are deployed on CMP systems. Due to the interference on shared resources, threads often demonstrate unbalanced progress wherein the lagging threads with slow progress are more critical to overall performance. In this paper, we propose a novel NUCA design called thread Criticality Assisted Replication and Migration (CARM). CARM exploits the runtime thread criticality information as hints to adjust the block replication and migration in NUCA. Specifically, CARM aims at boosting parallel application execution through prioritizing block replication and migration for critical threads. Full-system experimental results show that CARM reduces the execution time of a set of PARSEC workloads by 13.7 and 6.8 percent on average compared with the tradition D-NUCA and Re-NUCA respectively. Moreover, CARM also consumes much less energy compared with the evaluated schemes.
Jianhua Li 0003, Minming Li, Chun Jason Xue, Fanfan Shen
IEEE Trans. Computers1
2015 Compiler-Assisted Refresh Minimization for Volatile STT-RAM Cache
abstract
Spin-transfer torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features such as high storage density and ultra low leakage power. However, long write latency and high write energy are the two challenges for STT-RAM. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data losses resulting from volatility, refresh schemes have been proposed. However, refresh operations consume additional overhead. In this paper, we propose to significantly reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed to further reduce the number of refreshes. Experimental results show that, on average, the proposed methods can reduce the number of refresh operations by 84.2 percent, and reduce the dynamic energy consumption by 38.0 percent for volatile STT-RAM caches while incurring only 4.1 percent performance degradation.
Qing'an Li, Yanxiang He, Jianhua Li 0003, Liang Shi 0001, Yiran Chen 0001, Chun Jason Xue
IEEE Trans. Computers3
2014 Dual partitioning multicasting for high-performance on-chip networks
Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yinlong Xu 0001
J. Parallel Distributed Comput.1
2014 Thread Progress Aware Coherence Adaption for Hybrid Cache Coherence Protocols
abstract
For chip multiprocessor systems (CMPs), the interference on shared resources such as on-chip caches typically leads to unbalanced progress among threads. Because of the inherent synchronization primitives, such as barriers and locks, cores running fast threads have to waste precious cycles to wait for cores with slow progress, which leads to performance and energy inefficiency. For the purpose of improving performance and reducing energy consumption, this paper proposes to adapt the cache coherence policy for threads according to their delay-tolerant levels. Specifically, this paper proposes Thread progrEss Aware Coherence Adaption (TEACA) which utilizes the thread progress information as hints for coherence adaption. TEACA dynamically utilize the memory system statistics to estimate the progress of threads. Based on the estimated thread progress information, TEACA categorizes threads into leader threads and laggard threads. The thread categorization decisions are then leveraged for efficient coherence adaption on CMP systems supporting hybrid coherence protocols. Experimental results show that, on a 64-core CMP system, TEACA outperforms directory protocol in application execution time and a recently proposed hybrid protocol in both application execution time and energy dissipation.
Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2014 WCET-Aware Re-Scheduling Register Allocation for Real-Time Embedded Systems With Clustered VLIW Architecture
abstract
Worst-case execution time (WCET) is one of the most important metric in real-time embedded system design. For embedded systems with clustered very long instruction word (VLIW) architecture, register allocation, instruction scheduling, and cluster assignment are three key activities for code optimization, which have profound impact on WCET. At the same time, these three activities exhibit a phase ordering problem, i.e., independently performing register allocation, scheduling, and cluster assignment could have a negative effect on the other phases, thereby generating sub-optimal compiled code. In this paper, a compiler level optimization, namely WCET-aware re-scheduling register allocation, is proposed to achieve WCET minimization for real-time embedded systems with clustered VLIW architecture. The novelty of the proposed approach is that the effects of register allocation, instruction scheduling, and cluster assignment on the quality of generated code are taken into account for WCET minimization. These three compilation processes are integrated into a single phase to obtain a balanced result. The proposed technique is implemented in Trimaran 4.0. The experimental results show that the proposed technique can reduce WCET effectively, by 34% on average.
Yazhi Huang, Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Compiler-Assisted STT-RAM-Based Hybrid Cache for Energy Efficient Embedded Systems
abstract
Hybrid caches consisting of static RAM (SRAM) and spin-torque transfer (STT)-RAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most of the management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques involve additional access operations, and thus lead to extra overheads. In this paper, we propose two compilation-based approaches to improve the energy efficiency and performance of STT-RAM-based hybrid cache by reducing the migration overheads. The first approach, migration-aware data layout, is proposed to reduce the migrations by rearranging the data layout. The second approach, migration-aware cache locking, is proposed to reduce the migrations by locking migration-intensive memory blocks into SRAM part of hybrid cache. Furthermore, experiments show that these two methods can be combined to reduce more migrations. The reduction of migration overheads can improve the energy efficiency and performance of STT-RAM-based hybrid cache. Experimental results show that, combining these two methods, on average, the number of write operations on STT-RAM is reduced by 17.6%, the number of migrations is reduced by 38.9%, the total dynamic energy is reduced by 15.6%, and the total access latency is reduced by 13.8%.
Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Yanxiang He
IEEE Trans. Very Large Scale Integr. Syst.2
2014 A Unified Write Buffer Cache Management Scheme for Flash Memory
abstract
NAND flash memory has been widely adopted in embedded systems as secondary storage. However, the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read-and-write speed asymmetry, inability of in-place updates, and performance-harmful erase operations. While write buffer cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named expectation-based least recently used (ExLRU) is proposed to improve the performance of flash memory through effectively reducing the number of erase operations and write activities. Different from the previous works, ExLRU accurately maintains access history information in the WBC, based on which a novel cost model is constructed to select data with the minimum write cost to write to flash memory. An efficient ExLRU implementation with negligible overhead is developed. Simulation results show that ExLRU outperforms state-of-the-art WBC management schemes under various workloads.
Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue, Chengmo Yang, Xuehai Zhou
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Compiler-assisted refresh minimization for volatile STT-RAM cache
abstract
Spin-Transfer Torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features: high storage density and negligible leakage power. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data loss resulting from volatility, refresh schemes are proposed. However, refresh operations consume additional energy. In this paper, we propose to reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed. Experimental results show that, on average, the proposedmethods can reduce the number of refresh operations by 73.3%, and reduce the dynamic energy consumption by 27.6%.
Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yiran Chen 0001, Yanxiang He
ASP-DAC2
2013 Cache coherence enabled adaptive refresh for volatile STT-RAM
Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001
DATE1
2013 Low-energy volatile STT-RAM cache design using cache-coherence-enabled adaptive refresh
abstract
Spin-Torque Transfer RAM (STT-RAM) is a promising candidate for SRAM replacement because of its excellent features, such as fast read access, high density, low leakage power, and CMOS technology compatibility. However, wide adoption of STT-RAM as cache memories is impeded by its long write latency and high write power. Recent work proposed improving the write performance through relaxing the retention time of STT-RAM cells. The resultant volatile STT-RAM needs to be periodically refreshed to prevent data loss. When volatile STT-RAM is applied as the last-level cache (LLC) in chip multiprocessor (CMP) systems, frequent refresh operations could dissipate significant extra energy. In addition, refresh operations could severely conflict with normal read/write operations to degrade overall system performance. Therefore, minimizing the performance impact caused by refresh operations is crucial for the adoption of volatile STT-RAM. In this article, we propose Cache-Coherence-Enabled Adaptive Refresh (CCear) to minimize the number of refresh operations for volatile STT-RAM, adopted as the LLC for CMP systems. Specifically, CCear interacts with cache coherence protocol and cache management policy to minimize the number of refresh operations on volatile STT-RAM caches. Full-system simulation results show that CCear performs close to an ideal refresh policy with low overhead. Compared with state-of-the-art refresh policies, CCear simultaneously improves the system performance and reduces the energy consumption. Moreover, the performance of CCear could be further enhanced using small filter caches to accommodate the not-refreshed private STT-RAM blocks.
Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001, Wei Wang 0237
ACM Trans. Design Autom. Electr. Syst.1
2013 Cooperating Virtual Memory and Write Buffer Management for Flash-Based Storage Systems
abstract
Flash memory is becoming the preferred choice of secondary storage in mobile devices and embedded systems. The performance of Flash memory is dictated by asymmetric speeds of read and write, limited number of erase times, and the absence of in-place updates. To improve the performance of Flash-based storage systems, the write buffer has been provided in Flash memories recently. At the same time, new virtual memory management strategies have been proposed in recent studies that consider the characteristics of Flash memory. Currently, approaches on these two memory layers are considered separately, which fail to explore the full potential of these two layers. In this paper, we propose cooperative management schemes for virtual memory and write buffer to maximize the performance of Flash-memory-based systems. Management on virtual memory is designed to exploit write buffer status via reordering of the write sequences. The proposed write buffer management scheme works seamlessly with the proposed virtual memory management scheme. Experimental results show that significant improvement in I/O performance and reduction of the number of erase and write operations can be achieved compared to the state-of-art approaches.
Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Xuehai Zhou
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Task Allocation on Nonvolatile-Memory-Based Hybrid Main Memory
abstract
In this paper, we consider the task allocation problem on a hybrid main memory composed of nonvolatile memory (NVM) and dynamic random access memory (DRAM). Compared to the conventional memory technology DRAM, the emerging NVM has excellent energy performance since it consumes orders of magnitude less leakage power. On the other hand, most types of NVMs come with the disadvantages of much shorter write endurance and longer write latency as opposed to DRAM. By leveraging the energy efficiency of NVM and long write endurance of DRAM, this paper explores task allocation techniques on hybrid memory for multiple objectives such as minimizing the energy consumption, extending the lifetime, and minimizing the memory size. The contributions of this paper are twofold. First, we design the integer linear programming (ILP) formulations that can solve different objectives optimally. Then, we propose two sets of heuristic algorithms including three polynomial time offline heuristics and three online heuristics. Experiments show that compared to the optimal solutions generated by the ILP formulations, the offline heuristics can produce near-optimal results.
Wanyong Tian, Yingchao Zhao 0001, Liang Shi 0001, Qing'an Li, Jianhua Li 0003, Chun Jason Xue, Minming Li, Enhong Chen
IEEE Trans. Very Large Scale Integr. Syst.5
2012 MAC: migration-aware compilation for STT-RAM based hybrid cache in embedded systems
abstract
Hybrid caches consisting of both STT-RAM and SRAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most work on hybrid caches employs migration based strategies to dynamically move write-intensive data from STT-RAM to SRAM. Migrations require additional read and write operations for data movement and may lead to significant overheads. To address this issue, this paper proposes a Migration-Aware Compilation (MAC) approach to improve the energy efficiency and performance of STT-RAM based hybrid cache. By re-arranging data layout, the data access pattern in memory blocks is changed such that the number of migrations is reduced without any hardware modification. The reduction of migration overheads in turn improves energy efficiency and performance. The experimental results show that with the proposed approach, on average, the number of write operations on STT-RAM is reduced by 13.4%, the number of migrations is reduced by 16.1%, the total dynamic energy is reduced by 8.5%, and the total latency is reduced by 12.1%.
Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yanxiang He
ISLPED2
2012 Hybrid nonvolatile disk cache for energy-efficient and high-performance systems
abstract
NAND flash memory has been employed as disk cache in recent years. It has the advantages of high performance, low leakage power, and cost efficiency. However, flash memory's performance is limited by the inability of in-place updates, coarse access granularity, and a limited number of write/erase times. In this article, we propose a hybrid nonvolatile disk cache architecture for high-performance and energy-efficient systems, where the disk cache is implemented with a small-size phase change memory (PCM) and a large-size NAND flash memory. Compared with current flash memory-based disk cache, it has the following advantages. (1) System performance is improved as requests are carefully directed between PCM and flash memory; (2) the energy consumption of disk cache is substantially reduced with significant reduction of additional operations, such as garbage collections; (3) the efficiency of flash memory is improved with the reduction of write activities on flash memory; and (4) lifetime of NAND flash memory is increased with most of the write operations assigned to PCM, where PCM's lifetime is guaranteed to be longer than the lifetime of flash memory. Simulation results show that the proposed methods can substantially improve the system performance, energy consumption, and lifetime of the hybrid disk cache.
Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Xuehai Zhou
ACM Trans. Design Autom. Electr. Syst.2
2011 ExLRU: a unified write buffer cache management for flash memory
abstract
NAND flash memory has been widely adopted in embedded systems as secondary storage. Yet the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read and write speed asymmetry, inability of in-place update, and performance harmful erase operations. While Write Buffer Cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named Expectation-based LRU (ExLRU) is proposed to improve the performance of write operations while at the same time reducing the number of erase operations on flash memory. ExLRU accurately maintains access history information in WBC, based on which a new cost model is constructed to select the data with minimum write cost to be written to flash memory. An efficient ExLRU implementation with negligible hardware overhead is further developed. Simulation results show that ExLRU outperforms state-of-art WBC management schemes under various workloads.
Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Chengmo Yang, Xuehai Zhou
EMSOFT2
2011 STT-RAM based energy-efficiency hybrid cache for CMPs
abstract
Modern high performance Chip Multiprocessor (CMP) systems rely on large on-chip cache hierarchy. As technology scales down, the leakage power of present SRAM based cache gradually dominates the on-chip power consumption, which can severely jeopardize system performance. The emerging nonvolatile Spin Transfer Torque RAM (STT-RAM) is a promising candidate for large on-chip cache because of the ultra low leakage power. However, the write operations on STT-RAM suffer from considerably higher energy as well as longer latency compared with SRAM which will make STT-RAM in trouble for write-intensive workloads. In this paper, we propose to integrate SRAM with STT-RAM to construct a novel hybrid cache architecture for CMPs. We also propose dedicated microarchitectural mechanisms to make the hybrid cache robust to workloads with different write patterns. Extensive simulation results demonstrate that the proposed hybrid scheme is adaptive to variations of workloads. Overall power consumption is reduced by 37.1% and performance is improved by 23.6% on average compared with SRAM based static NUCA under the same area configuration.
Jianhua Li 0003, Chun Jason Xue, Yinlong Xu 0001
VLSI-SoC1
2010 LADPM: Latency-Aware Dual-Partition Multicast Routing for Mesh-Based Network-on-Chips
abstract
Networks-on-Chips (NoCs) provides an efficient architectural paradigm as interconnect for state-of-the-art Chip Multi-processors (CMPs). With the increasing development of novel applications in NoCs, one-to-many (multicast) or one-to-all (broadcast) communications are becoming universal and indispensable. The performance constraint metrics, such as power consumption and network latency, are often stringent on NoC systems. Without multicast support, the performance of traditional NoCs will be significantly degraded by such communications. In this paper, we propose Latency-Aware Dual-Partition Multicast (LADPM) routing for mesh-based on-chip networks to reduce packet latency and balance network load. A detailed wormhole router design is also presented for the proposed LADPM scheme. LADPM scheme can adaptively make routing decision based on the distribution of the destination nodes of the multicast traffic. Experimental results, implemented under a cycle-accurate simulator, show that compared with the best known multicast scheme RPM, LADPM reduces Energy-Delay Product by 25.4% on average. More importantly, in heavy traffic load networks, LADPM is a scalable solution.
Jianhua Li 0003, Chun Jason Xue, Yinlong Xu 0001
ICPADS1