EDBT 2026 Demo / reviewers in the wild / expert
Jianhua Li 0003
dblp:93/3389-3
· DBLP profile ↗
23ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-7086-7625ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fabdb: a low-latency fault-tolerant architecture based on dynamic bypass for network-on-chip
Zhuoxuan Ji, Jianhua Li 0003, Huaguo Liang |
J. Supercomput. | 3 |
| 2024 | F-Bypass: A Low-Power Network-on-Chip Design Utilizing Bypass to Improve Network ConnectivityabstractWith the development of transistor feature size to nanometer level, static power consumption has gradually become the main factor affecting the overall power consumption of network-on-chip (NoC). Power gating is an effective technology to reduce static power consumption, but it also brings new challenges, such as BET violation, wake-up latency and network connectivity. Therefore, a power gating method is needed to improve NoC performance and reduce static power consumption. This article proposes a low-power bypass method, namely Forwarding bypass (F-Bypass). First, F-Bypass adds bypass paths between all input and output ports and the network interface (NI) and connects the pop-up port and injection port in NI through the bypass path. When the router is powered off, F-Bypass performs wake-up-free packet transmission, which reduces the break-even time (BET) violation and cumulative wake-up latency while ensuring network connectivity. Secondly, this article adds the modified VC state table to NI so that the power-off router can perform normal traffic control. Finally, a new wake-up criterion is proposed, which can effectively avoid the frequent wake-up of power-off routers, and the detailed hardware implementation of F-Bypass is provided. The simulation results under integrated traffic load show that compared with the traditional scheme, the delay of F-Bypass is reduced by 2.2%, the throughput is increased by 13.1%, and the total static power consumption is reduced by 75.2%. Key performance indicators are superior to other solutions, and the increased area cost is moderate. Shuaijie Yuan, Jianhua Li 0003, Huaguo Liang |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2023 | A transparent virtual channel power gating method for on-chip network routers
Wu Zhou 0007, Jianhua Li 0003 |
Integr. | 3 |
| 2023 | Transit ring: bubble flow control for eliminating inter-ring communication congestion
Chenglong Sun, Qi Wang 0027, Jianhua Li 0003 |
J. Supercomput. | 5 |
| 2022 | WiGRUNT: WiFi-Enabled Gesture Recognition Using Dual-Attention NetworkabstractGestures constitute an important form of nonverbal communication where bodily actions are used for delivering messages alone or in parallel with spoken words. Recently, there exists an emerging trend of WiFi sensing-enabled gesture recognition due to its inherent merits like remote sensing, non-line-of-sight covering, and privacy-friendly. However, current WiFi-based approaches mainly reply on domain-specific training since they don’t know “where to look” and “when to look.” To this end, we propose WiGRUNT, a WiFi-enabled gesture recognition system using dual-attention network, to mimic how a keen human being intercepting a gesture regardless of the environment variations. The key insight is to train the network to dynamically focus on the domain-independent features of a gesture on the WiFi channel state information via a spatial-temporal dual-attention mechanism. WiGRUNT roots in a deep residual network (ResNet) backbone to evaluate the importance of spatial-temporal clues and exploit their inbuilt sequential correlations for fine-grained gesture recognition. We evaluate WiGRUNT on the open Widar3 dataset and show that it significantly outperforms its state-of-the-art rivals by achieving the best-ever performance in-domain or cross-domain. Yu Gu 0003, Xiang Zhang 0011, Yantong Wang, Meng Wang 0001, Huan Yan 0005, Yusheng Ji, Zhi Liu 0002, Jianhua Li 0003, Mianxiong Dong |
IEEE Trans. Hum. Mach. Syst. | 8 |
| 2019 | CPCA: An efficient wireless routing algorithm in WiNoC for cross path congestion awareness
Jianhua Li 0003, Chenglong Sun, Huaguo Liang, Gaoming Du |
Integr. | 3 |
| 2017 | Thread Criticality Assisted Replication and Migration for Chip Multiprocessor CachesabstractNon-Uniform Cache Architecture (NUCA) is a viable solution to mitigate the problem of large on-chip wire delay due to the rapid increase in the cache capacity of chip multiprocessors (CMPs). Through partitioning the last-level cache (LLC) into smaller banks connected by on-chip network, the access latency will exhibit non-uniform distribution. Various works have well explored the NUCA design, including block migration, block replication and block searching. However, all of the previous mechanisms designed for NUCA are thread-oblivious when multi-threaded applications are deployed on CMP systems. Due to the interference on shared resources, threads often demonstrate unbalanced progress wherein the lagging threads with slow progress are more critical to overall performance. In this paper, we propose a novel NUCA design called thread Criticality Assisted Replication and Migration (CARM). CARM exploits the runtime thread criticality information as hints to adjust the block replication and migration in NUCA. Specifically, CARM aims at boosting parallel application execution through prioritizing block replication and migration for critical threads. Full-system experimental results show that CARM reduces the execution time of a set of PARSEC workloads by 13.7 and 6.8 percent on average compared with the tradition D-NUCA and Re-NUCA respectively. Moreover, CARM also consumes much less energy compared with the evaluated schemes. Jianhua Li 0003, Minming Li, Chun Jason Xue, Fanfan Shen |
IEEE Trans. Computers | 1 |
| 2015 | Compiler-Assisted Refresh Minimization for Volatile STT-RAM CacheabstractSpin-transfer torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features such as high storage density and ultra low leakage power. However, long write latency and high write energy are the two challenges for STT-RAM. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data losses resulting from volatility, refresh schemes have been proposed. However, refresh operations consume additional overhead. In this paper, we propose to significantly reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed to further reduce the number of refreshes. Experimental results show that, on average, the proposed methods can reduce the number of refresh operations by 84.2 percent, and reduce the dynamic energy consumption by 38.0 percent for volatile STT-RAM caches while incurring only 4.1 percent performance degradation. Qing'an Li, Yanxiang He, Jianhua Li 0003, Liang Shi 0001, Yiran Chen 0001, Chun Jason Xue |
IEEE Trans. Computers | 3 |
| 2014 | Dual partitioning multicasting for high-performance on-chip networks
Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yinlong Xu 0001 |
J. Parallel Distributed Comput. | 1 |
| 2014 | Thread Progress Aware Coherence Adaption for Hybrid Cache Coherence ProtocolsabstractFor chip multiprocessor systems (CMPs), the interference on shared resources such as on-chip caches typically leads to unbalanced progress among threads. Because of the inherent synchronization primitives, such as barriers and locks, cores running fast threads have to waste precious cycles to wait for cores with slow progress, which leads to performance and energy inefficiency. For the purpose of improving performance and reducing energy consumption, this paper proposes to adapt the cache coherence policy for threads according to their delay-tolerant levels. Specifically, this paper proposes Thread progrEss Aware Coherence Adaption (TEACA) which utilizes the thread progress information as hints for coherence adaption. TEACA dynamically utilize the memory system statistics to estimate the progress of threads. Based on the estimated thread progress information, TEACA categorizes threads into leader threads and laggard threads. The thread categorization decisions are then leveraged for efficient coherence adaption on CMP systems supporting hybrid coherence protocols. Experimental results show that, on a 64-core CMP system, TEACA outperforms directory protocol in application execution time and a recently proposed hybrid protocol in both application execution time and energy dissipation. Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | WCET-Aware Re-Scheduling Register Allocation for Real-Time Embedded Systems With Clustered VLIW ArchitectureabstractWorst-case execution time (WCET) is one of the most important metric in real-time embedded system design. For embedded systems with clustered very long instruction word (VLIW) architecture, register allocation, instruction scheduling, and cluster assignment are three key activities for code optimization, which have profound impact on WCET. At the same time, these three activities exhibit a phase ordering problem, i.e., independently performing register allocation, scheduling, and cluster assignment could have a negative effect on the other phases, thereby generating sub-optimal compiled code. In this paper, a compiler level optimization, namely WCET-aware re-scheduling register allocation, is proposed to achieve WCET minimization for real-time embedded systems with clustered VLIW architecture. The novelty of the proposed approach is that the effects of register allocation, instruction scheduling, and cluster assignment on the quality of generated code are taken into account for WCET minimization. These three compilation processes are integrated into a single phase to obtain a balanced result. The proposed technique is implemented in Trimaran 4.0. The experimental results show that the proposed technique can reduce WCET effectively, by 34% on average. Yazhi Huang, Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Compiler-Assisted STT-RAM-Based Hybrid Cache for Energy Efficient Embedded SystemsabstractHybrid caches consisting of static RAM (SRAM) and spin-torque transfer (STT)-RAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most of the management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques involve additional access operations, and thus lead to extra overheads. In this paper, we propose two compilation-based approaches to improve the energy efficiency and performance of STT-RAM-based hybrid cache by reducing the migration overheads. The first approach, migration-aware data layout, is proposed to reduce the migrations by rearranging the data layout. The second approach, migration-aware cache locking, is proposed to reduce the migrations by locking migration-intensive memory blocks into SRAM part of hybrid cache. Furthermore, experiments show that these two methods can be combined to reduce more migrations. The reduction of migration overheads can improve the energy efficiency and performance of STT-RAM-based hybrid cache. Experimental results show that, combining these two methods, on average, the number of write operations on STT-RAM is reduced by 17.6%, the number of migrations is reduced by 38.9%, the total dynamic energy is reduced by 15.6%, and the total access latency is reduced by 13.8%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Yanxiang He |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | A Unified Write Buffer Cache Management Scheme for Flash MemoryabstractNAND flash memory has been widely adopted in embedded systems as secondary storage. However, the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read-and-write speed asymmetry, inability of in-place updates, and performance-harmful erase operations. While write buffer cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named expectation-based least recently used (ExLRU) is proposed to improve the performance of flash memory through effectively reducing the number of erase operations and write activities. Different from the previous works, ExLRU accurately maintains access history information in the WBC, based on which a novel cost model is constructed to select data with the minimum write cost to write to flash memory. An efficient ExLRU implementation with negligible overhead is developed. Simulation results show that ExLRU outperforms state-of-the-art WBC management schemes under various workloads. Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue, Chengmo Yang, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Compiler-assisted refresh minimization for volatile STT-RAM cacheabstractSpin-Transfer Torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features: high storage density and negligible leakage power. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data loss resulting from volatility, refresh schemes are proposed. However, refresh operations consume additional energy. In this paper, we propose to reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed. Experimental results show that, on average, the proposedmethods can reduce the number of refresh operations by 73.3%, and reduce the dynamic energy consumption by 27.6%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yiran Chen 0001, Yanxiang He |
ASP-DAC | 2 |
| 2013 | Cache coherence enabled adaptive refresh for volatile STT-RAM
Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001 |
DATE | 1 |
| 2013 | Low-energy volatile STT-RAM cache design using cache-coherence-enabled adaptive refreshabstractSpin-Torque Transfer RAM (STT-RAM) is a promising candidate for SRAM replacement because of its excellent features, such as fast read access, high density, low leakage power, and CMOS technology compatibility. However, wide adoption of STT-RAM as cache memories is impeded by its long write latency and high write power. Recent work proposed improving the write performance through relaxing the retention time of STT-RAM cells. The resultant volatile STT-RAM needs to be periodically refreshed to prevent data loss. When volatile STT-RAM is applied as the last-level cache (LLC) in chip multiprocessor (CMP) systems, frequent refresh operations could dissipate significant extra energy. In addition, refresh operations could severely conflict with normal read/write operations to degrade overall system performance. Therefore, minimizing the performance impact caused by refresh operations is crucial for the adoption of volatile STT-RAM. In this article, we propose Cache-Coherence-Enabled Adaptive Refresh (CCear) to minimize the number of refresh operations for volatile STT-RAM, adopted as the LLC for CMP systems. Specifically, CCear interacts with cache coherence protocol and cache management policy to minimize the number of refresh operations on volatile STT-RAM caches. Full-system simulation results show that CCear performs close to an ideal refresh policy with low overhead. Compared with state-of-the-art refresh policies, CCear simultaneously improves the system performance and reduces the energy consumption. Moreover, the performance of CCear could be further enhanced using small filter caches to accommodate the not-refreshed private STT-RAM blocks. Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001, Wei Wang 0237 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | Cooperating Virtual Memory and Write Buffer Management for Flash-Based Storage SystemsabstractFlash memory is becoming the preferred choice of secondary storage in mobile devices and embedded systems. The performance of Flash memory is dictated by asymmetric speeds of read and write, limited number of erase times, and the absence of in-place updates. To improve the performance of Flash-based storage systems, the write buffer has been provided in Flash memories recently. At the same time, new virtual memory management strategies have been proposed in recent studies that consider the characteristics of Flash memory. Currently, approaches on these two memory layers are considered separately, which fail to explore the full potential of these two layers. In this paper, we propose cooperative management schemes for virtual memory and write buffer to maximize the performance of Flash-memory-based systems. Management on virtual memory is designed to exploit write buffer status via reordering of the write sequences. The proposed write buffer management scheme works seamlessly with the proposed virtual memory management scheme. Experimental results show that significant improvement in I/O performance and reduction of the number of erase and write operations can be achieved compared to the state-of-art approaches. Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Task Allocation on Nonvolatile-Memory-Based Hybrid Main MemoryabstractIn this paper, we consider the task allocation problem on a hybrid main memory composed of nonvolatile memory (NVM) and dynamic random access memory (DRAM). Compared to the conventional memory technology DRAM, the emerging NVM has excellent energy performance since it consumes orders of magnitude less leakage power. On the other hand, most types of NVMs come with the disadvantages of much shorter write endurance and longer write latency as opposed to DRAM. By leveraging the energy efficiency of NVM and long write endurance of DRAM, this paper explores task allocation techniques on hybrid memory for multiple objectives such as minimizing the energy consumption, extending the lifetime, and minimizing the memory size. The contributions of this paper are twofold. First, we design the integer linear programming (ILP) formulations that can solve different objectives optimally. Then, we propose two sets of heuristic algorithms including three polynomial time offline heuristics and three online heuristics. Experiments show that compared to the optimal solutions generated by the ILP formulations, the offline heuristics can produce near-optimal results. Wanyong Tian, Yingchao Zhao 0001, Liang Shi 0001, Qing'an Li, Jianhua Li 0003, Chun Jason Xue, Minming Li, Enhong Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2012 | MAC: migration-aware compilation for STT-RAM based hybrid cache in embedded systemsabstractHybrid caches consisting of both STT-RAM and SRAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most work on hybrid caches employs migration based strategies to dynamically move write-intensive data from STT-RAM to SRAM. Migrations require additional read and write operations for data movement and may lead to significant overheads. To address this issue, this paper proposes a Migration-Aware Compilation (MAC) approach to improve the energy efficiency and performance of STT-RAM based hybrid cache. By re-arranging data layout, the data access pattern in memory blocks is changed such that the number of migrations is reduced without any hardware modification. The reduction of migration overheads in turn improves energy efficiency and performance. The experimental results show that with the proposed approach, on average, the number of write operations on STT-RAM is reduced by 13.4%, the number of migrations is reduced by 16.1%, the total dynamic energy is reduced by 8.5%, and the total latency is reduced by 12.1%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yanxiang He |
ISLPED | 2 |
| 2012 | Hybrid nonvolatile disk cache for energy-efficient and high-performance systemsabstractNAND flash memory has been employed as disk cache in recent years. It has the advantages of high performance, low leakage power, and cost efficiency. However, flash memory's performance is limited by the inability of in-place updates, coarse access granularity, and a limited number of write/erase times. In this article, we propose a hybrid nonvolatile disk cache architecture for high-performance and energy-efficient systems, where the disk cache is implemented with a small-size phase change memory (PCM) and a large-size NAND flash memory. Compared with current flash memory-based disk cache, it has the following advantages. (1) System performance is improved as requests are carefully directed between PCM and flash memory; (2) the energy consumption of disk cache is substantially reduced with significant reduction of additional operations, such as garbage collections; (3) the efficiency of flash memory is improved with the reduction of write activities on flash memory; and (4) lifetime of NAND flash memory is increased with most of the write operations assigned to PCM, where PCM's lifetime is guaranteed to be longer than the lifetime of flash memory. Simulation results show that the proposed methods can substantially improve the system performance, energy consumption, and lifetime of the hybrid disk cache. Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Xuehai Zhou |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | ExLRU: a unified write buffer cache management for flash memoryabstractNAND flash memory has been widely adopted in embedded systems as secondary storage. Yet the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read and write speed asymmetry, inability of in-place update, and performance harmful erase operations. While Write Buffer Cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named Expectation-based LRU (ExLRU) is proposed to improve the performance of write operations while at the same time reducing the number of erase operations on flash memory. ExLRU accurately maintains access history information in WBC, based on which a new cost model is constructed to select the data with minimum write cost to be written to flash memory. An efficient ExLRU implementation with negligible hardware overhead is further developed. Simulation results show that ExLRU outperforms state-of-art WBC management schemes under various workloads. Liang Shi 0001, Jianhua Li 0003, Chun Jason Xue, Chengmo Yang, Xuehai Zhou |
EMSOFT | 2 |
| 2011 | STT-RAM based energy-efficiency hybrid cache for CMPsabstractModern high performance Chip Multiprocessor (CMP) systems rely on large on-chip cache hierarchy. As technology scales down, the leakage power of present SRAM based cache gradually dominates the on-chip power consumption, which can severely jeopardize system performance. The emerging nonvolatile Spin Transfer Torque RAM (STT-RAM) is a promising candidate for large on-chip cache because of the ultra low leakage power. However, the write operations on STT-RAM suffer from considerably higher energy as well as longer latency compared with SRAM which will make STT-RAM in trouble for write-intensive workloads. In this paper, we propose to integrate SRAM with STT-RAM to construct a novel hybrid cache architecture for CMPs. We also propose dedicated microarchitectural mechanisms to make the hybrid cache robust to workloads with different write patterns. Extensive simulation results demonstrate that the proposed hybrid scheme is adaptive to variations of workloads. Overall power consumption is reduced by 37.1% and performance is improved by 23.6% on average compared with SRAM based static NUCA under the same area configuration. Jianhua Li 0003, Chun Jason Xue, Yinlong Xu 0001 |
VLSI-SoC | 1 |
| 2010 | LADPM: Latency-Aware Dual-Partition Multicast Routing for Mesh-Based Network-on-ChipsabstractNetworks-on-Chips (NoCs) provides an efficient architectural paradigm as interconnect for state-of-the-art Chip Multi-processors (CMPs). With the increasing development of novel applications in NoCs, one-to-many (multicast) or one-to-all (broadcast) communications are becoming universal and indispensable. The performance constraint metrics, such as power consumption and network latency, are often stringent on NoC systems. Without multicast support, the performance of traditional NoCs will be significantly degraded by such communications. In this paper, we propose Latency-Aware Dual-Partition Multicast (LADPM) routing for mesh-based on-chip networks to reduce packet latency and balance network load. A detailed wormhole router design is also presented for the proposed LADPM scheme. LADPM scheme can adaptively make routing decision based on the distribution of the destination nodes of the multicast traffic. Experimental results, implemented under a cycle-accurate simulator, show that compared with the best known multicast scheme RPM, LADPM reduces Energy-Delay Product by 25.4% on average. More importantly, in heavy traffic load networks, LADPM is a scalable solution. Jianhua Li 0003, Chun Jason Xue, Yinlong Xu 0001 |
ICPADS | 1 |