EDBT 2026 Demo / reviewers in the wild / expert
Jiwei Lu
dblp:70/1631
· DBLP profile ↗
8ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-0480-7237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 56% Processor architecture and microarchitecture · 18% Storage systems · 18% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 92% Compilers and program optimization · 8% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management › storage management › file systems
user-space file systems |
0.5 | 1 | 2021 | XFUSE: An Infrastructure for Running Filesystem Services in User Space · USENIX ATC 2021 |
Storage systems
file systems |
0.1 | 1 | 2021 | XFUSE: An Infrastructure for Running Filesystem Services in User Space · USENIX ATC 2021 |
Memory systems › non-volatile memory
magnetic random access memory |
0.1 | 1 | 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010 |
Memory systems
non-volatile memory |
0.1 | 1 | 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010 |
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM |
0.1 | 1 | 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010 |
Memory systems
cache |
0.1 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Memory systems › cache › prefetching
data prefetching |
0.1 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Processor architecture and microarchitecture › multithreading
helper thread prefetching |
0.1 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Processor architecture and microarchitecture
multithreading |
0.1 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Compilers and program optimization
dynamic optimization |
0.0 | 1 | 2003 | The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization System · MICRO 2003 |
Memory systems › cache
prefetching |
0.0 | 1 | 2003 | The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization System · MICRO 2003 |
Emerging computing paradigms
spintronics |
0.0 | 1 | 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Processor architecture and microarchitecture
dynamic optimization |
0.0 | 1 | 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005 |
Methods — techniques the papers use, named apart from their topics
static prefetching comparison · 0.1dynamic optimization · 0.1software scouting · 0.1dynamic binary optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | XFUSE: An Infrastructure for Running Filesystem Services in User Space
Qianbo Huai, Windsor W. Hsu, Jiwei Lu |
USENIX ATC | 3 |
| 2012 | Self-assembled multiferroic magnetic QCA structures for low power systemsabstractThe discovery of multiferroic materials has lead to a great interest in creating logic circuits which exploit both the magnetic and electrical properties of these materials. In this work we focus on self-assembled array structures of Magnetic Quantum Cellular Automata (MQCA) composed of multiferroic nanopillars which can be configured and clocked using solely electric fields. Furthermore, due to the switching nature of these nanopillars, the arrays can be reconfigured to implement read/write memory or multiple logic circuits in a similar fashion to FPGAs. Finally, we develop a SPICE model for the multiferroic nanopillars and demonstrate the functionality of the array. Mircea R. Stan, Mehdi Kabir, Jiwei Lu, Stuart A. Wolf |
ISCAS | 3 |
| 2011 | RAMA: a self-assembled multiferroic magnetic QCA for low power systemsabstractRecently, with the discovery of multiferroic materials, there has been a great interest in creating logic devices which exploit both magnetic and electric properties of these materials. This paper proposes a reconfigurable array of magnetic automata (RAMA) made of multiferroic nanopillars which can be operated using electric fields. Furthermore, due to the switching nature of these nanopillars, the array can be reconfigured to implement multiple logic circuits in a similar fashion to FPGAs. The paper proposes a compact model of the multiferroic switching mechanism which can be used to describe the behavior of the nanopillars in a circuit simulator. In addition, results from micromagnetic simulations of the MQCA bits indicate that it can operate with energy consumptions that are magnitudes lower than conventional CMOS technologies. Finally, the paper discusses the reliability of the nanopillar switching and suggests ways to optimize the error rates. Mehdi Kabir, Mircea R. Stan, Stuart A. Wolf, Ryan B. Comes, Jiwei Lu |
ACM Great Lakes Symposium on VLSI | 5 |
| 2010 | The Promise of Nanomagnetics and Spintronics for Future Logic and Universal MemoryabstractThis paper is both a review of some recent developments in the utilization of magnetism for applications to logic and memory and a description of some new innovations in nanomagnetics and spintronics. Nanomagnetics is primarily based on the magnetic interactions, while spintronics is primarily concerned with devices that utilize spin polarized currents. With the end of complementary metal-oxide-semiconductor (CMOS) in sight, nanomagnetics can provide a new paradigm for information process using the principles of magnetic quantum cellular automata (MQCA). This paper will review and describe these principles and then introduce a new nonlithographic method of producing reconfigurable arrays of MQCAs and/or storage bits that can be configured electrically. Furthermore, this paper will provide a brief description of magnetoresistive random access memory (MRAM), the first mainstream spintronic nonvolatile random access memory and project how far its successor spin transfer torque random access memory (STT-RAM) can go to provide a truly universal memory that can in principle replace most, if not all, semiconductor memories in the near future. For completeness, a description of an all-metal logic architecture based on magnetoresistive structures (transpinnor) will be described as well as some approaches to logic using magnetic tunnel junctions (MTJs). Stuart A. Wolf, Jiwei Lu, Mircea R. Stan, Eugene Chen, Daryl M. Treger |
Proc. IEEE | 2 |
| 2006 | Region Monitoring for Local Phase Detection in Dynamic Optimization SystemsabstractDynamic optimization relies on phase detection for two important functions (1) To detect change in code working set and (2) To detect change in performance characteristics that can affect optimization strategy. Current prototype runtime optimization systems (J. Lu et al) compare aggregate metrics like CPI over fixed time intervals to detect a change in working set and a change in performance. While simple and cost-effective, these metrics are sensitive to sampling rate and interval size. A phase detection scheme that computes performance metrics by aggregating the performance of individually optimized regions can be misled by some regions impacting aggregate metrics adversely. In this paper, we investigate the benefits and limitations of using aggregate metrics for phase detection, which we call global phase detection (GPD). We present a new model to detect change in working set and propose that the scope of phase detection be limited to within the candidate regions for optimization. By associating phase detection to individual regions we can isolate the effects of regions that are inherently unstable. This approach, which we call local phase detection (LPD), shows improved performance on several benchmarks even when global phase detection is not able to detect stable phases. Abhinav Das, Jiwei Lu, Wei-Chung Hsu |
CGO | 2 |
| 2005 | Performance of Runtime Optimization on BLASTabstractOptimization of a real world application BLAST is used to demonstrate the limitations of static and profile-guided optimizations and to highlight the potential of runtime optimization systems. We analyze the performance profile of this application to determine performance bottlenecks and evaluate the effect of aggressive compiler optimizations on BLAST. We find that applying common optimizations (e.g. O3) can degrade performance. Profile guided optimizations do not show much improvement across the board, as current implementations do not address critical performance bottlenecks in BLAST. In some cases, these optimizations lower performance significantly due to unexpected secondary effects of aggressive optimizations. We also apply runtime optimization to BLAST using the ADORE framework. ADORE is able to detect performance bottlenecks and deploy optimizations resulting in performance gains up to 58% on some queries using data cache prefetching. Abhinav Das, Jiwei Lu, Howard Chen 0002, Jinpyo Kim, Pen-Chung Yew, Wei-Chung Hsu, Dong-yuan Chen |
CGO | 2 |
| 2005 | Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP ProcessorabstractData prefetching via helper threading has been extensively investigated on simultaneous multi-threading (SMT) or virtual multi-threading (VMT) architectures. Although reportedly large cache latency can be hidden by helper threads at runtime, most techniques rely on hardware support to reduce context switch overhead between the main thread and helper thread as well as rely on static profile feedback to construct the help thread code. This paper develops a new solution by exploiting helper threaded prefetching through dynamic optimization on the latest UltraSPARC chip-multiprocessing (CMP) processor. Our experiments show that by utilizing the otherwise idle processor core, a single user-level helper thread is sufficient to improve the runtime performance of the main thread without triggering multiple thread slices. Moreover, since the multiple cores are physically decoupled in the CMP, contention introduced by helper threading is minimal. This paper also discusses several key technical challenges of building a lightweight dynamic optimization/software scouting system on the UltraSPARC/Solaris platform. Jiwei Lu, Abhinav Das, Wei-Chung Hsu, Santosh G. Abraham |
MICRO | 1 |
| 2003 | The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization SystemabstractTraditional software controlled data cache prefetching is often ineffective due to the lack of runtime cache miss and miss address information. To overcome this limitation, we implement runtime data cache prefetching in the dynamic optimization system ADORE (ADaptive Object code Reoptimization). Its performance has been compared with static software prefetching on the SPEC2000 benchmark suite. Runtime cache prefetching shows better performance. On an Itanium 2 based Linux workstation, it can increase performance by more than 20% over static prefetching on some benchmarks. For benchmarks that do not benefit from prefetching, the runtime optimization system adds only 1%-2% overhead. We have also collected cache miss profiles to guide static data cache prefetching in the ORC compiler. With that information the compiler can effectively avoid generating prefetches for loops that hit well in the data cache. Jiwei Lu, Howard Chen 0002, Wei-Chung Hsu, Bobbie Othmer, Pen-Chung Yew, Dong-yuan Chen |
MICRO | 1 |