Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiwei Lu

dblp:70/1631 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-0480-7237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 56% Processor architecture and microarchitecture · 18% Storage systems · 18%
Software engineering, system software, and programming languages
2 papers
Operating systems · 92% Compilers and program optimization · 8%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › resource management › storage management › file systems
user-space file systems
0.512021
XFUSE: An Infrastructure for Running Filesystem Services in User Space · USENIX ATC 2021
Storage systems
file systems
0.112021
XFUSE: An Infrastructure for Running Filesystem Services in User Space · USENIX ATC 2021
Memory systems › non-volatile memory
magnetic random access memory
0.112010
The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010
Memory systems
non-volatile memory
0.112010
The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.112010
The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010
Memory systems
cache
0.112005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005
Memory systems › cache › prefetching
data prefetching
0.112005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005
Processor architecture and microarchitecture › multithreading
helper thread prefetching
0.112005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005
Processor architecture and microarchitecture
multithreading
0.112005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005
Compilers and program optimization
dynamic optimization
0.012003
The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization System · MICRO 2003
Memory systems › cache
prefetching
0.012003
The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization System · MICRO 2003
Emerging computing paradigms
spintronics
0.012010
The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory · Proc. IEEE 2010
Processor architecture and microarchitecture
chip multiprocessor
0.012005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005
Processor architecture and microarchitecture
dynamic optimization
0.012005
Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor · MICRO 2005

Methods — techniques the papers use, named apart from their topics

static prefetching comparison · 0.1dynamic optimization · 0.1software scouting · 0.1dynamic binary optimization · 0.1
YearPublicationVenuePosition
2021 XFUSE: An Infrastructure for Running Filesystem Services in User Space
Qianbo Huai, Windsor W. Hsu, Jiwei Lu
USENIX ATC3
2012 Self-assembled multiferroic magnetic QCA structures for low power systems
abstract
The discovery of multiferroic materials has lead to a great interest in creating logic circuits which exploit both the magnetic and electrical properties of these materials. In this work we focus on self-assembled array structures of Magnetic Quantum Cellular Automata (MQCA) composed of multiferroic nanopillars which can be configured and clocked using solely electric fields. Furthermore, due to the switching nature of these nanopillars, the arrays can be reconfigured to implement read/write memory or multiple logic circuits in a similar fashion to FPGAs. Finally, we develop a SPICE model for the multiferroic nanopillars and demonstrate the functionality of the array.
Mircea R. Stan, Mehdi Kabir, Jiwei Lu, Stuart A. Wolf
ISCAS3
2011 RAMA: a self-assembled multiferroic magnetic QCA for low power systems
abstract
Recently, with the discovery of multiferroic materials, there has been a great interest in creating logic devices which exploit both magnetic and electric properties of these materials. This paper proposes a reconfigurable array of magnetic automata (RAMA) made of multiferroic nanopillars which can be operated using electric fields. Furthermore, due to the switching nature of these nanopillars, the array can be reconfigured to implement multiple logic circuits in a similar fashion to FPGAs. The paper proposes a compact model of the multiferroic switching mechanism which can be used to describe the behavior of the nanopillars in a circuit simulator. In addition, results from micromagnetic simulations of the MQCA bits indicate that it can operate with energy consumptions that are magnitudes lower than conventional CMOS technologies. Finally, the paper discusses the reliability of the nanopillar switching and suggests ways to optimize the error rates.
Mehdi Kabir, Mircea R. Stan, Stuart A. Wolf, Ryan B. Comes, Jiwei Lu
ACM Great Lakes Symposium on VLSI5
2010 The Promise of Nanomagnetics and Spintronics for Future Logic and Universal Memory
abstract
This paper is both a review of some recent developments in the utilization of magnetism for applications to logic and memory and a description of some new innovations in nanomagnetics and spintronics. Nanomagnetics is primarily based on the magnetic interactions, while spintronics is primarily concerned with devices that utilize spin polarized currents. With the end of complementary metal-oxide-semiconductor (CMOS) in sight, nanomagnetics can provide a new paradigm for information process using the principles of magnetic quantum cellular automata (MQCA). This paper will review and describe these principles and then introduce a new nonlithographic method of producing reconfigurable arrays of MQCAs and/or storage bits that can be configured electrically. Furthermore, this paper will provide a brief description of magnetoresistive random access memory (MRAM), the first mainstream spintronic nonvolatile random access memory and project how far its successor spin transfer torque random access memory (STT-RAM) can go to provide a truly universal memory that can in principle replace most, if not all, semiconductor memories in the near future. For completeness, a description of an all-metal logic architecture based on magnetoresistive structures (transpinnor) will be described as well as some approaches to logic using magnetic tunnel junctions (MTJs).
Stuart A. Wolf, Jiwei Lu, Mircea R. Stan, Eugene Chen, Daryl M. Treger
Proc. IEEE2
2006 Region Monitoring for Local Phase Detection in Dynamic Optimization Systems
abstract
Dynamic optimization relies on phase detection for two important functions (1) To detect change in code working set and (2) To detect change in performance characteristics that can affect optimization strategy. Current prototype runtime optimization systems (J. Lu et al) compare aggregate metrics like CPI over fixed time intervals to detect a change in working set and a change in performance. While simple and cost-effective, these metrics are sensitive to sampling rate and interval size. A phase detection scheme that computes performance metrics by aggregating the performance of individually optimized regions can be misled by some regions impacting aggregate metrics adversely. In this paper, we investigate the benefits and limitations of using aggregate metrics for phase detection, which we call global phase detection (GPD). We present a new model to detect change in working set and propose that the scope of phase detection be limited to within the candidate regions for optimization. By associating phase detection to individual regions we can isolate the effects of regions that are inherently unstable. This approach, which we call local phase detection (LPD), shows improved performance on several benchmarks even when global phase detection is not able to detect stable phases.
Abhinav Das, Jiwei Lu, Wei-Chung Hsu
CGO2
2005 Performance of Runtime Optimization on BLAST
abstract
Optimization of a real world application BLAST is used to demonstrate the limitations of static and profile-guided optimizations and to highlight the potential of runtime optimization systems. We analyze the performance profile of this application to determine performance bottlenecks and evaluate the effect of aggressive compiler optimizations on BLAST. We find that applying common optimizations (e.g. O3) can degrade performance. Profile guided optimizations do not show much improvement across the board, as current implementations do not address critical performance bottlenecks in BLAST. In some cases, these optimizations lower performance significantly due to unexpected secondary effects of aggressive optimizations. We also apply runtime optimization to BLAST using the ADORE framework. ADORE is able to detect performance bottlenecks and deploy optimizations resulting in performance gains up to 58% on some queries using data cache prefetching.
Abhinav Das, Jiwei Lu, Howard Chen 0002, Jinpyo Kim, Pen-Chung Yew, Wei-Chung Hsu, Dong-yuan Chen
CGO2
2005 Dynamic Helper Threaded Prefetching on the Sun UltraSPARC CMP Processor
abstract
Data prefetching via helper threading has been extensively investigated on simultaneous multi-threading (SMT) or virtual multi-threading (VMT) architectures. Although reportedly large cache latency can be hidden by helper threads at runtime, most techniques rely on hardware support to reduce context switch overhead between the main thread and helper thread as well as rely on static profile feedback to construct the help thread code. This paper develops a new solution by exploiting helper threaded prefetching through dynamic optimization on the latest UltraSPARC chip-multiprocessing (CMP) processor. Our experiments show that by utilizing the otherwise idle processor core, a single user-level helper thread is sufficient to improve the runtime performance of the main thread without triggering multiple thread slices. Moreover, since the multiple cores are physically decoupled in the CMP, contention introduced by helper threading is minimal. This paper also discusses several key technical challenges of building a lightweight dynamic optimization/software scouting system on the UltraSPARC/Solaris platform.
Jiwei Lu, Abhinav Das, Wei-Chung Hsu, Santosh G. Abraham
MICRO1
2003 The Performance of Runtime Data Cache Prefetching in a Dynamic Optimization System
abstract
Traditional software controlled data cache prefetching is often ineffective due to the lack of runtime cache miss and miss address information. To overcome this limitation, we implement runtime data cache prefetching in the dynamic optimization system ADORE (ADaptive Object code Reoptimization). Its performance has been compared with static software prefetching on the SPEC2000 benchmark suite. Runtime cache prefetching shows better performance. On an Itanium 2 based Linux workstation, it can increase performance by more than 20% over static prefetching on some benchmarks. For benchmarks that do not benefit from prefetching, the runtime optimization system adds only 1%-2% overhead. We have also collected cache miss profiles to guide static data cache prefetching in the ORC compiler. With that information the compiler can effectively avoid generating prefetches for loops that hit well in the data cache.
Jiwei Lu, Howard Chen 0002, Wei-Chung Hsu, Bobbie Othmer, Pen-Chung Yew, Dong-yuan Chen
MICRO1