VLDB 2026 Research / reviewers in the wild / expert
Jingfei Kong
dblp:88/5847
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-authorComputer networks · 1Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
GPUs and heterogeneous computing · 51% Memory systems · 42% Hardware reliability and fault tolerance · 6% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Network and information security
2 papers |
Hardware security and side channels · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware security and side channels › side-channel attack
cache side-channel attacks |
0.3 | 2 | 2013 | Architecting against Software Cache-Based Side-Channel Attacks · IEEE Trans. Computers 2013 Hardware-software integrated approaches to defend against software cache-based side channel attacks · HPCA 2009 |
Compilers and program optimization › accelerator compilation
GPU compiler |
0.3 | 2 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
GPUs and heterogeneous computing
GPU programming |
0.3 | 2 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
Compilers and program optimization › memory optimization
memory hierarchy optimization |
0.1 | 1 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 |
Compilers and program optimization › accelerator compilation
GPU compiler optimization |
0.1 | 1 | 2010 | A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
Compilers and program optimization › memory optimization
memory access optimization |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Memory systems › data locality
data reuse |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.1 | 1 | 2010 | A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
Hardware security and side channels
side-channel attack |
0.0 | 1 | 2013 | Architecting against Software Cache-Based Side-Channel Attacks · IEEE Trans. Computers 2013 |
Memory systems
cache design |
0.0 | 1 | 2013 | Architecting against Software Cache-Based Side-Channel Attacks · IEEE Trans. Computers 2013 |
Hardware reliability and fault tolerance › error correction
cache error correction |
0.0 | 1 | 2013 | Architecting against Software Cache-Based Side-Channel Attacks · IEEE Trans. Computers 2013 |
GPUs and heterogeneous computing › GPU memory access
memory coalescing |
0.0 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Memory systems
cache |
0.0 | 1 | 2009 | Hardware-software integrated approaches to defend against software cache-based side channel attacks · HPCA 2009 |
Memory systems
cache coherence |
0.0 | 1 | 2009 | Hardware-software integrated approaches to defend against software cache-based side channel attacks · HPCA 2009 |
Methods — techniques the papers use, named apart from their topics
preloading · 0.5software random permutation · 0.3hardware-software integrated defense · 0.3kernel generation · 0.3architecture-specific optimization · 0.3parallelism management · 0.2memory hierarchy optimization · 0.2data prefetching · 0.2auto-tuning · 0.2informing loads · 0.2software permutation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Architecting against Software Cache-Based Side-Channel AttacksabstractUsing cache-like architectural components including data caches, instruction caches, or branch target buffers as a side channel, software cache-based side-channel attacks are able to derive secret keys used in cryptographic operations through legitimate software activities. Existing software solutions are typically application specific and incur substantial performance overhead. Recent hardware proposals against attacks on data caches, although effective in reducing performance overhead, may still be vulnerable to advanced attacks. Furthermore, efficient defenses against attacks on other cache structures, including instruction caches and branch target buffers, are missing. In this paper, we propose hardware-software integrated approaches to defend against software cache-based attacks comprehensively. For attacks on data caches, we propose to use preloading, informing loads, and informing loads with software random permutation to secure the partition-locked cache (PLcache), the random permutation (RPcache) and regular caches, respectively. These approaches present different tradeoffs between hardware complexity and performance overhead. To defend against attacks on instruction caches, we show that the PLcache with preloading and the RPcache provide good protection. To defend against attacks based on branch target buffers, we propose to adopt a new update policy to eliminate potential information leaking. Our experiments show that the proposed schemes not only provide strong security protection but also incur small performance overhead. Jingfei Kong, Onur Aciiçmez, Jean-Pierre Seifert, Huiyang Zhou |
IEEE Trans. Computers | 1 |
| 2012 | A unified optimizing compiler framework for different GPGPU architecturesabstractThis article presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. The input to our compiler is a naïve GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler generates two kernels, one optimized for global memories and the other for texture memories. The proposed compilation process is effective for both AMD/ATI and NVIDIA GPUs. The experiments show that our optimized code achieves very high performance, either superior or very close to highly fine-tuned libraries. Yi Yang 0018, Ping Xiang, Jingfei Kong, Mike Mantor, Huiyang Zhou |
ACM Trans. Archit. Code Optim. | 3 |
| 2010 | Improving privacy and lifetime of PCM-based main memoryabstractPhase change memory (PCM) is a promising technology for computer memory systems. However, the non-volatile nature of PCM poses serious threats to computer privacy. The low programming endurance of PCM devices also limits the lifetime of PCM-based main memory (PRAM). In this paper, we first adopt counter-mode encryption for privacy protection and show that encryption significantly reduces the effectiveness of some previously proposed wear-leveling techniques for PRAM. To mitigate such adverse impact, we propose simple, yet effective extensions to the encryption scheme. In addition, we propose to reuse the encryption counters as age counters and to dynamically adjust the strength of error correction code (ECC) to extend the lifetime of PRAM. Our experiments show that our mechanisms effectively achieve privacy protection and lifetime extension for PRAM with very low performance overhead. Jingfei Kong, Huiyang Zhou |
DSN | 1 |
| 2010 | A GPGPU compiler for memory optimization and parallelism managementabstractThis paper presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou |
PLDI | 3 |
| 2010 | An optimizing compiler for GPGPU programs with input-data sharingabstractDeveloping high performance GPGPU programs is challenging for application developers since the performance is dependent upon how well the code leverages the hardware features of specific graphics processors. To solve this problem and relieve application developers of low-level hardware-specific optimizations, we introduce a novel compiler to optimize GPGPU programs. Our compiler takes a naive GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler then analyzes the code, identifies memory access patterns, and generates optimized code. The proposed compiler optimizations target at one category of scientific and media processing algorithms, which has the characteristics of input-data sharing when computing neighboring output pixels/elements. Many commonly used algorithms, such as matrix multiplication, convolution, etc., share such characteristics. For these algorithms, novel approaches are proposed to enforce memory coalescing and achieve effective data reuse. Data prefetching and hardware-specific tuning are also performed automatically with our compiler framework. The experimental results based on a set of applications show that our compiler achieves very high performance, either superior or very close to the highly fine-tuned library, NVIDIA CUBLAS 2.1. Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou |
PPoPP | 3 |
| 2009 | Hardware-software integrated approaches to defend against software cache-based side channel attacksabstractSoftware cache-based side channel attacks present serious threats to modern computer systems. Using caches as a side channel, these attacks are able to derive secret keys used in cryptographic operations through legitimate activities. Among existing countermeasures, software solutions are typically application specific and incur substantial performance overhead. Recent hardware proposals including the partition-locked cache (PLcache) and random-permutation cache (RPcache) (Wang and Lee, 2007), although very effective in reducing performance overhead while enhancing the security level, may still be vulnerable to advanced cache attacks. In this paper, we propose three hardware-software approaches to defend against software cache-based attacks - they present different tradeoffs between hardware complexity and performance overhead. First, we propose to use preloading to secure the PLcache. Second, we leverage informing loads, which is a lightweight architectural support originally proposed to improve memory performance, to protect the RPcache. Third, we propose novel software permutation to replace the random permutation hardware in the RPcache. This way, regular caches can be protected with hardware support for informing loads. In our experiments, we analyze various processor models for their vulnerability to cache attacks and demonstrate that even to the processor model that is most vulnerable to cache attacks, our proposed software-hardware integrated schemes provide strong security protection. Jingfei Kong, Onur Aciiçmez, Jean-Pierre Seifert, Huiyang Zhou |
HPCA | 1 |
| 2004 | An adaptive coordinated medium access control for wireless sensor networksabstractWe have developed adaptive coordinated medium access control (AC-MAC), a contention-based medium access control protocol for wireless sensor networks. To handle the load variations in some real-time sensor applications, ACMAC introduces the adaptive duty cycle scheme within the framework of sensor-MAC (S-MAC). The novelty of our protocol is that it improves latency and throughput under a wide range of traffic loads while remaining as energy-efficient as S-MAC. We illustrate such optimized trade-offs of AC-MAC via extensive simulations performed over wireless sensor networks. Our simulation results show that AC-MAC is as energy-efficient as S-MAC while its latency and throughput are always trying to follow the classic IEEE 802.11 MAC (no duty cycle), which outperform the S-MAC (fixed duty cycle), specially under the heavy load. Jing Ai, Jingfei Kong, Damla Turgut |
ISCC | 2 |