EDBT 2026 Demo / reviewers in the wild / expert
Feiqi Su
dblp:48/6381
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 46% Performance modeling and evaluation · 46% Distributed systems · 7% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.1 | 1 | 2009 | Modeling and Stack Simulation of CMP Cache Capacity and Accessibility · IEEE Trans. Parallel Distributed Syst. 2009 |
Memory systems › cache
chip multiprocessor cache |
0.1 | 1 | 2009 | Modeling and Stack Simulation of CMP Cache Capacity and Accessibility · IEEE Trans. Parallel Distributed Syst. 2009 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2009 | Modeling and Stack Simulation of CMP Cache Capacity and Accessibility · IEEE Trans. Parallel Distributed Syst. 2009 |
Performance modeling and evaluation
stack simulation |
0.1 | 1 | 2009 | Modeling and Stack Simulation of CMP Cache Capacity and Accessibility · IEEE Trans. Parallel Distributed Syst. 2009 |
Distributed systems › replication
data replication |
0.0 | 1 | 2009 | Modeling and Stack Simulation of CMP Cache Capacity and Accessibility · IEEE Trans. Parallel Distributed Syst. 2009 |
Methods — techniques the papers use, named apart from their topics
single-pass stack simulation · 0.1execution-driven simulation · 0.1abstract modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Directory Lookaside Table: Enabling scalable, low-conflict, many-core cache coherence directoryabstractMaintaining hardware cache coherence on future CMPs becomes increasingly important and difficult as the number of cores keeps accelerating in mainstream multicore chips. The simple snooping-bus coherence scheme is not suitable due to its limited scalability. The sparse coherence directory approach may incur extra cache invalidations due to a topological mismatch between the coherence directory and the directories of all cache modules. In this paper, we propose an innovative CMP coherence directory that has three important properties. First, the directory has a simple set-associative design with small associativity. The number of directory entries matches the total number of cache blocks. Second, an augmented Directory Lookaside Table (DLT) allows blocks to be displaced from their primary sets in the coherence directory for alleviating hot-set conflicts. Third, to avoid expensive presence bits, each copy of a block along with the located core ID occupies a separate directory entry. Performance evaluations based on multithreaded and multi-programmed workloads demonstrate significant advantages of the proposed CMP directory over directories with traditional set-associative or skewed associative designs. Xudong Shi 0003, Feiqi Su, Jih-Kwon Peir |
ICPADS | 2 |
| 2010 | Weak execution ordering - exploiting iterative methods on many-core GPUsabstractOn NVIDIA's many-core GPUs, there is no synchronization function among parallel thread blocks. When fine-granularity of data communication and synchronization is required for large-scale parallel programs executed by multiple thread blocks, frequent host synchronization are necessary, and they incur a significant overhead. We investigate a class of applications which uses a chaotic version of iterative methods to obtain numerical solutions for partial differential equations (PDE). Such a fast PDE solver is parallelized on GPUs with multiple thread blocks. In this parallel implementation, although frequent data communication is needed between adjacent thread blocks, a precise order of the data communication is not necessary. Separate communication threads are used for periodically exchanging the boundary values with adjacent thread blocks through the global memory. Since a precise order of the data communication is not required, the computation and the communication threads can be overlapped to alleviate the communication overhead. Performance measurements of two popular applications, Poisson image editing from computer graphics and shape from shading from computer vision, on Tesla C1060 show that a speedup of 4–5 times is achievable for both applications in comparison with the solution using host synchronization. Jianmin Chen, Feiqi Su, Jih-Kwon Peir, Jeff Ho, Lu Peng 0001 |
ISPASS | 3 |
| 2009 | Modeling and Stack Simulation of CMP Cache Capacity and AccessibilityabstractPerformance trade-offs between fast data access by local data replication and cache capacity maximization by global data sharing have been extensively studied for many-core Chip Multiprocessors (CMPs). Costly simulations over a wide spectrum of the design space are generally required to gain insight for a sound design. To lower the cost, we develop an abstract model for understanding the performance impact of data replication on CMP caches. To overcome the lack of real-time interactions among multiple cores in the model, we further develop an efficient single-pass stack simulation to study the performance of CMP cache organizations with various degrees of data replication. The global stack logically incorporates a shared stack and per-core private stacks; shared/private reuse (stack) distances can be collected in a single-pass simulation. With the reuse distances, one can calculate the performance of CMP cache organizations with various degrees of data replication. We verify both the model and the stack simulation against execution-driven simulations with commercial multithreaded workloads. The results show that the abstract model provides accurate information about performance trade-offs of data replication. The stack simulation accurately predicts the performance of various cache organizations with 2-9 percent error margins using only about 8 percent of the simulation time. Xudong Shi 0003, Feiqi Su, Jih-Kwon Peir, Ye Xia 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | Comparative evaluation of multi-core cache occupancy strategiesabstractIntelligent sharing cache space among multiple cores on a Chip Multiprocessor (CMP) has become an important research topic. There are many design options to trade off and many possible performance metrics to evaluate. It generally requires costly simulations to gain insights over a wide-spectrum of cache sharing and partitioning methods. In this paper, we use an efficient single-pass stack simulation method to understand the effectiveness of cache sharing through natural competition (i. e. shared cache) and through static or dynamic cache partitioning, such as equal partition, utility-based partition, etc. The results demonstrate that cache occupancy through natural competition favors the core with more frequently misses. It may take away the needed space for other cores and increase the overall miss ratio. Furthermore, we find that the existing cache partitioning schemes based on fixed-length portioning phases may not be optimal compared with the scheme that allows variable phase lengths. Feiqi Su, Xudong Shi 0003, Ye Xia 0001, Jih-Kwon Peir |
ICPADS | 1 |
| 2007 | Modeling and Single-Pass Simulation of CMP Cache Capacity and AccessibilityabstractThe future chip-multiprocessors (CMPs) with a large number of cores faces difficult issues in efficient utilizing on-chip storage space. Tradeoffs between data accessibility and effective on-chip capacity have been studied extensively. It requires costly simulations to understand a wide-spectrum of design spaces. In this paper, we first develop an abstract model for understanding the performance impact with respect to the degree of data replication. To overcome the lack of real-time interactions among multiple cores in the abstract model, we propose an efficient single-pass stack simulation method to study the performance of a variety of cache organizations on CMPs. The proposed global stack logically incorporates a shared stack and per-core private stacks to collect shared/private reuse (stack) distances for every memory reference in a single simulation pass. With the collected reuse distances, performance in terms of hits/misses and average memory access times can be calculated for multiple cache organizations. The basic stack simulation results can further derive other CMP cache organizations with various degrees of data replication. We verify both the modeling and the stack results against individual execution-driven simulations that consider realistic cache parameters and delays using a set of commercial multithreaded workloads. We also compare the simulation time saving with the stack simulation. The results show that stack simulation can accurately model the performance of various studied cache organizations with 2-9% error margins using only about 8% of the simulation time. The results also show that the effectiveness of various techniques for optimizing the CMP on-chip storage is closely related to the working sets of the workloads as well as the total cache sizes Xudong Shi 0003, Feiqi Su, Jih-Kwon Peir, Ye Xia 0001 |
ISPASS | 2 |
| 2006 | Overlapping dependent loads with addressless preloadabstractModern out-of-order processors with non-blocking caches exploit Memory-Level Parallelism (MLP) by overlapping cache misses in a wide instruction window. The exploitation of MLP, however, can be limited due to long-latency operations in producing the base address of a cache miss load. When the parent instruction is also a cache miss load, a serialization of the two loads must be enforced to satisfy the load-load data dependence.In this paper, we propose a mechanism that dynamically captures the load-load data dependences at runtime. A special Preload is issued in place of the dependent load without waiting for the parent load, thus effectively overlapping the two loads. The Preload provides necessary information for the memory controller to calculate the correct memory address upon the availability of the parent's data to eliminate any interconnect delay between the two loads. Performance evaluations based on SPEC2000 and Olden applications show that significant speedups up to 40% with an average of 16% are achievable using the Preload. In conjunction with other aggressive MLP exploitation methods, such as runahead execution, the Preload can make more significant improvement with an average of 22%. Xudong Shi 0003, Feiqi Su, Jih-Kwon Peir |
PACT | 3 |
| 2002 | A Mesh Watermarking Approach for Appearance AttributesabstractWe describe an algorithm to watermark appearance attributes, as well as the shape of the mesh. Appearance attributes are potential watermarking primitives and the watermarking approach for them can be generalized from that for the shape. The major challenge of generalization is that the watermarking for appearance attributes has more constraints. We focus on this challenge. Especially for the normal vector, we embed the watermark by modifying its orientation, not magnitude. Results show our scheme effectively improves the capacity and enhances the robustness of mesh watermarking. Liangjun Zhang, Ruofeng Tong 0001, Feiqi Su, Jinxiang Dong |
PG | 3 |