Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhizhong Tang

dblp:86/6700 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 96% Program analysis · 4%
Network and information security
1 paper
Network security · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 60% Reconfigurable computing and FPGAs · 40%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network security
intrusion detection and prevention
0.112009
A parameterized multilevel pattern matching architecture on FPGAs for network intrusion detection and prevention · Sci. China Ser. F Inf. Sci. 2009
Compilers and program optimization › instruction scheduling
software pipelining
0.132007
Single-dimension software pipelining for multidimensional loops · ACM Trans. Archit. Code Optim. 2007
GPMB - software pipelining branch-intensive loops · MICRO 1993
A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992
Compilers and program optimization › loop transformation
loop scheduling
0.112007
Single-dimension software pipelining for multidimensional loops · ACM Trans. Archit. Code Optim. 2007
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling
0.112007
Single-dimension software pipelining for multidimensional loops · ACM Trans. Archit. Code Optim. 2007
Processor architecture and microarchitecture
instruction-level parallelism
0.022007
Single-dimension software pipelining for multidimensional loops · ACM Trans. Archit. Code Optim. 2007
A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992
Reconfigurable computing and FPGAs
FPGA architecture
0.012009
A parameterized multilevel pattern matching architecture on FPGAs for network intrusion detection and prevention · Sci. China Ser. F Inf. Sci. 2009
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.011993
GPMB - software pipelining branch-intensive loops · MICRO 1993
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.011993
GPMB - software pipelining branch-intensive loops · MICRO 1993
Program analysis
static analysis
0.011993
GPMB - software pipelining branch-intensive loops · MICRO 1993
Compilers and program optimization
instruction scheduling
0.011992
A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.011992
A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992

Methods — techniques the papers use, named apart from their topics

pattern matching · 0.2FPGA · 0.2modulo scheduling · 0.1hyperplane scheduling · 0.1data dependence graph · 0.1static program analysis · 0.0branch prediction · 0.0
YearPublicationVenuePosition
2016 NestedMP: Enabling cache-aware thread mapping for nested parallel shared memory applications
Jiangzhou He, Zhizhong Tang
Parallel Comput.3
2014 OpenMDSP: Extending OpenMP to Program Multi-Core DSPs
Jiangzhou He, Wen-Guang Chen, Guangri Chen, Zhizhong Tang, Handong Ye
J. Comput. Sci. Technol.5
2011 OpenMDSP: Extending OpenMP to Program Multi-Core DSP
abstract
Multi-core Digital Signal Processors (DSP) are widely used in wireless telecommunication, core network transcoding, industrial control, and audio/video processing etc. Comparing with general purpose multi-processors, the multi-core DSPs normally have more complex memory hierarchy, such as on-chip core-local memory and non-cache-coherent shared memory. As a result, it is very challenging to write efficient multi-core DSP applications. The current approach to program multi-core DSPs is based on proprietary vendor SDKs, which only provides low-level, non-portable primitives. While it is acceptable to write coarse-grained task level parallel code with these SDKs, it is very tedious and error prone to write fine-grained data parallel code with them. We believe it is desired to have a high-level and portable parallel programming model for multi-core DSPs. In this paper, we propose Open MDSP, an extension of Open MP designed for multi-core DSPs. The goal of Open MDSP is to fill the gap between Open MP memory model and the memory hierarchy of multi-core DSPs. We propose three class of directives in Open MDSP: (1) data placement directives allow programmers to control the placement of global variables conveniently, (2) distributed array directives divide whole array into sections and promote them into core-local memory to improve performance, and (3) stream access directives promote big array into core-local memory section by section during a parallel loop's processing. We implement the compiler and runtime system for Open MDSP on Free Scale MSC8156. Benchmarking result shows that seven out of nine benchmarks achieve a speedup of more than 5 with 6 threads.
Jiangzhou He, Guangri Chen, Zhizhong Tang, Handong Ye
PACT5
2010 A Novel Memory Subsystem Evaluation Framework for Chip Multiprocessors
abstract
This paper presents a fast and cycle-accurate memory subsystem modeling and evaluating framework for Chip Multiprocessors (CMPs), called TSIM (Tsinghua SIMulator), which gives a flexible and extensible approach to evaluating architecture designs, models or algorithms, including the network-on-chip interconnection, cache hardware prefetcher, memory system protocol, replacement policy, etc. TSIM is trace-driven, adopting a dynamic binary instrumentation technique to generate the running trace information of applications on-the-fly. After receiving the trace information, TSIM will reappear the on-chip memory behaviors of applications. By introducing the concept of statistical meta metrics, TSIM separates the analysis stage from the simulation process per se, and this provides a great facilitation for a user to count and sample the performance metrics. Compared to the real cache system, TSIM achieves an accuracy of 90.66% at the average speed of 327 KIPS. Meanwhile, TSIM accelerates the simulating speed by almost 10 times, compared to the traditional cycle-accurate cache simulators. On the other hand, when TSIM is used to characterize the on-chip memory system behaviors of SPEC CPU 2000 benchmarks, experimental results about the on-chip memory behaviors are the same as others.
Fucen Zeng, Lin Qiao, Zhizhong Tang
HPCC4
2009 Efficient shared cache management through sharing-aware replacement and streaming-aware insertion policy
abstract
Multi-core processors with shared caches are now commonplace. However, prior works on shared cache management primarily focused on multi-programmed workloads. These schemes consider how to partition the cache space given that simultaneously-running applications may have different cache behaviors. In this paper, we examine policies for managing shared caches for running single multi-threaded applications. First, we show that the shared-cache miss rate can be significantly reduced by reserving a certain amount of space for shared data. Therefore, we modify the replacement policy to dynamically partition each set between shared and private data. Second, we modify the insertion policy to prevent streaming data (data not reused before eviction) from promoting to the MRU position. Finally, we use a low-overhead sampling mechanism to dynamically select the optimal policy. Compared to LRU policy, our scheme reduces the miss rate on average by 8.7% on 8MB caches and 20.1% on 16MB caches respectively.
Wenlong Li 0003, Changkyu Kim, Zhizhong Tang
IPDPS4
2009 Understanding the Memory Behavior of Emerging Multi-core Workloads
abstract
This paper characterizes the memory behavior on emerging RMS (recognition, mining, and synthesis) workloads for future multi-core processors. As multi-core processors proliferate across different application domains, and the number of on-die cores continues to increase, a key issue facing processor architects is the design of the on-die last level cache (LLC). In this paper, we explore the LLC design space for multi-threaded RMS workloads by examining the working set sizes, data sharing behavior, and spatial data locality. Our study reveals that these RMS workloads are memory intensive, have large working-set sizes greater than 16 MB on average, exhibit a significant amount of data sharing, about 47% on average, and show strong strided streaming access behavior with 77% of accesses in regular pattern. Based on the observations, we then investigate the potential cache architecture choices for future multi-core design. Our experiments show that for these workloads large DRAM caches can be useful to address their large working sets; e.g., a 128 MB DRAM cache can reduce the average L1 miss penalty by 18%; shared last level cache provides better cache performance than private cache; e.g., a 8 MB shared cache provides 25% performance improvement over a private one with the same total size; and stride based hardware prefetcher provides significant performance benefit by 25%. As a result, we suggest a memory hierarchy with a 128 MB DRAM cache, a 8 MB on-die SRAM shared cache and an 8-entry stride prefetcher to accommodate RMS workloads.
Junmin Lin, Wenlong Li 0003, Aamer Jaleel, Zhizhong Tang
ISPDC5
2009 A parameterized multilevel pattern matching architecture on FPGAs for network intrusion detection and prevention
Dongsheng Wang 0002, Zhizhong Tang
Sci. China Ser. F Inf. Sci.3
2008 Data Sharing Analysis of Emerging Parallel Media Mining Workloads
Wenlong Li 0003, Junmin Lin, Aamer Jaleel, Zhizhong Tang
HiPC5
2007 Single-dimension software pipelining for multidimensional loops
abstract
Traditionally, software pipelining is applied either to the innermost loop of a given loop nest or from the innermost loop to outer loops. This paper proposes a three-step approach, called single-dimension software pipelining (SSP) , to software pipeline a loop nest at an arbitrary loop level that has a rectangular iteration space and contains no sibling inner loops in it. The first step identifies the most profitable loop level for software pipelining in terms of initiation rate, data reuse potential, or any other optimization criteria. The second step simplifies the multidimensional data-dependence graph (DDG) of the selected loop level into a one-dimensional DDG and constructs a one-dimensional (1D) schedule. Based on the one-dimensional schedule, the third step derives a simple mapping function that specifies the schedule time for the operation instances in the multidimensional loop. The classical modulo scheduling is subsumed by SSP as a special case. SSP is also closely related to hyperplane scheduling, and, in fact, extends it to be resource constrained. We prove that SSP schedules are correct and at least as efficient as those schedules generated by traditional modulo scheduling methods. We extend SSP to schedule imperfect loop nests, which are most common at the instruction level. Multiple initiation intervals are naturally allowed to improve execution efficiency. Feasibility and correctness of our approach are verified by a prototype implementation in the ORC compiler for the IA-64 architecture, tested with loop nests from Livermore and SPEC2000 floating-point benchmarks. Preliminary experimental results reveal that, compared to modulo scheduling, software pipelining at an appropriate loop level results in significant performance improvement. Software pipelining is beneficial even with prior loop transformations.
Hongbo Rong, Zhizhong Tang, R. Govindarajan, Alban Douillet, Guang R. Gao
ACM Trans. Archit. Code Optim.2
2005 A Static Data Dependence Analysis Approach for Software Pipelining
Lin Qiao, Weitong Huang, Zhizhong Tang
NPC3
2005 A Dynamic Data Dependence Analysis Approach for Software Pipelining
Lin Qiao, Weitong Huang, Zhizhong Tang
NPC3
2005 Coping with Data Dependencies of Multi-dimensional Array References
Lin Qiao, Weitong Huang, Zhizhong Tang
NPC3
2004 Single-Dimension Software Pipelining for Multi-Dimensional Loops
abstract
Traditionally, software pipelining is applied either to the innermost loop of a given loop nest or from the innermost loop to outer loops. We propose a three-step approach, called single-dimension software pipelining (SSP), to software pipeline a loop nest at an arbitrary loop level. The first step identifies the most profitable loop level for software pipelining in terms of initiation rate or data reuse potential. The second step simplifies the multidimensional data-dependence graph (DDG) into a 1-dimensional DDG and constructs a 1-dimensional schedule for the selected loop level. The third step derives a simple mapping function which specifies the schedule time for the operations of the multidimensional loop, based on the 1-dimensional schedule. We prove that the SSP method is correct and at least as efficient as other modulo scheduling methods. We establish the feasibility and correctness of our approach by implementing it on the IA-64 architecture. Experimental results on a small number of loops show significant performance improvements over existing modulo scheduling methods that software pipeline a loop nest from the innermost loop.
Hongbo Rong, Zhizhong Tang, R. Govindarajan, Alban Douillet, Guang R. Gao
CGO2
2004 Increasing Software-Pipelined Loops in the Itanium-Like Architecture
Zhizhong Tang
ISPA4
1993 GPMB - software pipelining branch-intensive loops
abstract
Compile-time code transformations which expose instruction-level parallelism (ILP) typically take into account the constraints imposed by all execution scenarios in the program. However, there are additional opportunities to increase ILP along some execution sequences if the constraints from alternative execution sequences can be ignored. Traditionally, profile information has been used to identify important execution sequences for aggressive compiler optimization and scheduling. The paper presents a set of static program analysis heuristics used in the IMPACT compiler to identify execution sequences for aggressive optimization. The authors show that the static program analysis heuristics identify execution sequences without hazardous conditions that tend to prohibit compiler optimizations. As a result, the static program analysis approach often achieves optimization results comparable to profile information in spite of its inferior branch prediction accuracies. This observation makes a strong case for using static program analysis with or without profile information to facilitate aggressive compiler optimization and scheduling.>
Zhizhong Tang, Chihong Zhang, Bogong Su, Stanley Habib
MICRO1
1993 URPR-1: A single-chip VLIW architecture
Bogong Su, Jian Wang 0046, Zhizhong Tang, Chihong Zhang
Microprocess. Microprogramming3
1992 A VLIW architecture for optimal execution of branch-intensive loops
Bogong Su, Zhizhong Tang, Stanley Habib
MICRO3