EDBT 2026 Demo / reviewers in the wild / expert
Nan Wu 0003
dblp:58/2484-3
· DBLP profile ↗
20ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Reconfigurable computing and FPGAs · 40% GPUs and heterogeneous computing · 35% Parallel and multicore computing · 12% | |
| Computer graphics and multimedia
2 papers |
Image and video coding · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 3 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU computing › GPU video coding
GPU-accelerated video encoding |
0.1 | 1 | 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processor · ACM Multimedia 2011 |
Image and video coding › video coding standards
H.264/AVC |
0.1 | 1 | 2009 | Streaming HD H.264 encoder on programmable processors · ACM Multimedia 2009 |
Parallel and multicore computing
thread-level parallelism |
0.0 | 1 | 2012 | The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only) · FPGA 2012 |
Methods — techniques the papers use, named apart from their topics
programmable stream processor · 0.2block-based parallel processing · 0.2CUDA · 0.2stream programming · 0.2software optimization · 0.2partial dynamic reconfiguration · 0.1hardware/software partitioning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Parallel performance modeling of irregular applications in cell-centered finite volume methods over unstructured tetrahedral meshes
Johannes Langguth, Nan Wu 0003, Jun Chai, Xing Cai |
J. Parallel Distributed Comput. | 2 |
| 2014 | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs
Dafei Huang, Mei Wen, Changqing Xun, Dong Chen 0015, Xing Cai, Yuran Qiao, Nan Wu 0003, Chunyuan Zhang |
Euro-Par | 7 |
| 2014 | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node
Xinnan Dong, Jun Chai, Mei Wen, Nan Wu 0003, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
ICA3PP (2) | 5 |
| 2013 | ACF: Networks-on-Chip Deadlock Recovery with Accurate Detection and Elastic Credit
Nan Wu 0003, Yuran Qiao, Mei Wen, Chunyuan Zhang |
APPT | 1 |
| 2013 | On the GPU-CPU Performance Portability of OpenCL for 3D Stencil ComputationsabstractAlthough OpenCL programming provides full code portability between different hardware platforms, performance portability can be far from satisfactory. In this work, we use a set of representative 3D stencil computations to study OpenCL's performance portability between GPUs and CPUs. For each stencil computation, we have devised different implementations of the computational kernel function, all being 100% code-portable between the two architectures. The most straightforward and compact implementation gives satisfactory CPU performance but performs poorly on GPUs, because such an implementation hampers effective use of the GPU hardware. By injecting code complexity into the involved loop nests, we can create kernel functions that still have full code portability but with increased performance portability. It is found that spatial data blocking and register reuse can be beneficial for performance on both GPUs and CPUs, whereas use of OpenCL's local memory (and subsequent temporal blocking) may only have positive effects on GPUs. Huayou Su, Nan Wu 0003, Mei Wen, Chunyuan Zhang, Xing Cai |
ICPADS | 2 |
| 2013 | Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 5 |
| 2013 | Accelerating thread-intensive and explicit memory management programs with dynamic partial reconfiguration
Qianming Yang, Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 3 |
| 2012 | Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin SimulationsabstractIn cutting-edge CPU/GPU hybrid clusters, such as Tianhe-1A, the aggregate CPU computing capability may amount to up to 1/3 of the aggregate GPU computing capability. It thus goes without saying that the CPUs and GPUs should jointly carry out the computational work. However, to effectively and simultaneously use both the hardware components requires great care when developing the parallel implementations. The challenges include (1) finding a balanced division of the workload between the CPU and GPU sides, and (2) hiding various overheads by overlapping computations with CPU-GPU data transfers and/or MPI communications. We study these issues in the context of real-world sedimentary basin simulations. Numerical experiments show that an appropriately devised CPU-GPU hybrid implementation is able to handle a global mesh resolution of 131,072*131,072, and a double-precision rate of 62 TFlops is achieved by using 1024 GPUs and 12288 CPU cores on Tianhe-1A. Such an extreme computing capability will be of great importance for carrying out high-resolution and continental-scale stratigraphic simulations in future. Mei Wen, Huayou Su, Wenjie Wei, Nan Wu 0003, Xing Cai, Chunyuan Zhang |
CLUSTER | 4 |
| 2012 | The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only)abstractA uniform FPGA-based architecture, an efficient programming model and a simple mapping method are paramount for PPGA technology to be more widely accepted. This paper presents MASALA, a dynamically reconfigurable FPGA-based accelerator specifically for parallel programs written in thread-intensive and explicit memory management (TEMM) programming models. The system uses TEMM programming model to parallelize the demanding application, including decomposing the application into separate thread blocks, decoupling compute and data load/store etc. Hardware engines are included into the MASALA by using partial dynamic reconfigure modules, each of which encapsulates Thread Process Engine implementing the thread functionality in hardware. A data dispatching scheme is also included in MASALA to enable the explicit communication among multiple memory hierarchies such as between inter-hardware engines, the host processor and hardware engines. At last, the paper illustrates a Multi-FPGA prototype system of the presented architecture: MASALA-SX. A large synthetic aperture radar (SAR) image formatting experiment shows that the MASALA architecture facilitates the construction of a TEMM program accelerator by providing it with greater performance and less power consumption than current CPU platforms, but without sacrificing programmability, flexibility and scalability. Mei Wen, Nan Wu 0003, Qianming Yang, Chunyuan Zhang |
FPGA | 2 |
| 2012 | Extending BORPH for shared memory reconfigurable computersabstractWe extend BORPH for shared memory reconfigurable computers in this paper. BORPH is an operating system designed for FPGA based reconfigurable computers. BORPH introduced the concept of hardware process in contrast to software process. With our extension, hardware processes are supported to communicate with other processes based on shared memory. In our system, the program of hardware process is not just hardware design, but the software program running on embedded processor in FPGA. Our experiment shows the overhead of shared memory segments management is acceptable. And with independent virtual memory access, bandwidth of repeated shared memory access is high. Changqing Xun, Mei Wen, Nan Wu 0003, Chunyuan Zhang, Hayden Kwok-Hay So |
FPL | 3 |
| 2012 | A Parallel H.264 Encoder with CUDA: Mapping and EvaluationabstractEfficient mapping of a real-time HD video application to graphics hardware is challenging. Developers face the challenges of choosing the right parallelism model, balancing thread's process granularity between massive computing resources on the GPU, and partitioning tasks between the CPU and GPU. The paper illustrated the mapping approaches by a case of HD H.264 encoder based on X264 reference code and then evaluating it on state-of-the-art CPU and GPUs in depth. In the paper, we first split most of the computing task into Single-Instruction Multiple-Thread (SIMT) kernels, which are then chained intocertaininput/output data stream. Then we implementeda completed H.264 encoding on the computer unified device architecture (CUDA) platform. Finally, we present methods for exploiting multi-level parallelism and memory efficiency when mapping H.264 code, which we use to increase the efficiency of the execution on GPUs. Our experimental results show that computation efficiency of GPU and then real-time encoding performance are achieved with CUDA. Nan Wu 0003, Mei Wen, Huayou Su, Ju Ren 0002, Chunyuan Zhang |
ICPADS | 1 |
| 2011 | A Multilevel Parallel Intra Coding for H.264/AVC Based on CUDAabstractIn this paper, we propose a multilevel parallel intra coding for H.264/AVC based on computed unified device architecture (CUDA). The proposed parallel algorithm improves the parallelism between 4×4 blocks within a macro block (MB) by throwing off some inappreciable prediction modes. By partitioning a frame into multi-slice, the parallelism between MBs can be exploited. In addition, a scalable parallel method for kernels is introduced to improve the performance of the proposed intra coding. Experimental results show that, more than 20 times speedup can be achieved with the assistance of GPU. Moreover, the entire encoder can meet the real-time processing requirement for HDTV. Huayou Su, Nan Wu 0003, Chunyuan Zhang, Mei Wen, Ju Ren 0002 |
ICIG | 2 |
| 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processorabstractThis article presents an efficient software parallel CAVLC encoder based on programmable stream processors (Storm- SP16 and GPU). For static processor Storm SP16, a block-based 16 ways parallel CAVLC is presented with streaming processing. A component-oriented CAVLC encoder is proposed aiming at dynamic stream processor GPU. Experiments results show that, compared to the CPU version, more than 70 times of speedup can be obtained for the CAVLC based on Storm and over 50 times for GPU-based component-oriented CAVLC encoder. The throughput of the presented CAVLC encoder is more than 10 times higher over that of published software CAVLC encoders on DSP and multi-core platforms. Huayou Su, Chunyuan Zhang, Jun Chai, Mei Wen, Nan Wu 0003, Ju Ren 0002 |
ACM Multimedia | 5 |
| 2010 | Software Managed Instruction Scratchpad Memory Optimization in Stream Architecture Based on Hot Code Analysis of KernelsabstractStream processors, such as Imagine, GPGPUs, FT64 and MASA, typically uses software managed scratchpad instruction memory which improves performance and significantly reduces energy consumption. In this paper, we build a kernel-storage model to analyze the hot spot of kernels in stream programs. Based on the analysis, we define Kernel Hot Code and prove that scratchpad instruction memory should focus on the access efficiency of it. A methodology for finding Kernel Hot Code in the kernels of different structures is presented as well. In accordance with this method, we develop HOIS for Stream Architecture, which adopts a software managed scratchpad memory to store Kernel Hot Code, and uses a small hardware managed victim cache to store the Kernel Cool Code. HOIS is evaluated by measuring the performance of six applications on the MASA_S simulation platform. The results show that HOIS can achieve high efficiency in predictable applications with little performance loss. Yi He 0008, Ju Ren 0002, Mei Wen, Qianming Yang, Nan Wu 0003, Chunyuan Zhang |
DSD | 5 |
| 2009 | Cache streamization for high performance stream processorabstractDue to high bandwidth demand on memory system of stream applications, most of stream processors use software-managed streaming memory. However, this memory disadvantages ease of programming, compatibility, and supporting irregular stream access, which hinder the usage of stream processor in broader application domains. Meanwhile, hardware-managed coherent caches overcome these shortcomings of software-managed streaming memory with side-effect due to lack of supporting stream. For this problem, this paper developed a streamization cache whose performance is comparable to streaming memory but is more easy to use. The paper presents the motivation and details of our proposed design, including three stream-specific techniques for cache on data fetch policy, replacement policy and multi-client access. Moreover, a streamization cache instance is implemented in FT64, a 64-bit high performance stream processor. Based on a set of streaming application benchmark, the paper estimates the performance, power consumption and the area cost of the proposed architecture. Results show that these streamization techniques for cache are worthwhile. Nan Wu 0003, Mei Wen, Ju Ren 0002, Yi He 0008, Changqing Xun, Chunyuan Zhang |
HiPC | 1 |
| 2009 | Streaming HD H.264 encoder on programmable processorsabstractProgrammable processors have great advantage over dedicated ASIC design under intense time-to-market pressure. However, real-time encoding of high-definition (HD) H.264 video (up to 1080p) is a challenge to most existing programmable processors. On the other hand, model-based design is widely accepted in developing complex media program. Stream model, an emerging model-based programming method, shows surprising efficiency on many compute-intensive domains especially for media processing. On the basis, this paper proposes a set of streaming techniques for H.264 encoding, and then develops all of the code based on the X264 reference code. Our streaming H.264 encoder is a pure software implementation completely written in high-level language without special hardware/algorithm support. Real execution results show that our encoder achieves significant speedup over the original X264 encoder on various programmable architectures: on X86 CoreTM2 E8200 the speedup is 1.8x, on MIPS 4KEc the speedup is 3.7x, on TMS320 C6416 DSP the speedup is 5.5x, on stream processor STORM-SP16 G220 the speedup is 6.1x. Especially, on STORM processor, the streaming encoder achieves the performance of 30.6 frames per second for a 1080P HD sequence, satisfying the real-time requirement. These indicate that streaming is extremely efficient for this kind of media workload. Our work is also applicable for other media processing applications, and provides architecture insights into dedicated ASIC or FPGA HD H.264 encoders. Nan Wu 0003, Mei Wen, Ju Ren 0002, Huayou Su, Changqing Xun, Chunyuan Zhang |
ACM Multimedia | 1 |
| 2008 | Load scheduling: Reducing pressure on distributed register files for freeabstractIn this paper we describe load scheduling, a novel method that balances load among register files by residual resources. Load scheduling can reduce register pressure for clustered VLIW processors with distributed register files while not increasing VLIW scheduling length. We have implemented load scheduling in compiler for Imagine and FT64 stream processors. The result shows that the proposed technique effectively reduces the number of variables spilled to memory, and can even eliminate it. The algorithm presented in this paper is extremely efficient in embedded processor with limited register resource because it can improve registers utilization instead of increasing the requirement for the number of registers. Mei Wen, Nan Wu 0003, Maolin Guan, Chunyuan Zhang |
ASP-DAC | 2 |
| 2007 | FT64: Scientific Computing with Streams
Mei Wen, Nan Wu 0003, Chunyuan Zhang, Qianming Yang, Changqing Xun |
HiPC | 2 |
| 2005 | Multiple-Morphs Adaptive Stream Architecture
Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Comput. Sci. Technol. | 2 |
| 2004 | A Parallel Reed-Solomon Decoder on the Imagine Stream Processor
Mei Wen, Chunyuan Zhang, Nan Wu 0003, Li Li 0005 |
ISPA | 3 |