EDBT 2026 Demo / reviewers in the wild / expert
Wai Teng Tang
dblp:62/6991
· DBLP profile ↗
13ranked-venue papers
4as first author
1since 2021 · last 2025
0000-0002-6553-1270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 24% High-performance computing · 23% Emerging computing paradigms · 17% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
neuromorphic computing |
0.8 | 2 | 2020 | An FPGA-Based Hardware Emulator for Neuromorphic Chip With RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips · ACM Trans. Archit. Code Optim. 2019 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.7 | 3 | 2017 | Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPU · IEEE Trans. Parallel Distributed Syst. 2015 Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemes · SC 2013 |
Reconfigurable computing and FPGAs
FPGA-based emulation |
0.4 | 1 | 2020 | An FPGA-Based Hardware Emulator for Neuromorphic Chip With RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
GPUs and heterogeneous computing
GPU computing |
0.4 | 2 | 2015 | A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPU · IEEE Trans. Parallel Distributed Syst. 2015 Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemes · SC 2013 |
Memory systems › memory-efficient data structures
sparse matrix compression |
0.4 | 2 | 2015 | A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPU · IEEE Trans. Parallel Distributed Syst. 2015 Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemes · SC 2013 |
Memory systems
in-memory computing |
0.4 | 1 | 2019 | A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips · ACM Trans. Archit. Code Optim. 2019 |
Memory systems › emerging memory technologies
RRAM crossbar |
0.4 | 1 | 2019 | A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips · ACM Trans. Archit. Code Optim. 2019 |
Processor architecture and microarchitecture
many-core architecture |
0.3 | 1 | 2017 | Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.3 | 1 | 2017 | Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
Compilers and program optimization
code generation |
0.2 | 1 | 2015 | A Code Generation Framework for Targeting Optimized Library Calls for Multiple Platforms · IEEE Trans. Parallel Distributed Syst. 2015 |
Parallel and multicore computing › parallel programming models
directive-based programming |
0.2 | 1 | 2015 | A Code Generation Framework for Targeting Optimized Library Calls for Multiple Platforms · IEEE Trans. Parallel Distributed Syst. 2015 |
Electronic design automation
system-level simulation |
0.1 | 1 | 2019 | A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips · ACM Trans. Archit. Code Optim. 2019 |
Performance modeling and evaluation
performance tuning |
0.1 | 1 | 2017 | Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017 |
GPUs and heterogeneous computing › GPU programming
GPU code generation |
0.1 | 1 | 2015 | A Code Generation Framework for Targeting Optimized Library Calls for Multiple Platforms · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
iterative methods |
0.1 | 1 | 2015 | A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPU · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
performance optimization |
0.0 | 1 | 2013 | Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemes · SC 2013 |
Methods — techniques the papers use, named apart from their topics
linear algebraic operation recognition · 0.4directive-based compilation · 0.4RRAM crossbar modeling · 0.4bit-representation-optimized compression · 0.4cycle-accurate simulation · 0.4locality-aware block mapping · 0.3intrinsics · 0.3hybrid COO+CSR format · 0.3OpenCL · 0.32-d jagged partitioning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised Few-Shot Food Recognition With Intra-Class Variation and Inter-Class Similarity ModelingabstractFew-shot food recognition aims to first train a meta-model based on an extensive labeled dataset, and then adapt it to recognize novel food classes with limited labeled data. Although existing studies have achieved compelling success, they still heavily relied on a large number of labeled food data for training the initial meta-model. To save the annotation cost, we propose the unsupervised food recognition task, which aims to train a meta-model using only unlabeled food data. Due to the two challenges presented in food images: 1) high intra-class variations and 2) high inter-class similarity, directly applying existing unsupervised few-shot learning methods could result in sub-optimal results. Towards this end, we propose a novel framework,i.e., Unsupervsied Few-shot Food Recognition with Intra-class Variation and Inter-class Similarity (UFFR-IVIS). It consists of two key components: (1) dual diversity-injected support/query representation learning that introduces instance-level and representation-level diversities for the representation learning of support/query instance to model the characteristics of high intra-class variation; and (2) dual regularization-enhanced meta learning that designs two regularizations: auxiliary task-based intra-class regularization and similarity-guided inter-class regularization to regularize the intra-class variation and inter-class similarity modeling, respectively. Extensive experiments on two food datasets demonstrate the superiority of our UFFR-IVIS. Xuemeng Song, Wai Teng Tang, See-Kiong Ng, Liqiang Nie, Roger Zimmermann |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Automatic detection of anatomical landmarks in brain MR scanning using multi-task deep neural networks
Xulei Yang, Wai Teng Tang, Gabriel Tjio, Si Yong Yeo, Yi Su 0001 |
Neurocomputing | 2 |
| 2020 | An FPGA-Based Hardware Emulator for Neuromorphic Chip With RRAMabstractNeuromorphic chip with RRAM devices has been demonstrated as a promising computing platform for neural network-based applications. By directly mapping the weight matrices of neural networks onto RRAM-based crossbar arrays, high energy, and area efficiency can be achieved. However, the design of an RRAM-based neuromorphic chip faces many constraints due to the variability and limitations of RRAM. Simulation and emulation can help in the design of a neuromorphic chip prior to fabrication. However, software-based chip simulation on CPU is slow, especially for large-scale network-on-chip (NoC)-based chip design. In this paper, we present a hardware emulator on field-programmable gate array (FPGA) for an RRAM-based neuromorphic chip. Our emulator supports the emulation of static and dynamic variation of the RRAM-based crossbars used in the neural cores of a neuromorphic chip. Furthermore, an NoC is also implemented on FPGA to emulate the communication between the neural cores. Using the emulator, we show that effects, such as RRAM write and read noise and stuck-at faults affect the accuracy of an application on a neuromorphic chip. We also demonstrate the utility of the emulator in investigating NoC topologies, routing buffer depths, and neural core mappings. Tao Luo 0014, Chuping Qu, Matthew Kay Fei Lee, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | A System-Level Simulator for RRAM-Based Neuromorphic Computing ChipsabstractAdvances in non-volatile resistive switching random access memory (RRAM) have made it a promising memory technology with potential applications in low-power and embedded in-memory computing devices owing to a number of advantages such as low-energy consumption, low area cost and good scaling. There have been proposals to employ RRAM in architecting chips for neuromorphic computing and artificial neural networks where matrix-vector multiplication can be computed in the analog domain in a single timestep. However, it is challenging to employ RRAM devices in neuromorphic chips owing to the non-ideal behavior of RRAM. In this article, we propose a cycle-accurate and scalable system-level simulator that can be used to study the effects of using RRAM devices in neuromorphic computing chips. The simulator models a spatial neuromorphic chip architecture containing many neural cores with RRAM crossbars connected via a Network-on-Chip (NoC). We focus on system-level simulation and demonstrate the effectiveness of our simulator in understanding how non-linear RRAM effects such as stuck-at-faults (SAFs), write variability, and random telegraph noise (RTN) can impact an application’s behavior. By using our simulator, we show that RTN and write variability can have adverse effects on an application. Nevertheless, we show that these effects can be mitigated through proper design choices and the implementation of a write-verify scheme. Matthew Kay Fei Lee, Yingnan Cui, Thannirmalai Somu, Tao Luo 0014, Jun Zhou 0014, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh |
ACM Trans. Archit. Code Optim. | 6 |
| 2018 | Exploiting Sparsity to Accelerate Fully Connected Layers of CNN-Based Applications on Mobile SoCsabstractConvolutional neural networks (CNNs) are widely employed in many image recognition applications. With the proliferation of embedded and mobile devices, such applications are becoming commonplace on mobile devices. Network pruning is a commonly used strategy to reduce the memory and storage footprints of CNNs on mobile devices. In this article, we propose customized versions of the sparse matrix multiplication algorithm to speed up inference on mobile devices and make it more energy efficient. Specifically, we propose a Block Compressed Sparse Column algorithm and a bit-representation-based algorithm (BitsGEMM) that exploit sparsity to accelerate the fully connected layers of a network on the NVIDIA Jetson TK1 platform. We evaluate the proposed algorithms using real-world object classification and object detection applications. Experiments show that performance speedups can be achieved over the original baseline implementation using cuBLAS. On object detection CNNs, an average speedup of 1.82× is obtained over baseline cuBLAS in the fully connected layer of the VGG model, whereas on classification CNNs, an average speedup of 1.51× is achieved for the fully connected layer of the pruned-VGG model. Energy consumption reduction of 43--46% is also observed due to decreased computational and memory bandwidth demands. Xinfeng Xie, Dayou Du, Qian Li 0027, Yun Liang 0001, Wai Teng Tang, Zhongliang Ong, Mian Lu, Huynh Phung Huynh, Rick Siow Mong Goh |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | Scale-Free Sparse Matrix-Vector Multiplication on Many-Core ArchitecturesabstractSparse matrix-vector multiplication (SpMV) is one of the most important kernels for many applications. In this paper, we study the implementation of SpMV for scale-free matrices on many-core architectures including graphic processing units and Xeon Phi coprocessors. We first propose a hardware oblivious implementation for heterogeneous many-core processors using OpenCL. Our OpenCL implementation uses a novel SpMV format called hybrid COO+CSR (HCC), which employs 2-D jagged partitioning to balance the workload among a large number of cores and improve the data locality. Moreover, the OpenCL implementation is designed to be parametric, which allows systematic performance tuning. We conduct experiments to evaluate the efficiency of our hardware oblivious implementation. Experiments show that it achieves comparable performance to the Intel MKL and state-of-the-art OpenCL-based ViennaCL library implementation. Although the OpenCL implementation provides functional portability for heterogeneous systems, it fails to take advantage of the low-level architectural features. To further improve the performance, we propose a hardware conscious implementation using the native parallel programming language. We use the Xeon Phi platform as a case study. In our hardware conscious implementation, we ensure that the HCC format efficiently utilizes the vector process units on Xeon Phi by employing low-level intrinsics, and improve the overall performance through locality-aware block mapping, and intrablock tiling. Experiments using a wide range of representative scale-free matrices demonstrate that compared with the OpenCL-based hardware oblivious implementation, the hardware conscious implementation achieves 2.2× speedup on average. Compared with MKL, the hardware conscious implementation achieves 3.1× speedup on Xeon Phi. Yun Liang 0001, Wai Teng Tang, Ruizhe Zhao, Mian Lu, Huynh Phung Huynh, Rick Siow Mong Goh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Optimizing and auto-tuning scale-free sparse matrix-vector multiplication on Intel Xeon PhiabstractRecently, the Intel Xeon Phi coprocessor has received increasing attention in high performance computing due to its simple programming model and highly parallel architecture. In this paper, we implement sparse matrix vector multiplication (SpMV) for scale-free matrices on the Xeon Phi architecture and optimize its performance. Scale-free sparse matrices are widely used in various application domains, such as in the study of social networks, gene networks and web graphs. We propose a novel SpMV format called vectorized hybrid COO+CSR (VHCC). Our SpMV implementation employs 2D jagged partitioning, tiling and vectorized prefix sum computations to improve hardware resource utilization, and thus overall performance. As the achieved performance depends on the number of vertical panels, we also develop a performance tuning method to guide its selection. Experimental results demonstrate that our SpMV implementation achieves an average 3× speedup over Intel MKL for a wide range of scale-free matrices. Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang 0001, Huynh Phung Huyng, Xibai Li, Rick Siow Mong Goh |
CGO | 1 |
| 2015 | A Code Generation Framework for Targeting Optimized Library Calls for Multiple PlatformsabstractDirective-based programming approaches such as OpenMP and OpenACC have gained popularity due to their ease of programming. These programming models typically involve adding compiler directives to code sections such as loops in order to parallelize them for execution on multicore CPUs or GPUs. However, one problem with this approach is that existing compilers generate code directly from the annotated sections and do not make use of hardware-specific architectural features. As a result, the generated code is unable to fully exploit the capabilities of the underlying hardware. Alternatively, we propose a code generation framework in which linear algebraic operations in the annotated codes are recognized, extracted and mapped to optimized vendor-provided platform-specific library calls. We demonstrate that such an approach can result in better performance in the generated code compared to those which are generated by existing compilers. This is substantiated by experimental results on multicore CPUs and GPUs. Wen Jun Tan, Wai Teng Tang, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPUabstractSparse matrix-vector multiplication (SpMV) is an important kernel that is used in many iterative algorithms for solving scientific and engineering problems. One of the main challenges of SpMV is its memory-boundedness due to the low arithmetic intensity of the kernel. Although compression has been proposed previously to improve SpMV performance on CPUs, its use has not been demonstrated on the GPU because of the serial nature of many compression and decompression schemes. In this paper, we introduce a family of bit-representation-optimized (BRO) compression formats for representing sparse matrices on GPUs. The proposed formats - BRO-CSR, BRO-ELL and BRO-HYB, perform compression on index data and help to speed up SpMV on GPUs through the reduction of memory traffic. We also propose two other hybrid BRO formats which can potentially perform better than both HYB and BRO-HYB formats. Experimental results demonstrate that compared to uncompressed CSR and ELLPACK formats, our proposed compressed BRO-CSR and BRO-ELL formats are able to achieve average speedups of 2× and 1.4× respectively. Furthermore, we demonstrate that by using BRO-ELL, the preconditioned conjugate gradient method is able to achieve an average speedup of 1.3× over ELLPACK. Wai Teng Tang, Wen Jun Tan, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Optimizing and Auto-Tuning Iterative Stencil Loops for GPUs with the In-Plane MethodabstractStencils represent an important class of computations that are used in many scientific disciplines. Increasingly, many of the stencil computations in scientific applications are being offloaded to GPUs to improve running times. Since a large part of the simulation time is spent inside the stencil kernels, optimizing the kernel is therefore important in the context of achieving greater computation efficiencies and reducing simulation time. In this work, we proposed a novel in-plane method for stencil computations on GPUs and compared its performance with the conventional method implemented in the Nvidia SDK. We also implemented an auto-tuning framework for our method to select the optimal parameters for different GPU architectures. A performance model was developed for our proposed method, and is used to speed up the auto-tuning process. Our results show that a speedup of nearly 2× can be achieved compared to Nvidia's implementation. Wai Teng Tang, Wen Jun Tan, Ratna Krishnamoorthy, Yi Wen Wong, Shyh-Hao Kuo, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong |
IPDPS | 1 |
| 2013 | Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemesabstractThe sparse matrix-vector (SpMV) multiplication routine is an important building block used in many iterative algorithms for solving scientific and engineering problems. One of the main challenges of SpMV is its memory-boundedness. Although compression has been proposed previously to improve SpMV performance on CPUs, its use has not been demonstrated on the GPU because of the serial nature of many compression and decompression schemes. In this paper, we introduce a family of bit-representation-optimized (BRO) compression schemes for representing sparse matrices on GPUs. The proposed schemes, BRO-ELL, BRO-COO, and BRO-HYB, perform compression on index data and help to speed up SpMV on GPUs through reduction of memory traffic. Furthermore, we formulate a BRO-aware matrix reordering scheme as a data clustering problem and use it to increase compression ratios. With the proposed schemes, experiments show that average speedups of 1.5x compared to ELLPACK and HYB can be achieved for SpMV on GPUs. Wai Teng Tang, Wen Jun Tan, Rajarshi Ray 0001, Yi Wen Wong, Weiguang Chen, Shyh-Hao Kuo, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong |
SC | 1 |
| 2012 | Tulipse: A Visualization Framework for User-Guided Parallelization
Yi Wen Wong, Tomasz Dubrownik, Wai Teng Tang, Wen Jun Tan, Rubing Duan, Rick Siow Mong Goh, Shyh-Hao Kuo, Stephen John Turner, Weng-Fai Wong |
Euro-Par | 3 |
| 2012 | Automatic Refactoring of Legacy Fortran Code to the Array Slicing NotationabstractThere are many legacy Fortran programs still in use today, especially scientific codes which were written decades ago. Many of these codes use explicit DO-loops in programs that tend to clutter the code and make it harder to understand and maintain. Modern features of the Fortran language, such as the array slicing notation and introduction of commonly used intrinsic functions, go a long way in helping programmers write code that is easier to read and maintain. We introduce a refactoring tool that can help to transform code to make use of the array slicing notation and related intrinsic functions. Chandrasehar Rajaseharan, Wen Jun Tan, Wai Teng Tang, Stephen John Turner, Shyh-Hao Kuo, Rick Siow Mong Goh, Weng-Fai Wong |
ICPADS | 3 |