EDBT 2026 Demo / reviewers in the wild / expert
Tuan Ta
dblp:122/4949
· DBLP profile ↗
8ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0001-8961-0005ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EVE: Ephemeral Vector EnginesabstractThere has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead. Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady O. Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, Christopher Batten |
HPCA | 2 |
| 2022 | big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on ChipabstractSingle-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads. Tuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou, Eric Hall, Courtney Golden, Christopher Batten |
MICRO | 1 |
| 2021 | Power Delivery and Thermal-Aware Arm-Based Multi-Tier 3D Architectureabstract3D integration is becoming a cost-effective way to incorporate more CPU cores and memory to improve the performance of computing systems. Meanwhile, due to the higher power density, power delivery and thermal issues become more significant in multi-tier 3DICs. In this paper, we explore and evaluate multiple design options for an Arm Neoverse-based 3D architecture focusing on power and thermals at 7nm process and sub-10$\mu $m pitch. Using a rapid voltage-drop and thermal analysis methodology, we model a system with a 32-core CPU layer and up to 4 layers of system-level caches, and quantity the trade-offs between performance, cost, voltage-drop, and temperature. A 3-layer configuration shows a good balance with 17% IPC gain and 17% lower cost, while incurring 15mV worse voltage drop and 8.5°C higher temperature compared with 2D. Our studies suggest that the co-optimization of system architecture, technology, and physical design is key for high-performance 3D systems. Lingjun Zhu, Tuan Ta, Rossana Liu, Rahul Mathur, Shidhartha Das, Ankit Kaul, Alejandro Rico, Doug Joseph, Brian Cline, Sung Kyu Lim |
ISLPED | 2 |
| 2020 | Efficiently Supporting Dynamic Task Parallelism on Heterogeneous Cache-Coherent SystemsabstractManycore processors, with tens to hundreds of tiny cores but no hardware-based cache coherence, can offer tremendous peak throughput on highly parallel programs while being complexity and energy efficient. Manycore processors can be combined with a few high-performance big cores for executing operating systems, legacy code, and serial regions. These systems use heterogeneous cache coherence (HCC) with hardware-based cache coherence between big cores and software-centric cache coherence between tiny cores. Unfortunately, programming these heterogeneous cache-coherent systems to enable collaborative execution is challenging, especially when considering dynamic task parallelism. This paper seeks to address this challenge using a combination of light-weight software and hardware techniques. We provide a detailed description of how to implement a work-stealing runtime to enable dynamic task parallelism on heterogeneous cache-coherent systems. We also propose direct task stealing (DTS), a new technique based on user-level interrupts to bypass the memory system and thus improve the performance and energy efficiency of work stealing. Our results demonstrate that executing dynamic task-parallel applications on a 64-core system (4 big, 60 tiny) with complexity-effective HCC and DTS can achieve: $7 \times$ speedup over a single big core; $1.4 \times$ speedup over an area-equivalent eight bigcore system with hardware-based cache coherence; and 21% better performance and similar energy efficiency compared to a 64-core system (4 big, 60 tiny) with full-system hardware-based cache coherence. Moyang Wang, Tuan Ta, Christopher Batten |
ISCA | 2 |
| 2019 | A Specialized Concurrent Queue for Scheduling Irregular Workloads on GPUsabstractThe persistent thread model offers a viable solution for accelerating data-irregular workloads on Graphic Processing Units (GPUs). However, as the number of active threads increases, contention and retries on shared resources limit the efficiency of task scheduling among the persistent threads. To address this, we propose a highly scalable, non-blocking concurrent queue suitable for use as a GPU persistent thread task scheduler. The proposed concurrent queue has two novel properties: 1) The supporting enqueue/dequeue queue operations never suffer from retry overhead because the atomic operation does not fail and the queue empty exception has been refactored; and 2) The queue operates on an arbitrary number of queue entries for the same cost as a single entry. A proxy thread in each thread group performs all atomic operations on behalf of all threads in the group. These two novel properties substantially reduce thread contention caused by the GPU's lock-step Single Instruction Multiple Threads (SIMT) execution model. David Troendle, Tuan Ta, Byunghyun Jang |
ICPP | 2 |
| 2017 | Understanding the Impact of Fine-Grained Data Sharing and Thread Communication on Heterogeneous Workload DevelopmentabstractThe conventional OpenCL 1.x style CPU-GPU heterogeneous computing paradigm treats the CPU and GPU processors as loosely connected separate entities. At best each executes independent tasks, but, more commonly, the CPU idles while waiting for results from the GPU. No data-sharing and communications are allowed during kernel execution. This model limits the number of applications that can harness the tremendous computing power of two processors. OpenCL 2.x and compliant hardware introduce a new memory model that enables a new computing paradigm where the task-parallel CPU and data-parallel GPU are tightly coupled and can closely cooperate on shared data in a lock-based or non-blocking fashion. This new model maximizes hardware utilization and performance, and opens more applications to GPU acceleration. The most significant new OpenCL 2.x features are fine-grained data sharing and thread communication through shared virtual memory, CPU-GPU cache coherence support, and system-level atomics. However, few applications that can exploit the benefits of tightly coupled CPU-GPU heterogeneous processors have emerged. Programming and debugging in this new environment have proven challenging. The resulting lack of benchmark workloads have also left hardware architects uninformed. To facilitate truly heterogeneous workload development and hardware architecture research, this paper focuses on understanding the impact of fine-grained data sharing between the CPU and GPU for future heterogeneous workload development. To that goal, we identify three CPU-GPU cooperation paradigms, demonstrate their performance benefits on real hardware using both in-house and publicly available benchmarks, and profile their detailed behavior and characteristics using an architectural simulator. Our experiments demonstrate that truly heterogeneous implementations of our studied benchmarks outperform their corresponding conventional CPU or GPU versions by up to 59.5% and 36.6% respectively. We analyze thread contention problems, latency of synchronization operations and inter-cluster memory traffic in each cooperation paradigm using a timing architectural simulator. Tuan Ta, David Troendle, Xiaoqi Hu, Byunghyun Jang |
ISPDC | 1 |
| 2014 | Improving smartphone battery life utilizing device-to-device cooperative relays underlaying LTE networksabstractThe utility of smartphones has been limited to a great extent by their short battery life. In this work, we propose a new approach to prolonging smartphone battery life. We introduce the notions of valueless and valued battery, as being the available battery when the user does or does not have access to a power source, respectively. We propose a cooperative system where users with high battery level help carry the traffic of users with low battery level. Our scheme helps increase the amount of valued battery in the network, thus it reduces the chance of users running out of battery early. Our system can be realized in the form of a proximity service (ProSe) which utilizes a device-to-device (D2D) communication architecture underlaying LTE. We show through simulations that our system reduces the probability of cellular users running out of battery before their target usage time (probability of outage). Our simulator source code is made available to the public. Tuan Ta, John S. Baras |
ICC | 1 |
| 2012 | Wormhole detection using channel characteristicsabstractThe potential applications and pervasive nature of mobile ad-hoc networks (MANETs) has made them an attractive target for attackers. The wireless medium of communication coupled with constrained resources enable attacks which can be executed by a weak adversary. A wormhole is one such attack which poses considerable threat, particularly to routing protocols. In this paper, we devise a novel scheme for detecting a wormhole by utilizing the inherent symmetry of electromagnetic wave propagation in the wireless medium. We demonstrate the loss of this symmetry in case of a wormhole attack and propose a method to detect and flag the adversary. We modify the insecure neighborhood discovery to incorporate authentication. We further extend this scheme to a trust system with low overhead. Shalabh Jain, Tuan Ta, John S. Baras |
ICC | 2 |