Xulong Tang

dblp:66/10956 · DBLP profile ↗
← Back
69ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0002-3385-2053ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 51 · 4 first-author · 37 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Potential of Layer Freezing for DNN Training Efficiency
abstract
With the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training.
Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan
ACM Great Lakes Symposium on VLSI8
2026 Personalized Dance Synthesis Based on Physical and Cognitive Intensities
abstract
Dance-based exergames like Just Dance can be a fun way to boost your fitness and sharpen your mind. However, designing the dance routines requires expertise in modeling and animation. We introduce an augmented reality (AR) personalized dance generation framework that synthesizes dance routines according to specified physical and cognitive intensities. Our system utilizes a curated library of motion-capture dance segments, which are intelligently combined through an optimization process to meet user-defined intensity and cognitive goals. This optimization also ensures smooth transitions between movements for natural dance flow. Users can customize routines by specifying physical constraints or injuries. Implemented in a depth-camera-based exergame that provides real-time performance feedback, our framework was evaluated through experiments and user studies confirming its effectiveness in generating personalized routines with varying levels of physical and cognitive intensity.
Xulong Tang, Eun Yeo, Ruiyu Mao, Xiaohu Guo, Rawan Alghofaili
VR1
2025 A Computation and Energy Efficient Hardware Architecture for SSL Acceleration
abstract
In Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency.
Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou
ASP-DAC9
2025 Cascade: A Dependency-aware Efficient Training Framework for Temporal Graph Neural Network
abstract
Temporal graph neural networks (TGNN) have gained significant momentum in many real-world dynamic graph tasks. These models use graph changes (i.e., events) as inputs to update nodes' status vectors (i.e., memories), which are then exploited to assist predictions. Despite their improved accuracies, the efficiency of TGNN training is significantly limited due to the inherent temporal relationship between the input events. Although larger training batches can improve parallelism and speed up TGNN training, they lead to infrequent memory updates, which cause outdated information and reduced accuracy. This trade-off forces current methods to use small batches, resulting in high latency and underutilized hardware. To address this, we propose an efficient TGNN training framework, Cascade, to adaptively boost TGNN training parallelism based on nodes' spatial and temporal dependencies. Cascade adopts a topology-aware scheduler that includes as many spatial-independent events in the same batches. Moreover, it leverages node memories' similarities to break temporal dependencies on stabilized nodes, enabling it to pack more temporal-independent events in the same batches. Additionally, Cascade adaptively decides nodes' update frequencies based on runtime feedback. Compared to prior state-of-the-art TGNN training frameworks, our approach can averagely achieve 2.3x (up to 5.1x) speed up without jeopardizing the resulted models' accuracy.
Yue Dai 0005, Xulong Tang, Youtao Zhang
ASPLOS (2)2
2025 Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning
abstract
Tensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically finding high-performance programs for specific hardware. However, the search process is often inefficient, taking hours or even days to discover optimal programs due to the exploration mechanisms guided by an accurate but slow-learned cost model. Meanwhile, the learned cost model trained on one platform cannot seamlessly adapt online to another, which we call cross-platform online unawareness. In this work, we propose Pruner and MoA-Pruner. Pruner is a ''Draft-then-Verify'' exploration mechanism that accelerates the schedule search process. Instead of applying the complex learned cost model to all explored candidates, Pruner drafts small-scale potential candidates by introducing a naive Symbol-based Analyzer (draft model), then identifies the best candidates by the learned cost model. MoA-Pruner introduces a Momentum online Adaptation strategy to address the cross-platform online unawareness.
Jun Shi 0007, Minfan Zhao, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan
ASPLOS (2)10
2025 OASIS: Object-Aware Page Management for Multi-GPU Systems
abstract
The ever-growing need for high-performance computing has driven the popularity of employing multi-GPU systems. Modern multi-GPU systems employ unified virtual memory (UVM) to manage page placement and migration. However, the page management in UVM is application object agnostic. In this paper, we characterize the page access behaviors in relation to the application objects, and reveal that the beneficial page management policy varies according to (i) the different data objects within the same application, and (ii) the different execution phases of the same object. This motivates the need for dynamic and proactive page management in multi-GPU systems. To this end, we propose OASIS, which dynamically identifies object patterns during the execution and proactively determines the appropriate page management policies for these objects at runtime. Experimental results show that OASIS improves the performance over uniformly adopting on-touch migration, access counter-based migration, and duplication by an average of $\mathbf{6 4 \%}$, $\mathbf{3 5 \%}$, and $\mathbf{4 2 \%}$, respectively. Moreover, OASIS achieves a $\mathbf{1 2 \%}$ performance improvement over the state-of-the-art technique (i.e., GRIT) while having significantly lower design complexity.
Bingyao Li 0001, M. Tarek Ibn Ziad, Lieven Eeckhout, Jun Yang 0002, Aamer Jaleel, Xulong Tang
HPCA7
2025 STMC: Small-Tile Multiple-Copy Compilation for Reliable Measurement-Based Quantum Computing
abstract
Measurement-based Quantum Computing (MBQC) achieves universal quantum computing by applying measurements on the photonic architectures. While it has many advantages, such as long qubit decoherence time and strong scalability, the success rate of MBQC execution is constrained by imperfect photon control, measurement, and fusion operations. Both fusion failure and photon loss necessitate the re-execution of the entire quantum circuit, leading to significant overhead in terms of additional execution time and increased consumption of resource state layers. Recent studies mainly focus on mitigating fusion failures and little attention has been paid to photon loss. In this paper, we propose STMC (Small-Tile Multiple-Copy) compilation framework to reduce the re-execution overhead caused by both the fusion failure and photon loss. Specifically, STMC first transforms a quantum circuit into a fusion graph and partitions the fusion graph into subgraphs. Then, STMC generates compact subgraph mappings that are appropriate for the size of a subportion in the resource state layer, referred to as a tile. Finally, STMC employs multiple copies of each subgraph when mapping to tiles, duplicates the execution of tiles in parallel, and finishes the whole circuit execution in order. The experimental results demonstrate that STMC achieves an average execution time speedup of 65.68× for successfully executing the circuit under a 75% fusion success rate, compared to prior work. Additionally, STMC reduces the number of resource state layers by three orders of magnitude and decreases the number of resource states by an average of 36.40×.
Rongchao Dong, Zewei Mo, Yingheng Li, Aditya Pawar, Jun Yang 0002, Youtao Zhang, Xulong Tang
ICCAD7
2025 Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised Learning
abstract
Self-supervised learning (SSL) offers a compelling solution to the challenge of extensive labeled data requirements in traditional supervised learning. With the proven success of Vision Transformers (ViTs) in supervised tasks, there is increasing interest in adapting them for SSL frameworks. However, the high computational demands of SSL pose substantial challenges, particularly on resource-limited platforms like edge devices, despite its ability to achieve high accuracy without labeled data. Recent studies in supervised learning have shown that token pruning can reduce training costs by removing less informative tokens without compromising accuracy. However, SSL’s dual-branch encoders make traditional single-branch pruning strategies less effective, as they fail to account for the critical cross-branch similarity information, leading to reduced accuracy in SSL. To this end, we introduce SimPrune, a novel token pruning strategy designed for ViTs in SSL. SimPrune leverages cross-branch similarity information to efficiently prune tokens, retaining essential semantic information across dual branches. Additionally, we incorporate a difficulty-aware pruning strategy to further enhance SimPrune's effectiveness. Experimental results show that our proposed approach effectively reduces training computation while maintaining accuracy. Specifically, our approach offers 24\% savings in training costs compared to SSL baseline, without sacrificing accuracy.
Sheng Li 0019, Qitao Tan, Yue Dai 0005, Zhenglun Kong, Jun Liu 0075, Ao Li 0004, Ninghao Liu 0001, Yufei Ding 0001, Xulong Tang, Geng Yuan
ICLR10
2025 MemFreezing: A Novel Adversarial Attack on Temporal Graph Neural Networks under Limited Future Knowledge
abstract
Temporal graph neural networks (TGNN) have achieved significant momentum in many real-world dynamic graph tasks. While most existing TGNN attack methods assume worst-case scenarios where attackers have complete knowledge of the input graph, the assumption may not always hold in real-world situations, where attackers can, at best, access information about existing nodes and edges but not future ones after the attack. However, studying adversarial attacks under these constraints is crucial, as limited future knowledge can reveal TGNN vulnerabilities overlooked in idealized settings. Nevertheless, designing effective attacks in such scenarios is challenging: the evolving graph can weaken their impact and make it hard to affect unseen nodes. To address these challenges, we introduce MemFreezing, a novel adversarial attack framework that delivers long-lasting and spreading disruptions in TGNNs without requiring post-attack knowledge of the graph. MemFreezing strategically injects fake nodes or edges to push node memories into a stable “frozen state,” reducing their responsiveness to subsequent graph changes and limiting their ability to convey meaningful information. As the graph evolves, these affected nodes maintain and propagate their frozen state through their neighbors. Experimental results show that MemFreezing persistently degrades TGNN performance across various tasks, offering a more enduring adversarial strategy under limited future knowledge.
Yue Dai 0005, Xulong Tang, Youtao Zhang, Jun Yang 0002
ICML3
2025 CIExplorer: Microarchitecture-Aware Exploration for Tightly Integrated Custom Instruction
abstract
Extending existing architectures with customized instruction extensions is emerging to achieve high performance and energy efficiency for specific applications.Automated discovery of custom instructions (CIs) is well-studied nowadays, which requires exploring combinations of different types and quantities of operations, resulting in a vast search space.However, previous works typically use microarchitectureagnostic cost models, leading to suboptimal CIs that may degrade performance.They leverage graph isomorphism to reduce area overhead, but few of them consider its potential to benefit performance-oriented exploration.To this end, we present CIExplorer, a framework for adaptive CI exploration.
Qingcai Jiang, Jun Shi 0007, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan
ICS8
2025 Reinforcement Learning-Guided Graph State Generation in Photonic Quantum Computers
abstract
The photonic quantum computer (PQC) is an emerging and promising quantum computing paradigm that has gained momentum in recent years.In PQC, computations are executed by performing measurements on photons in graph states (i.e., a collection of entangled photons).The graph state generation process is fulfilled by applying a sequence of quantum gates to quantum emitters, referred to as the "generation sequence".In a generation sequence, i) the time required to complete the generation sequence, ii) the number of quantum emitters used, and iii) the number of CZ gates performed between emitters greatly affect the fidelity of the generated graph state.In this paper, we propose RLGS (Reinforcement Learningguided Graph State generation), a novel compilation framework to identify optimal generation sequences that optimize the three fidelity metrics.Experimental results show that RLGS achieves an average reduction in generation time of 31.1%,49.6%, and 57.5% for small, medium, and large graph states compared to the baseline.Additionally, the reductions in the number of quantum emitters are 13.9%, 16.7%, and 17.5%, whereas the reductions in the number of CZ gates are 37.7%, 53.4%, and 57.8%, respectively.
Yingheng Li, Yue Dai 0005, Aditya Pawar, Rongchao Dong, Jun Yang 0002, Youtao Zhang, Xulong Tang
ISCA7
2025 CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
abstract
Music-Driven Dance Generation seeks to create dance movements synchronized with music, playing a key role in applications like performance and gaming. While solo dance generation has seen progress, group dance generation remains underexplored. Although several methods have been proposed, existing approaches frequently fail to ensure spatial-temporal coherence, resulting in unrealistic and aesthetically unpleasing performances. To tackle the issue, we introduce CoheDancers, a novel framework for Music-Driven Interactive Group Dance Generation. CoheDancers aims to enhance group dance generation coherence by decomposing it into three key aspects: synchronization, naturalness, and fluidity. Correspondingly, we develop a Cycle Consistency based Dance Synchronization strategy to foster music-dance correspondences, an Auto-Regressive-based Exposure Bias Correction strategy to enhance the fluidity of the generated dances, and an Adversarial Training Strategy to augment the naturalness of the group dance output. Collectively, these strategies enable CoheDancers to produce highly coherent group dances with superior quality. Furthermore, to establish better benchmarks for Group Music2Dance, we construct the most diverse and comprehensive open-source dataset to date, I-Dancers, featuring rich dancer interactions, and create comprehensive evaluation metrics. Experimental evaluations on I-Dancers and other extant datasets substantiate that CoheDancers achieves unprecedented state-of-the-art performance. Code is available at https://github.com/XulongT/CoheDancers.
Kaixing Yang, Xulong Tang, Biao Qin, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan
ACM Multimedia2
2025 MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
abstract
Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underutilize genre conditioning, often treating it as auxiliary modifiers rather than core semantic drivers. This oversight compromises music-motion synchronization and disrupts dance genre continuity, particularly during complex rhythmic transitions, thereby leading to visually unsatisfactory effects. To address the challenge, we propose MEGADance, a novel architecture for music-driven 3D dance generation. By decoupling choreographic consistency into dance generality and genre specificity, MEGADance demonstrates significant dance quality and strong genre controllability. It consists of two stages: (1) High-Fidelity Dance Quantization Stage (HFDQ), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) and reconstructs them with kinematic-dynamic constraints, and (2) Genre-Aware Dance Generation Stage (GADG), which maps music into the latent representation by synergistic utilization of Mixture-of-Experts (MoE) mechanism with Mamba-Transformer hybrid backbone. Extensive experiments on the FineDance and AIST++ dataset demonstrate the state-of-the-art performance of MEGADance both qualitatively and quantitatively. Code is available at https://github.com/XulongT/MEGADance.
Kaixing Yang, Xulong Tang, Ziqiao Peng, Jun He 0008, Hongyan Liu 0002
NeurIPS2
2024 FMCC: Flexible Measurement-based Quantum Computation over Cluster State
abstract
Measurement-based quantum computing (MBQC) is a promising quantum computing paradigm that performs computation through "one-way" measurements on entangled quantum qubits. It is widely used in photonic quantum computing (PQC), where the computation is carried out on photonic cluster states (i.e., a 2-D mesh of entangled photons). In MBQC-based PQC, the cluster state depth (i.e., the length of one-way measurements) plays an important role in the overall execution time and circuit error. In this paper, we propose FMCC, a compilation framework that employs dynamic programming with heuristics to efficiently minimize the cluster state depth. Experimental results on six quantum applications show that FMCC achieves 51.7%, 57.4%, and 56.8% average depth reductions in small, medium, and large qubit counts compared to the state-of-the-art MBQC compilations.
Yingheng Li, Aditya Pawar, Zewei Mo, Youtao Zhang, Jun Yang 0002, Xulong Tang
ASPLOS (4)6
2024 QRCC: Evaluating Large Quantum Circuits on Small Quantum Computers through Integrated Qubit Reuse and Circuit Cutting
abstract
Quantum computing has recently emerged as a promising computing paradigm for many application domains. However, the size of quantum circuits that can be run with high fidelity is constrained by the limited quantity and quality of physical qubits. Recently proposed schemes, such as wire cutting and qubit reuse, mitigate the problem but produce sub-optimal results as they address the problem individually. In addition, gate cutting, an alternative circuit-cutting strategy that is suitable for circuits computing expectation values, has not been fully explored in the field.
Aditya Pawar, Yingheng Li, Zewei Mo, Yanan Guo 0002, Xulong Tang, Youtao Zhang, Jun Yang 0002
ASPLOS (4)5
2024 LOTUS: learning-based online thermal and latency variation management for two-stage detectors on edge devices
abstract
Two-stage object detectors exhibit high accuracy and precise localization, especially for identifying small objects that are favorable for various edge applications. However, the high computation costs associated with two-stage detection methods cause more severe thermal issues on edge devices, incurring dynamic runtime frequency change and thus large inference latency variations. Furthermore, the dynamic number of proposals in different frames leads to various computations over time, resulting in further latency variations. The significant latency variations of detectors on edge devices can harm user experience and waste hardware resources. To avoid thermal throttling and provide stable inference speed, we propose Lotus, a novel framework that is tailored for two-stage detectors to dynamically scale CPU and GPU frequencies jointly in an online manner based on deep reinforcement learning (DRL). To demonstrate the effectiveness of Lotus, we implement it on NVIDIA Jetson Orin Nano and Mi 11 Lite mobile platforms. The results indicate that Lotus can consistently and significantly reduce latency variation, achieve faster inference, and maintain lower CPU and GPU temperatures under various settings. Our code is available at [link].
Yifan Gong 0004, Yushu Wu, Zheng Zhan 0001, Pu Zhao 0001, Liangkai Liu, Chao Wu 0006, Xulong Tang, Yanzhi Wang 0001
DAC7
2024 FCM: A Fusion-aware Wire Cutting Approach for Measurement-based Quantum Computing
abstract
Measurement-based quantum computing (MBQC) is a promising quantum computing paradigm that carries out computation through one-way measurements on entangled photon qubits. Practical photonic hardware first generates a 2D mesh of resource states with each being a small number of entangled photon qubits and then exploits fusion operations to connect resource states to scale up the computation. Given that the fusion operation is highly error-prone, it is important to reduce the number of fusions for an MBQC circuit.
Zewei Mo, Yingheng Li, Aditya Pawar, Xulong Tang, Jun Yang 0002, Youtao Zhang
DAC4
2024 GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page Placement
abstract
Multi-GPU systems have become popular to cater to the growing demands for high parallelism and large memory capacity. However, the delivered performance is constrained by the non-uniform memory access (NUMA) overhead arising from data sharing and communication across multiple GPUs. Recent multi-GPUs employ unified virtual memory (UVM) to simplify the programming effort. In UVM-enabled multi-GPUs, three popular page placement schemes are adopted to mitigate the NUMA overheads: i) on-touch page migration, ii) access counter-based migration, and iii) page duplication. However, we observe that the preferred page placement scheme varies across i) different applications, ii) different pages of the same application, and iii) even different execution phases of a single page, making it challenging to find a “one-size-fits-all” page placement scheme. To this end, we propose GRIT, which dynamically and automatically determines the appropriate page placement schemes at runtime in a fine-grained manner to enhance multi-GPU performance and scalability. Experimental results indicate that GRIT achieves an average of 60%, 49%, and 29% performance improvements over uniformly adopting on-touch migration, access counter-based migration, and page duplication, respectively.
Bingyao Li 0001, Aamer Jaleel, Jun Yang 0002, Xulong Tang
HPCA5
2024 Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised Learning
abstract
Deep Neural Networks (DNNs), essential for diverse applications such as visual recognition and eldercare, often require a large amount of labeled data for training, making widespread deployment of DNNs a challenging task. Self-supervised learning (SSL) emerges as a promising approach, which leverages inherent patterns within data through diverse augmentations to train models without explicit labels. However, while SSL has shown notable advancements in accuracy, its high computation costs remain a daunting impediment, particularly for resource-constrained platforms. To address this problem, we introduce SimWnW, a similarity-based efficient self-supervised learning framework. By strategically removing less important regions in augmented images and feature maps, SimWnW not only reduces computation costs but also eliminates irrelevant features that might slow down the learning process, thereby accelerating model convergence. The experimental results show that SimWnW effectively reduces the amount of computation costs in self-supervised model training without compromising accuracy. Specifically, SimWnW yields up to 54\% and 51\% computation savings in training from scratch and transfer learning tasks, respectively.
Sheng Li 0019, Chao Wu 0006, Ao Li 0004, Yanzhi Wang 0001, Xulong Tang, Geng Yuan
ICLR5
2024 STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPU
abstract
NVIDIA's Multi-Instance GPU (MIG) technology enables partitioning GPU computing power and memory into sep-arate hardware instances, providing complete isolation including compute resources, caches, and memory. However, prior work identifies that MIG does not partition the last-level TLB (i.e., L3 TLB), which remains shared among all instances. To enhance TLB reach, NVIDIA GPUs reorganized the TLB structure with 16 sub-entries in each L3 TLB entry that have a one-to-one mapping to the address translations for 16 pages of size 64 KB located within the same 1 MB aligned range. Our comprehensive investigation of address translation efficiency in MIG identifies two main issues caused by L3 TLB sharing interference: (i) it results in performance degradation for co-running applications, and (ii) TLB sub-entries are not fully utilized before eviction. Based on this observation, we propose STAR to improve the utilization of TLB sub-entries through dynamic sharing of TLB entries across multiple base addresses. STAR evaluates TLB entries based on their sub-entry utilization to optimize address translation storage, dynamically adjusting between a shared and non-shared state to cater to current demand. We show that STAR improves overall performance by an average of 28.7% across various multi-tenant workloads.
Bingyao Li 0001, Lieven Eeckhout, Jun Yang 0002, Aamer Jaleel, Xulong Tang
MICRO7
2024 CoDancers: Music-Driven Coherent Group Dance Generation with Choreographic Unit
abstract
Dance and music are intimately interconnected, with group dance being a crucial part of dance artistry. Consequently, Music-Driven Group Dance Generation has been a fundamental and challenging task in various fields like education, art, and sports. However, existing methods fail to fully explore group dance coherence. Thus, we propose CoDancers, a novel and efficient retrieval-based music-driven group dance generation framework. CoDancers improves performance by decomposing group dance coherence into individual movement coherence and group interaction coherence for specialized design, incorporating a Spatial-Temporal Group Dance Blender block, a Acoustic-Semantic Music Miner block, and a Stereotype-Reducing Dance Generator block. Experimental results on the public dataset demonstrate the superiority of our method over existing baselines, achieving state-of-the-art performance. The code is available at https://github.com/XulongT/CoDancers.
Kaixing Yang, Xulong Tang, Ran Diao, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan
ICMR2
2024 BeatDance: A Beat-Based Model-Agnostic Contrastive Learning Framework for Music-Dance Retrieval
abstract
Dance and music are closely related forms of expression, with mutual retrieval between dance videos and music being a fundamental task in various fields like education, art, and sports. However, existing methods often suffer from unnatural generation effects or fail to fully explore the correlation between music and dance. To overcome these challenges, we propose BeatDance, a novel beat-based model-agnostic contrastive learning framework. BeatDance incorporates a Beat-Aware Music-Dance InfoExtractor, a Trans-Temporal Beat Blender, and a Beat-Enhanced Hubness Reducer to improve Music-Dance retrieval performance by utilizing the alignment between music beats and dance movements. We also introduce the Music-Dance (M-D) dataset, a large-scale collection of over 10,000 Music-Dance video pairs for training and testing. Experimental results on the M-D dataset demonstrate the superiority of our method over existing baselines, achieving state-of-the-art performance. The code and dataset are available at https://github.com/XulongT/BeatDance.
Kaixing Yang, Xukun Zhou, Xulong Tang, Ran Diao, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan
ICMR3
2023 Orchestrating Measurement-Based Quantum Computation over Photonic Quantum Processors
abstract
Quantum computing has rapidly evolved in recent years and has established its supremacy in many application domains. While matter-based qubit platforms such as superconducting qubits have received the most attention so far, there is a rising interest in photonic qubits lately, which show advantages in parallelism, speed, and scalability. Photonic qubits are best served by the paradigm of measurement-based quantum computation (MBQC). To deliver the promise of measurement-based photonic quantum computing (MBPQC), the photon cluster state depth and photon utilization are two of the most important metrics. However, little attention has been paid to optimizing the depth and utilization when mapping quantum circuits to the photon clusters. In this paper, we propose a compiler framework that achieves automatic and dynamic depth and utilization optimizations. Our approach consists of an MBPQC mapping mechanism that maps optimized measurement patterns on a cluster state and a cluster state pruning strategy that removes all possible redundancies without impacting the circuit functions. Experimental results on five quantum benchmark with three different qubit numbers indicate our approach achieves an average of 63.4% cluster depth reduction and 22.8% photon utilization improvements.
Yingheng Li, Aditya Pawar, Mohadeseh Azari, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002, Kaushik Parasuram Seshadreesan, Xulong Tang
DAC8
2023 Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUs
abstract
Unified Virtual Memory (UVM) is a promising feature in CPU-GPU heterogeneous systems that allows data structures to be accessed by both CPU and GPUs through unified pointers without explicit data copying. However, the delivered performance of UVM significantly relies on the efficiency of address translation. The current GPU thread block (TB) management is not aware of the translation process and heavily thrashes the per-streaming multiprocessor (SM) private Translation Look-ahead Buffers (TLBs). In this paper, we conduct a comprehensive characterization of 10 GPU benchmarks and quantify the translation reuses among the thread blocks. Our observation reveals that there exists substantial translation reuse within TBs rather than across the TBs. Moreover, the inter-TB interference significantly enlarges the intra-TB translation reuse distances. To this end, we propose a translation-aware TB scheduling and lightweight GPU L1 TLB partitioning to effectively mitigate the contention. Experimental results show that our proposed approach improves the L1 TLB hit rate, and this improvement translates to, on average, a 12.5% execution time reduction.
Bingyao Li 0001, Xulong Tang
DAC3
2023 EP-ORAM: Efficient NVM-Friendly Path Eviction for Ring ORAM in Hybrid Memory
abstract
Recent studies showed that only ORAM (oblivious RAM) can securely protect memory access patterns (i.e., data privacy) on modern computer systems. Ring ORAM is a promising ORAM protocol as it demands O(1) memory accesses for servicing each user memory request. However, Ring ORAM exhibits low memory utilization, i.e., its memory requirement is 4.8× of the protected user space. While adopting NVM (non-volatile memory) can alleviate the memory requirement, a simple implementation tends to introduce large performance degradation, preventing its adoption in practice.In this paper, we propose EP-ORAM, an NVM-friendly Ring ORAM implementation on DRAM/NVM hybrid memory. EP-ORAM is developed based on two key observations: (1) for tree-based Ring ORAM memory organization, saving bottom levels in NVM can dramatically reduce the DRAM memory requirement; (2) the tradeoffs among Ring ORAM operations expose design opportunities without security compromise. We, therefore, propose to save the bottom levels of the ORAM tree in NVM and shorten the path of EvictPath operation, which not only mitigates the number of NVM writes but also speeds up the execution. Our experimental results show that, under the design constraints of similar performance as the baseline that saves two bottom levels in NVM, EP-ORAM helps to save three levels in NVM, achieving 50% DRAM space reduction. In addition, EP-ORAM reduces the NVM writes by 15%.
Mehrnoosh Raoufi, Jun Yang 0002, Xulong Tang, Youtao Zhang
DAC3
2023 CEGMA: Coordinated Elastic Graph Matching Acceleration for Graph Matching Networks
abstract
The recently proposed Graph Matching Network models (GMNs) effectively improve the inference accuracy of graph similarity analysis tasks. GMNs often take graph pairs as input, embed nodes features, and match nodes between graphs for similarity analysis. While GMNs deliver high inference accuracy, the all-to-all node matching stage in GMNs introduces quadratic computing complexity with excessive memory accesses, resulting in significant computing and memory overhead that cannot be handled by existing approaches. In this paper, we propose the Coordinated Elastic Graph Matching Accelerator (CEGMA), a software and hardware co-design accelerator to address the challenges of GMNs. Specifically, by exploiting duplicate subgraphs in the input graphs, we develop an elastic matching filter to significantly reduce the quadratic computing overhead. By exploring the substantial data reuses oriented from accessing node features, we propose a cross-graph coordinator that fuses cross-graph similarity computing with intra-graph computing to enhance data locality. Experimental results show that, on average, CEGMA achieves 353× and 6.5× speedups in GMN computing compared to state-of-the-art GPU implementation and GNN accelerators, respectively.
Yue Dai 0005, Youtao Zhang, Xulong Tang
HPCA3
2023 Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding
abstract
Multi-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average.
Bingyao Li 0001, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang 0002, Xulong Tang
HPCA6
2023 AB-ORAM: Constructing Adjustable Buckets for Space Reduction in Ring ORAM
abstract
Ring ORAM (Oblivious RAM) is a secure primitive that mitigates the large performance degradation of ORAM through reduced online memory bandwidth demand, i.e., the number of memory accesses at servicing a real memory request. Ring ORAM requires 4× or more of the protected data space to enable the optimization and thus presents high capacity pressure on modern memory systems. While recent studies strive to reduce its space consumption through bucket compaction, the large space consumption remains a major design challenge for Ring ORAM.In this paper, we propose AB-ORAM to reduce the space capacity demand in Ring ORAM. AB-ORAM identifies two inefficient use of memory space in Ring ORAM: (i) accessed blocks hold useless data until the next reshuffle operation; and (ii) large buckets provide a diminishing performance benefit for tree levels close to the leaves. AB-ORAM then proposes two schemes to exploit the optimization opportunities, respectively. Specifically, it reclaims accessed blocks early by allocating them to buckets that need a reshuffle; and shrinks the bucket size for tree level close to the leaves for a better space/performance trade-off. We evaluate the proposed AB-ORAM design and compare it to the state-of-the-art. Our results show that AB-ORAM achieves an average of 36% space reduction over the state-of-the-art while introducing very low performance overhead.
Mehrnoosh Raoufi, Jun Yang 0002, Xulong Tang, Youtao Zhang
HPCA3
2023 FlexGM: An Adaptive Runtime System to Accelerate Graph Matching Networks on GPUs
abstract
GMNs (Graph Matching Networks) exploit recently developed GNNs (Graph Neural Networks) to analyze the similarity between two graphs. They are increasingly deployed in many application domains due to their improved inference accuracy. A GMN consists of two stages, i.e., node-embedding and node-matching stages. The node-matching stage matches node features from two graphs for similarity, which accounts for over 90% of the total execution time. However, it is challenging to accelerate GMNs on GPUs due to their diverse computing patterns for different graph inputs. For large graphs, the overhead comes mainly from the high computation overhead, which increases quadratically to the size of the graphs; for small graphs, the overhead comes from the low parallelism and resource utilization.In this paper, we propose FlexGM, a flexible runtime, to adaptively accelerate GMNs on GPUs. For large graphs, we exploit the massive computation redundancy in GMNs and develop a low-overhead deduplication module to mitigate the high computation overhead. For small graphs, we develop a unified matching module to optimize GPU hardware resource usage. An adaptive module manager is then developed to judiciously select beneficial optimization strategies. Experimental results show that the FlexGM system achieves 2.5× (up to 7.6 ×) average speedup over existing methods.
Yue Dai 0005, Xulong Tang, Youtao Zhang
ICCD2
2023 SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
Sheng Li 0019, Geng Yuan, Yue Dai 0005, Youtao Zhang, Yanzhi Wang 0001, Xulong Tang
ICLR6
2023 IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE Invalidations
abstract
Multi-GPU systems have emerged as a desirable platform to deliver high computing capabilities and large memory capacity to accommodate large dataset sizes. However, naively employing multi-GPU incurs non-scalable performance. One major reason is that execution efficiency suffers expensive address translations in multi-GPU systems. The data-sharing nature of GPU applications requires page migration between GPUs to mitigate non-uniform memory access overheads. Unfortunately, frequent page migration incurs substantial page table invalidation overheads to ensure translation coherence. A comprehensive investigation of multi-GPU address translation efficiency identifies two significant bottlenecks caused by page table invalidation requests: (i) increased latency for demand TLB miss requests and (ii) increased waiting latency for performing page migrations. Based on observations, we propose IDYLL, which reduces the number of page table invalidations by maintaining an “in-PTE" directory and reduces invalidation latency by batching multiple invalidation requests to exploit spatial locality. We show that IDYLL improves overall performance by 69.9% on average.
Bingyao Li 0001, Yanan Guo 0002, Aamer Jaleel, Jun Yang 0002, Xulong Tang
MICRO6
2023 SupeRBNN: Randomized Binary Neural Network Using Adiabatic Superconductor Josephson Devices
abstract
Adiabatic Quantum-Flux-Parametron (AQFP) is a superconducting logic with extremely high energy efficiency. By employing the distinct polarity of current to denote logic ‘0’ and ‘1’, AQFP devices serve as excellent carriers for binary neural network (BNN) computations. Although recent research has made initial strides toward developing an AQFP-based BNN accelerator, several critical challenges remain, preventing the design from being a comprehensive solution. In this paper, we propose SupeRBNN, an AQFP-based randomized BNN acceleration framework that leverages software-hardware co-optimization to eventually make the AQFP devices a feasible solution for BNN acceleration. Specifically, we investigate the randomized behavior of the AQFP devices and analyze the impact of crossbar size on current attenuation, subsequently formulating the current amplitude into the values suitable for use in BNN computation. To tackle the accumulation problem and improve overall hardware performance, we propose a stochastic computing-based accumulation module and a clocking scheme adjustment-based circuit optimization method. To effectively train the BNN models that are compatible with the distinctive characteristics of AQFP devices, we further propose a novel randomized BNN training solution that utilizes algorithm-hardware co-optimization, enabling simultaneous optimization of hardware configurations. In addition, we propose implementing batch normalization matching and the weight rectified clamp method to further improve the overall performance. We validate our SupeRBNN framework across various datasets and network architectures, comparing it with implementations based on different technologies, including CMOS, ReRAM, and superconducting RSFQ/ERSFQ. Experimental results demonstrate that our design achieves an energy efficiency of approximately 7.8 × 104 times higher than that of the ReRAM-based BNN framework while maintaining a similar level of model accuracy. Furthermore, when compared with superconductor-based counterparts, our framework demonstrates at least two orders of magnitude higher energy efficiency.
Zhengang Li 0001, Geng Yuan, Tomoharu Yamauchi, Masoud Zabihi, Yanyue Xie, Peiyan Dong, Xulong Tang, Nobuyuki Yoshikawa, Devesh Tiwari, Yanzhi Wang 0001, Olivia Chen
MICRO7
2022 You Already Have It: A Generator-Free Low-Precision DNN Training Framework Using Stochastic Rounding
Geng Yuan, Sung-En Chang, Qing Jin, Alec Lu, Yanyu Li, Yushu Wu, Zhenglun Kong, Yanyue Xie, Peiyan Dong, Minghai Qin, Xulong Tang, Zhenman Fang, Yanzhi Wang 0001
ECCV (12)12
2022 Q-GPU: A Recipe of Optimizations for Quantum Circuit Simulation Using GPUs
abstract
In recent years, quantum computing has undergone significant developments and has established its supremacy in many application domains. While quantum hardware is accessible to the public through the cloud environment, a robust and efficient quantum circuit simulator is necessary to investigate the constraints and foster quantum computing development, such as quantum algorithm development and quantum device architecture exploration. In this paper, we observe that most of the publicly available quantum circuit simulators (e.g., QISKit from IBM, QDK from Microsoft, and Qsim-Cirq from Google) suffer from slow simulation and poor scalability when the number of qubits increases. To this end, we systematically investigate the deficiencies in quantum circuit simulation (QCS) and propose Q-GPU, a framework that leverages GPUs with comprehensive optimizations to allow efficient and scalable QCS. Specifically, Q-GPU features i) proactive state amplitude transfer, ii) zero state amplitude pruning, iii) delayed qubit involvement, and iv) lossless nonzero state amplitude compression. Experimental results across nine representative quantum circuits indicate that Q-GPU significantly reduces the execution time of the state-of-the-art GPU-based QCS by 71.89% (3.55× speedup). Q-GPU also outperforms the state-of-the-art OpenMP CPU implementation, the Google Qsim-Cirq simulator, and the Microsoft QDK simulator by 1.49×, 2.02×, and 10.82×, respectively.
Yilun Zhao 0002, Yanan Guo 0002, Amanda Dumi, Devin M. Mulvey, Shiv Upadhyay, Youtao Zhang, Kenneth D. Jordan, Jun Yang 0002, Xulong Tang
HPCA10
2022 Fine-Granular Computation and Data Layout Reorganization for Improving Locality
abstract
While data locality and cache performance have been investigated in great depth by prior research (in the context of both high-end systems and embedded/mobile systems), one of the important characteristics of prior approaches is that they transform loop and/or data space (e.g., array layout) as a whole. Unfortunately, such coarse-grain approaches bring three critical issues. First, they implicitly assume that all parts of a given array would equally benefit from the identified data layout transformation. Second, they also assume that a given loop transformation would have the same locality impact on an entire data array. Third and more importantly, such coarse-grain approaches are local by their nature and difficult to achieve globally optimal executions. Motivated by these drawbacks of existing code and data space reorganization/optimization techniques, this paper proposes to determine multiple loop transformation matrices for each loop nest in the program and multiple data layout transformations for each array accessed by the program, in an attempt to exploit data locality at a finer granularity. It leverages bipartite graph matching and extends the proposed fine-granular integrated loop-layout strategy to a multicore setting as well. Our experimental results show that the proposed approach significantly improves the data locality and outperforms existing schemes - 9.1% average performance improvement in single-threaded executions and 11.5% average improvement in multi-threaded executions over the state-of-the-art.
Mahmut T. Kandemir, Xulong Tang, Jagadish Kotra, Mustafa Karaköy
ICCAD2
2022 Enhancing GPU Performance via Neighboring Directory Table Based Inter-TLB Sharing
abstract
Modern discrete GPUs support Unified Virtual Memory (UVM), simplifying GPU programming. However, UVM entails address translation on each memory access, which introduces expensive performance overhead during address translation. In this work, we select various workloads and conduct experiments on GPU performance. Our investigation shows that many workloads have low L1 TLB hit ratios of less than 40% on average. Even for a particular workload, the hit ratio is as low as 15%, which leads to significant performance degradation. Through further analysis, we find that a lot of common entries exist between neighboring private L1 TLBs, showing clear inter-TLB sharing behavior. To leverage the sharing, we propose a Neighboring Directory table based hardware scheme, named NeiDty. In NeiDty, L1 TLBs can probe physical addresses from neighboring L1 TLBs through a lightweight interconnect network. And NeiDty uses neighboring directory tables to keep track of the shared entries among neighboring L1-TLBs. In addition, we find it better to update address translation after two consecutive neighboring TLB hits than one hit. We run eight typical workloads with Gem5-GPU, and the results show that NeiDty increases the average hit ratio of L1 TLB TLB by 14% and improves the average performance by 10%.
Yajuan Du, Xulong Tang
ICCD5
2022 Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse Training
abstract
Recently, sparse training has emerged as a promising paradigm for efficient deep learning on edge devices. The current research mainly devotes the efforts to reducing training costs by further increasing model sparsity. However, increasing sparsity is not always ideal since it will inevitably introduce severe accuracy degradation at an extremely high sparsity level. This paper intends to explore other possible directions to effectively and efficiently reduce sparse training costs while preserving accuracy. To this end, we investigate two techniques, namely, layer freezing and data sieving. First, the layer freezing approach has shown its success in dense model training and fine-tuning, yet it has never been adopted in the sparse training domain. Nevertheless, the unique characteristics of sparse training may hinder the incorporation of layer freezing techniques. Therefore, we analyze the feasibility and potentiality of using the layer freezing technique in sparse training and find it has the potential to save considerable training costs. Second, we propose a data sieving method for dataset-efficient training, which further reduces training costs by ensuring only a partial dataset is used throughout the entire training process. We show that both techniques can be well incorporated into the sparse training algorithm to form a generic framework, which we dub SpFDE. Our extensive experiments demonstrate that SpFDE can significantly reduce training costs while preserving accuracy from three dimensions: weight sparsity, layer freezing, and dataset sieving. Our code and models will be released.
Geng Yuan, Yanyu Li, Sheng Li 0019, Zhenglun Kong, Sergey Tulyakov, Xulong Tang, Yanzhi Wang 0001, Jian Ren 0005
NeurIPS6
2022 An efficient segmented quantization for graph neural networks
Yue Dai 0005, Xulong Tang, Youtao Zhang
CCF Trans. High Perform. Comput.2
2022 Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization Framework
abstract
Efficient deployment of Deep Neural Networks (DNNs) on edge devices (i.e., FPGAs and mobile platforms) is very challenging, especially under a recent witness of the increasing DNN model size and complexity. Model compression strategies, including weight quantization and pruning, are widely recognized as effective approaches to significantly reduce computation and memory intensities, and have been implemented in many DNNs on edge devices. However, most state-of-the-art works focus on ad hoc optimizations, and there lacks a thorough study to comprehensively reveal the potentials and constraints of different edge devices when considering different compression strategies. In this article, we qualitatively and quantitatively compare the energy efficiency of FPGA-based and mobile-based DNN executions using mobile GPU and provide a detailed analysis. Based on the observations obtained from the analysis, we propose a unified optimization framework using block-based pruning to reduce the weight storage and accelerate the inference speed on mobile devices and FPGAs, achieving high hardware performance and energy-efficiency gain while maintaining accuracy.
Geng Yuan, Peiyan Dong, Mengshu Sun, Wei Niu 0002, Zhengang Li 0001, Yuxuan Cai 0001, Yanyu Li, Jun Liu 0075, Weiwen Jiang, Xue Lin 0001, Bin Ren 0002, Xulong Tang, Yanzhi Wang 0001
ACM Trans. Embed. Comput. Syst.12
2022 Automatic Mapping of the Best-Suited DNN Pruning Schemes for Real-Time Mobile Acceleration
abstract
Weight pruning is an effective model compression technique to tackle the challenges of achieving real-time deep neural network (DNN) inference on mobile devices. However, prior pruning schemes have limited application scenarios due to accuracy degradation, difficulty in leveraging hardware acceleration, and/or restriction on certain types of DNN layers. In this article, we propose a general, fine-grained structured pruning scheme and corresponding compiler optimizations that are applicable to any type of DNN layer while achieving high accuracy and hardware inference performance. With the flexibility of applying different pruning schemes to different layers enabled by our compiler optimizations, we further probe into the new problem of determining the best-suited pruning scheme considering the different acceleration and accuracy performance of various pruning schemes. Two pruning scheme mapping methods—one -search based and the other is rule based—are proposed to automatically derive the best-suited pruning regularity and block size for each layer of any given DNN. Experimental results demonstrate that our pruning scheme mapping methods, together with the general fine-grained structured pruning scheme, outperform the state-of-the-art DNN optimization framework with up to 2.48 \( \times \) and 1.73 \( \times \) DNN inference acceleration on CIFAR-10 and ImageNet datasets without accuracy loss.
Yifan Gong 0004, Geng Yuan, Zheng Zhan 0001, Wei Niu 0002, Zhengang Li 0001, Pu Zhao 0001, Yuxuan Cai 0001, Sijia Liu 0001, Bin Ren 0002, Xue Lin 0001, Xulong Tang, Yanzhi Wang 0001
ACM Trans. Design Autom. Electr. Syst.11
2021 YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
abstract
The rapid development and wide utilization of object detection techniques have aroused attention on both accuracy and speed of object detectors. However, the current state-of-the-art object detection works are either accuracy-oriented using a large model but leading to high latency or speed-oriented using a lightweight model but sacrificing accuracy. In this work, we propose YOLObile framework, a real-time object detection on mobile devices via compression-compilation co-design. A novel block-punched pruning scheme is proposed for any kernel size. To improve computational efficiency on mobile devices, a GPU-CPU collaborative scheme is adopted along with advanced compiler-assisted optimizations. Experimental results indicate that our pruning scheme achieves 14x compression rate of YOLOv4 with 49.0 mAP. Under our YOLObile framework, we achieve 17 FPS inference speed using GPU on Samsung Galaxy S20. By incorporating our proposed GPU-CPU collaborative scheme, the inference speed is increased to 19.1 FPS, and outperforms the original YOLOv4 by 5x speedup. Source code is at: https://github.com/nightsnack/YOLObile.
Yuxuan Cai 0001, Hongjia Li 0003, Geng Yuan, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001
AAAI6
2021 A Compression-Compilation Co-Design Framework Towards Real-Time Object Detection on Mobile Devices
abstract
The rapid development and wide utilization of object detection techniques have aroused requirements for both accuracy and speed of object detectors. In this work, we propose a compression-compilation co-design framework to achieve real-time YOLOv4 inference on mobile devices. We propose a novel fine-grained structured pruning, which maintain high accuracy while achieving high hardware parallelism. Our pruned YOLOv4 achieves 48.9 mAP and 17 FPS inference speed on an off-the-shelf Samsung Galaxy S20 smartphone, which is 5.5x faster than the original state-of-the-art detector YOLOv4.
Yuxuan Cai 0001, Geng Yuan, Hongjia Li 0003, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001
AAAI6
2021 Towards a Secure Integrated Heterogeneous Platform via Cooperative CPU/GPU Encryption
abstract
Nowadays, emerging integrated heterogeneous platforms play major roles to host autonomous systems. However, the security issue that comes with such heterogeneous architectures has not been thoroughly explored and imposes great threats and vulnerabilities to these systems. We set out to explore the security issues for the heterogeneous architectures and the corresponding mitigation mechanisms. We investigate the side-channel timing attack in a modern integrated CPU/GPU platform and propose a CPU/GPU co-encryption mechanism CoENC to mitigate the timing attack to provide a secure platform for autonomous systems. Evaluations demonstrate CoENC can effectively enhance the security 29~44 times compared to the baseline with an extra 14%~31% latency overhead.
Rujia Wang, Zihang Jiang, Xulong Tang, Shouyi Yin, Yang Hu 0001
ATS4
2021 ScaleDNN: Data Movement Aware DNN Training on Multi-GPU
abstract
Training Deep Neural Networks (DNNs) models is a time-consuming process that requires immense amount of data and computation. To this end, GPUs are widely adopted to accelerate the training process. However, the delivered training performance rarely scales with the increase in the number of GPUs. The major reason behind this is the large amount of data movement that prevents the system from providing the GPUs with the required data in a timely fashion. In this paper, we propose ScaleDNN, a framework that systematically and comprehensively investigates and optimizes data-parallel training on two types of multi-GPU systems (PCIe-based and NVLink-based). Specifically, ScaleDNN performs: i) CPU-centric input batch splitting, ii) mini-batch data pre-loading, and iii) model parameter compression to effectively a) reduce the data movement between the CPU and multiple GPUs, and b) hide the data movement overheads by overlapping the data transfer with the GPU computation. Our experimental results show that ScaleDNN achieves up to 39.38%, with an average of 17.96% execution time saving over modern data parallelism on PCIe-based multi-GPU system. The corresponding execution time reduction on NVLink-based multi-GPU system is up to 19.20% with an average of 10.26%.
Weizheng Xu, Ashutosh Pattnaik, Geng Yuan, Yanzhi Wang 0001, Youtao Zhang, Xulong Tang
ICCAD6
2021 Automated Runtime-Aware Scheduling for Multi-Tenant DNN Inference on GPU
abstract
With the fast development of deep neural networks (DNNs), many real-world applications are adopting multiple models to conduct compound tasks, such as co-running classification, detection, and segmentation models on autonomous vehicles. Such multi-tenant DNN inference cases greatly exacerbate the computational complexity and call for comprehensive collaboration for graph-level operator scheduling, runtime-level resource awareness, as well as hardware scheduler support. However, the current scheduling support for such multi-tenant inference is still relatively backward. In this work, we propose a resource-aware scheduling framework for efficient multi-tenant DNN inference on GPU, which automatically coordinates DNN computing in different execution levels. Leveraging the unified scheduling intermediate representation and the automated ML-based searching algorithm, optimal schedules could be generated to wisely adjust model concurrency and interleave DNN model operators, maintaining a continuously balanced resource utilization across the entire inference process, and eventually improving the runtime efficiency. Experiments show that we could consistently achieve$1.3\times\sim 1.7\times$speed-up, comparing to regular DNN runtime libraries (e.g., CuDNN, TVM) and particular concurrent scheduling methods (e.g., NVIDIA Multi-Stream).
Fuxun Yu, Shawn Bray, Di Wang 0003, Longfei Shangguan, Xulong Tang, Xiang Chen 0010
ICCAD5
2021 Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB Design
abstract
In recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively.
Bingyao Li 0001, Jieming Yin, Youtao Zhang, Xulong Tang
MICRO4
2021 Characterizing AI Model Inference Applications Running in the SGX Environment
abstract
Intel Software Guard Extensions (SGX) is a set of extensions built into Intel CPUs for the trusted computation. It creates a hardware-assisted secure container, within which programs are protected from data leakage and data manipulations by privileged software and hypervisors. With the trend that more and more machine learning based programs are moving to cloud computing, SGX can be used in cloud-based Machine Learning applications to protect user data from malicious privileged programs.However, applications running in SGX suffer from several overheads, including frequent context switching, memory page encryption/decryption, and memory page swapping, which significantly degrade the execution efficiency. In this paper, we aim to i) comprehensively explore the execution of general AI applications running on SGX, ii) systematically characterize the data reuses at both page granularity and cacheline granularity, and iii) provide optimization insights for efficient deployment of machine learning based applications on SGX. To the best of our knowledge, our work is the first to study machine learning applications on SGX and explore the potential of data reuses to reduce the runtime overheads in SGX.
Shixiong Jing, Qinkun Bao, Pei Wang 0007, Xulong Tang, Dinghao Wu
NAS4
2021 Fluid: a framework for approximate concurrency via controlled dependency relaxation
abstract
In this work, we introduce the Fluid framework, a set of language, compiler and runtime extensions that allow for the expression of regions within which dataflow dependencies can be approximated in a disciplined manner. Our framework allows the eager execution of dependent tasks before their inputs have finalized in order to capitalize on situations where an eagerly-consumed input has a high probability of sufficiently resembling the value or structure of the final value that would have been produced in a conservative/precise execution schedule. We introduce controlled access to the early consumption of intermediate values and provide hooks for user-specified quality assurance mechanisms that can automatically enforce re-execution of eagerly-executed tasks if their output values do not meet heuristic expectations. Our experimental analysis indicates that the fluidized versions of the applications bring 22.2% average execution time improvements, over their original counterparts, under the default values of our fluidization parameters. The Fluid approach is largely orthogonal to approaches that aim to reduce the task effort itself and we show that utilizing the Fluid framework can yield benefits for both originally precise and originally approximate versions of computation.
Huaipan Jiang, Haibo Zhang 0005, Xulong Tang, Vineetha Govindaraj, Jack Sampson, Mahmut T. Kandemir, Danfeng Zhang
PLDI3
2021 Distance-in-time versus distance-in-space
abstract
Cache behavior is one of the major factors that influence the performance of applications. Most of the existing compiler techniques that target cache memories focus exclusively on reducing data reuse distances in time (DIT). However, current manycore systems employ distributed on-chip caches that are connected using an on-chip network. As a result, a reused data element/block needs to travel over this on-chip network, and the distance to be traveled -- reuse distance in space (DIS) -- can be as influential in dictating application performance as reuse DIT. This paper represents the first attempt at defining a compiler framework that accommodates both DIT and DIS. Specifically, it first classifies data reuses into four groups: G1: (low DIT, low DIS), G2: (high DIT, low DIS), G3: (low DIT, high DIS), and G4: (high DIT, high DIS). Then, observing that reuses in G1 represent the ideal case and there is nothing much to be done in computations in G4, it proposes a "reuse transfer" strategy that transfers select reuses between G2 and G3, eventually, transforming each reuse to either G1 or G4. Finally, it evaluates the proposed strategy using a set of 10 multithreaded applications. The collected results reveal that the proposed strategy reduces parallel execution times of the tested applications between 19.3% and 33.3%.
Mahmut T. Kandemir, Xulong Tang, Hui Zhao 0013, Jihyun Ryoo, Mustafa Karaköy
PLDI2
2021 Compiler support for near data computing
abstract
Recent works from both hardware and software domains offer various optimizations that try to take advantage of near data computing (NDC) opportunities. While the results from these works indicate performance improvements of various magnitudes, the existing literature lacks a detailed quantification of the potential of NDC and analysis of compiler optimizations on tapping into that potential. This paper first presents an analysis of the NDC potential when executing multithreaded applications on manycore platforms. It then presents two compiler schemes designed to take advantage of NDC. The first of these schemes try to increase the amount of computation that can be performed in a hardware component, whereas the second compiler strategy strikes a balance between optimizing NDC and exploiting data reuse, by being more selective on when to perform NDC (even if the opportunity presents itself) and how. The collected experimental results on a 5×5 manycore system reveal that our first and second compiler schemes improve the overall performance of our multithreaded applications by, respectively, 22.5% and 25.2%, on average. Furthermore, these two compiler schemes are only 6.8% and 4.1% worse than an oracle scheme that makes the best near data computing decisions for each and every computation.
Mahmut T. Kandemir, Jihyun Ryoo, Xulong Tang, Mustafa Karaköy
PPoPP3
2021 Work in Progress: Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization Framework
abstract
Efficient deployment of Deep Neural Networks (DNNs) on edge devices (i.e., FPGAs and mobile platforms) is very challenging, especially under a recent witness of the increasing DNN model size and complexity. Although various optimization approaches have been proven to be effective in many DNNs on edge devices, most state-of-the-art work focuses on ad-hoc optimizations, and there lacks a thorough study to comprehensively reveal the potentials and constraints of different edge devices when considering different optimizations. In this paper, we qualitatively and quantitatively compare the energyefficiency of FPGA-based and mobile-based DNN executions, and provide detailed analysis.
Geng Yuan, Peiyan Dong, Mengshu Sun, Wei Niu 0002, Zhengang Li 0001, Yuxuan Cai 0001, Jun Liu 0075, Weiwen Jiang, Xue Lin 0001, Bin Ren 0002, Xulong Tang, Yanzhi Wang 0001
RTAS11
2021 Algorithm-hardware Co-design of Attention Mechanism on FPGA Devices
abstract
Multi-head self-attention (attention mechanism) has been employed in a variety of fields such as machine translation, language modeling, and image processing due to its superiority in feature extraction and sequential data analysis. This is benefited from a large number of parameters and sophisticated model architecture behind the attention mechanism. To efficiently deploy attention mechanism on resource-constrained devices, existing works propose to reduce the model size by building a customized smaller model or compressing a big standard model. A customized smaller model is usually optimized for the specific task and needs effort in model parameters exploration. Model compression reduces model size without hurting the model architecture robustness, which can be efficiently applied to different tasks. The compressed weights in the model are usually regularly shaped (e.g. rectangle) but the dimension sizes vary (e.g. differs in rectangle height and width). Such compressed attention mechanism can be efficiently deployed on CPU/GPU platforms as their memory and computing resources can be flexibly assigned with demand. However, for Field Programmable Gate Arrays (FPGAs), the data buffer allocation and computing kernel are fixed at run time to achieve maximum energy efficiency. After compression, weights are much smaller and different in size, which leads to inefficient utilization of FPGA on-chip buffer. Moreover, the different weight heights and widths may lead to inefficient FPGA computing kernel execution. Due to the large number of weights in the attention mechanism, building a unique buffer and computing kernel for each compressed weight on FPGA is not feasible. In this work, we jointly consider the compression impact on buffer allocation and the required computing kernel during the attention mechanism compressing. A novel structural pruning method with memory footprint awareness is proposed and the associated accelerator on FPGA is designed. The experimental results show that our work can compress Transformer (an attention mechanism based model) by 95x. The developed accelerator can fully utilize the FPGA resource, processing the sparse attention mechanism with the run-time throughput performance of 1.87 Tops in ZCU102 FPGA.
Xinyi Zhang 0001, Yawen Wu, Peipei Zhou 0001, Xulong Tang, Jingtong Hu
ACM Trans. Embed. Comput. Syst.4
2020 Enhancing Address Translations in Throughput Processors via Compression
abstract
Efficient memory sharing among multiple compute engines plays an important role in shaping the overall application performance on CPU-GPU heterogeneous platforms. Unified Virtual Memory (UVM) is a promising feature that allows globally-visible data structures and pointers such that the GPU can access the physical memory space on the CPU side, and take advantage of the host OS paging mechanism without explicit programmer effort. However, a key requirement for the guaranteed performance is effective hardware support of address translation. Particularly, we observe that GPU execution suffers from high TLB miss rates in a UVM environment, especially for irregular and/or memory-intensive applications. In this paper, we propose simple yet effective compression mechanisms for address translations to improve GPU TLB hit rates. Specifically, we explore and leverage the TLB compressibility during the execution of GPU applications to design efficient address translation compression with minimal runtime overhead. Experimental results across 22 applications indicate that our proposed approach significantly improves GPU TLB hit rates, which translate to 12% average performance improvement. Particularly, for 16 irregular and/or memory-intensive applications, the performance improvements achieved reach up to 69.2%, with an average of 16.3%.
Xulong Tang, Weizheng Xu, Mahmut T. Kandemir, Rami G. Melhem, Jun Yang 0002
PACT1
2020 Enabling Latency-Aware Data Initialization for Integrated CPU/GPU Heterogeneous Platform
abstract
Nowadays, driven by the needs of autonomous driving and edge intelligence, integrated CPU/GPU heterogeneous platform has gained significant attention from both academia and industry. As the representative series, NVIDIA Jetson family perform well in terms of computation capability, power consumption, and mobile size. Even so, the integrated heterogeneous platform only contains one limited physical memory, which is shared by the CPU and GPU cores and can be the performance bottleneck of the mobile/edge applications. On the other hand, with the unified memory (UM) model introduced in GPU programming, not only the memory allocation is significantly reduced, which mitigates the memory bottleneck of the integrated platforms but also the memory management and programming are simplified. However, as a programming legacy, the UM model still follows the conventional copy-then-execute model, initializing data on the CPU side after allocating memory. This legacy programming mode not only causes significant initialization latency but also slows the execution of the following kernel. In this article, we propose a framework to enable the latency-aware data initialization on the integrated heterogeneous platform. The framework not only includes three data initialization modes, the CPU initialization, GPU initialization, and hybrid initialization, but also utilizes an affinity estimation model to wisely decide the best initialization mode for an application such that the initialization latency performance of the application can be optimized. We evaluate our design on NVIDIA TX2 and AGX platforms. The results demonstrate that the framework can accurately select a data initialization mode for a given application to significantly reduce the initialization latency. We envision this latency-aware data initialization framework being adopted in a full-version of autonomous solution (e.g., Autoware) in the future.
Zihang Jiang, Zhen Wang 0019, Xulong Tang, Cong Liu 0005, Shouyi Yin, Yang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Architecture-Centric Bottleneck Analysis for Deep Neural Network Applications
abstract
The ever-growing complexity and popularity of machine learning and deep learning applications have motivated an urgent need of effective and efficient support for these applications on contemporary computing systems. In this paper, we thoroughly analyze the various DNN algorithms on three widely used architectures (CPU, GPU, and Xeon Phi). The DNN algorithms we choose for evaluation include i) Unet - for biomedical image segmentation, based on Convolutional Neural Network (CNN), ii) NMT - for neural machine translation based on Recurrent Neural Network (RNN), iii) ResNet-50, and iv) DenseNet - both for image processing based on CNNs. The ultimate goal of this paper is to answer four fundamental questions: i) whether the different DNN networks exhibit similar behavior on a given execution platform? ii) whether, across different platforms, a given DNN network exhibits different behaviors? iii) for the same execution platform and the same DNN network, whether different execution phases have different behaviors? and iv) are the current major general-purpose platforms tuned sufficiently well for different DNN algorithms? Motivated by these questions, we conduct an in-depth investigation of running DNN applications on modern systems. Specifically, we first identify the most time-consuming functions (hotspot functions) across different networks and platforms. Next, we characterize performance bottlenecks and discuss them in detail. Finally, we port selected hotspot functions to a cycle-accurate simulator, and use the results to direct architectural optimizations to better support DNN applications.
Jihyun Ryoo, Mengran Fan, Xulong Tang, Huaipan Jiang, Meena Arunachalam, Sharada Naveen, Mahmut T. Kandemir
HiPC3
2019 Opportunistic computing in GPU architectures
abstract
Data transfer overhead between computing cores and memory hierarchy has been a persistent issue for von Neumann architectures and the problem has only become more challenging with the emergence of manycore systems. A conceptually powerful approach to mitigate this overhead is to bring the computation closer to data, known as Near Data Computing (NDC). Recently, NDC has been investigated in different flavors for CPU-based multicores, while the GPU domain has received little attention. In this paper, we present a novel NDC solution for GPU architectures with the objective of minimizing on-chip data transfer between the computing cores and Last-Level Cache (LLC). To achieve this, we first identify frequently occurring Load-Compute-Store instruction chains in GPU applications. These chains, when offloaded to a compute unit closer to where the data resides, can significantly reduce data movement. We develop two offloading techniques, called LLC-Compute and Omni-Compute. The first technique, LLC-Compute, augments the LLCs with computational hardware for handling the computation offloaded to them. The second technique (Omni-Compute) employs simple bookkeeping hardware to enable GPU cores to compute instructions offloaded by other GPU cores. Our experimental evaluations on nine GPGPU workloads indicate that the LLC-Compute technique provides, on an average, 19% performance improvement (IPC), 11% performance/watt improvement, and 29% reduction in on-chip data movement compared to the baseline GPU design. The Omni-Compute design boosts these benefits to 31%, 16% and 44%, respectively.
Ashutosh Pattnaik, Xulong Tang, Onur Kayiran, Adwait Jog, Asit K. Mishra, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
ISCA2
2019 Co-optimizing memory-level parallelism and cache-level parallelism
abstract
Minimizing cache misses has been the traditional goal in optimizing cache performance using compiler based techniques. However, continuously increasing dataset sizes combined with large numbers of cache banks and memory banks connected using on-chip networks in emerging manycores/accelerators makes cache hit–miss latency optimization as important as cache miss rate minimization. In this paper, we propose compiler support that optimizes both the latencies of last-level cache (LLC) hits and the latencies of LLC misses. Our approach tries to achieve this goal by improving the parallelism exhibited by LLC hits and LLC misses. More specifically, it tries to maximize both cache-level parallelism (CLP) and memory-level parallelism (MLP). This paper presents different incarnations of our approach, and evaluates them using a set of 12 multithreaded applications. Our results indicate that (i) optimizing MLP first and CLP later brings, on average, 11.31% performance improvement over an approach that already minimizes the number of LLC misses, and (ii) optimizing CLP first and MLP later brings 9.43% performance improvement. In comparison, balancing MLP and CLP brings 17.32% performance improvement on average.
Xulong Tang, Mahmut T. Kandemir, Mustafa Karaköy, Meenakshi Arunachalam
PLDI1
2018 Quantifying and Optimizing Data Access Parallelism on Manycores
abstract
The following topics are dealt with: storage management; cache storage; pattern clustering; cloud computing; optimisation; flash memories; resource allocation; scheduling; parallel processing; data mining.
Jihyun Ryoo, Orhan Kislal, Xulong Tang, Mahmut T. Kandemir
MASCOTS3
2018 Enhancing computation-to-core assignment with physical location information
abstract
Going beyond a certain number of cores in modern architectures requires an on-chip network more scalable than conventional buses. However, employing an on-chip network in a manycore system (to improve scalability) makes the latencies of the data accesses issued by a core non-uniform. This non-uniformity can play a significant role in shaping the overall application performance. This work presents a novel compiler strategy which involves exposing architecture information to the compiler to enable an optimized computation-to-core mapping. Specifically, we propose a compiler-guided scheme that takes into account the relative positions of (and distances between) cores, last-level caches (LLCs) and memory controllers (MCs) in a manycore system, and generates a mapping of computations to cores with the goal of minimizing the on-chip network traffic. The experimental data collected using a set of 21 multi-threaded applications reveal that, on an average, our approach reduces the on-chip network latency in a 6×6 manycore system by 38.4% in the case of private LLCs, and 43.8% in the case of shared LLCs. These improvements translate to the corresponding execution time improvements of 10.9% and 12.7% for the private LLC and shared LLC based systems, respectively.
Orhan Kislal, Jagadish Kotra, Xulong Tang, Mahmut T. Kandemir, Myoungsoo Jung
PLDI3
2017 POSTER: Location-Aware Computation Mapping for Manycore Processors
abstract
Employing an on-chip network in a manycore system (to improve scalability) makes the latencies of data accesses issued by a core non-uniform, which significant impact application performance. This paper presents a compiler strategy which involves exposing architecture information to the compiler to enable optimized computation-to-core mapping. Our scheme takes into account the relative positions of (and distances between) cores, last-level caches (LLCs) and memory controllers (MCs) in a manycore system, and generates a mapping of computations to cores with the goal of minimizing the on-chip network traffic. Our experiments of 12 multi-threaded applications reveal that, on average, our approach reduces the on-chip network latency in a 6x6 manycore system by 49.5% in the case of private LLCs and 52.7% in the case of shared LLCs. These improvements translate to the corresponding execution time improvements of 14.8% and 15.2% for the private LLC and shared LLC based systems.
Orhan Kislal, Jagadish Kotra, Xulong Tang, Mahmut T. Kandemir, Myoungsoo Jung
PACT3
2017 Controlled Kernel Launch for Dynamic Parallelism in GPUs
abstract
Dynamic parallelism (DP) is a promising feature for GPUs, which allows on-demand spawning of kernels on the GPU without any CPU intervention. However, this feature has two major drawbacks. First, the launching of GPU kernels can incur significant performance penalties. Second, dynamically-generated kernels are not always able to efficiently utilize the GPU cores due to hardware-limits. To address these two concerns cohesively, we propose SPAWN, a runtime framework that controls the dynamically-generated kernels, thereby directly reducing the associated launch overheads and queuing latency. Moreover, it allows a better mix of dynamically-generated and original (parent) kernels for the scheduler to effectively hide the remaining overheads and improve the utilization of the GPU resources. Our results show that, across 13 benchmarks, SPAWN achieves 69% and 57% speedup over the flat (non-DP) implementation and baseline DP, respectively.
Xulong Tang, Ashutosh Pattnaik, Huaipan Jiang, Onur Kayiran, Adwait Jog, Sreepathi Pai, Mohamed Assem Ibrahim, Mahmut T. Kandemir, Chita R. Das
HPCA1
2017 DEMM: A Dynamic Energy-Saving Mechanism for Multicore Memories
abstract
Since main memory system contributes to a large and increasing fraction of server/datacenter energy consumption, there have been several efforts to reduce its power and energy consumption. DVFS schemes have been used to reduce the memory power, but they come with a performance penalty. In this work, we propose DEMM, an OS-based, high performance DVFS mechanism that reduces memory power by dynamically scaling individual memory channel frequencies/voltages. Our strategy also involves clustering the running applications based on their sensitivities to memory latency, and assigning memory channels to the application clusters. We introduce a new metric called Discrete Misses per Kilo Cycle (DMPKC) to capture the performance sensitivities of the applications to memory frequency modulation. DEMM allows us to save power in the memory system with negligible impact on performance. We demonstrate around 25% savings in the memory system energy and 10% savings in the total system energy, with only a 4% loss in workload performance.
Akbar Sharifi, Wei Ding 0008, Diana R. Guttman, Hui Zhao 0013, Xulong Tang, Mahmut T. Kandemir, Chita R. Das
MASCOTS5
2017 Data movement aware computation partitioning
abstract
Data access costs dominate the execution times of most parallel applications and they are expected to be even more important in the future. To address this, recent research has focused on Near Data Processing (NDP) as a new paradigm that tries to bring computation to data, instead of bringing data to computation (which is the norm in conventional computing). This paper explores the potential of compiler support in exploiting NDP in the context of emerging manycore systems. To that end, we propose a novel compiler algorithm that partitions the computations in a given loop nest into subcomputations and schedules the resulting subcomputations on different cores with the goal of reducing the distance-to-data on the on-chip network. An important characteristic of our approach is that it exploits NDP while taking advantage of data locality. Our experiments with 12 multithreaded applications running on a state-of-the-art commercial manycore system indicate that the proposed compiler-based approach significantly reduces data movements on the on-chip network by taking advantage of NDP, and these benefits lead to an average execution time improvement of 18.4%.
Xulong Tang, Orhan Kislal, Mahmut T. Kandemir, Mustafa Karaköy
MICRO1
2016 μC-States: Fine-grained GPU Datapath Power Management
abstract
To improve the performance of Graphics Processing Units (GPUs) beyond simply increasing core count, architects are recently adopting a scale-up approach: the peak throughput and individual capabilities of the GPU cores are increasing rapidly. This big-core trend in GPUs leads to various challenges, including higher static power consumption and lower and imbalanced utilization of the datapath components of a big core. As we show in this paper, two key problems ensue: (1) the lower and imbalanced datapath utilization can waste power as an application does not always utilize all portions of the big core datapath, and (2) the use of big cores can lead to application performance degradation in some cases due to the higher memory system contention caused by the more memory requests generated by each big core.
Onur Kayiran, Adwait Jog, Ashutosh Pattnaik, Rachata Ausavarungnirun, Xulong Tang, Mahmut T. Kandemir, Gabriel H. Loh, Onur Mutlu, Chita R. Das
PACT5
2016 Scheduling Techniques for GPU Architectures with Processing-In-Memory Capabilities
abstract
Processing data in or near memory (PIM), as opposed to in conventional computational units in a processor, can greatly alleviate the performance and energy penalties of data transfers from/to main memory. Graphics Processing Unit (GPU) architectures and applications, where main memory bandwidth is a critical bottleneck, can benefit from the use of PIM. To this end, an application should be properly partitioned and scheduled to execute on either the main, powerful GPU cores that are far away from memory or the auxiliary, simple GPU cores that are close to memory (e.g., in the logic layer of 3D-stacked DRAM).
Ashutosh Pattnaik, Xulong Tang, Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Chita R. Das
PACT2
2016 Improving bank-level parallelism for irregular applications
abstract
Observing that large multithreaded applications with irregular data access patterns exhibit very low memory bank-level parallelism (BLP) during their execution, we propose a novel loop iteration scheduling strategy built upon the inspector-executor paradigm. A unique characteristic of this strategy is that it considers both bank-level parallelism (from an inter-core perspective) and bank reuse (from an intra-core perspective) in a unified framework. Its primary goal is to improve bank-level parallelism, and bank reuse is taken into account only if doing so does not hurt bank-level parallelism. Our experiments with this strategy using eight application programs on both a simulator and a real multicore system show an average BLP improvement of 46.8% and an average execution time reduction of 18.3%.
Xulong Tang, Mahmut T. Kandemir, Praveen Yedlapalli, Jagadish Kotra
MICRO1
2015 Optimizing off-chip accesses in multicores
abstract
In a network-on-chip (NoC) based manycore architecture, an off-chip data access (main memory access) needs to travel through the on-chip network, spending considerable amount of time within the chip (in addition to the memory access latency). In addition, it contends with on-chip (cache) accesses as both use the same NoC resources. In this paper, focusing on data-parallel, multithreaded applications, we propose a compiler-based off-chip data access localization strategy, which places data elements in the memory space such that an off-chip access traverses a minimum number of links (hops) to reach the memory controller that handles this access. This brings three main benefits. First, the network latency of off-chip accesses gets reduced; second, the network latency of on-chip accesses gets reduced; and finally, the memory latency of off-chip accesses improves, due to reduced queue latencies. We present an experimental evaluation of our optimization strategy using a set of 13 multithreaded application programs under both private and shared last-level caches. The results collected emphasize the importance of optimizing the off-chip data accesses.
Wei Ding 0008, Xulong Tang, Mahmut T. Kandemir, Emre Kultursay
PLDI2
2015 Memory Row Reuse Distance and its Role in Optimizing Application Performance
abstract
Continuously increasing dataset sizes of large-scale applications overwhelm on-chip cache capacities and make the performance of last-level caches (LLC) increasingly important. That is, in addition to maximizing LLC hit rates, it is becoming equally important to reduce LLC miss latencies. One of the critical factors that influence LLC miss latencies is row-buffer locality (i.e., the fraction of LLC misses that hit in the large buffer attached to a memory bank). While there has been a plethora of recent works on optimizing row-buffer performance, to our knowledge, there is no study that quantifies the full potential of row-buffer locality and impact of maximizing it on application performance.
Mahmut T. Kandemir, Hui Zhao 0013, Xulong Tang, Mustafa Karaköy
SIGMETRICS3
2012 FlexBFS: a parallelism-aware implementation of breadth-first search on GPU
abstract
In this paper, we present FlexBFS, a parallelism-aware implementation for breadth-first search on GPU. Our implementation can adjust the computation resources according to the feedback of available parallelism dynamically. We also optimized our program in three ways: (1)a simplified two-level queue management,(2)a combined kernel strategy and (3)a high-degree vertices specialization approach. Our experimental results show that it can achieve 3~20 times speedup against the fastest serial version, and can outperform the TBB based multi-threading CPU version and the previous most effective GPU version on all types of input graphs.
Gu Liu, Hong An, Wenting Han, Xuechao Wei, Xulong Tang
PPoPP8