Huayou Su

dblp:24/7455 · DBLP profile ↗
← Back
33ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-3587-0917ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AMXStencil: Boosting the Performance of Stencil Computations on AMX-Powered CPUs via Fusion
Zitong An, Kangkang Chen, Huayou Su, Jinwei Xu, Xi Yang 0020
APPT3
2026 DOA: Dataflow Optimization for Attention on Multi-core DSPs with Three-Level Memory Hierarchy
Zhiquan Lai, Shun Ouyang, Zhaoning Zhang 0001, Menghan Jia, Huayou Su, Dongsheng Li 0001
APPT8
2026 JSQKV: Joint Sparsification and Quantization for KV-Cache Compression and Decode Acceleration
Xiaoli Gong, Huayou Su, Qingxia Chen, Jin Zhang 0003
APPT4
2026 SparX: Cache-Aware Sparse Matrix Multiplication Kernels on AMX-Based CPUs
Dongsheng Li 0001, Huayou Su
Euro-Par (2)3
2026 Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUs
abstract
Graph Neural Networks (GNNs) are becoming increasingly popular in graph data processing due to their excellent performance in feature extraction on graph datasets. Compared to GPUs, CPUs are more widely accessible and serve as a practical platform for GNN inference. However, achieving efficient GNN execution on CPUs remains a challenge. We first comprehensively evaluate and quantitatively analyze the performance of GNN inference on multi-core CPUs using the state-of-the-art frameworks, identifying four key performance bottlenecks: inefficient sparse computation, poor data locality, workload imbalance, and inefficient General Matrix Multiplication (GEMM). To tackle these issues, we introduce a set of joint optimizations. Specifically, for the aggregation phase, we propose three optimizations: a register padding and tiling Graph Sparse-dense Matrix Multiplication (GSpMM) algorithm that leverages the computation capability of long vector processing units on modern multi-core CPUs, a destination node-oriented indexes reorganization to enhance data locality, and a boundary buffer-based method to balance the workloads. Additionally, for the update phase, we develop an efficient bias fusion GEMM algorithm, tailored for the irregular matrices. We evaluate the proposed optimizations extensively with three popular GNN models on three typical multi-core CPU platforms. Experimental results on Intel, AMD, and ARM platforms show that our optimizations outperform the state-of-the-art GNN framework DGL by an average factor of 2.41×, 1.58×, and 2.04× (up to 4.75×, 2.70×, and 3.55×), respectively. Compared to PyG, our implementations achieve an average speedup of 1.70×, 1.86×, and 2.44×, respectively.
Kangkang Chen, Huayou Su, Xi Yang 0020, Zitong An, Yong Dou, Dongsheng Li 0001
ACM Trans. Archit. Code Optim.2
2026 Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-Scale MoE Models
abstract
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budgets with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Experts (MoE) has drawn significant attention as it can scale models to extra-large sizes with a near-stable computation budget. However, inefficient distributed training of large-scale MoE models hinders their broader application. Specifically, a considerable dynamic load imbalance occurs among devices during training, significantly reducing throughput. Several load-balancing works have been proposed to address the challenge. System-level solutions draw more attention for their hardware affinity and non-disruption of model convergence compared to algorithm-level ones. However, they are troubled by high communication costs and poor communication-computation overlap. To address these challenges, we propose a systematic load-balancing method, Pro-Prophet, which consists of a planner and a scheduler for efficient parallel training of large-scale MoE models. To adapt to the dynamic load imbalance, we have profiled training statistics and utilized them to design Pro-Prophet. For lower communication volume, the Pro-Prophet planner determines a series of lightweight load-balancing strategies and efficiently searches for a communication-efficient one for training based on the statistics. For sufficient overlapping of communication and computation, the Pro-Prophet scheduler schedules the data-dependent operations based on the statistics and operation characteristics, further improving the training throughput. We conduct extensive experiments in various clusters and MoE models. The results indicate that Pro-Prophet achieves up to 2.66x speedup on MoE-GPT models compared to two popular MoE frameworks, namely Deepspeed-MoE and FasterMoE. Furthermore, Pro-Prophet has demonstrated a load-balancing improvement of up to 11.01x and speedups on modern MoE models up to 1.22x compared to a representative load-balancing work, FasterMoE.
Zhiquan Lai, Dongsheng Li 0001, Ke-shi Ge, Huayou Su
IEEE Trans. Parallel Distributed Syst.8
2025 Heterogeneous Feature-Aware Graph Neural Network for Tabular Data
Hongxiao Fei, Jinqi Hu, Tingxuan Chen, Huayou Su, Zanqun Liu
DASFAA (3)5
2025 REAMP: A Redundancy Elimination System for AMP-GNN Acceleration
Yongquan Fu, Huayou Su
ICIC (21)3
2025 ORE: An Offline Redundancy Elimination System for GNN Acceleration
Yongquan Fu, Huayou Su
ICIC (22)3
2024 Prism: Decomposing Program Semantics for Code Clone Detection through Compilation
abstract
Code clone detection (CCD) is of critical importance in software engineering, while semantic similarity is a key evaluation factor for CCD. The embedding technique, which represents an object using a numerical vector, is utilized to generate code representations, where code snippets with similar semantics (clone pairs) should have similar vectors. However, due to the diversity and flexibility of high-level program languages, the code representation of clone pairs may be inconsistent. Assembly code provides the program execution trace and can normalize the diversity of high-level languages in terms of the program behavior semantics. After revisiting the assembly language, we find that different assembly codes can align with the computational logic and memory access patterns of cloned pairs. Therefore, the use of multiple assembly languages can capture the behavior semantics to enhance the understanding of programs. Thus, we propose Prism, a new method for code clone detection fusing behavior semantics from multiple architecture assembly code, which directly captures multilingual domains' syntax and semantic information. Additionally, we introduce a multi-feature fusion strategy that leverages global information interaction to expand the representation space. This fusion process allows us to capture the complementary information from each feature and leverage the relationships between them to create a more expressive representation of the code. After testing the OJClone dataset, the Prism model exhibited exceptional performance with precision and recall scores of 0.999 and 0.999, respectively.
Haoran Li 0014, Siqian Wang, Weihong Quan, Xiaoli Gong, Huayou Su, Jin Zhang 0003
ICSE5
2024 SyncIntellects: Orchestrating LLM Inference with Progressive Prediction and QoS-Friendly Control
abstract
Large Language Models (LLMs) have shown impressive capabilities, especially in the realm of Human-Machine Chat Systems. Nevertheless, these models entail significant computational expenses, particularly when generating tokens. As a remedy to enhance system throughput and hardware utilization, batch scheduling is commonly adopted. This method involves initiating a batch of inference requests concurrently and then waiting for their completion. A significant challenge encountered with task-batching is the need to group requests with similar response lengths. However, accurately predicting response length proves to be a daunting task, and the inherent variability in response length leads to suboptimal resource utilization.In this paper, we introduce SyncIntellects, a framework designed to orchestrate Large Language Model (LLM) Inference with fine-grained response length prediction and Quality of Service (QoS)-Friendly length control. Specifically, SyncIntellects enhances response length prediction by leveraging embedding information during token generation through a transformer-based model. Subsequently, a dynamic response length controller based on Prompt Engineering techniques is employed to ensure alignment of response lengths without compromising the QoS of the responses. We have implemented SyncIntellects and seamlessly integrated it with a chatbot engine based on the llama2 7B model. We conduct comprehensive experiments on an NVIDIA A100-based testbed, and the results demonstrate a significant reduction in latency by 17.76% on average, along with an increase in throughput by 9.34%.
Xue Lin 0006, Peining Yue, Haoran Li 0014, Jin Zhang 0003, Baoyu Fan, Huayou Su, Xiaoli Gong
IWQoS7
2023 Optimizing GNN Inference Processing on Very Long Vector Processor
Kangkang Chen, Huayou Su, Chaorun Liu
ICA3PP (6)2
2022 An Efficient Transformer Inference Engine on DSP
Kangkang Chen, Huayou Su, Chaorun Liu, Xiaoli Gong
ICA3PP2
2021 Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement Learning
abstract
The communication among agents is important for Multi-Agent Reinforcement Learning (MARL). In this work, we propose GraphComm, a method makes use of the relation-ships among agents for MARL communication. GraphComm takes the explicit relations (e.g., agent types), which can be provided through some knowledge background, into account to better model the relationships among agents. Besides explicit relations, GraphComm considers implicit relations, which are formed by agent interactions. GraphComm use Graph Neural Networks (GNNs) to model the relational information, and use GNNs to assist the learning of agent communication. We show that GraphComm can obtain better results than state-of-the-art methods on the challenging StarCraft II unit micromanagement tasks through extensive experimental evaluation.
Yongquan Fu, Huayou Su, Hengyue Pan, Peng Qiao, Yong Dou, Cheng Wang 0003
ICASSP3
2021 Beyond AP: a new evaluation index for multiclass classification task accuracy
Kaifang Zhang, Huayou Su, Yong Dou
Appl. Intell.2
2021 Multilevel parallelism optimization of stencil computations on SIMDlized NUMA architectures
Kaifang Zhang, Huayou Su, Yong Dou
J. Supercomput.2
2020 Learning Network Representation Through Reinforcement Learning
abstract
Network Representation Learning embeds each node in a network into a low-dimensional real-value vector which can be used for downstream tasks such as link prediction and recommendation. Many existing approaches use unsupervised or (semi-)supervised methods to explore the network topology and learn representations from it. In contrast, we propose, reinforcement learning network representations (RLNet), which explores the idea of using reinforcement learning to learn to explore the network and to obtain network representations. Based on reward signals, RLNet learns an actor which uses a policy to determine the network navigation actions. RLNet uses node representations to parameterize its policy, and the representations are learned together with the policy. Through experiments based on multiple datasets, we show that RLNet can obtain promising results in link prediction tasks.
Yongquan Fu, Adele Lu Jia, Huayou Su, Chengsong Wang, Yong Dou
ICASSP4
2020 A High-Throughput LDPC Decoder Based on GPUs for 5G New Radio
abstract
In this paper, we propose a GPU-based QC-LDPC decoder for 5G New Radio(NR). Different from existing LDPC decoders based on GPUs, our decoder achieves high throughput when decoding LDPC codes with high code rates. Moreover, we implement the shortening and puncturing techniques which are exploited by 5G NR. The decoding algorithm Min-Sum approximation algorithm(MSA) is optimized to implement efficient parallel decoding on the GPU. In order to save the on-chip and the off-chip bandwidth, we propose the two-level quantization scheme and implement data packing on the GPU. We also analyse the optimum thread assignment for different code rates based on our implementation. By using the optimum settings on the GPU, the decoding throughput achieves 1.38 Gbps in the case of (2080, 1760), r=5/6 on Nvidia RTX 2080Ti.
Rongchun Li, Hengyue Pan, Huayou Su, Yong Dou
ISCC4
2020 DWS-MKL: Depth-width-scaling multiple kernel learning for data classification
Tingting Wang 0011, Huayou Su, Junbao Li
Neurocomputing2
2019 Author Disambiguation through Adversarial Network Representation Learning
abstract
Many persons share with the same name. Distinguishing different persons with the same name is important but challenging. Albeit much work has been proposed for author disambiguation, most of them do not adequately consider the heterogeneous relationships among authors and papers. In our work, ambiguous names and their related information, such as papers, conferences, titles, abstracts, etc., are constructed into a heterogeneous network which consists of different edge types. To fully incorporate all the information of the constructed network, we use Generative Adversarial Networks (GAN) to learn the network representation of the heterogeneous network. Although GAN has been used in many fields such as image generation, it hasn't been used to obtain representations for the heterogeneous network. As far as we know, our work is the first work which use adversarial training to learn heterogeneous network representation. After the representations are learned, they are partitioned into different groups each representing distinct authors. After extensive experiments on three major author disambiguation datasets, we demonstrate that our method outperforms several state-of-the-art baselines in author disambiguation problem.
Liwen Peng, Dongsheng Li 0001, Yongquan Fu, Huayou Su
IJCNN6
2019 A Skewness-Aware Matrix Factorization Approach for Mesh-Structured Cloud Services
abstract
Online cloud services need to fulfill clients' requests scalably and fast. State-of-the-art cloud services are increasingly deployed as a distributed service mesh. Service to service communication is frequent in the mesh. Unfortunately, problematic events may occur between any pair of nodes in the mesh, therefore, it is vital to maximize the network visibility. A state-of-the-art approach is to model pairwise RTTs based on a latent factor model represented as a low-rank matrix factorization. A latent factor corresponds to a rank-1 component in the factorization model, and is shared by all node pairs. However, different node pairs usually experience a skewed set of hidden factors, which should be fully considered in the model. In this paper, we propose a skewness-aware matrix factorization method named SMF. We decompose the matrix factorization into basic units of rank-one latent factors, and progressively combine rank-one factors for different node pairs. We present a unifying framework to automatically and adaptively select the rank-one factors for each node pair, which not only preserves the low rankness of the matrix model, but also adapts to skewed network latency distributions. Over real-world RTT data sets, SMF significantly improves the relative error by a factor of 0.2 x to 10 x, converges fast and stably, and compactly captures fine-grained local and global network latency structures.
Yongquan Fu, Dongsheng Li 0001, Pere Barlet-Ros, Chun Huang 0006, Zhen Huang 0006, Huayou Su
IEEE/ACM Trans. Netw.7
2018 Deep Discriminative Clustering Network
abstract
Deep clustering aims to cluster unlabeled data by embedding them into a subspace based on deep model. The key challenge of deep clustering is to learn discriminative representations for input data with high dimensions. In this paper, we present a deep discriminative clustering network for clustering the real-world images. We use a convolutional auto-encoder stacked with a softmax layer to predict clustering assignments. To learn a discriminative representations, the proposed approach adds discriminative loss as embedded regularization with relative entropy minimization. With the discriminative loss, the network can not only produce clustering assignments, but also learn discriminative features by reducing intra-cluster distance and increasing inter-cluster distance. We evaluate the proposed method on three datasets: MNIST-full, YTF and FRGC-v2.0. We outperform state-of-the-art results on MNIST-full and FRGC-v2.0 and achieve competitive result on YTF. The source code has been made publicly available at https://github.com/shaoxuying/DeepDiscriminativeClusteringNetwork.
Xuying Shaol, Ke-shi Ge, Huayou Su, Lei Luo 0002, Baoyun Peng, Dongsheng Li 0001
IJCNN3
2017 Efficient parallel implementation of a density peaks clustering algorithm on graphics processing unit
abstract
The density peak (DP) algorithm has been widely used in scientific research due to its novel and effective peak density-based clustering approach. However, the DP algorithm uses each pair of data points several times when determining cluster centers, yielding high computational complexity. In this paper, we focus on accelerating the time-consuming density peaks algorithm with a graphics processing unit (GPU). We analyze the principle of the algorithm to locate its computational bottlenecks, and evaluate its potential for parallelism. In light of our analysis, we propose an efficient parallel DP algorithm targeting on a GPU architecture and implement this parallel method with compute unified device architecture (CUDA), called the ‘CUDA-DP platform’. Specifically, we use shared memory to improve data locality, which reduces the amount of global memory access. To exploit the coalescing accessing mechanism of GPU, we convert the data structure of the CUDA-DP program from array of structures to structure of arrays. In addition, we introduce a binary search-and-sampling method to avoid sorting a large array. The results of the experiment show that CUDA-DP can achieve a 45-fold acceleration when compared to the central processing unit based density peaks implementation.
Ke-shi Ge, Huayou Su, Dongsheng Li 0001, Xicheng Lu
Frontiers Inf. Technol. Electron. Eng.2
2015 An analytical GPU performance model for 3D stencil computations from the angle of data traffic
Huayou Su, Xing Cai, Mei Wen, Chunyuan Zhang
J. Supercomput.1
2013 On the GPU-CPU Performance Portability of OpenCL for 3D Stencil Computations
abstract
Although OpenCL programming provides full code portability between different hardware platforms, performance portability can be far from satisfactory. In this work, we use a set of representative 3D stencil computations to study OpenCL's performance portability between GPUs and CPUs. For each stencil computation, we have devised different implementations of the computational kernel function, all being 100% code-portable between the two architectures. The most straightforward and compact implementation gives satisfactory CPU performance but performs poorly on GPUs, because such an implementation hampers effective use of the GPU hardware. By injecting code complexity into the involved loop nests, we can create kernel functions that still have full code portability but with increased performance portability. It is found that spatial data blocking and register reuse can be beneficial for performance on both GPUs and CPUs, whereas use of OpenCL's local memory (and subsequent temporal blocking) may only have positive effects on GPUs.
Huayou Su, Nan Wu 0003, Mei Wen, Chunyuan Zhang, Xing Cai
ICPADS1
2013 Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang
J. Supercomput.2
2012 Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin Simulations
abstract
In cutting-edge CPU/GPU hybrid clusters, such as Tianhe-1A, the aggregate CPU computing capability may amount to up to 1/3 of the aggregate GPU computing capability. It thus goes without saying that the CPUs and GPUs should jointly carry out the computational work. However, to effectively and simultaneously use both the hardware components requires great care when developing the parallel implementations. The challenges include (1) finding a balanced division of the workload between the CPU and GPU sides, and (2) hiding various overheads by overlapping computations with CPU-GPU data transfers and/or MPI communications. We study these issues in the context of real-world sedimentary basin simulations. Numerical experiments show that an appropriately devised CPU-GPU hybrid implementation is able to handle a global mesh resolution of 131,072*131,072, and a double-precision rate of 62 TFlops is achieved by using 1024 GPUs and 12288 CPU cores on Tianhe-1A. Such an extreme computing capability will be of great importance for carrying out high-resolution and continental-scale stratigraphic simulations in future.
Mei Wen, Huayou Su, Wenjie Wei, Nan Wu 0003, Xing Cai, Chunyuan Zhang
CLUSTER2
2012 Parallelization Design of Irregular Algorithms of Video Processing on GPUs
abstract
In this paper, we present the parallelization design consideration for irregular algorithms of video processing on GPUs. Enrich parallelism can be exploited by scheduling the processing order or making a tradeoff between performance and parallelism for irregular algorithms (such as CAVLC and deblocking filter). We implement a component-oriented CAVLC encoder and a direction-oriented deblocking filter on GPUs. The experiment results show that, compared with the implementation on CPU, the optimized parallel methods achieve high performance in term of speedup ratio from 63 to 44, relatively for deblocking filter and CAVLC. It shows that the rich parallelism is one of the most important factors to gain high performance for irregular algorithms based on GPUs. In addition, it seems that for some irregular kernels, the number of SM of GPU is more important to the performance than the computation capability.
Huayou Su, Jun Chai, Mei Wen, Ju Ren 0002, Chunyuan Zhang
ICME1
2012 A Parallel H.264 Encoder with CUDA: Mapping and Evaluation
abstract
Efficient mapping of a real-time HD video application to graphics hardware is challenging. Developers face the challenges of choosing the right parallelism model, balancing thread's process granularity between massive computing resources on the GPU, and partitioning tasks between the CPU and GPU. The paper illustrated the mapping approaches by a case of HD H.264 encoder based on X264 reference code and then evaluating it on state-of-the-art CPU and GPUs in depth. In the paper, we first split most of the computing task into Single-Instruction Multiple-Thread (SIMT) kernels, which are then chained intocertaininput/output data stream. Then we implementeda completed H.264 encoding on the computer unified device architecture (CUDA) platform. Finally, we present methods for exploiting multi-level parallelism and memory efficiency when mapping H.264 code, which we use to increase the efficiency of the execution on GPUs. Our experimental results show that computation efficiency of GPU and then real-time encoding performance are achieved with CUDA.
Nan Wu 0003, Mei Wen, Huayou Su, Ju Ren 0002, Chunyuan Zhang
ICPADS3
2012 Improving Performance of GPU Specific OpenCL Program on CPUs
abstract
OpenCL provides unified programming interface for various parallel computing platforms. The OpenCL framework manifests good functional portability, the programs can be run on platforms supporting OpenCL programming without any modification. However, most of the OpenCL programs are optimized for massively parallel processors, such as GPU, it's hard to achieve good performance on general multi-core processors without sophisticate modification to the GPU specific OpenCL programs. The major reason is the immense gap between CPU and GPU architecture. In this paper, we evaluate the performance portability of OpenCL programs between CPU and GPU, and analyse the reasons why GPU specific OpenCL programs are not fit for CPU. Based on the profiling, we proposed three optimization strategies for improving performance of GPU specific OpenCL programs on CPU, including increasing the granularity of task partition, optimizing the usage of memory hierarchy and block-based data accessing. In addition, we applied the proposed techniques on several benchmarks. The experimental results show that the performance of the optimized OpenCL programs achieve high performance in terms of speedup ratio from 2 to 4 on CPUs, when compared with their corresponding GPU specific ones.
Qiang Lan, Changqing Xun, Mei Wen, Huayou Su, Chunyuan Zhang
PDCAT4
2011 A Multilevel Parallel Intra Coding for H.264/AVC Based on CUDA
abstract
In this paper, we propose a multilevel parallel intra coding for H.264/AVC based on computed unified device architecture (CUDA). The proposed parallel algorithm improves the parallelism between 4×4 blocks within a macro block (MB) by throwing off some inappreciable prediction modes. By partitioning a frame into multi-slice, the parallelism between MBs can be exploited. In addition, a scalable parallel method for kernels is introduced to improve the performance of the proposed intra coding. Experimental results show that, more than 20 times speedup can be achieved with the assistance of GPU. Moreover, the entire encoder can meet the real-time processing requirement for HDTV.
Huayou Su, Nan Wu 0003, Chunyuan Zhang, Mei Wen, Ju Ren 0002
ICIG1
2011 High-efficient software parallel CAVLC encoder based on programmable stream processor
abstract
This article presents an efficient software parallel CAVLC encoder based on programmable stream processors (Storm- SP16 and GPU). For static processor Storm SP16, a block-based 16 ways parallel CAVLC is presented with streaming processing. A component-oriented CAVLC encoder is proposed aiming at dynamic stream processor GPU. Experiments results show that, compared to the CPU version, more than 70 times of speedup can be obtained for the CAVLC based on Storm and over 50 times for GPU-based component-oriented CAVLC encoder. The throughput of the presented CAVLC encoder is more than 10 times higher over that of published software CAVLC encoders on DSP and multi-core platforms.
Huayou Su, Chunyuan Zhang, Jun Chai, Mei Wen, Nan Wu 0003, Ju Ren 0002
ACM Multimedia1
2009 Streaming HD H.264 encoder on programmable processors
abstract
Programmable processors have great advantage over dedicated ASIC design under intense time-to-market pressure. However, real-time encoding of high-definition (HD) H.264 video (up to 1080p) is a challenge to most existing programmable processors. On the other hand, model-based design is widely accepted in developing complex media program. Stream model, an emerging model-based programming method, shows surprising efficiency on many compute-intensive domains especially for media processing. On the basis, this paper proposes a set of streaming techniques for H.264 encoding, and then develops all of the code based on the X264 reference code. Our streaming H.264 encoder is a pure software implementation completely written in high-level language without special hardware/algorithm support. Real execution results show that our encoder achieves significant speedup over the original X264 encoder on various programmable architectures: on X86 CoreTM2 E8200 the speedup is 1.8x, on MIPS 4KEc the speedup is 3.7x, on TMS320 C6416 DSP the speedup is 5.5x, on stream processor STORM-SP16 G220 the speedup is 6.1x. Especially, on STORM processor, the streaming encoder achieves the performance of 30.6 frames per second for a 1080P HD sequence, satisfying the real-time requirement. These indicate that streaming is extremely efficient for this kind of media workload. Our work is also applicable for other media processing applications, and provides architecture insights into dedicated ASIC or FPGA HD H.264 encoders.
Nan Wu 0003, Mei Wen, Ju Ren 0002, Huayou Su, Changqing Xun, Chunyuan Zhang
ACM Multimedia5