Youbo Mao

dblp:425/3661 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0006-0542-1033ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 33% Hardware accelerators and domain-specific architectures · 33% Embedded and real-time systems · 33%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing
0.912025
FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile Devices · ACM Multimedia 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile Devices · ACM Multimedia 2025
Embedded and real-time systems › on-device inference
mobile inference
0.912025
FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile Devices · ACM Multimedia 2025
Image and video coding › image decoding
JPEG decoding
0.312025
FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile Devices · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

multi-core processing · 1.7hybrid parallel pipeline · 1.7frequency domain inference · 1.7
YearPublicationVenuePosition
2025 FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile Devices
abstract
With the increasing popularity of image and video analysis on mobile devices, high-throughput image inference has become essential. However, current mobile deep learning frameworks face key bottlenecks: high computational load in JPEG image recognition and low processor efficiency, which limit overall image processing throughput. To address these issues, this paper proposes the FCG framework (Frequency Domain model for CPU and GPU), a mobile JPEG inference framework based on frequency domain data and a hybrid parallel architecture that enables high-throughput inference for JPEG-encoded images on mobile devices. FCG decouples JPEG decoding from model inference by discarding the traditional RGB decoding process and retaining only the Huffman decoding. This decoding step is further accelerated through multi-core processing, significantly reducing the computational burden and latency during preprocessing. In light of the characteristics of frequency domain data and the heterogeneous CPU/GPU processors on mobile devices, FCG reconstructs the deep learning model to ensure recognition accuracy while optimizing resource utilization. By effectively allocating tasks and combining parallel and sequential execution, FCG optimizes processor resource utilization to achieve high throughput and low latency. FCG outperforms the state-of-the-art NN-Stretch by reducing latency by 36%. It also achieves significant throughput improvements—3.6x, 3.3x, and 2.8x—on CPU, GPU, and CPU+GPU configurations, respectively, compared to sequential inference systems. Additionally, FCG reduces power consumption by 56%, 35%, and 43% in these configurations.
Youbo Mao, Ziyang Kang, Jiyao Chen, Zenglin Yang, Zhijun Li 0002
ACM Multimedia1