Dongjie Tang

dblp:268/1475 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 gCom: Fine-grained Compressors in Graphics Memory of Mobile GPU
abstract
Today, GPUs significantly boost rendering performance. However, the high memory requirements limit their use, especially on low-end mobile platforms. Compression techniques have been widely adopted to reduce memory consumption but face two primary issues when applied to mobile GPUs: (1) low repetition ratio caused by small raw data sizes and concurrency, and (2) low locality caused by unpredictable rendering behaviors. These two limitations result in a low compression ratio when compressors are applied to low-end mobile devices. This article introduces gCom , a fine-grained rendering compressor accelerated by GPUs. To improve the compression ratio, gCom incorporates the following innovations. First, unlike other compression techniques that use frames or tiles as basic processing units, gCom is the first to employ a fine-grained processing unit (i.e., the color channel), enhancing repetition amplification without increasing raw data. Second, gCom introduces two key features— Hierarchical Delta and Channel Decorrelator —which maximize the locality of adjacent channels and reduce raw data size. Third, to maintain the original GPU throughput, gCom revolutionizes the Golomb-Rice algorithm and proposes a new compression approach, the Parallel-Oriented Golomb-Rice algorithm, enabling parallel execution of both decompression and compression processes. The entire design of gCom utilizes only idle resources and existing commands on mobile GPUs, thus keeping purchasing costs low. To date, gCom has improved the channel locality by nearly 50%. The best compression achievement received by gCom has reached around 20%.
Dongjie Tang, Yun Wang 0039, Yicheng Gu, Fangxin Liu, Zhengwei Qi
ACM Trans. Archit. Code Optim.1
2025 Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture
Yun Wang 0039, Bing Deng, Xia Jiang, Xuyan Hu, Dongjie Tang, Randy Xu, Yijin Sun, Zhengwei Qi
IEEE Trans. Serv. Comput.5
2024 CARE: Cloudified Android With Optimized Rendering Platform
abstract
Due to the excellent rendering capabilities, GPUs are mainstream accelerators in the Cloud-rendering industry. However, current Cloud-rendering systems suffer from a CPU-GPU workload imbalance that not only degrades application performance but also causes a significant waste of GPU resources. Recent proposals (such as API-forwarding and c-GPU) for improving CPU-GPU balance are promising but fail to solve system-resource redundancy issues (i.e., each instance tends to occupy all resources, exceeding its requirements). Such behavior will increase CPU load and lower effective GPU utilization. To demonstrate the severity of the issue, we evaluated real-world applications and results show that in most cases, nearly 50% of resources are useless. To solve this problem, we present CARE, the first framework intended to reduce the system-level redundancy by cloudifying the system from monolithic to Cloud-native. To allow users to configure required services, CARE puts forward a functional unit calledConfigurable Android (CA). To allow multiple instances to share certain types of resources, CARE innovatesSharing Resource (SR). To reduce the unused services, CARE introducesPruning Resources (PR). To further alleviate the CPU pressure and achieve CPU-GPU balance, we propose rShare, a system aiming at enhancing CPU effective utilization and increasing Android instance density of the Cloud-rendering platform. Based on Kubernetes, rShare divides all the CPUs into non-overlapping shared CPU pools, allocates instances to pools within milliseconds, and dynamically migrates them by tracking their QoS status. So far, CARE primarily focuses on Android systems and can handle 60 heavyweight instances (e.g., KOG (King of Glory)) on Intel SG1. rShare can apply instance allocation within milliseconds and increase the platform density by 39.4%.
Yuxin Xiang, Dongjie Tang, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Cathy Bao, Yicheng Gu, Zhengwei Qi, Haibing Guan
IEEE Trans. Multim.2
2023 rShare: Alleviating long startup on the Cloud-rendering platform through de-systemization
Dongjie Tang, Marc Mao, Cathy Bao, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Yun Wang 0039, Zhengwei Qi, Haibing Guan, Xiaojie Cao
J. Syst. Archit.1
2021 CARE: Cloudified Android OSes on the Cloud Rendering
abstract
GPUs have become ubiquitous in the Cloud-rendering areas due to the outstanding rendering performance. However, many existing Cloud-rendering systems suffer from low GPU utilization caused by the CPU bottleneck. Recent proposals (e.g., API-forwarding and c-GPU) for GPU-usage optimization are promising but fail to address the system-resource redundancy issues (i.e., each instance tends to occupy all the system resources exceeding their requirements), leading to unnecessary CPU consumption and lowering GPU utilization. We conducted an experiment by testing real-world applications on the percentage of unused resources to demonstrate the severity of this issue. Nearly 50% of resources are unused.
Dongjie Tang, Cathy Bao, Qiming Shi, Marc Mao, Randy Xu, Linsheng Li, Mohammad R. Haghighat, Zhengwei Qi, Haibing Guan
ACM Multimedia1
2021 gRemote: Cloud rendering on GPU resource pool based on API-forwarding
Dongjie Tang, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan
J. Syst. Archit.1
2020 gRemote: API-Forwarding Powered Cloud Rendering
abstract
Traditional GPU resource allocation approaches, widely adopted in today's data centers, only focus on the server-side functions while ignoring the client-side. These approaches waste client-side hardware resources. To solve this problem, remote API-forwarding architectures appear. Through running applications on the client-side, remote API-forwarding architectures offload some workloads to the client. However, many remote API-forwarding systems suffer from one big issue: shared-resource interference, stemming from two reasons: (a) GPU resource racing caused by resource overuse for a single client, and (b) CPU resource racing caused by resource shortage among clients. This paper presents gRemote, an open-source GPU-remoting system that can address this issue. To mitigate the CPU resource shortage, gRemote improves CPU configurations by expanding CPU resources from the server-side to both server- and client-side. To maintain the reasonable GPU usage for individual tasks, we innovate a new resource-sharing mechanism called GPU throttle. gRemote supports 1,228 OpenGL commands with around 10% shared-resource interference.
Dongjie Tang, Yun Wang 0039, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan
HPDC1