Aodong Chen

dblp:365/4729 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 40% Parallel and multicore computing · 40% GPUs and heterogeneous computing · 20%
Artificial intelligence
2 papers
Video understanding and tracking · 88% Deep learning architectures and training · 12%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › deep video understanding › video reasoning
story video understanding
1.012026
StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset · Int. J. Comput. Vis. 2026
Computer vision › Video understanding and tracking
video question answering
1.012026
StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset · Int. J. Comput. Vis. 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN inference accelerator
0.912025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025
GPUs and heterogeneous computing
GPU computing
0.912025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025
Parallel and multicore computing › parallel query processing
operator parallelism
0.912025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025
Hardware accelerators and domain-specific architectures › dataflow optimization
operator scheduling
0.912025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025
Parallel and multicore computing
parallel scheduling
0.912025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025
Machine learning › Deep learning architectures and training › neural network inference
DNN inference
0.312025
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs · IEEE Trans. Computers 2025

Methods — techniques the papers use, named apart from their topics

operator fusion · 1.7CUDA graphs · 1.7CUDA streams · 0.9CUDA Streams · 0.9
YearPublicationVenuePosition
2026 StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
Zhengqian Wu, Zhixian Liu, Aodong Chen, Jingyang Zhang, Ruizhe Li 0004, Hanlin Ge, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001
Int. J. Comput. Vis.3
2025 Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
abstract
GPUs have become thedefactohardware devices for accelerating Deep Neural Network (DNN) inference workloads. However, the conventionalsequential execution mode of DNN operatorsin mainstream deep learning frameworks cannot fully utilize GPU resources, even with the operator fusion enabled, due to the increasing complexity of model structures and a greater diversity of operators. Moreover, theinadequate operator launch orderin parallelized execution scenarios can lead to GPU resource wastage and unexpected performance interference among operators. In this paper, we proposeOpara, a resource- and interference-aware DNNOperatorparallel scheduling framework to accelerate DNN inference on GPUs. Specifically,Oparafirst employsCUDA StreamsandCUDA Graphtoparallelizethe execution of multiple operators automatically. To further expedite DNN inference,Oparaleverages the resource demands of operators to judiciously adjust the operator launch order on GPUs, overlapping the execution of compute-intensive and memory-intensive operators. We implement and open source a prototype ofOparabased on PyTorch in anon-intrusivemanner. Extensive prototype experiments with representative DNN and Transformer-based models demonstrate thatOparaoutperforms the default sequentialCUDA Graphin PyTorch and the state-of-the-art operator parallelism systems by up to$1.68\boldsymbol{\times}$and$1.29\boldsymbol{\times}$, respectively, yet with acceptable runtime overhead.
Aodong Chen, Fei Xu 0009, Li Han 0001, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers1