Shangkun Sun

dblp:321/3659 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-2922-9349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Visual content generation and editing · 36% Multimedia systems and quality of experience · 24% Image and video processing · 21%
Artificial intelligence
3 papers
Video understanding and tracking · 34% 3D vision · 24% Deep learning architectures and training · 24%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › motion estimation
optical flow
1.022025
StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences · NeurIPS 2024
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow · ICCV 2025
Computer vision › Video understanding and tracking
long video understanding
0.912025
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow · ICCV 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow · ICCV 2025
Visual content generation and editing › video editing
text-driven video editing
0.912025
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment · AAAI 2025
Visual content generation and editing
video editing
0.912025
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment · AAAI 2025
Multimedia systems and quality of experience
video quality assessment
0.912025
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment · AAAI 2025
Computer vision › Video understanding and tracking
spatio-temporal modeling
0.812024
StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences · NeurIPS 2024
Image and video coding › video compression
learned video compression
0.712023
OpenDMC: An Open-Source Library and Performance Evaluation for Deep-learning-based Multi-frame Compression · ACM Multimedia 2023
Performance modeling and evaluation
benchmarking
0.712023
OpenDMC: An Open-Source Library and Performance Evaluation for Deep-learning-based Multi-frame Compression · ACM Multimedia 2023
Machine learning › Deep learning architectures and training
convolutional neural network
0.612022
SKFlow: Learning Optical Flow with Super Kernels · NeurIPS 2022
Computer vision › 3D vision
motion estimation
0.612022
SKFlow: Learning Optical Flow with Super Kernels · NeurIPS 2022
Computer vision › 3D vision › motion estimation
optical flow
0.612022
SKFlow: Learning Optical Flow with Super Kernels · NeurIPS 2022
Machine learning › Deep learning architectures and training › convolutional neural network › receptive field
receptive field design
0.612022
SKFlow: Learning Optical Flow with Super Kernels · NeurIPS 2022
Image and video coding
quality assessment
0.312025
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment · AAAI 2025
Multimedia systems and quality of experience
subjective quality assessment
0.312025
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment · AAAI 2025

Methods — techniques the papers use, named apart from their topics

temporal granularity optimization · 1.7optical flow · 1.7motion token pruning · 1.7integrative spatiotemporal coherence · 1.5in-batch multi-frame pipeline · 1.5global temporal regressor · 1.5rate-distortion optimization · 1.3mean opinion score · 0.9human-aligned metric · 0.9super kernels · 0.6depthwise convolution · 0.6conical connections · 0.6
YearPublicationVenuePosition
2025 VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
abstract
Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this, we introduce VE-Bench, a benchmark suite tailored to the assessment of text-driven video editing. This suite includes VE-Bench DB, a video quality assessment (VQA) database for video editing. VE-Bench DB encompasses a diverse set of source videos featuring various motions and subjects, along with multiple distinct editing prompts, editing results from 8 different models, and the corresponding Mean Opinion Scores (MOS) from 24 human annotators. Based on VE-Bench DB, we further propose VE-Bench QA, a quantitative human-aligned measurement for the text-driven video editing task. In addition to the aesthetic, distortion, and other visual quality indicators that traditional VQA methods emphasize, VE-Bench QA focuses on the text-video alignment and the relevance modeling between source and edited videos. It introduces a new assessment network for video editing that attains superior performance in alignment with human preferences.To the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and proposes an effective subjective-aligned quantitative metric for this domain. All models, data, and code will be publicly available to the community.
Shangkun Sun, Songlin Fan, Wenxu Gao, Wei Gao 0003
AAAI1
2025 Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
abstract
Long-form video understanding has always been a challenging problem due to the significant redundancy in both temporal and spatial contents. This challenge is further exacerbated by the limited context length of Multimodal Large Language Models (MLLMs). To address this issue, many previous works have attempted to extract key video information, where the "key" is typically semantic-aware and heavily dependent on the CLIP model as prior. In this paper, we propose Flow4Agent, a novel framework that pioneeringly incorporates motion priors from optical flow to facilitate LLM-based long video understanding. Flow4Agent mitigates the redundancy in long videos at both temporal and spatial levels through two core modules: Temporal Granularity Optimization (TGO) adaptively refines framelevel hierarchies, which first leverages coarse flow priors to group similar visual contents and then applies semantic priors to filter out highly irrelevant scene information. Motion Token Pruning (MTP) further refines the intra-frame visual representations, pruning high-redundancy video tokens using fine-grained optical flow information. Extensive experiments demonstrate that our Flow4Agent outperforms existing methods across a wide range of video MLLM benchmarks, especially for hour-level video understanding tasks, achieving 64.7% on Video-MME, 71.4% on MLVU and 60.4% on LongVideoBench.
Ruyang Liu, Shangkun Sun, Wei Gao 0003, Ge Li 0002
ICCV2
2024 StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences
abstract
Prior multi-frame optical flow methods typically estimate flow repeatedly in a pair-wise manner, leading to significant computational redundancy. To mitigate this, we implement a Streamlined In-batch Multi-frame (SIM) pipeline, specifically tailored to video inputs to minimize redundant calculations. It enables the simultaneous prediction of successive unidirectional flows in a single forward pass, boosting processing speed by 44.43% and reaching efficiencies on par with two-frame networks. Moreover, we investigate various spatiotemporal modeling methods for optical flow estimation within this pipeline. Notably, we propose a simple yet highly effective parameter-efficient Integrative spatiotemporal Coherence (ISC) modeling method, alongside a lightweight Global Temporal Regressor (GTR) to harness temporal cues. The proposed ISC and GTR bring powerful spatiotemporal modeling capabilities and significantly enhance accuracy, including in occluded areas, while adding modest computations to the SIM pipeline. Compared to the baseline, our approach, StreamFlow, achieves performance enhancements of 15.45% and 11.37% on the Sintel clean and final test sets respectively, with gains of 15.53% and 10.77% on occluded regions and only a 1.11% rise in latency. Furthermore, StreamFlow exhibits state-of-the-art cross-dataset testing results on Sintel and KITTI, demonstrating its robust cross-domain generalization capabilities. The code is available [here](https://github.com/littlespray/StreamFlow).
Shangkun Sun, Huaxia Li, Thomas H. Li, Wei Gao 0003
NeurIPS1
2024 Closing the Gap Between Theory and Practice During Alternating Optimization for GANs
abstract
Synthesizing high-quality and diverse samples is the main goal of generative models. Despite recent great progress in generative adversarial networks (GANs), mode collapse is still an open problem, and mitigating it will benefit the generator to better capture the target data distribution. This article rethinks alternating optimization in GANs, which is a classic approach to training GANs in practice. We find that the theory presented in the original GANs does not accommodate this practical solution. Under the alternating optimization manner, the vanilla loss function provides an inappropriate objective for the generator. This objective forces the generator to produce the output with the highest discriminative probability of the discriminator, which leads to mode collapse in GANs. To address this problem, we introduce a novel loss function for the generator to adapt to the alternating optimization nature. When updating the generator by the proposed loss function, the reverse Kullback-Leibler divergence between the model distribution and the target distribution is theoretically optimized, which encourages the model to learn the target distribution. The results of extensive experiments demonstrate that our approach can consistently boost model performance on various datasets and network structures.
Yuanqi Chen, Shangkun Sun, Ge Li 0002, Wei Gao 0003, Thomas H. Li
IEEE Trans. Neural Networks Learn. Syst.2
2023 OpenDMC: An Open-Source Library and Performance Evaluation for Deep-learning-based Multi-frame Compression
abstract
Video streaming has become an essential component of our everyday routines. Nevertheless, video data imposes a significant strain on data usage, demanding substantial bandwidth and storage resources for effective transmission. To suit explosively increasing video transmission and storage requirements, deep-learning-based video compression has developed rapidly in the past few years. New methods have mushroomed in order to achieve better Rate-Distortion (RD) performance. However, the absence of an algorithm library that can effectively sort, classify, and conduct extensive benchmark testing on existing algorithms remains a challenge. In this paper, we present an open-source algorithm library called OpenDMC, which integrates a variety of end-to-end video compression methods in cross-platform environments. We provide comprehensive descriptions of the algorithms used in the library, including their contributions and implementation details. We perform a thorough benchmarking test to evaluate the performance of the algorithms. We meticulously compare and analyze each algorithm based on various metrics, including RD performance, running time, and GPU memory usage. The open-source library for OpenDMC is available at https://openi.pcl.ac.cn/OpenDMC/.
Wei Gao 0003, Shangkun Sun, Huiming Zheng, Yuyang Wu, Yongchi Zhang
ACM Multimedia2
2022 SKFlow: Learning Optical Flow with Super Kernels
abstract
Optical flow estimation is a classical yet challenging task in computer vision. One of the essential factors in accurately predicting optical flow is to alleviate occlusions between frames. However, it is still a thorny problem for current top-performing optical flow estimation methods due to insufficient local evidence to model occluded areas. In this paper, we propose the Super Kernel Flow Network (SKFlow), a CNN architecture to ameliorate the impacts of occlusions on optical flow estimation. SKFlow benefits from the super kernels which bring enlarged receptive fields to complement the absent matching information and recover the occluded motions. We present efficient super kernel designs by utilizing conical connections and hybrid depth-wise convolutions. Extensive experiments demonstrate the effectiveness of SKFlow on multiple benchmarks, especially in the occluded areas. Without pre-trained backbones on ImageNet and with a modest increase in computation, SKFlow achieves compelling performance and ranks $\textbf{1st}$ among currently published methods on the Sintel benchmark. On the challenging Sintel clean and final passes (test), SKFlow surpasses the best-published result in the unmatched areas ($7.96$ and $12.50$) by $9.09\%$ and $7.92\%$. The code is available at https://github.com/littlespray/SKFlow.
Shangkun Sun, Yuanqi Chen, Yu Zhu 0006, Guodong Guo
NeurIPS1