Yuchong Sun

dblp:206/8045 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
14since 2021 · last 2025
0009-0004-6559-5620ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
abstract
Imitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building alive virtual characters. Previous studies aim to predict gaze trajectories when humans are free-viewing an image, searching for required targets, or looking for clues to answer questions in an image. While these tasks focus on visual-centric scenarios, humans move their gaze also along with audio signal inputs in more common scenarios. To fill this gap, we introduce a new task that predicts human gaze trajectories in a visual scene with synchronized audio inputs and provide a new dataset containing 20k gaze points from 8 subjects. To effectively integrate audio information and simulate the dynamic process of human gaze motion, we propose a novel learning framework called EyEar (Eye moving while Ear listening) based on physics-informed dynamics, which considers three key factors to predict gazes: eye inherent motion tendency, vision salient attraction, and audio semantic attraction. We also propose a probability density score to overcome the high individual variability of gaze trajectories, thereby improving the stabilization of optimization and the reliability of the evaluation. Experimental results show that EyEar outperforms all the baselines in the context of all evaluation metrics, thanks to the proposed components in the learning model.
Xin Cheng 0008, Yuchong Sun, Ruihua Song, Hao Sun 0002, Denghao Zhang
AAAI3
2025 MuKA: Multimodal Knowledge Augmented Visual Information-Seeking
abstract
The visual information-seeking task aims to answer visual questions that require external knowledge, such as “On what date did this building officially open?”. Existing methods using retrieval-augmented generation framework primarily rely on textual knowledge bases to assist multimodal large language models (MLLMs) in answering questions. However, the text-only knowledge can impair information retrieval for the multimodal query of image and question, and also confuse MLLMs in selecting the most relevant information during generation. In this work, we propose a novel framework MuKA which leverages a multimodal knowledge base to address these limitations. Specifically, we construct a multimodal knowledge base by automatically pairing images with text passages in existing datasets. We then design a fine-grained multimodal interaction to effectively retrieve multimodal documents and enrich MLLMs with both retrieved texts and images. MuKA outperforms state-of-the-art methods by 38.7% and 15.9% on the InfoSeek and E-VQA benchmark respectively, demonstrating the importance of multimodal knowledge in enhancing both retrieval and answer generation.
Lianghao Deng, Yuchong Sun, Shizhe Chen, Ruihua Song
COLING2
2025 ETVA: Evaluation of Text-to-Video Alignment via Fine-Grained Question Generation and Answering
abstract
Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. First, a multi-agent system parses prompts into semantic scene graphs to generate atomic questions. Then we design a knowledge-augmented multi-stage reasoning framework for question answering, where an auxiliary LLM first retrieves relevant common-sense knowledge (e.g., physical laws), and then video LLM answers the generated questions through a multi-stage reasoning mechanism. Extensive experiments demonstrate that ETVA achieves a Spearman's correlation coefficient of 58.47, showing a much higher correlation with human judgment than existing metrics which attain only 31.0. We also construct a comprehensive benchmark specifically designed for text-to-video alignment evaluation, featuring 2k diverse prompts and 12k atomic questions spanning 10 categories. Through a systematic evaluation of 15 existing text-to-video models, we identify their key capabilities and limitations, paving the way for next-generation T2V generation.
Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Kieran Liu, Ruihua Song
ICCV3
2025 Uncovering Personality Traits via Multimodal LLM for Personalized Image Emotion Analysis
abstract
Human emotion induced by images is strongly linked to individual personalities. Most existing works focus on analyzing the dominant emotions, i.e., the emotion that most viewers have for an image, leaving personalized image emotion analysis less explored. In this paper, we propose MLLM-PIEA, a framework based on Multimodal Large Language Models (MLLMs) for Personalized Image Emotion Analysis. To better represent the personalities of different viewers, we propose using an MLLM to uncover personality traits from the viewers’ experience data. These personality traits are summarized as structured descriptions and then used to augment another MLLM for emotion analysis. Experimental results show that our method brings a significant relative improvement of 28.8% over the baseline method.
Jianzhang Gao, Hao Pu, Yuchong Sun, Ruihua Song
ICME3
2025 JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
abstract
This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFusion module for spatio-temporal audio-video fusion and synchrony-aware learnable queries to bridge a pretrained JAV-DiT generator. This design enables temporally coherent video-audio understanding and generation from multimodal instructions. We design an effective three-stage training pipeline consisting of multimodal pretraining, audio-video fine-tuning, and large-scale instruction-tuning, to progressively build multimodal comprehension and generation from existing vision-language models. For instruction tuning, we construct JavisInst-Omni, a high-quality instruction dataset with over 200K GPT-4o-curated audio-video-text dialogues that cover diverse and multi-level comprehension and generation scenarios. On JAV comprehension and generation benchmarks, our experiments show that JavisGPT outperforms existing MLLMs, particularly in complex and temporally synchronized settings.
Kai Liu 0023, Jungang Li, Yuchong Sun, Shengqiong Wu, Jianzhang Gao, Daoan Zhang, Wei Zhang 0090, Sheng Jin 0002, Sicheng Yu, Geng Zhan, Jiayi Ji, Fan Zhou 0007, Shuicheng Yan, Hao Fei 0001, Tat-Seng Chua
NeurIPS3
2025 ReGA: Reasoning and Grounding Decoupled GUI Navigation Agents
Feiyue Ni, Yanchu Guan, Yuchong Sun, Dong Wang 0062, Chenyi Zhuang, Jinjie Gu, Ruihua Song
NLPCC (1)3
2024 Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models
abstract
Yuchong Sun, Che Liu, Kun Zhou, Jinwen Huang, Ruihua Song, Xin Zhao, Fuzheng Zhang, Di Zhang, Kun Gai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuchong Sun, Kun Zhou 0002, Jinwen Huang, Ruihua Song, Wayne Xin Zhao, Di Zhang 0026, Kun Gai
ACL (1)1
2024 ViCo: Engaging Video Comment Generation with Human Preference Rewards
Yuchong Sun, Bei Liu 0001, Xu Chen 0017, Ruihua Song, Jianlong Fu
MMAsia1
2023 CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment
Hongwei Xue, Yuchong Sun, Bei Liu 0001, Jianlong Fu, Ruihua Song, Houqiang Li, Jiebo Luo 0001
ICLR2
2023 TeViS: Translating Text Synopses to Video Storyboards
abstract
A video storyboard is a roadmap for video creation which consists of shot-by-shot images to visualize key plots in a text synopsis. Creating video storyboards, however, remains challenging which not only requires cross-modal association between high-level texts and images but also demands long-term reasoning to make transitions smooth across shots. In this paper, we propose a new task called Text synopsis to Video Storyboard (TeViS) which aims to retrieve an ordered sequence of images as the video storyboard to visualize the text synopsis. We construct a MovieNet-TeViS dataset based on the public MovieNet dataset [17]. It contains 10K text synopses each paired with keyframes manually selected from corresponding movies by considering both relevance and cinematic coherence. To benchmark the task, we present strong CLIP-based baselines and a novel VQ-Trans model. VQ-Trans first encodes text synopsis and images into a joint embedding space and uses vector quantization (VQ) to improve the visual representation. Then, it auto-regressively generates a sequence of visual features for retrieval and ordering. Experimental results demonstrate that VQ-Trans significantly outperforms prior methods and the CLIP-based baselines. Nevertheless, there is still a large gap compared to human performance suggesting room for promising future work. The code and data are available at: https://ruc-aimind.github.io/projects/TeViS/
Xu Gu 0003, Yuchong Sun, Feiyue Ni, Shizhe Chen, Xihua Wang 0002, Ruihua Song
ACM Multimedia2
2023 Going Beyond Closed Sets: A Multimodal Perspective for Video Emotion Analysis
Hao Pu, Yuchong Sun, Ruihua Song, Xu Chen 0017, Hao Jiang 0022, Zhao Cao
PRCV (6)2
2023 Expanding the Horizons: Exploring Further Steps in Open-Vocabulary Segmentation
Xihua Wang 0002, Yuchong Sun, Ruihua Song
PRCV (10)4
2022 Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
abstract
We study joint video and language (VL) pretraining to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that high-resolution videos and diversified semantics can significantly improve cross-modality learning. In this paper, we propose a novel High-resolution and Diversified VIdeo-LAnguage pre-training model (HD-VILA) for many visual tasks. In particular, we collect a large dataset with two distinct properties: 1) the first high-resolution dataset including 371.5k hours of 720p videos, and 2) the most diversified dataset covering 15 popular YouTube categories. To enable VL pre-training, we jointly optimize the HD-VILA model by a hybrid Transformer that learns rich spatiotemporal features, and a multimodal Transformer that enforces interactions of the learned video features with diversified texts. Our pre-training model achieves new state-of-the-art results in 10 VL understanding tasks and 2 more novel text-to-visual generation tasks. For example, we outperform SOTA models with relative increases of 40.4% R@1 in zero-shot MSR-VTT text-to-video retrieval task, and 55.4% in high-resolution dataset LSMDC. The learned VL embedding is also effective in generating visually pleasing and semantically relevant results in text-to-visual editing and super-resolution tasks.
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu 0001, Huan Yang 0005, Jianlong Fu, Baining Guo
CVPR4
2022 Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning
abstract
Large-scale video-language pre-training has shown significant improvement in video-language understanding tasks. Previous studies of video-language pretraining mainly focus on short-form videos (i.e., within 30 seconds) and sentences, leaving long-form video-language pre-training rarely explored. Directly learning representation from long-form videos and language may benefit many long-formvideo-language understanding tasks. However, it is challenging due to the difficulty of modeling long-range relationships and the heavy computational burden caused by more frames. In this paper, we introduce a Long-Form VIdeo-LAnguage pre-training model (LF-VILA) and train it on a large-scale long-form video and paragraph dataset constructed from an existing public dataset. To effectively capturethe rich temporal dynamics and to better align video and language in an efficient end-to-end manner, we introduce two novel designs in our LF-VILA model. We first propose a Multimodal Temporal Contrastive (MTC) loss to learn the temporal relation across different modalities by encouraging fine-grained alignment between long-form videos and paragraphs. Second, we propose a Hierarchical Temporal Window Attention (HTWA) mechanism to effectively capture long-range dependency while reducing computational cost in Transformer. We fine-tune the pre-trained LF-VILA model on seven downstream long-form video-language understanding tasks of paragraph-to-video retrieval and long-form video question-answering, and achieve new state-of-the-art performances. Specifically, our model achieves 16.1% relative improvement on ActivityNet paragraph-to-video retrieval task and 2.4% on How2QA task, respectively. We release our code, dataset, and pre-trained models at https://github.com/microsoft/XPretrain.
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 0001, Huan Yang 0005, Jianlong Fu
NeurIPS1
2017 High-speed driver for SiC MOSFET based on class-E inverter
abstract
This paper presents a high-speed resonant driver for silicon carbide (SiC) MOSFET, which is based on the class-E inverter. Gate-to-source resistance and input capacitance of SiC MOSFET are included in the resonant circuit of the driver for avoiding the driving-signal strain. Additionally, clamp diodes are added to the output filter in the proposed driver, which makes a square-waveform-like driving signal. The design example along with experimental waveforms is shown in this paper. It was possible in the laboratory experiments to drive the SiC MOSFET at 7 MHz and 13 MHz frequencies with satisfying the zero-voltage switching condition, which showed validity and effectiveness of the proposed driver and its design procedure.
Yuchong Sun, Ryoko Sugano, Xiuqin Wei, Takashi Hikihara, Hiroo Sekiya
ISCAS1