Enxin Song

dblp:353/2190 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 43% Video understanding and tracking · 32% Deep learning architectures and training · 11%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 77% Multimedia analysis and retrieval · 23%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
video-language model
1.822026
MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024
Computer vision › Video understanding and tracking › video question answering
long-form video question answering
1.012026
MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.912025
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis · ICLR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Bringing RNNs Back to Efficient Open-Ended Video Understanding · ICCV 2025
Computer vision › Vision and language
video captioning
0.912025
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis · ICLR 2025
Computer vision › Video understanding and tracking
long video understanding
0.812024
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024
Multimedia analysis and retrieval
video understanding
0.312025
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025
Natural language and speech › Question answering and dialogue systems
long-context memory
0.212024
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024

Methods — techniques the papers use, named apart from their topics

transformer · 2.5token merging · 1.7masked image modeling · 1.7large multimodal model · 1.7vision-question matching · 1.0training-free adaptation · 1.0memory consolidation · 1.0recurrent neural network · 0.9memory mechanism · 0.8
YearPublicationVenuePosition
2026 MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific vision tasks. Yet, existing methods either employ complex spatial-temporal modules or rely heavily on additional perception models to extract temporal features for video understanding, performing well only on short videos. For long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges. Leveraging the hierarchical memory structure of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination, we propose MovieChat within a training-free memory consolidation mechanism to overcome these challenges, which transfers dense frames from short-term memory into sparse tokens in long-term memory by temporally merging adjacent frames. We lift pre-trained large multi-modal models for understanding long videos without additional trainable modules, employing a zero-shot approach. Additionally, in our new version, MovieChat+, we design an enhanced training-free vision-question matching-based memory consolidation mechanism to better anchor predictions to relevant visual content. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1 K benchmark with 1 K long video, 2 K temporal grounding labels, and 14 K manual annotations.
Enxin Song, Wenhao Chai, Tian Ye 0001, Jenq-Neng Hwang, Xi Li 0001, Gaoang Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai, Xuexiang Wen, Tian Ye 0001, Gaoang Wang
ICCV2
2025 Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
abstract
We present Meissonic, which elevates non-autoregressive text-to-image Masked Image Modeling (MIM) to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing methods in generating high-quality, high-resolution images. Extensive experiments validate Meissonic’s capabilities, demonstrating its potential as a new standard in text-to-image synthesis.
Jinbin Bai, Tian Ye 0001, Wei Chow, Enxin Song, Xiangtai Li, Zhen Dong 0003, Lei Zhu 0003, Shuicheng Yan
ICLR4
2025 AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
abstract
Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality.
Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Saining Xie, Christopher D. Manning
ICLR2
2024 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat.
Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang
CVPR1