VLDB 2026 Research / reviewers in the wild / expert
Enxin Song
dblp:353/2190
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Vision and language · 43% Video understanding and tracking · 32% Deep learning architectures and training · 11% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 77% Multimedia analysis and retrieval · 23% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
video-language model |
1.8 | 2 | 2026 | MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024 |
Computer vision › Video understanding and tracking › video question answering
long-form video question answering |
1.0 | 1 | 2026 | MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.9 | 1 | 2025 | Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 1 | 2025 | Bringing RNNs Back to Efficient Open-Ended Video Understanding · ICCV 2025 |
Computer vision › Vision and language
video captioning |
0.9 | 1 | 2025 | AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025 |
Visual content generation and editing › image generation
text-to-image generation |
0.9 | 1 | 2025 | Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis · ICLR 2025 |
Computer vision › Video understanding and tracking
long video understanding |
0.8 | 1 | 2024 | MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024 |
Multimedia analysis and retrieval
video understanding |
0.3 | 1 | 2025 | AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems
long-context memory |
0.2 | 1 | 2024 | MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.5token merging · 1.7masked image modeling · 1.7large multimodal model · 1.7vision-question matching · 1.0training-free adaptation · 1.0memory consolidation · 1.0recurrent neural network · 0.9memory mechanism · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MovieChat+: Question-Aware Sparse Memory for Long Video Question AnsweringabstractRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific vision tasks. Yet, existing methods either employ complex spatial-temporal modules or rely heavily on additional perception models to extract temporal features for video understanding, performing well only on short videos. For long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges. Leveraging the hierarchical memory structure of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination, we propose MovieChat within a training-free memory consolidation mechanism to overcome these challenges, which transfers dense frames from short-term memory into sparse tokens in long-term memory by temporally merging adjacent frames. We lift pre-trained large multi-modal models for understanding long videos without additional trainable modules, employing a zero-shot approach. Additionally, in our new version, MovieChat+, we design an enhanced training-free vision-question matching-based memory consolidation mechanism to better anchor predictions to relevant visual content. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1 K benchmark with 1 K long video, 2 K temporal grounding labels, and 14 K manual annotations. Enxin Song, Wenhao Chai, Tian Ye 0001, Jenq-Neng Hwang, Xi Li 0001, Gaoang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai, Xuexiang Wen, Tian Ye 0001, Gaoang Wang |
ICCV | 2 |
| 2025 | Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image SynthesisabstractWe present Meissonic, which elevates non-autoregressive text-to-image Masked Image Modeling (MIM) to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing methods in generating high-quality, high-resolution images. Extensive experiments validate Meissonic’s capabilities, demonstrating its potential as a new standard in text-to-image synthesis. Jinbin Bai, Tian Ye 0001, Wei Chow, Enxin Song, Xiangtai Li, Zhen Dong 0003, Lei Zhu 0003, Shuicheng Yan |
ICLR | 4 |
| 2025 | AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkabstractVideo detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality. Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Saining Xie, Christopher D. Manning |
ICLR | 2 |
| 2024 | MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingabstractRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat. Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang |
CVPR | 1 |