Songhao Han

dblp:330/4484 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 32% Language models and text generation · 24% Motion planning and robot control · 21%
Computer graphics and multimedia
1 paper
Audio and music processing · 67% Multimedia analysis and retrieval · 33%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Computer vision › Video understanding and tracking
frame selection
0.912025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning
0.912025
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025
Robotics › Motion planning and robot control › motion planning › manipulation planning
long-horizon manipulation
0.912025
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought
0.912025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Robotics › Motion planning and robot control
robot planning
0.912025
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025
Computer vision › Video understanding and tracking
video question answering
0.912025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Computer vision › Video understanding and tracking › deep video understanding
video reasoning
0.912025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025
Audio and music processing
music generation
0.712023
Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023
Multimedia analysis and retrieval › audio-visual analysis
video-music correspondence
0.712023
Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023
Audio and music processing › music generation
video-to-music generation
0.712023
Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023
Natural language and speech › Language models and text generation
instruction tuning
0.312025
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025
Machine learning › Generative modeling
multimodal generation
0.212023
Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023

Methods — techniques the papers use, named apart from their topics

retrieval-based evaluation · 1.3music prior modeling · 1.3vision-language-action controller · 0.9vision-language model · 0.9semantic-aware redundancy reduction · 0.9hierarchical framework · 0.9GPT-4o · 0.9
YearPublicationVenuePosition
2025 VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
abstract
The advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual annotations with insufficient granularity or automatic construction methods with redundant frame-by-frame analysis, limiting their scalability and effectiveness for complex reasoning. To address these challenges, we introduce VideoEspresso, a novel dataset that features VideoQA pairs preserving essential spatial details and temporal coherence, along with multimodal annotations of intermediate reasoning steps. Our construction pipeline employs a semantic-aware method to reduce redundancy, followed by generating QA pairs using GPT-4o. We further develop video Chain-of-Thought (CoT) annotations to enrich reasoning processes, guiding GPT-4o in extracting logical relationships from QA pairs and video content. To exploit the potential of high-quality VideoQA pairs, we propose a Hybrid LVLMs Collaboration framework, featuring a Frame Selector and a two-stage instruction fine-tuned reasoning LVLM. This framework adaptively selects core frames and performs CoT reasoning using multimodal evidence. Evaluated on our proposed benchmark with 14 tasks against 9 popular LVLMs, our method outperforms existing baselines on most tasks, demonstrating superior video reasoning capabilities. Our code and dataset have been released at: https://github.com/hshjerry/VideoEspresso
Songhao Han, Wei Huang 0042, Hairong Shi, Le Zhuo, Xiu Su, Xiaojuan Qi 0001, Yue Liao, Si Liu 0001
CVPR1
2025 RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
abstract
Recent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs’ strengths in semantic reasoning and long-horizon planning. These System 2 capabilities—characterized by deliberative, goal-directed thinking—remain underexplored due to the limited temporal scale and structural complexity of current benchmarks. To address this gap, we introduce RoboCerebra, a benchmark for evaluating high-level reasoning in long-horizon robotic manipulation. RoboCerebra includes: (1) a large-scale simulation dataset with extended task horizons and diverse subtask sequences in household environments; (2) a hierarchical framework combining a high-level VLM planner with a low-level vision-language-action (VLA) controller; and (3) an evaluation protocol targeting planning, reflection, and memory through structured System 1–System 2 interaction. The dataset is constructed via a top-down pipeline, where GPT generates task instructions and decomposes them into subtask sequences. Human operators execute the subtasks in simulation, yielding high-quality trajectories with dynamic object variations. Compared to prior benchmarks, RoboCerebra features significantly longer action sequences and denser annotations. We further benchmark state-of-the-art VLMs as System 2 modules and analyze their performance across key cognitive dimensions, advancing the development of more capable and generalizable robotic planners.
Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang 0004, Chen Gao 0005, Shuicheng Yan, Si Liu 0001
NeurIPS1
2024 Mask-Enhanced Segment Anything Model for Tumor Lesion Semantic Segmentation
Hairong Shi, Songhao Han, Shaofei Huang 0001, Yue Liao, Guanbin Li, Xiangxing Kong, Xiaomu Wang, Si Liu 0001
MICCAI (8)2
2023 Video Background Music Generation: Dataset, Method and Evaluation
abstract
Music is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it requires music-video datasets, efficient architectures for video-to-music generation, and reasonable metrics, none of which currently exist. To close this gap, we introduce a complete recipe including dataset, benchmark model, and evaluation metric for video background music generation. We present SymMV, a video and symbolic music dataset with various musical annotations. To the best of our knowledge, it is the first video-music dataset with rich musical annotations. We also propose a benchmark video background music generation framework named V-MusProd, which utilizes music priors of chords, melody, and accompaniment along with video-music relations of semantic, color, and motion features. To address the lack of objective metrics for video-music correspondence, we design a retrieval-based metric VMCP built upon a powerful video-music representation learning model. Experiments show that with our dataset, V-MusProd outperforms the state-of-the-art method in both music quality and correspondence with videos. We believe our dataset, benchmark model, and evaluation metric will boost the development of video background music generation. Our dataset and code are available at https://github.com/zhuole1025/SymMV.
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang 0002, Si Liu 0001
ICCV7