EDBT 2026 Demo / reviewers in the wild / expert
Songhao Han
dblp:330/4484
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 32% Language models and text generation · 24% Motion planning and robot control · 21% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 67% Multimedia analysis and retrieval · 33% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Computer vision › Video understanding and tracking
frame selection |
0.9 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning |
0.9 | 1 | 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025 |
Robotics › Motion planning and robot control › motion planning › manipulation planning
long-horizon manipulation |
0.9 | 1 | 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought |
0.9 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Robotics › Motion planning and robot control
robot planning |
0.9 | 1 | 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025 |
Computer vision › Video understanding and tracking
video question answering |
0.9 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Computer vision › Video understanding and tracking › deep video understanding
video reasoning |
0.9 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation · NeurIPS 2025 |
Audio and music processing
music generation |
0.7 | 1 | 2023 | Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023 |
Multimedia analysis and retrieval › audio-visual analysis
video-music correspondence |
0.7 | 1 | 2023 | Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023 |
Audio and music processing › music generation
video-to-music generation |
0.7 | 1 | 2023 | Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.3 | 1 | 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection · CVPR 2025 |
Machine learning › Generative modeling
multimodal generation |
0.2 | 1 | 2023 | Video Background Music Generation: Dataset, Method and Evaluation · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
retrieval-based evaluation · 1.3music prior modeling · 1.3vision-language-action controller · 0.9vision-language model · 0.9semantic-aware redundancy reduction · 0.9hierarchical framework · 0.9GPT-4o · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame SelectionabstractThe advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual annotations with insufficient granularity or automatic construction methods with redundant frame-by-frame analysis, limiting their scalability and effectiveness for complex reasoning. To address these challenges, we introduce VideoEspresso, a novel dataset that features VideoQA pairs preserving essential spatial details and temporal coherence, along with multimodal annotations of intermediate reasoning steps. Our construction pipeline employs a semantic-aware method to reduce redundancy, followed by generating QA pairs using GPT-4o. We further develop video Chain-of-Thought (CoT) annotations to enrich reasoning processes, guiding GPT-4o in extracting logical relationships from QA pairs and video content. To exploit the potential of high-quality VideoQA pairs, we propose a Hybrid LVLMs Collaboration framework, featuring a Frame Selector and a two-stage instruction fine-tuned reasoning LVLM. This framework adaptively selects core frames and performs CoT reasoning using multimodal evidence. Evaluated on our proposed benchmark with 14 tasks against 9 popular LVLMs, our method outperforms existing baselines on most tasks, demonstrating superior video reasoning capabilities. Our code and dataset have been released at: https://github.com/hshjerry/VideoEspresso Songhao Han, Wei Huang 0042, Hairong Shi, Le Zhuo, Xiu Su, Xiaojuan Qi 0001, Yue Liao, Si Liu 0001 |
CVPR | 1 |
| 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation EvaluationabstractRecent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs’ strengths in semantic reasoning and long-horizon planning. These System 2 capabilities—characterized by deliberative, goal-directed thinking—remain underexplored due to the limited temporal scale and structural complexity of current benchmarks. To address this gap, we introduce RoboCerebra, a benchmark for evaluating high-level reasoning in long-horizon robotic manipulation. RoboCerebra includes: (1) a large-scale simulation dataset with extended task horizons and diverse subtask sequences in household environments; (2) a hierarchical framework combining a high-level VLM planner with a low-level vision-language-action (VLA) controller; and (3) an evaluation protocol targeting planning, reflection, and memory through structured System 1–System 2 interaction. The dataset is constructed via a top-down pipeline, where GPT generates task instructions and decomposes them into subtask sequences. Human operators execute the subtasks in simulation, yielding high-quality trajectories with dynamic object variations. Compared to prior benchmarks, RoboCerebra features significantly longer action sequences and denser annotations. We further benchmark state-of-the-art VLMs as System 2 modules and analyze their performance across key cognitive dimensions, advancing the development of more capable and generalizable robotic planners. Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang 0004, Chen Gao 0005, Shuicheng Yan, Si Liu 0001 |
NeurIPS | 1 |
| 2024 | Mask-Enhanced Segment Anything Model for Tumor Lesion Semantic Segmentation
Hairong Shi, Songhao Han, Shaofei Huang 0001, Yue Liao, Guanbin Li, Xiangxing Kong, Xiaomu Wang, Si Liu 0001 |
MICCAI (8) | 2 |
| 2023 | Video Background Music Generation: Dataset, Method and EvaluationabstractMusic is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it requires music-video datasets, efficient architectures for video-to-music generation, and reasonable metrics, none of which currently exist. To close this gap, we introduce a complete recipe including dataset, benchmark model, and evaluation metric for video background music generation. We present SymMV, a video and symbolic music dataset with various musical annotations. To the best of our knowledge, it is the first video-music dataset with rich musical annotations. We also propose a benchmark video background music generation framework named V-MusProd, which utilizes music priors of chords, melody, and accompaniment along with video-music relations of semantic, color, and motion features. To address the lack of objective metrics for video-music correspondence, we design a retrieval-based metric VMCP built upon a powerful video-music representation learning model. Experiments show that with our dataset, V-MusProd outperforms the state-of-the-art method in both music quality and correspondence with videos. We believe our dataset, benchmark model, and evaluation metric will boost the development of video background music generation. Our dataset and code are available at https://github.com/zhuole1025/SymMV. Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang 0002, Si Liu 0001 |
ICCV | 7 |