EDBT 2026 Demo / reviewers in the wild / expert
Sen Jia 0003
dblp:35/3232-3
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-8570-2172ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Video understanding and tracking · 50% Vision and language · 26% 3D vision · 12% | |
| Computer graphics and multimedia
1 paper |
Rendering · 77% Visualization and visual analytics · 23% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | Multiple Human Motion Understanding · AAAI 2026 |
Computer vision › 3D vision › 3d generation
3d scene generation |
0.9 | 1 | 2025 | Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking › human action analysis › action understanding
action counting |
0.9 | 1 | 2025 | MoCount: Motion-Based Repetitive Action Counting · ACM Multimedia 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation |
0.9 | 1 | 2025 | Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction tuning |
0.9 | 1 | 2025 | Human Motion Instruction Tuning · CVPR 2025 |
Computer vision › Video understanding and tracking › human action analysis › action understanding › action counting
repetitive action counting |
0.9 | 1 | 2025 | MoCount: Motion-Based Repetitive Action Counting · ACM Multimedia 2025 |
Rendering
scene graph |
0.9 | 1 | 2025 | Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking › motion analysis
human motion analysis |
0.6 | 2 | 2026 | Multiple Human Motion Understanding · AAAI 2026 Human Motion Instruction Tuning · CVPR 2025 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis
human behavior analysis |
0.3 | 1 | 2025 | Human Motion Instruction Tuning · CVPR 2025 |
Visualization and visual analytics
spatial ability |
0.3 | 1 | 2025 | Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 1.9training-free generation · 1.7graph representation · 1.7social-temporal learner · 1.0sparse spatial-temporal module · 0.9multimodal large language model · 0.9motion encoder · 0.93d motion estimation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiple Human Motion UnderstandingabstractWe introduce LLaMMo (Large Language and Multi-Person Motion Assistant), the first instruction-tuning multimodal framework tailored for multi-human motion analysis. LLaMMo incorporates a novel human-centric and social-temporal learner that models and fuses both intra-person dynamics and inter-person dependencies, yielding robust, context-aware representations of complex group behaviors while maintaining low computational overhead. To support LLaMMo, we construct LLaVerse, a large-scale dataset with fine-grained manual annotations covering diverse multi-person activities spanning daily social interaction and professional team sports. Built on top of LLaVerse, we also propose LLaMI-Bench, a dedicated benchmark for evaluating multi-human behavior understanding across motion and video modalities. Extensive experiments demonstrate that LLaMMo consistently outperforms baselines in understanding multi-person interactions under low-latency settings, with notable gains in both social and sport-specific contexts. Lei Li 0050, Sen Jia 0003, Jenq-Neng Hwang |
AAAI | 2 |
| 2025 | Human Motion Instruction TuningabstractThis paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang |
CVPR | 2 |
| 2025 | Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory BalanceabstractRecent advances in generative modeling enable neural networks to generate weights without relying on gradient-based optimization. However, current methods are limited by issues of over-coupling and long-horizon. The former tightly binds weight generation with task-specific objectives, thereby limiting the flexibility of the learned optimizer. The latter leads to inefficiency and low accuracy during inference, caused by the lack of local constraints. In this paper, we propose Lo-Hp, a decoupled two-stage weight generation framework that enhances flexibility through learning various optimization policies. It adopts a hybrid-policy sub-trajectory balance objective, which integrates on-policy and off-policy learning to capture local optimization policies. Theoretically, we demonstrate that learning solely local optimization policies can address the long-horizon issue while enhancing the generation of global optimal weights. In addition, we validate Lo-Hp’s superior accuracy and inference efficiency in tasks that require frequent weight updates, such as transfer learning, few-shot learning, domain generalization, and large language model adaptation. Yunchuan Guan, Yu Liu 0040, Ke Zhou 0001, Sen Jia 0003, Zhiqi Shen 0001, Tao Chen 0030, Jenq-Neng Hwang, Lei Li 0050 |
ECAI | 5 |
| 2025 | MoCount: Motion-Based Repetitive Action CountingabstractExisting action counting methods typically rely on pixel-based changes within videos, leading to high computational redundancy and low accuracy due to the limited spatial sensitivity. To address these challenges, we introduce MoCount, the first framework that leverages 3D motion representations for counting tasks. MoCount significantly reduces computational overhead and improves counting accuracy, benefiting from the simplicity of motion representation and strong spatial sensitivity. Specifically, we utilize a motion estimator to convert video subjects into 3D motion data. A motion encoder, combined with a Sparse Spatial-Temporal module, is then applied to extract robust human body representations, yielding precise counting results. Extensive experiments on the RepCount and UCFRep datasets show that MoCount achieves state-of-the-art performance, reducing inference latency by approximately 2-3 times compared to existing video counting models. These advantages position MoCount as a leading solution for real-world action counting applications. Ruocheng Gu, Sen Jia 0003, Yule Ma, Jinqin Zhong, Jenq-Neng Hwang, Lei Li 0050 |
ACM Multimedia | 2 |
| 2025 | Graph Canvas for Controllable 3D Scene GenerationabstractSpatial intelligence is fundamental to AI systems that interact with the physical world, particularly in 3D scene generation and spatial comprehension. Current layout generation in 3D scene synthesis remains highly complex, often constrained by predefined datasets and limited dynamic adaptation to changing spatial relationships. In this paper, we propose GraphCanvas3D, a flexible, query-driven framework for controllable 3D scene generation. Unlike traditional methods that require retraining and predefined input masks for modifications, GraphCanvas3D provides a training-free solution supporting the generation of diverse scenes-both indoor and outdoor-through free manipulation of objects and scene elements. Our framework employs hierarchical, graph-driven scene descriptions, representing spatial elements as graph nodes and establishing coherent relationships among objects in 3D environments. The decoupled object representation enables flexible, on-the-fly scene adjustments and dynamic, customizable scene creation. Experimental results and user studies demonstrate that GraphCanvas3D improves usability, adaptability, and generalization across various 3D scene generation tasks, offering a powerful tool for scalable and diverse scene synthesis. Sen Jia 0003, Jingzhe Shi, Can Jin, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050 |
ACM Multimedia | 3 |