Sen Jia 0003

dblp:35/3232-3 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-8570-2172ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 50% Vision and language · 26% 3D vision · 12%
Computer graphics and multimedia
1 paper
Rendering · 77% Visualization and visual analytics · 23%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Multiple Human Motion Understanding · AAAI 2026
Computer vision › 3D vision › 3d generation
3d scene generation
0.912025
Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025
Computer vision › Video understanding and tracking › human action analysis › action understanding
action counting
0.912025
MoCount: Motion-Based Repetitive Action Counting · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation
0.912025
Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction tuning
0.912025
Human Motion Instruction Tuning · CVPR 2025
Computer vision › Video understanding and tracking › human action analysis › action understanding › action counting
repetitive action counting
0.912025
MoCount: Motion-Based Repetitive Action Counting · ACM Multimedia 2025
Rendering
scene graph
0.912025
Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.622026
Multiple Human Motion Understanding · AAAI 2026
Human Motion Instruction Tuning · CVPR 2025
Computer vision › Video understanding and tracking › video analytics › behavior analysis
human behavior analysis
0.312025
Human Motion Instruction Tuning · CVPR 2025
Visualization and visual analytics
spatial ability
0.312025
Graph Canvas for Controllable 3D Scene Generation · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

instruction tuning · 1.9training-free generation · 1.7graph representation · 1.7social-temporal learner · 1.0sparse spatial-temporal module · 0.9multimodal large language model · 0.9motion encoder · 0.93d motion estimation · 0.9
YearPublicationVenuePosition
2026 Multiple Human Motion Understanding
abstract
We introduce LLaMMo (Large Language and Multi-Person Motion Assistant), the first instruction-tuning multimodal framework tailored for multi-human motion analysis. LLaMMo incorporates a novel human-centric and social-temporal learner that models and fuses both intra-person dynamics and inter-person dependencies, yielding robust, context-aware representations of complex group behaviors while maintaining low computational overhead. To support LLaMMo, we construct LLaVerse, a large-scale dataset with fine-grained manual annotations covering diverse multi-person activities spanning daily social interaction and professional team sports. Built on top of LLaVerse, we also propose LLaMI-Bench, a dedicated benchmark for evaluating multi-human behavior understanding across motion and video modalities. Extensive experiments demonstrate that LLaMMo consistently outperforms baselines in understanding multi-person interactions under low-latency settings, with notable gains in both social and sport-specific contexts.
Lei Li 0050, Sen Jia 0003, Jenq-Neng Hwang
AAAI2
2025 Human Motion Instruction Tuning
abstract
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction.
Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang
CVPR2
2025 Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
abstract
Recent advances in generative modeling enable neural networks to generate weights without relying on gradient-based optimization. However, current methods are limited by issues of over-coupling and long-horizon. The former tightly binds weight generation with task-specific objectives, thereby limiting the flexibility of the learned optimizer. The latter leads to inefficiency and low accuracy during inference, caused by the lack of local constraints. In this paper, we propose Lo-Hp, a decoupled two-stage weight generation framework that enhances flexibility through learning various optimization policies. It adopts a hybrid-policy sub-trajectory balance objective, which integrates on-policy and off-policy learning to capture local optimization policies. Theoretically, we demonstrate that learning solely local optimization policies can address the long-horizon issue while enhancing the generation of global optimal weights. In addition, we validate Lo-Hp’s superior accuracy and inference efficiency in tasks that require frequent weight updates, such as transfer learning, few-shot learning, domain generalization, and large language model adaptation.
Yunchuan Guan, Yu Liu 0040, Ke Zhou 0001, Sen Jia 0003, Zhiqi Shen 0001, Tao Chen 0030, Jenq-Neng Hwang, Lei Li 0050
ECAI5
2025 MoCount: Motion-Based Repetitive Action Counting
abstract
Existing action counting methods typically rely on pixel-based changes within videos, leading to high computational redundancy and low accuracy due to the limited spatial sensitivity. To address these challenges, we introduce MoCount, the first framework that leverages 3D motion representations for counting tasks. MoCount significantly reduces computational overhead and improves counting accuracy, benefiting from the simplicity of motion representation and strong spatial sensitivity. Specifically, we utilize a motion estimator to convert video subjects into 3D motion data. A motion encoder, combined with a Sparse Spatial-Temporal module, is then applied to extract robust human body representations, yielding precise counting results. Extensive experiments on the RepCount and UCFRep datasets show that MoCount achieves state-of-the-art performance, reducing inference latency by approximately 2-3 times compared to existing video counting models. These advantages position MoCount as a leading solution for real-world action counting applications.
Ruocheng Gu, Sen Jia 0003, Yule Ma, Jinqin Zhong, Jenq-Neng Hwang, Lei Li 0050
ACM Multimedia2
2025 Graph Canvas for Controllable 3D Scene Generation
abstract
Spatial intelligence is fundamental to AI systems that interact with the physical world, particularly in 3D scene generation and spatial comprehension. Current layout generation in 3D scene synthesis remains highly complex, often constrained by predefined datasets and limited dynamic adaptation to changing spatial relationships. In this paper, we propose GraphCanvas3D, a flexible, query-driven framework for controllable 3D scene generation. Unlike traditional methods that require retraining and predefined input masks for modifications, GraphCanvas3D provides a training-free solution supporting the generation of diverse scenes-both indoor and outdoor-through free manipulation of objects and scene elements. Our framework employs hierarchical, graph-driven scene descriptions, representing spatial elements as graph nodes and establishing coherent relationships among objects in 3D environments. The decoupled object representation enables flexible, on-the-fly scene adjustments and dynamic, customizable scene creation. Experimental results and user studies demonstrate that GraphCanvas3D improves usability, adaptability, and generalization across various 3D scene generation tasks, offering a powerful tool for scalable and diverse scene synthesis.
Sen Jia 0003, Jingzhe Shi, Can Jin, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050
ACM Multimedia3