Yuke Lou

dblp:330/4468 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-5165-6251ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 100%
Artificial intelligence
3 papers
Reinforcement learning · 43% Generative modeling · 28% Video understanding and tracking · 14%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation · ICCV 2025
Computer animation and physical simulation
human-scene interaction
0.912025
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation · ICCV 2025
Computer animation and physical simulation › procedural animation
behavioral animation
0.812024
CBIL: Collective Behavior Imitation Learning for Fish from Real Videos · ACM Trans. Graph. 2024
Machine learning › Reinforcement learning
imitation learning
0.712023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Computer animation and physical simulation › gesture generation
co-speech gesture generation
0.612022
Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings · ACM Trans. Graph. 2022
Computer animation and physical simulation
motion synthesis
0.612022
Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings · ACM Trans. Graph. 2022
Natural language and speech › Language models and text generation › text generation › story generation
script generation
0.312025
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation · ICCV 2025
Computer vision › Video understanding and tracking › video analytics › behavior analysis
animal behavior analysis
0.212024
CBIL: Collective Behavior Imitation Learning for Fish from Real Videos · ACM Trans. Graph. 2024
Computer vision › Video understanding and tracking
human motion prediction
0.212023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023
Robotics › Autonomous driving › trajectory prediction
multi-person motion prediction
0.212023
Social Motion Prediction with Cognitive Hierarchies · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 1.7self-supervised learning · 1.5masked video autoencoder · 1.5adversarial imitation learning · 1.5generative adversarial imitation learning · 0.7cognitive hierarchy framework · 0.7behavioral cloning · 0.7rhythm-based segmentation · 0.6hierarchical neural embedding · 0.6contrastive learning · 0.6
YearPublicationVenuePosition
2025 SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Wenjia Wang 0009, Liang Pan, Zhiyang Dou, Jidong Mei, Zhouyingcheng Liao, Yuke Lou, Yifan Wu 0039, Lei Yang 0045, Jingbo Wang 0003, Taku Komura
ICCV6
2024 CBIL: Collective Behavior Imitation Learning for Fish from Real Videos
abstract
Reproducing realistic collective behaviors presents a captivating yet formidable challenge. Traditional rule-based methods rely on hand-crafted principles, limiting motion diversity and realism in generated collective behaviors. Recent imitation learning methods learn from data but often require ground-truth motion trajectories and struggle with authenticity, especially in high-density groups with erratic movements. In this paper, we present a scalable approach, Collective Behavior Imitation Learning (CBIL), for learning fish schooling behavior directly from videos , without relying on captured motion trajectories. Our method first leverages Video Representation Learning, in which a Masked Video AutoEncoder (MVAE) extracts implicit states from video inputs in a self-supervised manner. The MVAE effectively maps 2D observations to implicit states that are compact and expressive for following the imitation learning stage. Then, we propose a novel adversarial imitation learning method to effectively capture complex movements of the schools of fish, enabling efficient imitation of the distribution of motion patterns measured in the latent space. It also incorporates bio-inspired rewards alongside priors to regularize and stabilize training. Once trained, CBIL can be used for various animation tasks with the learned collective motion priors. We further show its effectiveness across different species. Finally, we demonstrate the application of our system in detecting abnormal fish behavior from in-the-wild videos.
Yifan Wu 0039, Zhiyang Dou, Yuko Ishiwaka, Shun Ogawa, Yuke Lou, Wenping Wang 0001, Lingjie Liu, Taku Komura
ACM Trans. Graph.5
2023 Social Motion Prediction with Cognitive Hierarchies
abstract
Humans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired framework. We present Wusi, a 3D multi-person motion dataset under the context of team sports, which features intense and strategic human interactions and diverse pose distributions. By reformulating the problem from a multi-agent reinforcement learning perspective, we incorporate behavioral cloning and generative adversarial imitation learning to boost learning efficiency and generalization. Furthermore, we take into account the cognitive aspects of the human social action planning process and develop a cognitive hierarchy framework to predict strategic human social interactions. We conduct comprehensive experiments to validate the effectiveness of our proposed dataset and approach.
Wentao Zhu 0004, Jason Qin, Yuke Lou, Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Yizhou Wang 0001
NeurIPS3
2023 Edit-History Vis: An Interactive Visual Exploration and Analysis on Wikipedia Edit History
abstract
We propose Edit-History Vis, a visual analytics system designed to facilitate interactive exploration on Wikipedia edit history at a fine-grained level. The examination of detailed changes in Wikipedia articles is crucial for understanding how authors’ perspectives vary and conflict during the collaborative editing process. However, it is challenging to reveal the details while preserving the heterogeneous attributes of revisions, namely the time, content, and editor. The Edit-History Vis system integrates editor and textual changes of revisions by utilizing a force-directed revision graph that groups revisions based on standpoints. Through this revision graph, users can identify and analyze editing events such as edit wars, vandalism, repair, and normal updates. The effectiveness of the system in analyzing the edit history is validated through a qualitative comparison with prior work and a quantitative rating from a user study.
Yuhan Guo 0004, Qin Han, Yuke Lou, Yiming Wang 0009, Can Liu 0004, Xiaoru Yuan
PacificVis3
2022 Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
abstract
Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to difficulties in mining the clear rhythm and semantics due to the complex yet subtle harmony between speech and gestures. We present a novel co-speech gesture synthesis method that achieves convincing results both on the rhythm and semantics. For the rhythm, our system contains a robust rhythm-based segmentation pipeline to ensure the temporal coherence between the vocalization and gestures explicitly. For the gesture semantics, we devise a mechanism to effectively disentangle both low- and high-level neural embeddings of speech and motion based on linguistic theory. The high-level embedding corresponds to semantics, while the low-level embedding relates to subtle variations. Lastly, we build correspondence between the hierarchical embeddings of the speech and the motion, resulting in rhythm- and semantics-aware gesture synthesis. Evaluations with existing objective metrics, a newly proposed rhythmic metric, and human feedback show that our method outperforms state-of-the-art systems by a clear margin.
Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, Libin Liu 0002
ACM Trans. Graph.3