Kai Ruan

dblp:160/6087 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%
Artificial intelligence
1 paper
Generative modeling · 77% Video understanding and tracking · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation
motion synthesis
1.922026
X-MoGen: Unified Motion Generation Across Humans and Animals · AAAI 2026
AniMo: Species-Aware Model for Text-Driven Animal Motion Generation · CVPR 2025
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation
1.012026
X-MoGen: Unified Motion Generation Across Humans and Animals · AAAI 2026
Machine learning › Generative modeling › motion generation
text-driven motion generation
0.912025
AniMo: Species-Aware Model for Text-Driven Animal Motion Generation · CVPR 2025
Computer vision › Video understanding and tracking › motion analysis
motion capture analysis
0.312025
AniMo: Species-Aware Model for Text-Driven Animal Motion Generation · CVPR 2025

Methods — techniques the papers use, named apart from their topics

spatiotemporal encoding · 1.7motion tokenization · 1.7masked modeling · 1.7masked motion modeling · 1.0graph variational autoencoder · 1.0
YearPublicationVenuePosition
2026 X-MoGen: Unified Motion Generation Across Humans and Animals
abstract
Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representation and improved generalization. However, morphological differences across species remain a key challenge, often compromising motion plausibility. To address this, we propose X-MoGen, the first unified framework for cross-species text-driven motion generation covering both humans and animals. X-MoGen adopts a two-stage architecture. First, a conditional graph variational autoencoder learns canonical T-pose priors, while an autoencoder encodes motion into a shared latent space regularized by morphological loss. In the second stage, we perform masked motion modeling to generate motion embeddings conditioned on textual descriptions. During training, a morphological consistency module is employed to promote skeletal plausibility across species. To support unified modeling, we construct UniMo4D, a large-scale dataset of 115 species and 119k motion sequences, which integrates human and animal motions under a shared skeletal topology for joint training. Extensive experiments on UniMo4D demonstrate that X-MoGen outperforms state-of-the-art methods on both seen and unseen species.
Kai Ruan, Liyang Qian, Guo Zhi Zhi, Gaoang Wang
AAAI2
2025 AniMo: Species-Aware Model for Text-Driven Animal Motion Generation
abstract
Text-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal ecology, and biomechanics. Animal motion modeling presents unique challenges due to species diversity, varied morphological structures, and different behavioral patterns in response to similar textual descriptions. To address these challenges, we propose AniMo for text-driven animal motion generation. AniMo consists of two stages: motion tokenization and text-to-motion generation. In the motion tokenization stage, we encode motions using a joint-aware spatiotemporal encoder with species-aware feature modulation, enabling the model to adapt to diverse skeletal structures across species. In the text-to-motion generation stage, we employ masked modeling to jointly learn the mapping from textual descriptions to motion tokens. Additionally, we introduce AniMo4D, a large-scale dataset containing 78,149 motion sequences and 185,435 textual descriptions across 114 animal species. Experimental results show that AniMo achieves superior performance on both the AniMo4D and AnimalML3D datasets, effectively capturing diverse morphological structures and behavioral patterns across animal species.
Kai Ruan, Gaoang Wang
CVPR2
2025 Few-Shot-Learning-Like Neural Dynamics for Time-Dependent Multilinear $\mathcal {M}$-Tensor Equation
abstract
In recent years, many discrete neural dynamics models are presented based on continuous models to solve the multilinear tensor equation. However, these existing discrete models all depend on numerical algorithms, such as Euler difference formula and Taylor-type difference formula, which may suffer from the problem of fixed selections with limited feasible parameters. In this article, a few-shot-learning-like neural dynamics (FLLND) model is constructed to find the solution to the time-dependent multilinear tensor equation (TMTE), which opens a new road in constructing the discrete computing model from its continuous counterpart. Specifically, to keep the consistency and better generalization of the constructed model, a few-shot-learning-like method is leveraged to learn parameters from a small dataset. Then, theoretical analyses are conducted to demonstrate the convergence and robustness of the constructed FLLND model in solving the TMTE problem. Finally, several TMTE examples are provided to illustrate the effectiveness and practicality of the FLLND model.
Kai Ruan, Huanmei Wu, Xin Ma 0008
IEEE Trans. Ind. Informatics2
2017 Improved image segmentation method based on morphological reconstruction
Yanpeng Wu, Xiaoqi Peng, Kai Ruan, Zhikun Hu
Multim. Tools Appl.3