Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ho Yin Au

dblp:324/4684 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-5588-4243ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 97% Multimedia analysis and retrieval · 3%
Artificial intelligence
4 papers
3D vision · 52% Generative modeling · 40% Video understanding and tracking · 8%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.222026
Deep Compositional Phase Diffusion for Long Motion Sequence Generation · NeurIPS 2025
SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control · AAAI 2026
Computer animation and physical simulation › motion synthesis
human motion synthesis
1.012026
SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control · AAAI 2026
Computer animation and physical simulation
motion control
1.012026
SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control · AAAI 2026
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation
1.012026
SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control · AAAI 2026
Computer vision › 3D vision › motion capture
multi-person motion capture
0.912025
Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint Cloud · IEEE Trans. Vis. Comput. Graph. 2025
Computer vision › 3D vision › multi-view geometry › triangulation
multi-view triangulation
0.912025
Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint Cloud · IEEE Trans. Vis. Comput. Graph. 2025
Computer animation and physical simulation › motion synthesis › motion composition
compositional motion generation
0.912025
Deep Compositional Phase Diffusion for Long Motion Sequence Generation · NeurIPS 2025
Computer animation and physical simulation
motion synthesis
0.912025
Deep Compositional Phase Diffusion for Long Motion Sequence Generation · NeurIPS 2025
Computer animation and physical simulation
music-driven dance generation
0.612022
ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic Graph · ACM Multimedia 2022
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.312025
Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint Cloud · IEEE Trans. Vis. Comput. Graph. 2025
Machine learning › Generative modeling
motion generation
0.212022
ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic Graph · ACM Multimedia 2022
Multimedia analysis and retrieval › audio-visual learning
music-dance alignment
0.212022
ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic Graph · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

gradient-based optimization · 2.0data augmentation · 2.0controlnet · 2.0agglomerative clustering · 2.0structural SVM · 1.1dynamic motion graph · 1.1cross-modality embedding · 1.1transformer · 0.9attention mechanism · 0.9
YearPublicationVenuePosition
2026 SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control
abstract
Traditional text-to-motion frameworks often lack precise control, and existing approaches based on joint keyframe locations provide only positional guidance, making it challenging and unintuitive to specify body part orientations and motion timing. To address these limitations, we introduce the Salient Orientation Symbolic (SOS) script, a programmable symbolic framework for specifying body part orientations and motion timing at keyframes. We further propose an automatic SOS extraction pipeline that employs temporally-constrained agglomerative clustering for frame saliency detection and a Saliency-based Masking Scheme (SMS) to generate sparse, interpretable SOS scripts directly from motion data. Moreover, we present the SOSControl framework, which treats the available orientation symbols in the sparse SOS script as salient and prioritizes satisfying these constraints during motion generation. By incorporating SMS-based data augmentation and gradient-based iterative optimization, the framework enhances alignment with user-specified constraints. Additionally, it employs a ControlNet-based ACTOR-PAE Decoder to ensure smooth and natural motion outputs. Extensive experiments demonstrate that the SOS extraction pipeline generates human-interpretable scripts with symbolic annotations at salient keyframes, while the SOSControl framework outperforms existing baselines in motion quality, controllability, and generalizability with respect to motion timing and body part orientation control.
Ho Yin Au, Junkun Jiang, Jie Chen 0026
AAAI1
2025 Deep Compositional Phase Diffusion for Long Motion Sequence Generation
abstract
Recent research on motion generation has shown significant progress in generating semantically aligned motion with singular semantics. However, when employing these models to create composite sequences containing multiple semantically generated motion clips, they often struggle to preserve the continuity of motion dynamics at the transition boundaries between clips, resulting in awkward transitions and abrupt artifacts. To address these challenges, we present Compositional Phase Diffusion, which leverages the Semantic Phase Diffusion Module (SPDM) and Transitional Phase Diffusion Module (TPDM) to progressively incorporate semantic guidance and phase details from adjacent motion clips into the diffusion process. Specifically, SPDM and TPDM operate within the latent motion frequency domain established by the pre-trained Action-Centric Motion Phase Autoencoder (ACT-PAE). This allows them to learn semantically important and transition-aware phase information from variable-length motion clips during training. Experimental results demonstrate the competitive performance of our proposed framework in generating compositional motion sequences that align semantically with the input conditions, while preserving phase transitional continuity between preceding and succeeding motion clips. Additionally, motion inbetweening task is made possible by keeping the phase parameter of the input motion sequences fixed throughout the diffusion process, showcasing the potential for extending the proposed framework to accommodate various application scenarios. Codes are available at https://github.com/asdryau/TransPhase.
Ho Yin Au, Jie Chen 0026, Junkun Jiang, Jingyu Xiang
NeurIPS1
2025 Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint Cloud
abstract
Multi-person motion capture over sparse angular observations is a challenging problem under interference from both self- and mutual-occlusions. Existing works produce accurate 2D joint detection, however, when these are triangulated and lifted into 3D, available solutions all struggle in selecting the most accurate candidates and associating them to the correct joint type and target identity. As such, in order to fully utilize all accurate 2D joint location information, we propose to independently triangulate between all same-typed 2D joints from all camera views regardless of their target ID, forming the Joint Cloud. Joint Cloud consist of both valid joints lifted from the same joint type and target ID, as well as falsely constructed ones that are from different 2D sources. These redundant and inaccurate candidates are processed over the proposed Joint Cloud Selection and Aggregation Transformer (JCSAT) involving three cascaded encoders which deeply explore the trajectile, skeletal structural, and view-dependent correlations among all 3D point candidates in the cross-embedding space. An Optimal Token Attention Path (OTAP) module is proposed which subsequently selects and aggregates informative features from these redundant observations for the final prediction of human motion. To demonstrate the effectiveness of JCSAT, we build and publish a new multi-person motion capture dataset BUMocap-X with complex interactions and severe occlusions. Comprehensive experiments over the newly presented as well as benchmark datasets validate the effectiveness of the proposed framework, which outperforms all existing state-of-the-art methods, especially under challenging occlusion scenarios.
Junkun Jiang, Jie Chen 0026, Ho Yin Au, Wei Xue 0002, Yike Guo
IEEE Trans. Vis. Comput. Graph.3
2024 Motion Part-Level Interpolation and Manipulation over Automatic Symbolic Labanotation Annotation
abstract
Motion sequencing is a crucial process in creating smooth and natural animations by arranging individual motion sequences based on desired action scripts. Existing methods either rely on carefully engineered key-frame libraries or implicitly encoded latent phase manifolds for sequential interpolation and manipulation. However, ensuring smooth and natural transitions becomes challenging when dealing with complex and diverse actions, and the manipulation flexibility is limited to the frame level. In this study, we introduce a novel motion sequencing framework centered around Labanotation. The framework leverages automatically annotated Labanotation for explicit representation of motion elements to the body-part level. The proposed Laban Masked Autoencoder (LBN-MAE) is able to directly complete, interpolate and translate Laban symbols into natural 3D trajectories. Our framework offers a compact and descriptive representation of motion, enabling precise motion control and reediting. Comparative evaluations against both conventional and state-of-the-art learning-based methods validate the effectiveness of our proposed framework.
Junkun Jiang, Ho Yin Au, Jie Chen 0026, Jingyu Xiang
IJCNN2
2022 ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic Graph
abstract
To generate dance that temporally and aesthetically matches the music is a challenging problem, as the following factors need to be considered. First, the aesthetic styles and messages conveyed by the motion and music should be consistent. Second, the beats of the generated motion should be locally aligned to the musical features. And finally, basic choreomusical rules should be observed, and the motion generated should be diverse. To address these challenges, we propose ChoreoGraph, which choreographs high-quality dance motion for a given piece of music over a Dynamic Graph. A data-driven learning strategy is proposed to evaluate the aesthetic style and rhythmic connections between music and motion in a progressively learned cross-modality embedding space. The motion sequences will be beats-aligned based on the music segments and then incorporated as nodes of a Dynamic Motion Graph. Compatibility factors such as the style and tempo consistency, motion context connection, action completeness, and transition smoothness are comprehensively evaluated to determine the node transition in the graph. We demonstrate that our repertoire-based framework can generate motions with aesthetic consistency and robustly extensible in diversity. Both quantitative and qualitative experiment results show that our proposed model outperforms other baseline models.
Ho Yin Au, Jie Chen 0026, Junkun Jiang, Yike Guo
ACM Multimedia1