Junkun Jiang

dblp:247/1318 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-7478-001XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control
abstract
Traditional text-to-motion frameworks often lack precise control, and existing approaches based on joint keyframe locations provide only positional guidance, making it challenging and unintuitive to specify body part orientations and motion timing. To address these limitations, we introduce the Salient Orientation Symbolic (SOS) script, a programmable symbolic framework for specifying body part orientations and motion timing at keyframes. We further propose an automatic SOS extraction pipeline that employs temporally-constrained agglomerative clustering for frame saliency detection and a Saliency-based Masking Scheme (SMS) to generate sparse, interpretable SOS scripts directly from motion data. Moreover, we present the SOSControl framework, which treats the available orientation symbols in the sparse SOS script as salient and prioritizes satisfying these constraints during motion generation. By incorporating SMS-based data augmentation and gradient-based iterative optimization, the framework enhances alignment with user-specified constraints. Additionally, it employs a ControlNet-based ACTOR-PAE Decoder to ensure smooth and natural motion outputs. Extensive experiments demonstrate that the SOS extraction pipeline generates human-interpretable scripts with symbolic annotations at salient keyframes, while the SOSControl framework outperforms existing baselines in motion quality, controllability, and generalizability with respect to motion timing and body part orientation control.
Ho Yin Au, Junkun Jiang, Jie Chen 0026
AAAI2
2026 Part-Level Semantic Fusion for Sketch-Based 3D Voxel Reconstruction
abstract
Reconstructing 3D shapes from monocular freehand sketches is challenging due to their fragmented structures, varying line thickness, and discontinuity. These characteristics cause ambiguity, making it difficult for existing methods to extract sufficient geometric feature information to distinguish subtle shape variations and internal details of the object contours depicted in the sketches, resulting in poor overall reconstruction quality. To address these challenges, we introduce the Part-level Semantic Fusion (PSFusion) module, which combines local units of image features with global feature guidance to enhance the representation of complex geometric structures and subtle contour variations. This approach reduces ambiguity in complex edges and local details, improving shape preservation and edge contour accuracy. Additionally, we propose the coarse-to-fine 3D Decoder consisting of one 3D convolutional network as a coarse-grained regressor and one Shuffle-UNet-based fine-grained refiner, to capture feature dependencies across spatial and channel dimensions. Shuffle operations facilitate information exchange among sub-features, enhancing structural differentiation and cross-feature dependency modelling. This significantly improves the handling of subtle textures and complex intersection boundaries. Extensive experiments on three public benchmarks show that our method outperforms baseline approaches, as demonstrated by both quantitative and qualitative results.
Fei Wang 0056, Yanlong Pan, Junkun Jiang, Dazhi Jiang, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2025 Efficient Real-Time Fine-Grained Action Recognition over a Progressive and Hierarchical Classification Framework
abstract
Real-Time fine-grained action recognition (AR) presents significant challenges in resource-constrained environments with strict accuracy requirements. This paper proposes an efficient real-time AR system that utilizes a progressive hierarchical classification framework to achieve high accuracy while minimizing computational demands. The system utilizes the YOLO model for initial single-frame classification, enabling precise identification of alarming actions with a high recall rate to facilitate timely alerts. Subsequently, a second-tier recognizer that relies on spatiotemporal features is applied for fine-grained AR of identified alarming actions. To enhance recognition accuracy, we introduce a hierarchical classification model where actions are grouped based on semantic and kinematic similarity, followed by further classification within each group. Additionally, we implement a multi-threaded scheduling pipeline that ensures prompt alarms with reasonable loading time for precise AR. Experimental results demonstrate that our system effectively balances computational efficiency with recognition accuracy, making it suitable for real-time deployment in resource-constrained settings.
Shuwen Niu, Junkun Jiang, Jie Chen 0026
ISCAS2
2025 Deep Compositional Phase Diffusion for Long Motion Sequence Generation
abstract
Recent research on motion generation has shown significant progress in generating semantically aligned motion with singular semantics. However, when employing these models to create composite sequences containing multiple semantically generated motion clips, they often struggle to preserve the continuity of motion dynamics at the transition boundaries between clips, resulting in awkward transitions and abrupt artifacts. To address these challenges, we present Compositional Phase Diffusion, which leverages the Semantic Phase Diffusion Module (SPDM) and Transitional Phase Diffusion Module (TPDM) to progressively incorporate semantic guidance and phase details from adjacent motion clips into the diffusion process. Specifically, SPDM and TPDM operate within the latent motion frequency domain established by the pre-trained Action-Centric Motion Phase Autoencoder (ACT-PAE). This allows them to learn semantically important and transition-aware phase information from variable-length motion clips during training. Experimental results demonstrate the competitive performance of our proposed framework in generating compositional motion sequences that align semantically with the input conditions, while preserving phase transitional continuity between preceding and succeeding motion clips. Additionally, motion inbetweening task is made possible by keeping the phase parameter of the input motion sequences fixed throughout the diffusion process, showcasing the potential for extending the proposed framework to accommodate various application scenarios. Codes are available at https://github.com/asdryau/TransPhase.
Ho Yin Au, Jie Chen 0026, Junkun Jiang, Jingyu Xiang
NeurIPS3
2025 Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint Cloud
abstract
Multi-person motion capture over sparse angular observations is a challenging problem under interference from both self- and mutual-occlusions. Existing works produce accurate 2D joint detection, however, when these are triangulated and lifted into 3D, available solutions all struggle in selecting the most accurate candidates and associating them to the correct joint type and target identity. As such, in order to fully utilize all accurate 2D joint location information, we propose to independently triangulate between all same-typed 2D joints from all camera views regardless of their target ID, forming the Joint Cloud. Joint Cloud consist of both valid joints lifted from the same joint type and target ID, as well as falsely constructed ones that are from different 2D sources. These redundant and inaccurate candidates are processed over the proposed Joint Cloud Selection and Aggregation Transformer (JCSAT) involving three cascaded encoders which deeply explore the trajectile, skeletal structural, and view-dependent correlations among all 3D point candidates in the cross-embedding space. An Optimal Token Attention Path (OTAP) module is proposed which subsequently selects and aggregates informative features from these redundant observations for the final prediction of human motion. To demonstrate the effectiveness of JCSAT, we build and publish a new multi-person motion capture dataset BUMocap-X with complex interactions and severe occlusions. Comprehensive experiments over the newly presented as well as benchmark datasets validate the effectiveness of the proposed framework, which outperforms all existing state-of-the-art methods, especially under challenging occlusion scenarios.
Junkun Jiang, Jie Chen 0026, Ho Yin Au, Wei Xue 0002, Yike Guo
IEEE Trans. Vis. Comput. Graph.1
2024 Exploring Latent Cross-Channel Embedding for Accurate 3d Human Pose Reconstruction in a Diffusion Framework
abstract
Monocular 3D human pose estimation poses significant challenges due to the inherent depth ambiguities that arise during the reprojection process from 2D to 3D. Conventional approaches that rely on estimating an over-fit projection matrix struggle to effectively address these challenges and often result in noisy outputs. Recent advancements in diffusion models have shown promise in incorporating structural priors to address reprojection ambiguities. However, there is still ample room for improvement as these methods often overlook the exploration of correlation between the 2D and 3D jointlevel features. In this study, we propose a novel cross-channel embedding framework that aims to fully explore the correlation between joint-level features of 3D coordinates and their 2D projections. In addition, we introduce a context guidance mechanism to facilitate the propagation of joint graph attention across latent channels during the iterative diffusion process. To evaluate the effectiveness of our proposed method, we conduct experiments on two benchmark datasets, namely Human3.6M and MPI-INF-3DHP. Our results demonstrate a significant improvement in terms of reconstruction accuracy compared to state-of-the-art methods. The code for our method will be made available online for further reference.
Junkun Jiang, Jie Chen 0026
ICASSP1
2024 Motion Part-Level Interpolation and Manipulation over Automatic Symbolic Labanotation Annotation
abstract
Motion sequencing is a crucial process in creating smooth and natural animations by arranging individual motion sequences based on desired action scripts. Existing methods either rely on carefully engineered key-frame libraries or implicitly encoded latent phase manifolds for sequential interpolation and manipulation. However, ensuring smooth and natural transitions becomes challenging when dealing with complex and diverse actions, and the manipulation flexibility is limited to the frame level. In this study, we introduce a novel motion sequencing framework centered around Labanotation. The framework leverages automatically annotated Labanotation for explicit representation of motion elements to the body-part level. The proposed Laban Masked Autoencoder (LBN-MAE) is able to directly complete, interpolate and translate Laban symbols into natural 3D trajectories. Our framework offers a compact and descriptive representation of motion, enabling precise motion control and reediting. Comparative evaluations against both conventional and state-of-the-art learning-based methods validate the effectiveness of our proposed framework.
Junkun Jiang, Ho Yin Au, Jie Chen 0026, Jingyu Xiang
IJCNN1
2022 ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic Graph
abstract
To generate dance that temporally and aesthetically matches the music is a challenging problem, as the following factors need to be considered. First, the aesthetic styles and messages conveyed by the motion and music should be consistent. Second, the beats of the generated motion should be locally aligned to the musical features. And finally, basic choreomusical rules should be observed, and the motion generated should be diverse. To address these challenges, we propose ChoreoGraph, which choreographs high-quality dance motion for a given piece of music over a Dynamic Graph. A data-driven learning strategy is proposed to evaluate the aesthetic style and rhythmic connections between music and motion in a progressively learned cross-modality embedding space. The motion sequences will be beats-aligned based on the music segments and then incorporated as nodes of a Dynamic Motion Graph. Compatibility factors such as the style and tempo consistency, motion context connection, action completeness, and transition smoothness are comprehensively evaluated to determine the node transition in the graph. We demonstrate that our repertoire-based framework can generate motions with aesthetic consistency and robustly extensible in diversity. Both quantitative and qualitative experiment results show that our proposed model outperforms other baseline models.
Ho Yin Au, Jie Chen 0026, Junkun Jiang, Yike Guo
ACM Multimedia3
2022 A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token Completion
abstract
Multi-person motion capture can be challenging due to ambiguities caused by severe occlusion, fast body movement, and complex interactions. Existing frameworks build on 2D pose estimations and triangulate to 3D coordinates via reasoning the appearance, trajectory, and geometric consistencies among multi-camera observations. However, 2D joint detection is usually incomplete and with wrong identity assignments due to limited observation angle, which leads to noisy 3D triangulation results. To overcome this issue, we propose to explore the short-range autoregressive characteristics of skeletal motion using transformer. First, we propose an adaptive, identity-aware triangulation module to reconstruct 3D joints and identify the missing joints for each identity. To generate complete 3D skeletal motion, we then propose a Dual-Masked Auto-Encoder (D-MAE) which encodes the joint status with both skeletal-structural and temporal position encoding for trajectory completion. D-MAE's flexible masking and encoding mechanism enable arbitrary skeleton definitions to be conveniently deployed under the same framework. In order to demonstrate the proposed model's capability in dealing with severe data loss scenarios, we contribute a high-accuracy and challenging motion capture dataset of multi-person interactions with severe occlusion. Evaluations on both benchmark and our new dataset demonstrate the efficiency of our proposed model, as well as its advantage against the other state-of-the-art methods.
Junkun Jiang, Jie Chen 0026, Yike Guo
ACM Multimedia1
2019 SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional Networks
abstract
Parsing sketches via semantic segmentation is attractive but challenging, because (i) free-hand drawings are abstract with large variances in depicting objects due to different drawing styles and skills; (ii) distorting lines drawn on the touchpad make sketches more difficult to be recognized; (iii) the high-performance image segmentation via deep learning technologies needs enormous annotated sketch datasets during the training stage. In this paper, we propose a Sketch-target deep FCN Segmentation Network(SFSegNet) for automatic free-hand sketch segmentation, labeling each sketch in a single object with multiple parts. SFSegNet has an end-to-end network process between the input sketches and the segmentation results, composed of 2 parts: (i) a modified deep Fully Convolutional Network(FCN) using a reweighting strategy to ignore background pixels and classify which part each pixel belongs to; (ii) affine transform encoders that attempt to canonicalize the shaking strokes. We train our network with the dataset that consists of 10,000 annotated sketches, to find an extensively applicable model to segment stokes semantically in one ground truth. Extensive experiments are carried out and segmentation results show that our method outperforms other state-of-the-art networks.
Junkun Jiang, Ruomei Wang 0001, Shujin Lin, Fei Wang 0056
IJCNN1