VLDB 2026 Research / reviewers in the wild / expert
Qiongjie Cui
dblp:232/2538
· DBLP profile ↗
30ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-8078-6706ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing global coherence and hand-level detail: Frequency-decomposed whole-body human motion prediction
Delong Yang, Linda Ma, Qiongjie Cui |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Adaptive Interaction Network for Human Motion Prediction During Human-Robot CollaborationabstractHuman motion prediction during human-robot collaboration is critical for achieving safe and efficient interactions in shared environments. Unlike previous studies that have primarily focused on humans and objects, we focus on the robot-aware human motion prediction task, which explicitly models the influence of robots on human motion. This task presents new challenges, such as capturing human-robot heterogeneity and modeling complex spatial interactions. To address these issues, we develop an adaptive interaction network (AINet) model that consists of two branches: a primary branch that predicts future human motion and an auxiliary branch that estimates robot trajectories. The two branches are jointly optimized and coupled via a Local-Global Spatial Interaction (LGSI) Module, which effectively captures fine-grained and global contextual dependencies between human and robot motion sequences. Then, we introduce an Adaptive Weighted Aggregation (AWA) Module to dynamically fuse motion features using context-dependent weights, thereby increasing adaptability across diverse scenarios. Furthermore, we adopt a coarse-to-fine prediction strategy, in which a coarse pseudo-trajectory is first predicted, followed by the refinement of detailed human poses based on that trajectory. Extensive experiments on three challenging datasets reveal that our proposed method achieves superior performance, validating its effectiveness in human-robot collaboration scenarios. Our code is available at: https://github.com/LyTingHub/AINet. Mengyuan Liu 0001, Yangting Lin, Qiongjie Cui |
IEEE Trans. Image Process. | 3 |
| 2025 | Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesabstractRecent advances in human motion prediction (HMP) have shifted focus from isolated motion data to integrating human-scene correlations. In particular, the latest methods leverage human gaze points, using their spatial coordinates to indicate intent—where a person might move within a 3D environment. Despite promising trajectory results, these methods often produce inaccurate poses by overlooking the semantic implications of gaze, specifically the affordances of observed objects, which indicate the possible interactions. To address this, we propose GAP3DS, an affordance-aware HMP model that utilizes gaze-informed object affordances to improve HMP in complex 3D environments. GAP3DS incorporates a gaze-guided affordance learner to identify relevant objects in the scene and infer their affordances based on human gaze, thus contextualizing future human-object interactions. This affordance information, enriched with visual features and gaze data, conditions the generation of multiple human-object interaction poses, which are subsequently decoded into final motion predictions. Extensive experiments on two datasets demonstrate that GAP3DS outperforms state-of-the-art methods in both trajectory and pose accuracy, producing more physically consistent and contextually grounded predictions. For more details and code, please refer to the project page. Zhenyu Lou, Qiongjie Cui |
CVPR | 5 |
| 2025 | OcSplats: Rendering Occluded Humans with Prior KnowledgeabstractThe task of reconstructing and rendering moving humans from monocular videos, particularly when occlusions are present, is fraught with difficulty due to insufficient visual information. Existing approaches struggle with two primary issues in delivering complete and high-fidelity rendering: the reliance on precise geometry constraints, often failing to account for occluded body parts, and the insufficient observation of unseen body parts, resulting in inconsistent reconstructions. To address these limitations, we introduce OcSplats, a deformable 3D gaussian splatting based method tailored for rendering humans in highly occluded scenarios using prior knowledge. OcSplats designs a human body prior-based geometry completion module to recover the occluded human geometry, ensuring complete reconstruction and rendering. Additionally, OcSplats employs a multi-view diffusion prior to regularize human reconstruction pipeline at novel camera poses beyond those in the occluded monocular video. We evaluate OcSplats on the ZJU-MoCap dataset and challenging OcMotion sequences, and experimental results demonstrate that OcSplats significantly outperforms existing state-of-the-art methods in rendering occluded humans. Jie Zhang 0002, Qiongjie Cui, Xulei Yang, Na Zhao 0004 |
ICME | 2 |
| 2025 | How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic SegmentationabstractLiDAR-based 3D panoptic segmentation often struggles with the inherent sparsity of data from LiDAR sensors, which makes it challenging to accurately recognize distant or small objects. Recently, a few studies have sought to overcome this challenge by integrating LiDAR inputs with camera images, leveraging the rich and dense texture information provided by the latter. While these approaches have shown promising results, they still face challenges, such as misalignment during data augmentation and the reliance on post-processing steps. To address these issues, we propose Image-Assists-LiDAR (IAL), a novel multi-modal 3D panoptic segmentation framework. In IAL, we first introduce a modality-synchronized data augmentation strategy, PieAug, to ensure alignment between LiDAR and image inputs from the start. Next, we adopt a transformer decoder to directly predict panoptic segmentation results. To effectively fuse LiDAR and image features into tokens for the decoder, we design a Geometric-guided Token Fusion (GTF) module. Additionally, we leverage the complementary strengths of each modality as priors for query initialization through a Prior-based Query Generation (PQG) module, enhancing the decoder’s ability to generate accurate instance masks. Our IAL framework achieves state-of-the-art performance compared to previous multi-modal 3D panoptic segmentation methods on two widely used benchmarks. Code and models are publicly available at https://github.com/IMPL-Lab/IAL.git. Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao 0004 |
ICML | 2 |
| 2025 | Toward Physically Stable Motion Generation: A New Paradigm of Human Pose RepresentationabstractIn machine learning, generating realistic human motion is paramount for a range of applications that require lifelike movements. Traditional methods have often overlooked the adherence to physical principles, leading to motion sequences that exhibit unrealistic behaviors such as foot sliding, penetration, and floating. These issues are particularly pronounced in complex tasks like dance choreography, which demand a higher degree of fidelity and realism. To address these challenges, we introduce RF-Rotation, a novel approach to human pose representation that strategically repositions the root joint of the SMPL model to align with both feet, while representing other joints through recursive bone rotations. It not only aligns more closely with the natural dynamics of human movement but also integrates an advanced contact predictor to ascertain the ground contact status of both feet, thereby preventing physically implausible movements on feet. We note that RF-Rotation is compatible with any motion generation tasks, including dance choreography, text-to-motion synthesis, and motion prediction, and can be seamlessly integrated into existing frameworks without modifications. Extensive experiments across three distinct tasks demonstrate the superior performance of RF-Rotation in enhancing the realism and stability of generated motion sequences. This method can significantly reduce foot sliding, floating, and penetration issues, without affecting computational efficiency, underscores its potential to set new standards in human motion generation. Qiongjie Cui, Zhenyu Lou, Zhenbo Song, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Hierarchical Motion-Enhanced Matching Framework for Few-Shot Action RecognitionabstractFew-Shot Action Recognition (FSAR) aims to recognize novel class action with limited annotated training data from the same class. Most FSAR methods subconsciously follow the few-shot image classification solutions by solely focusing on appearance-level matching between support and query videos, such as part-level matching, frame-level matching, and segment-level matching. However, these methods, almost always, have two main limitations: 1) generally ignore the relationship among these part-, frame- and segment-level features and 2) may mismatch the same class actions under fast-term and slow-term dynamics. To this end, we present a novel Hierarchical Motion-enhanced Matching (HM${^{2}}$) framework to hierarchically learn the relation-aware multi-modal features, and jointly promote the multi-modal matching, including appearance-level matching on segments, frames, and parts, as well as the motion-level matching on dynamics. Specifically, we first propose a new Hierarchical Tokenizer (HT) to learn multi-modal features, namely utilizing a hierarchical Transformer to learn appearance-level features, along with a Slow-Fast Aware Motion (SFAM) strategy to learn motion-level features covering fast- and slow-term dynamics. Next, we propose a new Relation-aware Matcher (RM) to match the multi-modal features, by leveraging a Hierarchical Relational Graph Convolutional Network (H-RGCN) to capture the relationship among these appearance-level features. Further, a Dual Sample-to-Class Matching (DSCM) strategy is proposed to measure the bidirectional similarities among appearance- and motion-modal features by sample-to-class matching and class-to-sample matching. Extensive experiments on four golden FSAR datasets demonstrate significant performance improvements of HM${^{2}}$compared with the state-of-the-art methods. Hailiang Gao, Guosen Xie, Rui Yan 0010, Qiongjie Cui, Hongyu Qu, Xiangbo Shu |
IEEE Trans. Multim. | 4 |
| 2025 | MossVLN: Memory-Observation Synergistic System for Continuous Vision-Language NavigationabstractNavigating in continuous environments with vision-language cues presents critical challenges, particularly in the accuracy of waypoint prediction and the quality of navigation decision-making. Traditional methods, which predominantly rely on spatial data from depth images or straightforward RGB-depth integrations, frequently encounter difficulties in environments where waypoints share similar spatial characteristics, leading to erroneous navigational outcomes. Additionally, the capacity for effective navigation decisions is often hindered by the inadequacies of traditional topological maps and the issue of uneven data sampling. In response, this paper introduces a robust memory-observation synergistic vision-language navigation framework to substantially enhance the navigation capabilities of agents operating in continuous environments. We present an advanced observation-driven waypoint predictor that effectively utilizes spatial data and integrates aligned visual and textual cues to significantly improve the accuracy of waypoint predictions within complex real-world scenarios. Additionally, we develop a strategic memory-observation planning approach that leverages memory panoramic environmental data and detailed current observation information, enabling more informed and precise navigation decisions. Our framework sets new performance benchmarks on the VLN-CE dataset, achieving a 60.25% Success Rate (SR) and a 50.89% Path Length Score (SPL) on the R2R-CE dataset's unseen validation splits. Furthermore, when adapted to a discrete environment, our model also shows exceptional performance on the R2R dataset, achieving a 74% SR and a 64% SPL on the unseen validation split. The code is available athttps://github.com/OpenMICG/MossVLN. Ting Yu 0002, Qiongjie Cui, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Expressive Forecasting of 3D Whole-Body Human MotionsabstractHuman motion forecasting, with the goal of estimating future human behavior over a period of time, is a fundamental task in many real-world applications. However, existing works typically concentrate on foretelling the major joints of the human body without considering the delicate movements of the human hands. In practical applications, hand gesture plays an important role in human communication with the real world, and expresses the primary intention of human beings. In this work, we are the first to formulate whole-body human pose forecasting task, which jointly predicts future both body and gesture activities. Correspondingly, we propose a novel Encoding-Alignment-Interaction (EAI) framework that aims to predict both coarse (body joints) and fine-grained (gestures) activities collaboratively, enabling expressive and cross-facilitated forecasting of 3D whole-body human motions. Specifically, our model involves two key constituents: cross-context alignment (XCA) and cross-context interaction (XCI). Considering the heterogeneous information within the whole-body, XCA aims to align the latent features of various human components, while XCI focuses on effectively capturing the context interaction among the human components. We conduct extensive experiments on a newly-introduced large-scale benchmark and achieve state-of-the-art performance. The code is public for research purposes at https://github.com/Dingpx/EAI. Pengxiang Ding, Qiongjie Cui |
AAAI | 2 |
| 2024 | GCNext: Towards the Unity of Graph Convolutions for Human Motion PredictionabstractThe past few years has witnessed the dominance of Graph Convolutional Networks (GCNs) over human motion prediction. Various styles of graph convolutions have been proposed, with each one meticulously designed and incorporated into a carefully-crafted network architecture. This paper breaks the limits of existing knowledge by proposing Universal Graph Convolution (UniGC), a novel graph convolution concept that re-conceptualizes different graph convolutions as its special cases. Leveraging UniGC on network-level, we propose GCNext, a novel GCN-building paradigm that dynamically determines the best-fitting graph convolutions both sample-wise and layer-wise. GCNext offers multiple use cases, including training a new GCN from scratch or refining a preexisting GCN. Experiments on Human3.6M, AMASS, and 3DPW datasets show that, by incorporating unique module-to-network designs, GCNext yields up to 9x lower computational cost than existing GCN methods, on top of achieving state-of-the-art performance. Our code is available at https://github.com/BradleyWang0416/GCNext. Xinshun Wang, Qiongjie Cui, Chen Chen 0001 |
AAAI | 2 |
| 2024 | Multimodal Sense-Informed Forecasting of 3D Human MotionsabstractPredicting future human pose is a fundamental application for machine intelligence, which drives robots to plan their behavior and paths ahead of time to seamlessly accomplish human-robot collaboration in real-world 3D scenarios. Despite encouraging results, existing approaches rarely consider the effects of the external scene on the motion sequence, leading to pronounced artifacts and physical implausibilities in the predictions. To address this limitation, this work introduces a novel multi-modal sense-informed motion prediction approach, which conditions high-fidelity generation on two modal information: external 3D scene, and internal human gaze, and is able to recognize their salience for future human activity. Furthermore, the gaze information is regarded as the human intention, and combined with both motion and scene features, we construct a ternary intention-aware attention to supervise the generation to match where the human wants to reach. Meanwhile, we introduce semantic coherence-aware attention to explicitly distinguish the salient point clouds and the underlying ones, to ensure a reasonable interaction of the generated sequence with the 3D scene. On two real-world benchmarks, the proposed method achieves state-of-the-art performance both in 3D human pose and trajectory prediction. More detailed results are available on the page: https://sites.google.com/view/cvpr2024sif3d. Zhenyu Lou, Qiongjie Cui |
CVPR | 2 |
| 2024 | Forecasting of 3D Whole-Body Human Poses with Grasping ObjectsabstractIn the context of computer vision and human-robot interaction, forecasting 3D human poses is crucial for understanding human behavior and enhancing the predictive capabilities of intelligent systems. While existing methods have made significant progress, they often focus on predicting major body joints, overlooking fine-grained gestures and their interaction with objects. Human hand movements, particularly during object interactions, play a pivotal role and provide more precise expressions of human poses. This work fills this gap and introduces a novel paradigm: forecasting 3D whole-body human poses with a focus on grasping objects. This task involves predicting activities across all joints in the body and hands, encompassing the complexities of internal heterogeneity and external interactivity. To tackle these challenges, we also propose a novel approach: C3HOST, cross-context cross-modal consolidation for 3D whole-body pose forecasting, effectively handles the complexities of internal heterogeneity and external interactivity. C3HOST involves distinct steps, including the heterogeneous content encoding and alignment, and cross-modal feature learning and interaction. These enable us to predict activities across all body and hand joints, ensuring high-precision whole-body human pose prediction, even during object grasping. Extensive experiments on two benchmarks demonstrate that our model significantly enhances the accuracy of whole-body human motion prediction. The project page is available at https://sites.google.com/view/c3host. Haitao Yan, Qiongjie Cui, Jiexin Xie, Shijie Guo |
CVPR | 2 |
| 2024 | Human Motion Forecasting in Dynamic Domain Shifts: A Homeostatic Continual Test-Time Adaptation Framework
Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084 |
ECCV (31) | 1 |
| 2024 | Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion PredictionabstractDiverse human motion prediction (HMP) is a fundamental application in computer vision that has recently attracted considerable interest. Prior methods primarily focus on the stochastic nature of human motion, while neglecting the specific impact of external environment, leading to the pronounced artifacts in prediction when applied to real-world scenarios. To fill this gap, this work introduces a novel task: predicting diverse human motion within real-world 3D scenes. In contrast to prior works, it requires harmonizing the deterministic constraints imposed by the surrounding 3D scenes with the stochastic aspect of human motion. For this purpose, we propose DiMoP3D, a diverse motion prediction framework with 3D scene awareness, which leverages the 3D point cloud and observed sequence to generate diverse and high-fidelity predictions. DiMoP3D is able to comprehend the 3D scene, and determines the probable target objects and their desired interactive pose based on the historical motion. Then, it plans the obstacle-free trajectory towards these interested objects, and generates diverse and physically-consistent future motions. On top of that, DiMoP3D identifies deterministic factors in the scene and integrates them into the stochastic modeling, making the diverse HMP in realistic scenes become a controllable stochastic generation process. On two real-captured benchmarks, DiMoP3D has demonstrated significant improvements over state-of-the-art methods, showcasing its effectiveness in generating diverse and physically-consistent motion predictions within real-world 3D environments. Zhenyu Lou, Qiongjie Cui, Zhenbo Song, Luoming Zhang, Huaxia Li |
NeurIPS | 2 |
| 2023 | Meta-Auxiliary Learning for Adaptive Human Pose PredictionabstractPredicting high-fidelity future human poses, from a historically observed sequence, is crucial for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on external datasets and then directly apply it to all test samples, emerge as the dominant solution to solve this issue. Despite encouraging progress, they remain non-optimal, as the unique properties (e.g., motion style, rhythm) of a specific sequence cannot be adapted. More generally, once encountering out-of-distributions, the predicted poses tend to be unreliable. Motivated by this observation, we propose a novel test-time adaptation framework that leverages two self-supervised auxiliary tasks to help the primary forecasting network adapt to the test sequence. In the testing phase, our model can adjust the model parameters by several gradient updates to improve the generation quality. However, due to catastrophic forgetting, both auxiliary tasks typically have a low ability to automatically present the desired positive incentives for the final prediction performance. For this reason, we also propose a meta-auxiliary learning scheme for better adaptation. Extensive experiments show that the proposed approach achieves higher accuracy and more realistic visualization. Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084 |
AAAI | 1 |
| 2023 | Test-time Personalizable Forecasting of 3D Human PosesabstractCurrent motion forecasting approaches typically train a deep end-to-end model from the source domain data, and then apply it directly to target subjects. Despite promising results, they remain non-optimal, due to privacy considerations, the test person and his/her natural properties (e.g., behavioral trait) are typically unseen in training. In this case, the source pre-trained model has a low ability to adapt to these out-of-source characteristics, resulting in an unreliable prediction. To tackle this issue, we propose a novel helper-predictor test-time personalization approach (H/P-TTP), which allows for a generalizable representation of out-of-source subjects to gain more realistic predictions. Concretely, the helper is preceded by explicit and implicit augmenters, where the former yields noisy sequences to improve robustness, while the latter is to generate novel-domain data with an adversarial learning paradigm. Then, the domain-generalizable learning is achieved where the helper can extract cross-subject invariant-knowledge to update the predictor. At test time, given a new person, the predictor is able to be further optimized to empower personalized capabilities to the specific properties. Extensive experiments show that with H/P-TTP, the existing models are significantly improved for various unseen subjects. The project page is available at https://sites.google.com/view/hp-ttp. Qiongjie Cui, Huaijiang Sun, Jianfeng Lu 0003, Bin Li 0084, Hongwei Yi |
ICCV | 1 |
| 2023 | Learning Snippet-to-Motion Progression for Skeleton-based Human Motion PredictionabstractExisting Graph Convolutional Networks to achieve human motion prediction largely adopt a one-step scheme, which output the prediction straight from history input, failing to exploit human motion patterns. We observe that human motions have transitional patterns and can be split into snippets representative of each transition. Each snippet can be reconstructed from its starting and ending poses referred to as the transitional poses. We propose a snippet-to-motion multi-stage framework that breaks motion prediction into sub-tasks easier to accomplish. Each sub-task integrates three modules: transitional pose prediction, snippet reconstruction, and snippet-to-motion prediction. Specifically, we propose to first predict only the transitional poses. Then we use them to reconstruct the corresponding snippets, obtaining a close approximation to the true motion sequence. Finally we refine them to produce the final prediction output. To implement the network, we propose a novel unified graph modeling, which allows for direct and effective feature propagation compared to existing approaches which rely on separate space-time modeling. Extensive experiments on Human 3.6M, CMU Mocap and 3DPW datasets verify the effectiveness of our method which achieves state-of-the-art performance. Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001 |
MMAsia | 2 |
| 2023 | Graph-Guided MLP-Mixer for Skeleton-Based Human Motion PredictionabstractIn recent years, Graph Convolutional Networks (GCNs) have been widely used in human motion prediction, but their performance remains unsatisfactory. Recently, MLP-Mixer, initially developed for vision tasks, has been leveraged into human motion prediction as a promising alternative to GCNs, which achieves both better performance and better efficiency than GCNs. Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001 |
MMAsia | 2 |
| 2022 | Overlooked Poses Actually Make Sense: Distilling Privileged Knowledge for Human Motion Prediction
Xiaoning Sun, Qiongjie Cui, Huaijiang Sun, Bin Li 0084, Jianfeng Lu 0003 |
ECCV (5) | 2 |
| 2021 | Towards Accurate 3D Human Motion Prediction From Incomplete ObservationsabstractPredicting accurate and realistic future human poses from historically observed sequences is a fundamental task in the intersection of computer vision, graphics, and artificial intelligence. Recently, continuous efforts have been devoted to addressing this issue, which has achieved remarkable progress. However, the existing work is seriously limited by complete observation, that is, once the historical motion sequence is incomplete (with missing values), it can only produce unexpected predictions or even deformities. Furthermore, due to inevitable reasons such as occlusion and the lack of equipment precision, the incompleteness of motion data occurs frequently, which hinders the practical application of current algorithms.In this work, we first notice this challenging problem, i.e., how to generate high-fidelity human motion predictions from incomplete observations. To solve it, we propose a novel multi-task graph convolutional network (MTGCN). Specifically, the model involves two branches, in which the primary task is to focus on forecasting future 3D human actions accurately, while the auxiliary one is to repair the missing value of the incomplete observation. Both of them are integrated into a unified framework to share the spatio-temporal representation, which improves the final performance of each collaboratively. On three large-scale datasets, for various data missing scenarios in the real world, extensive experiments demonstrate that our approach is consistently superior to the state-of-the-art methods in which the missing values from incomplete observations are not explicitly analyzed. Qiongjie Cui, Huaijiang Sun |
CVPR | 1 |
| 2021 | Deep Human Dynamics PriorabstractMotion capture (MoCap) technology aims to provide an accurate record of human motion, with specific potentials in activity analysis, human behavior understanding, as well as multimedia industries of animation production and special effects movies. However, because of joint occlusion and limitation of equipment precision, the raw motion data are often damaged, which severely hinders its downstream applications. The latest method relies on deep neural networks to reconstruct the underlying complete motion from the degraded observation, achieving remarkable results. Unfortunately, due to the non-enumerability of human motion, the trained model from large-scale training data often fails to comprehensively cover incomputable action categories, which may lead to a sharp decline in the performance of deep learning-based methods. To handle these limitations, we propose an untrained deep generative model, in which Graph Convolutional Networks (GCNs) are utilized to efficiently capture complicated topological relationships of human joints. We show that the untrained GCN architecture with randomly-initialized weights is sufficient to extract some low-level statistics for human motion reconstruction without any training process. Notably, the performance of our approach is comparable to that of those trained models, while its application is not restricted by the availability of training data or a pre-trained network. Moreover, the proposed model even surpasses the state-of-the-art methods when encountering unprecedented samples in the human action database, regardless of the tasks of human motion recovery and gap-filling problem. Qiongjie Cui, Huaijiang Sun, Yue Kong, Xiaoning Sun |
ACM Multimedia | 1 |
| 2021 | Efficient human motion prediction using temporal convolutional generative adversarial network
Qiongjie Cui, Huaijiang Sun, Yue Kong, Yanmeng Li |
Inf. Sci. | 1 |
| 2021 | R-CTSVM+: Robust capped L1-norm twin support vector machine with privileged information
Yanmeng Li, Huaijiang Sun, Wenzhu Yan, Qiongjie Cui |
Inf. Sci. | 4 |
| 2021 | A dual-stream framework guided by adaptive Gaussian maps for interactive image segmentation
Zongyuan Ding, Tao Wang 0020, Quan-Sen Sun, Qiongjie Cui, Fuhua Chen |
Knowl. Based Syst. | 4 |
| 2020 | Learning Dynamic Relationships for 3D Human Motion Predictionabstract3D human motion prediction, i.e., forecasting future sequences from given historical poses, is a fundamental task for action analysis, human-computer interaction, machine intelligence. Recently, the state-of-the-art method assumes that the whole human motion sequence involves a fully-connected graph formed by links between each joint pair. Although encouraging performance has been made, due to the neglect of the inherent and meaningful characteristics of the natural connectivity of human joints, unexpected results may be produced. Moreover, such a complicated topology greatly increases the training difficulty. To tackle these issues, we propose a deep generative model based on graph networks and adversarial learning. Specifically, the skeleton pose is represented as a novel dynamic graph, in which natural connectivities of the joint pairs are exploited explicitly, and the links of geometrically separated joints can also be learned implicitly. Notably, in the proposed model, the natural connection strength is adaptively learned, whereas, in previous schemes, it was constant. Our approach is evaluated on two representations (i.e., angle-based, position-based) from various large-scale 3D skeleton benchmarks (e.g., H3.6M, CMU, 3DPW MoCap). Extensive experiments demonstrate that our approach achieves significant improvements against existing baselines in accuracy and visualization. Code will be available at https://github.com/cuiqiongjie/LDRGCN. Qiongjie Cui, Huaijiang Sun |
CVPR | 1 |
| 2020 | A novel local region-based active contour model for image segmentation using Bayes theorem
Guo Cao, Tao Wang 0020, Qiongjie Cui, Bisheng Wang |
Inf. Sci. | 4 |
| 2020 | Efficient human motion recovery using bidirectional attention network
Qiongjie Cui, Huaijiang Sun, Yue Kong |
Neural Comput. Appl. | 1 |
| 2019 | A Deep Bi-directional Attention Network for Human Motion RecoveryabstractHuman motion capture (mocap) data, recording the movement of markers attached to specific joints, has gradually become the most popular solution of animation production. However, the raw motion data are often corrupted due to joint occlusion, marker shedding and the lack of equipment precision, which severely limits the performance in real-world applications. Since human motion is essentially a sequential data, the latest methods resort to variants of long short-time memory network (LSTM) to solve related problems, but most of them tend to obtain visually unreasonable results. This is mainly because these methods hardly capture long-term dependencies and cannot explicitly utilize relevant context, especially in long sequences. To address these issues, we propose a deep bi-directional attention network (BAN) which can not only capture the long-term dependencies but also adaptively extract relevant information at each time step. Moreover, the proposed model, embedded attention mechanism in the bi-directional LSTM (BLSTM) structure at the encoding and decoding stages, can decide where to borrow information and use it to recover corrupted frame effectively. Extensive experiments on CMU database demonstrate that the proposed model consistently outperforms other state-of-the-art methods in terms of recovery accuracy and visualization. Qiongjie Cui, Huaijiang Sun, Yue Kong |
IJCAI | 1 |
| 2019 | Nonlocal low-rank regularization for human motion recovery based on similarity analysis
Qiongjie Cui, Beijia Chen, Huaijiang Sun |
Inf. Sci. | 1 |
| 2019 | Robust low-rank kernel multi-view subspace clustering based on the Schatten p-norm and correntropy
Huaijiang Sun, Zhenwen Ren, Qiongjie Cui, Yanmeng Li |
Inf. Sci. | 5 |