EDBT 2026 Demo / reviewers in the wild / expert
Ju Dai
dblp:209/7205
· DBLP profile ↗
42ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0002-9397-8539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 28 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ECLIPSE: Continuous Alpha Field Modulation for Zero-Shot Educational Facial Expression Recognition
Yixiao Xu, Yulian Sheng, Junxuan Bai, Feng Zhou 0007, Ju Dai, JunJun Pan |
ICIC (19) | 5 |
| 2026 | LiteNeRFAvatar: A lightweight NeRF with local feature learning for dynamic human avatar
JunJun Pan, Junxuan Bai, Ju Dai |
Pattern Recognit. | 4 |
| 2026 | CTNet: Color transformation network for low-light image enhancement
Lidong Xie, Runmin Cong, Ju Dai, Wenhan Yang, JunJun Pan |
Pattern Recognit. | 3 |
| 2026 | EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent DiffusionabstractSpeech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies. Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032 |
Vis. Comput. | 2 |
| 2026 | Multi-level fusion tokens for enhanced self-supervised skeleton-based action recognition
Kaida Ning, Feng Zhou 0007, JunJun Pan, Hongwen Xu, Ju Dai |
Vis. Comput. | 6 |
| 2026 | CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery
Feng Zhou 0007, Shuang Ji, Pei Shen, Ju Dai, JunJun Pan, Yukun Lai, Paul L. Rosin |
Vis. Comput. | 4 |
| 2026 | Mask-aware tri-modal learning for indoor 3D object detection
Feng Zhou 0007, Kaida Ning, JunJun Pan, Jin Li 0068, Ju Dai |
Vis. Comput. | 6 |
| 2026 | AgeStyle: Dynamic age-guided motion transfer in virtual realityabstractAge significantly influences human motor patterns, yet existing virtual reality (VR) systems lack dynamic modelling of these variations. This paper introduces AgeStyle, a versatile framework that integrates age-guided style selection with motion style transfer to convert user-uploaded videos into interactive 3D motion models. Utilizing 2D joint detection and 3D pose estimation, AgeStyle constructs motion representations enhanced by a CLIP-driven Cross-Attention module, capturing the distinct traits of different age groups—child flexibility, adult efficiency, and elderly stability. Our system enables real-time switching between motion styles and perspectives through voice commands, offering an immersive exploration of age-related movements. Quantitative experiments on the XIA dataset demonstrate AgeStyle’s competitive performance in both content preservation and style consistency, achieving average CC and SC++ scores of 7.4 and 14.8, respectively. AgeStyle represents a meaningful advancement in VR character design, with broad potential applications in education, healthcare, rehabilitation, and interactive entertainment. For the demo, please refer to https://youtu.be/eo7Shy0Ukps . The source code of AgeStyle is available at https://github.com/codeozzz/ageStyle . Feng Zhou 0007, Ju Dai, Sen-Zhe Xu 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2025 | Position-Aware Guided Point Cloud Completion with CLIP ModelabstractPoint cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods. Feng Zhou 0007, Ju Dai, Lei Li 0050, Junliang Xing |
AAAI | 3 |
| 2025 | Human Motion Instruction TuningabstractThis paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang |
CVPR | 6 |
| 2025 | Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationabstractIn 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module—Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git. Ju Dai, Xin Zhao 0025, Feng Zhou 0007, JunJun Pan, Lei Li 0050 |
CVPR | 2 |
| 2025 | Chat-Driven 3D Human Pose and Shape Editing with Large Language ModelsabstractGenerating and creating humanoid 3D models has received increasing attention recently due to its fundamental support for many high-level 3D applications. Although automatic 3D pose and shape reconstruction methods have achieved promising results, there are still some failure cases due to self-occlusions, viewpoint changes, and the complexity of human pose articulations. In this paper, we propose a novel way to leverage Large Language Models (LLMs) to interactively reconstruct human pose and shape based on a Skinned Multi-Person Linear (SMPL) model. We construct a mapping table to fine-tune an LLM, enabling it to understand user inputs better and output the positional information of joint points. Additionally, a simple neural network is adopted to regress the shape cues of the SMPL. We demonstrate a gallery of results of numerous poses and shapes. We validate our method via numerical evaluations, user studies, and comparisons to manually posed characters and previous work. Feng Zhou 0007, Ju Dai, Mengxiao Zhu 0004, Yongmei Zhang, Yukun Lai, Paul L. Rosin |
ICASSP | 3 |
| 2025 | Hierarchical Proxy Learning for Cloth-Changing Person Re-IdentificationabstractCloth-Changing person Re-Identification (CC-ReID) depends significantly on learning discriminative features under the cloth-changing scenario. It is quite challenging due to the large intra-person variance and small inter-person variance caused by clothes changing. To address these issues, in this work we propose a Hierarchical Proxy Learning (HPL) framework to extract clothes-irrelevant and person-invariant features. Specifically, we employ person labels as the main proxy. Instead of leveraging clothing labels as sub proxy, we further propose a clustering-based automatic sub-proxy mining scheme. More specifically, we first construct a person-aware Main Proxy Learning (MPL) to improve the separability of different persons. Then, a Sub Proxy Learning (SPL) is constructed to enhance the intra-person compactness. Finally, a Sub-to-Main Proxy Learning (S2MPL) is proposed to promote the cooperation between the main proxies and sub proxies. In addition, to weed out the negative effect of clothes, we propose a Sample Balance and Diversity (SBD) module, which balances the number of sub proxies in a mini-batch and utilizes semantic guidance to enrich the diversity of clothes, simultaneously. Extensive experiments on two public CC-ReID datasets demonstrate the superiority of our proposed method over most state-of-the-art methods. Chenyang Yu, Xuehu Liu, Ju Dai, Huchuan Lu |
ICASSP | 3 |
| 2025 | AU-Blendshape for Fine-Grained Stylized 3D Facial Expression ManipulationabstractWhile 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-grained facial expression manipulation across identities. AUBlendSet is a blendshape data collection based on 32 standard facial action units (AUs) across 500 identities, along with an additional set of facial postures annotated with detailed AUs. Based on AUBlendSet, we propose AUBlendNet to learn AU-Blendshape basis vectors for different character styles. AUBlendNet predicts, in parallel, the AU-Blendshape basis vectors of the corresponding style for a given identity mesh, thereby achieving stylized 3D emotional facial manipulation. We comprehensively validate the effectiveness of AUBlendSet and AUBlendNet through tasks such as stylized facial expression manipulation, speech-driven emotional facial animation, and emotion recognition data augmentation. Through a series of qualitative and quantitative experiments, we demonstrate the potential and importance of AUBlendSet and AUBlendNet in 3D facial animation tasks. To the best of our knowledge, AUBlendSet is the first dataset, and AUBlendNet is the first network for continuous 3D facial expression manipulation for any identity through facial AUs. Our source code is available at https://github.com/wslh852/AUBlendNet.git. Ju Dai, Feng Zhou 0007, Kaida Ning, Lei Li 0050, JunJun Pan |
ICCV | 2 |
| 2025 | Learning Hanzi Character Through VR-Based Mortise-Tenon
Conglin Ma, Sen-Zhe Xu 0001, Ju Dai, Jie Liu 0022, Feng Zhou 0007 |
ICXR | 4 |
| 2025 | BACH: Bi-Stage Data-Driven Piano Performance Animation for Controllable Hand MotionabstractABSTRACT This paper presents a novel framework for generating piano performance animations using a two‐stage deep learning model. By using discrete musical score data, the framework transforms sparse control signals into continuous, natural hand motions. Specifically, in the first stage, by incorporating musical temporal context, the keyframe predictor is leveraged to learn keyframe motion guidance. Meanwhile, the second stage synthesizes smooth transitions between these keyframes via an inter‐frame sequence generator. Additionally, a Laplacian operator‐based motion retargeting technique is introduced, ensuring that the generated animations can be adapted to different digital human models. We demonstrate the effectiveness of the system through an audiovisual multimedia application. Our approach provides an efficient, scalable method for generating realistic piano animations and holds promise for broader applications in animation tasks driven by sparse control signals. Jihui Jiao, Ju Dai, JunJun Pan |
Comput. Animat. Virtual Worlds | 3 |
| 2025 | Motion In-Betweening via Recursive Keyframe PredictionabstractABSTRACT Motion in‐betweening is a flexible and efficient technique for generating 3‐dimensional animations. In this paper, we propose a keyframe‐driven method that effectively addresses the pose ambiguity issue and achieves robust in‐betweening performance. We introduce a keyframe‐driven synthesis framework. At each recursion, the key poses at both ends keep predicting the new one at the midpoint. The recursive breakdown reduces motion ambiguities by simplifying the in‐betweening sequence as the integration of short clips. The hybrid positional encoding scales the hidden states to adapt to long‐ and short‐term dependencies. Additionally, we employ a temporal refinement network to capture the local motion relationships, thereby enhancing the consistency of the predicted pose sequence. Through comprehensive evaluations that include both quantitative and qualitative comparisons, the proposed model demonstrates its competitiveness in prediction accuracy and in‐betweening flexibility. Ju Dai, Junxuan Bai, JunJun Pan |
Comput. Animat. Virtual Worlds | 2 |
| 2025 | Motion Editing for Quadruped Characters via Latent Frequency EmbeddingabstractThe accurate and diversified generation of motion sequences for virtual characters poses both an enticing and challenging task within the domain of 3D animation and game content production. To achieve a natural and realistic full-body motion, the movements of virtual characters must adhere to a set of constraints, promoting reliable and seamless pose-changing. This study presents a two-stage model specifically designed to learn Inverse Kinematics (IK) constraints from the representative quadruped character poses. In the first stage, we employ frequency analysis to decompose motion poses into the base-level and style-level components. The base-level content encapsulates the global correlations in the dataset, while the style-level variation centers on distinguishing the local attributes in similar data elements. In order to construct data correlations among poses, we embed the decomposed pose feature into a latent space in the second stage. The kernel matrix of the embedding, which is refined from the original joint angles to the decomposed representation and the IK constraints, creates a more compact distribution of the pose similarity and also guarantees a plausible sampling result with certain IK constraints. Moreover, new motions from the edited IK constraints can also be generated by proposing a searching strategy to adapt to our latent embedding. Experimental results reveal that our method is competitive with the state-of-the-art synthetic approaches in terms of accuracy, highlighting our considerable potential for high efficiency in the animation production. JunJun Pan, Ju Dai, Yang Gao 0032, Junxuan Bai, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Diffusion model with temporal constraint for 3D human pose estimation
Zhangmeng Chen, Ju Dai, JunJun Pan, Feng Zhou 0007 |
Vis. Comput. | 2 |
| 2025 | Dual-path spatio-temporal Mamba for skeleton-based action recognition
Ju Dai, Feng Zhou 0007, JunJun Pan, Hongwen Xu |
Vis. Comput. | 2 |
| 2024 | AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh RecoveryabstractEstimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regression to accurately and efficiently predict 3D joint locations and reconstruct the hand mesh. We adopt a pooling attention module that learns to focus on relevant regions in the input image to extract better features for handling occlusions, while greatly reducing the computational cost. The multi-scale 2D heatmaps provide spatial constraints to guide the 3D vertex predictions. By exploiting the complementary strengths of sparse 2D supervision and dense mesh regression, our method accurately reconstructs hand meshes with realistic details. Extensive experiments on standard benchmarks demonstrate that the proposed method efficiently improves the performance of 3D hand pose estimation and mesh recovery. The reproducible recipes are available at https://github.com/SDiannn/AHRNET-Heatmap. Feng Zhou 0007, Pei Shen, Ju Dai, Yukun Lai, Paul L. Rosin |
ICASSP | 3 |
| 2024 | Enhancing Real-Time Fluid Simulations with Lagrangian Methods and GPU-Based Techniques in Unity
Yanrui Sun, Feng Zhou 0007, Ju Dai |
ICXR | 3 |
| 2024 | Foot-constrained spatial-temporal transformer for keyframe-based complex motion synthesisabstractAbstract Keyframe‐based motion synthesis holds significant effects in games and movies. Existing methods for complex motion synthesis often require secondary post‐processing to eliminate foot sliding to yield satisfied motions. In this paper, we analyze the cause of the sliding issue attributed to the mismatch between root trajectory and motion postures. To address the problem, we propose a novel end‐to‐end Spatial‐Temporal transformer network conditioned on foot contact information for high‐quality keyframe‐based motion synthesis. Specifically, our model mainly compromises a spatial‐temporal transformer encoder and two decoders to learn motion sequence features and predict motion postures and foot contact states. A novel constrained embedding, which consists of keyframes and foot contact constraints, is incorporated into the model to facilitate network learning from diversified control knowledge. To generate matched root trajectory with motion postures, we design a differentiable root trajectory reconstruction algorithm to construct root trajectory based on the decoder outputs. Qualitative and quantitative experiments on the public LaFAN1, Dance, and Martial Arts datasets demonstrate the superiority of our method in generating high‐quality complex motions compared with state‐of‐the‐arts. Ju Dai, Junxuan Bai, Zhangmeng Chen, JunJun Pan |
Comput. Animat. Virtual Worlds | 2 |
| 2024 | DGFormer: Dynamic graph transformer for 3D human pose estimation
Zhangmeng Chen, Ju Dai, Junxuan Bai, JunJun Pan |
Pattern Recognit. | 2 |
| 2024 | Self-Supervised Lightweight Depth Estimation in Endoscopy Combining CNN and TransformerabstractIn recent years, an increasing number of medical engineering tasks, such as surgical navigation, pre-operative registration, and surgical robotics, rely on 3D reconstruction techniques. Self-supervised depth estimation has attracted interest in endoscopic scenarios because it does not require ground truth. Most existing methods depend on expanding the size of parameters to improve their performance. There, designing a lightweight self-supervised model that can obtain competitive results is a hot topic. We propose a lightweight network with a tight coupling of convolutional neural network (CNN) and Transformer for depth estimation. Unlike other methods that use CNN and Transformer to extract features separately and then fuse them on the deepest layer, we utilize the modules of CNN and Transformer to extract features at different scales in the encoder. This hierarchical structure leverages the advantages of CNN in texture perception and Transformer in shape extraction. In the same scale of feature extraction, the CNN is used to acquire local features while the Transformer encodes global information. Finally, we add multi-head attention modules to the pose network to improve the accuracy of predicted poses. Experiments demonstrate that our approach obtains comparable results while effectively compressing the model parameters on two datasets. Zhuoyue Yang, JunJun Pan, Ju Dai |
IEEE Trans. Medical Imaging | 3 |
| 2024 | MPMNet: A Data-Driven MPM Framework for Dynamic Fluid-Solid InteractionabstractHigh-accuracy, high-efficiency physics-based fluid-solid interaction is essential for reality modeling and computer animation in online games or real-time Virtual Reality (VR) systems. However, the large-scale simulation of incompressible fluid and its interaction with the surrounding solid environment is either time-consuming or suffering from the reduced time/space resolution due to the complicated iterative nature pertinent to numerical computations of involved Partial Differential Equations (PDEs). In recent years, we have witnessed significant growth in exploring a different, alternative data-driven approach to addressing some of the existing technical challenges in conventional model-centric graphics and animation methods. This article showcases some of our exploratory efforts in this direction. One technical concern of our research is to address the central key challenge of how to best construct the numerical solver effectively and how to best integrate spatiotemporal/dimensional neural networks with the available MPM's pressure solvers. In particular, we devise the MPMNet, a hybrid data-driven framework supporting the popular and powerful MPM, to combine the comprehensive properties of MPM in numerically handling physical behaviors ranging from fluid to deformable solids and the high efficiency of data-driven models. At the architectural level, our MPMNet comprises three primary components: A data processing module to describe the physical properties by way of the input fields; A deep neural network group to learn the spatiotemporal features; And an iterative refinement process to continue to reduce possible numerical errors. The goal of these special technical developments is to aim at involved numerical acceleration while preserving physical accuracy, realizing efficient and accurate fluid-solid interactions in a data-driven fashion. The extensive experimental results verify that our MPMNet can tremendously speed up the computation compared with the popular numerical methods as the complexity of interaction scenes increases while better retaining the numerical accuracy. Jin Li 0068, Yang Gao 0032, Ju Dai, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Free editing of Shape and Texture with Deformable Net for 3D Caricature Generation
Yuanyuan lin, Ju Dai, JunJun Pan, Feng Zhou 0007, Junxuan Bai |
Vis. Comput. | 2 |
| 2023 | GFENet: Group-Free Enhancement Network for Indoor Scene 3D Object Detection
Feng Zhou 0007, Ju Dai, JunJun Pan, Mengxiao Zhu 0004, Xingquan Cai, Chen Wang 0043 |
CGI (3) | 2 |
| 2023 | KD-Former: Kinematic and dynamic coupled transformer network for 3D human motion prediction
Ju Dai, Junxuan Bai, Feng Zhou 0007, JunJun Pan |
Pattern Recognit. | 1 |
| 2023 | A lightweight pose estimation network with multi-scale receptive field
Shuo Li 0001, Ju Dai, Zhangmeng Chen, JunJun Pan |
Vis. Comput. | 2 |
| 2022 | Attribute-Decomposable Motion Compression Network for 3D MoCap DataabstractMotion Capture (MoCap) data is one type of fundamental asset for the digital entertainment. The progressively increasing 3D applications make MoCap data compression unprecedentedly important. In this paper, we propose an end-to-end attribute-decomposable motion compression network using the AutoEncoder architecture. Specifically, the algorithm consists of an LSTM-based encoder-decoder for compression and decompression. The encoder module decomposes human motion into multiple uncorrelated semantic attributes, including action content, arm space, and motion mirror. The decoder module is responsible for reconstructing vivid motion based on the decomposed high-level characteristics. Our method is computationally efficient with powerful compression ability, outperforming the state-of-the-art methods in terms of compression rate and compression error. Furthermore, our model can generate new motion data given a combination of different motion attributes while existing methods have no such capability. Zengming Chen, Junxuan Bai, Ju Dai |
DCC | 3 |
| 2022 | MCGNet: Multi-Level Context-aware and Geometric-aware Network for 3D Object DetectionabstractHough voting based on PointNet++ [1] is effective against 3D object detection, which has been verified by VoteNet [2], H3DNet [3], etc. However, we find there is still room for improvements in two aspects. The first is that most existing methods ignores the particular significance of different format inputs and geometric primitives for predicting object proposals. The second is that the feature extracted by PointNet++ overlooks contextual information about each object. In this paper, to tackle the above issues, we introduce MCGNet to learn multi-level geometric-aware and scale-aware contextual information for 3D object detection. Specifically, our network mainly consists of the baseline module based on H3DNet, geometric-aware module, and context-aware module. The baseline module feeding with four-types inputs (Point, Edge, Surface, and Line) concentrates on extracting diversified geometric primitives, i.e., BB centers, BB face centers, and BB edge centers. The geometric-aware module is proposed to learn the different contributions among the four-types feature maps and the three geometric primitives. The context-aware module aims to establish long-range dependencies features for either four-types feature maps or three geometric primitives. Extensive experiments on two large datasets with real 3D scans, SUN RGB-D and ScanNet datasets, demonstrate that our method is effective against 3D object detection. Keng Chen, Feng Zhou 0007, Ju Dai, Pei Shen, Xingquan Cai, Fengquan Zhang |
ICIP | 3 |
| 2022 | Robust AUV Visual Loop-Closure Detection Based on Variational Autoencoder NetworkabstractThe visual loop-closure detection for autonomous underwater vehicles (AUVs) is a key component to reduce the drift error accumulated in simultaneous localization and mapping tasks. However, due to viewpoint changes, textureless images, and fast-moving objects, the loop closure detection in dramatically changing underwater environments remains a challenging problem to traditional geometric methods. Inspired by strong feature learning ability of deep neural networks, we propose an underwater loop-closure detection method based on a variational autoencoder network in this article. Our proposed method can learn effective image representations to deal with the challenges caused by dynamic underwater environments. Specifically, the proposed network is an unsupervised method, which avoids the difficulty and cost of labeling a great quantity of underwater data. Also included is a semantic object segmentation module, which is utilized to segment the underwater environments and assign weights to objects in order to alleviate the impact of fast-moving objects. Furthermore, an underwater image description scheme is used to enable efficient access to geometric and object-level semantic information, which helps to build a robust and real-time system in dramatically changing underwater scenarios. Finally, we test the proposed system under complex underwater environments and get a recall rate of 92.31% in the tested environments. Yangyang Wang 0005, Xiaorui Ma, Jie Wang 0003, Shilong Hou, Ju Dai, Dongbing Gu, Hongyu Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Neurophysiological and Subjective Analysis of VR Emotion Induction ParadigmabstractThe ecological validity of emotion-inducing scenarios is essential for emotion research. In contrast to the classical passive induction paradigm, immersive VR fully engages the psychological and physiological components of the subject, which is considered an ecologically valid paradigm for studying emotion. Several studies investigate the emotional responses to different VR tasks or games using subjective scales. However, little research regards VR as an eliciting material, especially when systematically analyzing emotional processes in VR from a neurophysiological perspective. To fill this gap and scientifically evaluate VR's ability to be used as an active method for emotion elicitation, we investigate the dynamic relationship between explicit information (subjective evaluations) and implicit information (objective neurophysiological data). A total of 28 participants are enlisted to watch eight VR videos while their SAM/IPQ scores and EEG data are recorded simultaneously. In ecologically valid scenarios, the subjective results demonstrate that VR has significant advantages for evoking emotion in arousal-valence. This conclusion is backed by our examination of objective neurophysiological evidence that VR videos effectively induce high-arousal emotions. In addition, we obtain features of critical channels and frequency oscillations associated with emotional valence, thereby validating previous research in more lifelike circumstances. In particular, we discover hemispheric asymmetry in the occipital region under high and low emotional arousal, which adds to our understanding of neural features and the dynamics of emotional arousal. As a result, we successfully integrate EEG and VR to demonstrate that VR is more pragmatic for evoking natural feelings and is beneficial for emotional research. Our research has set a precedent for new methodologies of using VR induction paradigms to acquire a more reliable explanation of affective computing. JunJun Pan, Yang Gao 0032, Yang Shen 0009, Ju Dai, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Diverse Dance Synthesis via Keyframes with Transformer ControllersabstractAbstract Existing keyframe‐based motion synthesis mainly focuses on the generation of cyclic actions or short‐term motion, such as walking, running, and transitions between close postures. However, these methods will significantly degrade the naturalness and diversity of the synthesized motion when dealing with complex and impromptu movements , e.g., dance performance and martial arts. In addition, current research lacks fine‐grained control over the generated motion, which is essential for intelligent human‐computer interaction and animation creation. In this paper, we propose a novel keyframe‐based motion generation network based on multiple constraints, which can achieve diverse dance synthesis via learned knowledge. Specifically, the algorithm is mainly formulated based on the recurrent neural network (RNN) and the Transformer architecture. The backbone of our network is a hierarchical RNN module composed of two long short‐term memory (LSTM) units, in which the first LSTM is utilized to embed the posture information of the historical frames into a latent space, and the second one is employed to predict the human posture for the next frame. Moreover, our framework contains two Transformer‐based controllers, which are used to model the constraints of the root trajectory and the velocity factor respectively, so as to better utilize the temporal context of the frames and achieve fine‐grained motion control. We verify the proposed approach on a dance dataset containing a wide range of contemporary dance. The results of three quantitative analyses validate the superiority of our algorithm. The video and qualitative experimental results demonstrate that the complex motion sequences generated by our algorithm can achieve diverse and smooth motion transitions between keyframes, even for long‐term synthesis. JunJun Pan, Junxuan Bai, Ju Dai |
Comput. Graph. Forum | 4 |
| 2021 | EmoDescriptor: A hybrid feature for emotional classification in dance movementsabstractAbstract Similar to language and music, dance performances provide an effective way to express human emotions. With the abundance of the motion capture data, content‐based motion retrieval and classification have been fiercely investigated. Although researchers attempt to interpret body language in terms of human emotions, the progress is limited by the scarce 3D motion database annotated with emotion labels. This article proposes a hybrid feature for emotional classification in dance performances. The hybrid feature is composed of an explicit feature and a deep feature. The explicit feature is calculated based on the Laban movement analysis, which considers the body, effort, shape, and space properties. The deep feature is obtained from latent representation through a 1D convolutional autoencoder. Eventually, we present an elaborate feature fusion network to attain the hybrid feature that is almost linearly separable. The abundant experiments demonstrate that our hybrid feature is superior to the separate features for the emotional classification in dance performances. Junxuan Bai, Rong Dai, Ju Dai, JunJun Pan |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | Dynamic imposter based online instance matching for person search
Ju Dai, Huchuan Lu, Hongyu Wang 0001 |
Pattern Recognit. | 1 |
| 2019 | Video Person Re-Identification by Temporal Residual LearningabstractIn this paper, we propose a novel feature learning framework for video person re-identification (re-ID). The proposed framework largely aims to exploit the adequate temporal information of video sequences and tackle the poor spatial alignment of moving pedestrians. More specifically, for exploiting the temporal information, we design a temporal residual learning (TRL) module to simultaneously extract the generic and specific features of consecutive frames. The TRL module is equipped with two bi-directional LSTM (BiLSTM), which are respectively responsible to describe a moving person in different aspects, providing complementary information for better feature representations. To deal with the poor spatial alignment in video re- ID datasets, we propose a spatial-temporal transformer network (ST2N) module. Transformation parameters in the ST2N module are learned by leveraging the high-level semantic information of the current frame as well as the temporal context knowledge from other frames. The proposed ST2N module with less learnable parameters allows effective person alignments under significant appearance changes. Extensive experimental results on the largescale MARS, PRID2011, ILIDS-VID and SDU-VID datasets demonstrate that the proposed method achieves consistently superior performance and outperforms most of the very recent state-of-the-art methods. Ju Dai, Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | A Bi-Directional Message Passing Model for Salient Object DetectionabstractRecent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detection. In this paper, we propose a novel bi-directional message passing model to integrate multi-level features for salient object detection. At first, we adopt a Multi-scale Context-aware Feature Extraction Module (MCFEM) for multi-level feature maps to capture rich context information. Then a bi-directional structure is designed to pass messages between multi-level features, and a gate function is exploited to control the message passing rate. We use the features after message passing, which simultaneously encode semantic information and spatial details, to predict saliency maps. Finally, the predicted results are efficiently combined to generate the final saliency map. Quantitative and qualitative experiments on five benchmark datasets demonstrate that our proposed model performs favorably against the state-of-the-art methods under different evaluation metrics. Lu Zhang 0053, Ju Dai, Huchuan Lu, You He 0002, Gang Wang 0012 |
CVPR | 2 |
| 2018 | Cross-view semantic projection learning for person re-identification
Ju Dai, Ying Zhang 0021, Huchuan Lu, Hongyu Wang 0001 |
Pattern Recognit. | 1 |
| 2009 | Waveform Analysis Mathematic Based on Ultrasonic Displacement MeasurementabstractThe measurement principle and sensor structure of a new-style time-difference method ultrasonic flow meter are presented. Software waveform analyses arithmetic based on FPGA and a bran-new echo datum mark position criterion are put forward. Exact measurement time can be helpful for improving system precision. In the end, realization project is simply described. Ju Dai |
IAS | 2 |