VLDB 2026 Research / reviewers in the wild / expert
JunJun Pan
dblp:11/7528 · also Junjun Pan
· DBLP profile ↗
76ranked-venue papers
17as first author
53since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 13 first-author · 36 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Correcting False Alarms from Unseen: Adapting Graph Anomaly Detectors at Test TimeabstractGraph anomaly detection (GAD), which aims to detect outliers in graph-structured data, has received increasing research attention recently. However, existing GAD methods assume identical training and testing distributions, which is rarely valid in practice. In real-world scenarios, unseen but normal samples may emerge during deployment, leading to a normality shift that degrades the performance of GAD models trained on the original data. Through empirical analysis, we reveal that the degradation arises from (1) semantic confusion, where unseen normal samples are misinterpreted as anomalies due to their novel patterns, and (2) aggregation contamination, where the representations of seen normal nodes are distorted by unseen normals through message aggregation. While retraining or fine-tuning GAD models could be a potential solution to the above challenges, the high cost of model retraining and the difficulty of obtaining labeled data often render this approach impractical in real-world applications. To bridge the gap, we proposed a lightweight and plug-and-play Test-time adaptation framework for correcting Unseen Normal pattErns (TUNE) in GAD. To address semantic confusion, a graph aligner is employed to align the shifted data to the original one at the graph attribute level. Moreover, we utilize the minimization of representation-level shift as a supervision signal to train the aligner, which leverages the estimated aggregation contamination as a key indicator of normality shift. Extensive experiments on 10 real-world datasets demonstrate that TUNE significantly enhances the generalizability of pre-trained GAD models to both synthetic and real unseen normal patterns. JunJun Pan, Yixin Liu 0001, Chuan Zhou 0001, Alan Wee-Chung Liew, Shirui Pan |
AAAI | 1 |
| 2026 | Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly DetectionabstractJunjun Pan, Yixin Liu, Rui Miao, Kaize Ding, Yu Zheng, Quoc Viet Hung Nguyen, Alan Wee-Chung Liew, Shirui Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. JunJun Pan, Yixin Liu 0001, Rui Miao 0003, Kaize Ding, Yu Zheng 0013, Nguyen Quoc Viet Hung, Alan Wee-Chung Liew, Shirui Pan |
ACL (1) | 1 |
| 2026 | ECLIPSE: Continuous Alpha Field Modulation for Zero-Shot Educational Facial Expression Recognition
Yixiao Xu, Yulian Sheng, Junxuan Bai, Feng Zhou 0007, Ju Dai, JunJun Pan |
ICIC (19) | 6 |
| 2026 | LiteNeRFAvatar: A lightweight NeRF with local feature learning for dynamic human avatar
JunJun Pan, Junxuan Bai, Ju Dai |
Pattern Recognit. | 1 |
| 2026 | CTNet: Color transformation network for low-light image enhancement
Lidong Xie, Runmin Cong, Ju Dai, Wenhan Yang, JunJun Pan |
Pattern Recognit. | 5 |
| 2026 | A Unified Viscoelastic Solver for Multiphase Fluid Simulation Based on a Mixture ModelabstractFluid simulation is a central topic in computer graphics, encompassing a wide range of methodologies for modeling Newtonian, non-Newtonian, and viscoelastic behaviors across both single-phase and multiphase settings. Existing single-phase frameworks have achieved high visual fidelity, yet multiphase simulations remain limited in accurately capturing complex phase interactions, particularly under high-viscosity-ratio or viscoelastic conditions. To address these challenges, we develop a unified multiphase viscoelastic formulation capable of handling diverse fluid types-including Newtonian, shear-dependent non-Newtonian, and viscoelastic flows-within a single consistent framework. The formulation extends mixture-model approaches through a multi-mode conformation tensor representation, which enhances numerical stability via phase-level stress corrections and efficiently captures a broad spectrum of rheological behaviors. Compared with existing techniques, our framework achieves improved momentum-mass consistency and numerical stability, maintaining physically plausible results across wide viscosity ranges, advancing the state of the art in multiphase viscoelastic fluid simulation. Long Shen, Yalan Zhang, Steffen Frey, Alexandru C. Telea, Jirí Kosinka, JunJun Pan, Xiaokun Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032 |
Vis. Comput. | 6 |
| 2026 | Multi-level fusion tokens for enhanced self-supervised skeleton-based action recognition
Kaida Ning, Feng Zhou 0007, JunJun Pan, Hongwen Xu, Ju Dai |
Vis. Comput. | 4 |
| 2026 | CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery
Feng Zhou 0007, Shuang Ji, Pei Shen, Ju Dai, JunJun Pan, Yukun Lai, Paul L. Rosin |
Vis. Comput. | 5 |
| 2026 | Mask-aware tri-modal learning for indoor 3D object detection
Feng Zhou 0007, Kaida Ning, JunJun Pan, Jin Li 0068, Ju Dai |
Vis. Comput. | 4 |
| 2025 | A Label-free Heterophily-guided Approach for Unsupervised Graph Fraud DetectionabstractGraph fraud detection (GFD) has rapidly advanced in protecting online services by identifying malicious fraudsters. Recent supervised GFD research highlights that heterophilic connections between fraudster and user greatly impacts detection performance, where the fraudsters tend to camouflage themselves by building more connections to benign users. Despite their promising performance, their label reliance limits its application in unsupervised scenarios; Additionally, accurately capturing complex and diverse heterophily patterns without labels poses a further challenge. Therefore, we propose a Heterophily-guided Unsupervised Graph fraud dEtection approach (HUGE) for unsupervised GFD, which contains two essential components: a heterophily estimation module and an alignment-based fraud detection module. In the heterophily estimation module, we design a novel unsupervised heterophily metric called HALO, which captures the critical graph properties for GFD, enabling its outstanding ability to estimate heterophily with attributes. In the alignment-based fraud detection module, we develop a joint MLP-GNN architecture with ranking loss and asymmetric alignment loss. The ranking loss aligns the predicted fraud score with the relative order of HALO, providing an extra robustness guarantee by comparing heterophily between non-adjacent nodes. Moreover, the asymmetric alignment loss effectively utilizes structural information to alleviate the feature-smooth effects. Extensive experiments on six datasets demonstrate that HUGE consistently outperforms competitors, showcasing its effectiveness and robustness. JunJun Pan, Yixin Liu 0001, Xin Zheng 0008, Yizhen Zheng, Alan Wee-Chung Liew, Fuyi Li, Shirui Pan |
AAAI | 1 |
| 2025 | NAT-3D: A Non-Rigid Approach for Accurate 3D Tracking of Basic Surgical Instrumentsabstract3D tracking technology serves as an essential enabler in numerous domains, most notably within digital healthcare, where it supports critical procedures such as surgical navigation, trajectory planning, and high-fidelity surgical simulation, etc. However, due to the slender, textureless, and reflective characteristics of basic surgical instruments, coupled with the complexity of their diverse motion patterns, existing systems and methods face significant limitations in achieving accurate 3D tracking of these instruments. We propose NAT-3D, an effective and efficient non-rigid 3D tracking approach based on multi-modal region estimation, integrating kinematic structural constraints and nonrigid constraints. This method enables accurate tracking and dynamic mapping of a variety of basic surgical instruments, including scalpels, scissors, clamps, and forceps, without the need for additional markers. The tracking covers various movement modes, including rigid motion, local mechanical motion, and non-rigid composite motion. Extensive experiments demonstrate that our method outperforms previous algorithms in terms of robustness, accuracy, applicability, and real-time performance. In addition, we introduce a novel dataset for Basic Surgical Instrument Tracking (BSIT), which can serve as a benchmark for future related research. Zihan Deng, Jier Zhang, Yushan Pan, JunJun Pan, Zhijie Xu |
BIBM | 5 |
| 2025 | Attention-Enhanced 3D Craniomaxillofacial Anatomical Landmark Detection Based on Projection
Yuyou Zhong, Xi Fu, Ruilin Zhao, Feng Niu, JunJun Pan |
CGI (2) | 9 |
| 2025 | Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature GenerationabstractGeneralized Zero-Shot Learning (GZSL) addresses the challenge of classifying unseen classes in the presence of seen classes by leveraging semantic attributes to bridge the gap for unseen classes. However, in image based disease classification, such as glioma sub-typing, distinguishing between classes using image semantic attributes can be challenging. To address this challenge, we introduce a novel GZSL method that eliminates the dependency on semantic information. Specifically, we propose that the primary of most classification in clinic is risk stratification, and classes are inherently ordered rather than purely categorical. Based on this insight, we present an inter-class feature augmentation (IFA) module, where distributions of different classes are ordered by their risk levels in a learned feature space using pre-defined joint conditional Gaussian distribution model. This ordering enables the generation of unseen class features through feature mixing of adjacent seen classes, effectively transforming the zero-shot learning problem into a supervised learning task. Our method eliminates the need for explicit semantic information, avoiding the cross-modal alignment between visual and semantic features. Moreover, the IFA module for GZSL requires no structural modifications to the existing classification models. In the experiment, both in-house and public datasets are used to evaluate our method across different tasks, including glioma subtyping, Alzheimer’s disease (AD) classification and diabetic retinopathy classification. Experimental results demonstrate that our method outperforms the state-of-the-art GZSL methods with statistical significance. Libiao Chen, Dong Nie, JunJun Pan, Zhenyu Tang 0002 |
CVPR | 3 |
| 2025 | Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationabstractIn 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module—Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git. Ju Dai, Xin Zhao 0025, Feng Zhou 0007, JunJun Pan, Lei Li 0050 |
CVPR | 5 |
| 2025 | AU-Blendshape for Fine-Grained Stylized 3D Facial Expression ManipulationabstractWhile 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-grained facial expression manipulation across identities. AUBlendSet is a blendshape data collection based on 32 standard facial action units (AUs) across 500 identities, along with an additional set of facial postures annotated with detailed AUs. Based on AUBlendSet, we propose AUBlendNet to learn AU-Blendshape basis vectors for different character styles. AUBlendNet predicts, in parallel, the AU-Blendshape basis vectors of the corresponding style for a given identity mesh, thereby achieving stylized 3D emotional facial manipulation. We comprehensively validate the effectiveness of AUBlendSet and AUBlendNet through tasks such as stylized facial expression manipulation, speech-driven emotional facial animation, and emotion recognition data augmentation. Through a series of qualitative and quantitative experiments, we demonstrate the potential and importance of AUBlendSet and AUBlendNet in 3D facial animation tasks. To the best of our knowledge, AUBlendSet is the first dataset, and AUBlendNet is the first network for continuous 3D facial expression manipulation for any identity through facial AUs. Our source code is available at https://github.com/wslh852/AUBlendNet.git. Ju Dai, Feng Zhou 0007, Kaida Ning, Lei Li 0050, JunJun Pan |
ICCV | 6 |
| 2025 | Iterative Foundation-Dedicated Learning: Optimized Key Frames, Prompts and Memories for Semi-supervised Segmentation
Ziman Yin, Dong Nie, Shuo Li 0001, JunJun Pan, Zhenyu Tang 0002 |
MICCAI (8) | 4 |
| 2025 | BACH: Bi-Stage Data-Driven Piano Performance Animation for Controllable Hand MotionabstractABSTRACT This paper presents a novel framework for generating piano performance animations using a two‐stage deep learning model. By using discrete musical score data, the framework transforms sparse control signals into continuous, natural hand motions. Specifically, in the first stage, by incorporating musical temporal context, the keyframe predictor is leveraged to learn keyframe motion guidance. Meanwhile, the second stage synthesizes smooth transitions between these keyframes via an inter‐frame sequence generator. Additionally, a Laplacian operator‐based motion retargeting technique is introduced, ensuring that the generated animations can be adapted to different digital human models. We demonstrate the effectiveness of the system through an audiovisual multimedia application. Our approach provides an efficient, scalable method for generating realistic piano animations and holds promise for broader applications in animation tasks driven by sparse control signals. Jihui Jiao, Ju Dai, JunJun Pan |
Comput. Animat. Virtual Worlds | 4 |
| 2025 | Motion In-Betweening via Recursive Keyframe PredictionabstractABSTRACT Motion in‐betweening is a flexible and efficient technique for generating 3‐dimensional animations. In this paper, we propose a keyframe‐driven method that effectively addresses the pose ambiguity issue and achieves robust in‐betweening performance. We introduce a keyframe‐driven synthesis framework. At each recursion, the key poses at both ends keep predicting the new one at the midpoint. The recursive breakdown reduces motion ambiguities by simplifying the in‐betweening sequence as the integration of short clips. The hybrid positional encoding scales the hidden states to adapt to long‐ and short‐term dependencies. Additionally, we employ a temporal refinement network to capture the local motion relationships, thereby enhancing the consistency of the predicted pose sequence. Through comprehensive evaluations that include both quantitative and qualitative comparisons, the proposed model demonstrates its competitiveness in prediction accuracy and in‐betweening flexibility. Ju Dai, Junxuan Bai, JunJun Pan |
Comput. Animat. Virtual Worlds | 4 |
| 2025 | Anchor-Based Multiview Subspace Clustering With Anchor-wise and Class-wise AlignmentsabstractMultiview subspace clustering has shown promising performance in multimedia and data mining applications. However, its employment in large-scale datasets is limited due to its quadratic or even cubic computational complexity. The anchor graph strategy, which selects a few important samples (anchors) to represent the whole data for different views, has been introduced to address this challenge. These methods rely on a heuristic assumption that the correspondence and class structures between the sets of anchors across different views are the same. This assumption ignores the difference in the ordering of anchors with respect to their associated classes and the number of anchors belonging to the same class from different views. As a result, this can lead to unsatisfactory clustering results due to incorrect anchorwise and classwise alignments. To tackle this issue, this article proposes an anchor-based multiview subspace clustering with anchorwise and classwise alignments (AMCA2) method. Specifically, the proposed method simultaneously aligns and fuses multiple anchor graphs anchor wisely and class wisely via learning permutation matrices and utilizing the Hadamard product. To further enhance the clustering performance of AMCA2, we propose a novel anchor selection method called kernel anchor selection (KAS) to select more representative anchors. Extensive experiments on ten benchmark datasets are conducted to show the superiority and effectiveness of AMCA2over the state-of-the-art methods. Ye Liu 0014, Hongshan Pu, JunJun Pan, Michael Kwok-Po Ng, Hongmin Cai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Motion Editing for Quadruped Characters via Latent Frequency EmbeddingabstractThe accurate and diversified generation of motion sequences for virtual characters poses both an enticing and challenging task within the domain of 3D animation and game content production. To achieve a natural and realistic full-body motion, the movements of virtual characters must adhere to a set of constraints, promoting reliable and seamless pose-changing. This study presents a two-stage model specifically designed to learn Inverse Kinematics (IK) constraints from the representative quadruped character poses. In the first stage, we employ frequency analysis to decompose motion poses into the base-level and style-level components. The base-level content encapsulates the global correlations in the dataset, while the style-level variation centers on distinguishing the local attributes in similar data elements. In order to construct data correlations among poses, we embed the decomposed pose feature into a latent space in the second stage. The kernel matrix of the embedding, which is refined from the original joint angles to the decomposed representation and the IK constraints, creates a more compact distribution of the pose similarity and also guarantees a plausible sampling result with certain IK constraints. Moreover, new motions from the edited IK constraints can also be generated by proposing a searching strategy to adapt to our latent embedding. Experimental results reveal that our method is competitive with the state-of-the-art synthetic approaches in terms of accuracy, highlighting our considerable potential for high efficiency in the animation production. JunJun Pan, Ju Dai, Yang Gao 0032, Junxuan Bai, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Diffusion model with temporal constraint for 3D human pose estimation
Zhangmeng Chen, Ju Dai, JunJun Pan, Feng Zhou 0007 |
Vis. Comput. | 3 |
| 2025 | Enhanced material point method with affine projection stabilizer for efficient hyperelastic simulations
Siyan Zhu, JunJun Pan |
Vis. Comput. | 3 |
| 2025 | Dual-path spatio-temporal Mamba for skeleton-based action recognition
Ju Dai, Feng Zhou 0007, JunJun Pan, Hongwen Xu |
Vis. Comput. | 4 |
| 2024 | An augmented reality laparoscopic colectomy navigation system
Junyang Lu, JunJun Pan |
BIBM | 5 |
| 2024 | Foot-constrained spatial-temporal transformer for keyframe-based complex motion synthesisabstractAbstract Keyframe‐based motion synthesis holds significant effects in games and movies. Existing methods for complex motion synthesis often require secondary post‐processing to eliminate foot sliding to yield satisfied motions. In this paper, we analyze the cause of the sliding issue attributed to the mismatch between root trajectory and motion postures. To address the problem, we propose a novel end‐to‐end Spatial‐Temporal transformer network conditioned on foot contact information for high‐quality keyframe‐based motion synthesis. Specifically, our model mainly compromises a spatial‐temporal transformer encoder and two decoders to learn motion sequence features and predict motion postures and foot contact states. A novel constrained embedding, which consists of keyframes and foot contact constraints, is incorporated into the model to facilitate network learning from diversified control knowledge. To generate matched root trajectory with motion postures, we design a differentiable root trajectory reconstruction algorithm to construct root trajectory based on the decoder outputs. Qualitative and quantitative experiments on the public LaFAN1, Dance, and Martial Arts datasets demonstrate the superiority of our method in generating high‐quality complex motions compared with state‐of‐the‐arts. Ju Dai, Junxuan Bai, Zhangmeng Chen, JunJun Pan |
Comput. Animat. Virtual Worlds | 6 |
| 2024 | DGFormer: Dynamic graph transformer for 3D human pose estimation
Zhangmeng Chen, Ju Dai, Junxuan Bai, JunJun Pan |
Pattern Recognit. | 4 |
| 2024 | Self-Supervised Lightweight Depth Estimation in Endoscopy Combining CNN and TransformerabstractIn recent years, an increasing number of medical engineering tasks, such as surgical navigation, pre-operative registration, and surgical robotics, rely on 3D reconstruction techniques. Self-supervised depth estimation has attracted interest in endoscopic scenarios because it does not require ground truth. Most existing methods depend on expanding the size of parameters to improve their performance. There, designing a lightweight self-supervised model that can obtain competitive results is a hot topic. We propose a lightweight network with a tight coupling of convolutional neural network (CNN) and Transformer for depth estimation. Unlike other methods that use CNN and Transformer to extract features separately and then fuse them on the deepest layer, we utilize the modules of CNN and Transformer to extract features at different scales in the encoder. This hierarchical structure leverages the advantages of CNN in texture perception and Transformer in shape extraction. In the same scale of feature extraction, the CNN is used to acquire local features while the Transformer encodes global information. Finally, we add multi-head attention modules to the pose network to improve the accuracy of predicted poses. Experiments demonstrate that our approach obtains comparable results while effectively compressing the model parameters on two datasets. Zhuoyue Yang, JunJun Pan, Ju Dai |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Multimodal Physiological Analysis of Impact of Emotion on Cognitive Control in VRabstractCognitive control is often perplexing to elucidate and can be easily influenced by emotions. Understanding the individual cognitive control level is crucial for enhancing VR interaction and designing adaptive and self-correcting VR/AR applications. Emotions can reallocate processing resources and influence cognitive control performance. However, current research has primarily emphasized the impact of emotional valence on cognitive control tasks, neglecting emotional arousal. In this study, we comprehensively investigate the influence of emotions on cognitive control based on the arousal-valence model. A total of 26 participants are recruited, inducing emotions through VR videos with high ecological validity and then performing related cognitive control tasks. Leveraging physiological data including EEG, HRV, and EDA, we employ classification techniques such as SVM, KNN, and deep learning to categorize cognitive control levels. The experiment results demonstrate that high-arousal emotions significantly enhance users' cognitive control abilities. Utilizing complementary information among multi-modal physiological signal features, we achieve an accuracy of 84.52% in distinguishing between high and low cognitive control. Additionally, time-frequency analysis results confirm the existence of neural patterns related to cognitive control, contributing to a better understanding of the neural mechanisms underlying cognitive control in VR. Our research indicates that physiological signals measured from both the central and autonomic nervous systems can be employed for cognitive control classification, paving the way for novel approaches to improve VR/AR interactions. JunJun Pan, Yang Gao 0032, Hong Qin 0001, Yang Shen 0009 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Efficient frictional contacts for soft body dynamics via ADMM
Siyan Zhu, Xiao Zhai, Aimin Hao, JunJun Pan |
Vis. Comput. | 6 |
| 2024 | Free editing of Shape and Texture with Deformable Net for 3D Caricature Generation
Yuanyuan lin, Ju Dai, JunJun Pan, Feng Zhou 0007, Junxuan Bai |
Vis. Comput. | 3 |
| 2023 | GFENet: Group-Free Enhancement Network for Indoor Scene 3D Object Detection
Feng Zhou 0007, Ju Dai, JunJun Pan, Mengxiao Zhu 0004, Xingquan Cai, Chen Wang 0043 |
CGI (3) | 3 |
| 2023 | PREM: A Simple Yet Effective Approach for Node-Level Graph Anomaly DetectionabstractNode-level graph anomaly detection (GAD) plays a critical role in identifying anomalous nodes from graph-structured data in various domains such as medicine, social networks, and e-commerce. However, challenges have arisen due to the diversity of anomalies and the dearth of labeled data. Existing methodologies - reconstruction-based and contrastive learning - while effective, often suffer from efficiency issues, stemming from their complex objectives and elaborate modules. To improve the efficiency of GAD, we introduce a simple method termed PREprocessing and Matching (PREM for short). Our approach streamlines GAD, reducing time and memory consumption while maintaining powerful anomaly detection capabilities. Comprising two modules - a pre-processing module and an ego-neighbor matching module - PREM eliminates the necessity for message-passing propagation during training, and employs a simple contrastive loss, leading to considerable reductions in training time and memory usage. Moreover, our method demonstrated robustness and effectiveness in five datasets. Notably, when validated on the ACM dataset, PREM achieved a 5% improvement in AUC, a 9-fold increase in training speed, and sharply reduce memory usage compared to the most efficient baseline. JunJun Pan, Yixin Liu 0001, Yizhen Zheng, Shirui Pan |
ICDM | 1 |
| 2023 | Multi-user upper limb rehabilitation training system integrating social interaction
Hui Liang 0004, Shiqing Liu, JunJun Pan, Yazhou Zhang 0001, Xiaohang Dong |
Comput. Graph. | 4 |
| 2023 | Virtual emotional gestures to assist in the examination of the mental health of the deaf-mutesabstractAbstract The particular characteristics of deaf‐mutes make them more likely to have mental health problems. Due to their particular way of communication, it is more difficult for them to deal with mental health problems than ordinary people. Nowadays, those psychologists who are good at sign language are in short supply, and remote assistance cannot achieve satisfactory results. Therefore, a library of virtual emotional gestures based on electroencephalogram (EEG) was established and a prototype system for mental health examination of deaf‐mutes was proposed, which help deaf‐mutes identify their psychological problems in time and assist medical staff to examine the psychological problems encountered by deaf‐mutes. In addition, the virtual library of emotional gestures is established with the assistance of the chief physician from a 3A hospital in Henan Province. More importantly, the later experiments demonstrate the applicability of this virtual system. Hui Liang 0004, JunJun Pan, Jialin Fu, Xiangwen Pang |
Comput. Animat. Virtual Worlds | 4 |
| 2023 | Virtual scene generation promotes shadow puppet art conservationabstractAbstract As an ancient performing art, shadow puppetry is a treasure of Chinese art. However, with the development of society, shadow puppetry becomes less well‐known among the young generation. To preserve and further spread this traditional culture, the digitization of shadow puppets is playing an increasingly important role in shadow puppetry conservation. Despite this, the spread of shadow puppetry culture in modern times still faces many hindrances. In digitalized shadow puppet art, the virtual scenes in certain degree determine the artistic effect of shadow puppet performance. The commonly used method for digital shadow puppet scene construction is via artificially created models which are then placed in corresponding positions. Obviously, this is a cumbersome and time‐consuming task. Therefore, a semantic‐based scene generation method for digital shadow puppet performance scene is proposed in this paper. According to this method, the key information is extracted from the descriptive text using the Chinese text segmentation technology. Meanwhile, we generate semantic scene graphs and search the corresponding shadow puppet models in the model library to construct the virtual scenes of digital shadow puppet performance. In the evaluation experiment, we invited 30 volunteers (10 female and 20 male) who had been exposed to traditional shadow puppet play in their daily lives. As suggested by the experimental results, the digital shadow puppet performance scene generated in this paper exhibit advantages of convenient use and high availability, which largely enhance the effect of digital shadow puppet performance. It should nonetheless be noted that it's not easy to extract the spatial relations from complicated texts, which inevitably limits the effectiveness of our scene generation method. The research work of this paper aims to promote digital shadow puppet technology and provide insights for the inheritance and conservation of traditional shadow puppet art. Hui Liang 0004, Xiaohang Dong, JunJun Pan |
Comput. Animat. Virtual Worlds | 3 |
| 2023 | Tucker network: Expressive power and comparison
Ye Liu 0014, JunJun Pan, Michael Kwok-Po Ng |
Neural Networks | 2 |
| 2023 | KD-Former: Kinematic and dynamic coupled transformer network for 3D human motion prediction
Ju Dai, Junxuan Bai, Feng Zhou 0007, JunJun Pan |
Pattern Recognit. | 6 |
| 2023 | Separable Quaternion Matrix Factorization for Polarization ImagesabstractAbstract. A transverse wave is a wave in which the particles are displaced perpendicular to the direction of the wave’s advance. Examples of transverse waves include ripples on the surface of water and light waves. Polarization is one of the primary properties of transverse waves. Analysis of polarization states can reveal valuable information about the sources. In this paper, we propose a separable low-rank quaternion linear mixing model for polarized signals: we assume each column of the source factor matrix equals a column of the polarized data matrix and refer to the corresponding problem as separable quaternion matrix factorization (SQMF). We discuss some properties of the matrix that can be decomposed by SQMF. To determine the source factor matrix in quaternion space, we propose a heuristic algorithm called quaternion successive projection algorithm (QSPA) inspired by the successive projection algorithm. To guarantee the effectiveness of QSPA, a new normalization operator is proposed for the quaternion matrix. We use a block coordinate descent algorithm to compute nonnegative activation matrix in real number space. We test our method on the applications of polarization image representation and spectro-polarimetric imaging unmixing to verify its effectiveness. JunJun Pan, Michael Kwok-Po Ng |
SIAM J. Imaging Sci. | 1 |
| 2023 | A lightweight pose estimation network with multi-scale receptive field
Shuo Li 0001, Ju Dai, Zhangmeng Chen, JunJun Pan |
Vis. Comput. | 4 |
| 2023 | Metaverse Virtual Social Center for the Elderly Communication During the Social DistancingabstractThe lack of social activities in the elderly for physical reasons can make them feel lonely and prone to depression. With the spread of COVID-19, it is difficult for the elderly to conduct the few social activities stably, causing the elderly to be more lonely. The metaverse is a virtual world that mirrors reality. It allows the elderly to get rid of the constraints of reality and perform social activities stably and continuously, providing new ideas for alleviating the loneliness of the elderly. Through the analysis of the needs of the elderly, a virtual social center framework for the elderly was proposed in this study. Besides, a prototype system was designed according to the framework. The elderly can socialize in virtual reality with metaverse-related technologies and human-computer interaction tools. Additionally, a test was jointly conducted with the chief physician of the geriatric rehabilitation department of a tertiary hospital. The results demonstrated that the mental state of the elderly who had used the virtual social center was significantly better than that of the elderly who had not used it. Thus, virtual social centers alleviated loneliness and depression in older adults. Virtual social centers can help the elderly relieve loneliness and depression when the global epidemic is normalizing and the population is aging. Hence, they have promotion value Hui Liang 0004, Jiupeng Li, JunJun Pan, Yazhou Zhang 0001, Xiaohang Dong |
Virtual Real. Intell. Hardw. | 4 |
| 2022 | Lane Detection Transformer Based on Multi-frame Horizontal and Vertical Attention and Visual Transformer Module
Yunchao Gu, JunJun Pan |
ECCV (39) | 4 |
| 2022 | A semantic-driven generation of 3D Chinese opera performance scenesabstractAbstract The emergence of digital opera has enriched the stage performance of Chinese opera and expanded its dissemination means. However, the modern spread of traditional Chinese opera still faces hindrances. Digital opera performances require the generation of virtual scenes of the stages and characters. However, traditional virtual scene generation requires workers to build 3D models using modeling software and incorporate them into the performance scene. This article proposes a semantic‐based generation method for Chinese opera performance scenes. First, we analyze the scene description scripts to understand the elements in Chinese opera virtual scenes. The prior probability is subsequently used to learn the model placement rules in the opera scene model. A digital scene suitable for Chinese opera performance is then generated. The final results show that the method can generate natural and receptive opera digital performance scenes. This article's research ideas and achievements are conducive to the promotion of development of Chinese digital opera technology. They possess substantial significance to the inheritance and development of traditional Chinese opera art. Hui Liang 0004, Xiaohang Dong, JunJun Pan, Jingyue Zhang, Ruicong Wang |
Comput. Animat. Virtual Worlds | 4 |
| 2022 | Intelligent recognition of portrait sketch components for child autism assessmentabstractAbstract For autistic children with slow language function, it is a classic and easy way to understand the development of their cognitive ability through simple painting experiments. Due to the lack of professional evaluators for painting assessment of autistic children, this article research and implement an intelligent assessment system for child autism through recognition of portrait sketch components. A portrait sketch database is constructed with the sample size of 30,400 after data expansion. The data consists of two formats: stroke vector sequence and 2D image. Then we propose a joint model coupled with LSTM and CNN features to automatically segment the portrait sketch components. It can perform better segmentation for exaggerated proportion and incomplete components samples. Finally, according to the evaluation criteria of the painting, we design an assessment model for child autism. The experiments are conducted in cooperation with relevant rehabilitation institutions to verify the effectiveness of the system. The analysis results show that our painting assessment system has a good ability to identify autistic tendencies. It can accurately evaluate children's autistic tendencies through “draw‐a‐man” experiments. Yang Shen 0009, Zhangmeng Chen, Hui Liang 0004, JunJun Pan |
Comput. Animat. Virtual Worlds | 7 |
| 2022 | Unsupervised Hyperspectral Pansharpening by Ratio Estimation and Residual Attention NetworkabstractMost deep learning-based hyperspectral pansharpening methods use the hyperspectral images (HSIs) as the ground truth. Training samples are usually obtained by blurring and downsampling the panchromatic image and HSI. However, the blurring and downsampling operation lose much spatial and spectral information. As a result, the model parameters trained by these reduced-resolution samples are unsuitable for fusing full-resolution images. To tackle this problem, we propose an unsupervised hyperspectral pansharpening method via ratio estimation (RE) and residual attention network (RE-RANet). The spatial and spectral information of the fused image are derived from the original panchromatic and HSI rather than reduced-resolution images. At first, we generate the initial ratio image using the ratio enhancement method. The initial ratio image is fine-tuned by the residual attention network (RANet) to generate a multichannel ratio image. Then, we inject the multichannel ratio image that contains spatial detail information into the HSI. Finally, the generated hyperspectral image is constrained by the spatial constraint loss and the spectral constraint loss. Experiments on the EO-1 and Chikusei datasets verify the effectiveness of the proposed method. Compared with other state-of-the-art approaches, our method performs well in qualitative visual effects and quantitative evaluation indicators. Jinyan Nie, Qizhi Xu, JunJun Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Hyperspectral Image Classification Based on Multiscale Spectral-Spatial Deformable NetworkabstractImage classification plays a fundamental role in hyperspectral image (HSI) analysis. Since the mixed pixels of the urban areas are generally more complex than other areas, the following two problems remain to be considered while dealing with urban HSI classification: 1) due to the fact that the spectral feature of different mixed pixels varies greatly in the same class, HSI classification of urban area is susceptible to the representativeness of the training samples and 2) since the urban area is densely packed with objects of different size, the comprehensive use of the spatial and spectral features to classify the objects is a difficult problem. To tackle these problems, HSI classification based on multiscale spectral–spatial deformable network (S2-DNet) is proposed. First, a$k$-means clustering method is adopted to cluster spectrum of each class, and representative samples are selected from the spectrum after clustering to reduce the impact of intraclass variation. Second, a spectral–spatial joint network is designed to extract the low-level features, including spectral features and spatial features. Third, the deformable network is introduced to extract high-level features of the object. Experimental results demonstrated that the proposed method outperformed the state-of-the-art methods on two widely used HSI data sets. Jinyan Nie, Qizhi Xu, JunJun Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Neurophysiological and Subjective Analysis of VR Emotion Induction ParadigmabstractThe ecological validity of emotion-inducing scenarios is essential for emotion research. In contrast to the classical passive induction paradigm, immersive VR fully engages the psychological and physiological components of the subject, which is considered an ecologically valid paradigm for studying emotion. Several studies investigate the emotional responses to different VR tasks or games using subjective scales. However, little research regards VR as an eliciting material, especially when systematically analyzing emotional processes in VR from a neurophysiological perspective. To fill this gap and scientifically evaluate VR's ability to be used as an active method for emotion elicitation, we investigate the dynamic relationship between explicit information (subjective evaluations) and implicit information (objective neurophysiological data). A total of 28 participants are enlisted to watch eight VR videos while their SAM/IPQ scores and EEG data are recorded simultaneously. In ecologically valid scenarios, the subjective results demonstrate that VR has significant advantages for evoking emotion in arousal-valence. This conclusion is backed by our examination of objective neurophysiological evidence that VR videos effectively induce high-arousal emotions. In addition, we obtain features of critical channels and frequency oscillations associated with emotional valence, thereby validating previous research in more lifelike circumstances. In particular, we discover hemispheric asymmetry in the occipital region under high and low emotional arousal, which adds to our understanding of neural features and the dynamics of emotional arousal. As a result, we successfully integrate EEG and VR to demonstrate that VR is more pragmatic for evoking natural feelings and is beneficial for emotional research. Our research has set a precedent for new methodologies of using VR induction paradigms to acquire a more reliable explanation of affective computing. JunJun Pan, Yang Gao 0032, Yang Shen 0009, Ju Dai, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | SCSF-Net: Single Class Scale Fixed Network for Object Detection in Optical Remote Sensing Images on Limited HardwareabstractThe detection of objects such as vehicle, airplane and ship is a fundamental problem in optical remote-sensing(ORS) image process. Despite a great success has achieved by migrating nature image detection methods to the remote sensing field, some challenges in hardware limit environments still remain to be solved, e.g., space-borne hardware and UAV-borne hardware. We proposed a low-computational network by digging several prior knowledge in the remote sensing field. By focusing on certain ground sample distance(gsd) and single target class, the proposed method gains high performance with only less than 1% parameters and less than 1% computation used comparing with the state-of-the-art detection method. Detection result on public available vehicle dataset demonstrates the effectiveness of the proposed method. Meanwhile, the ship and airplane detection results of two private datasets are also shown. Our vehicle detection code on limited hardware is now available at https://github.com/minghuicode/scsf-detector. Qingpeng Li, JunJun Pan, Yunchao Gu |
IGARSS | 3 |
| 2021 | Diabetic Retinopathy Grading Base on Contrastive Learning and Semi-supervised Learning
Yunchao Gu, JunJun Pan, Zhong Zhou |
ISBRA | 3 |
| 2021 | Diverse Dance Synthesis via Keyframes with Transformer ControllersabstractAbstract Existing keyframe‐based motion synthesis mainly focuses on the generation of cyclic actions or short‐term motion, such as walking, running, and transitions between close postures. However, these methods will significantly degrade the naturalness and diversity of the synthesized motion when dealing with complex and impromptu movements , e.g., dance performance and martial arts. In addition, current research lacks fine‐grained control over the generated motion, which is essential for intelligent human‐computer interaction and animation creation. In this paper, we propose a novel keyframe‐based motion generation network based on multiple constraints, which can achieve diverse dance synthesis via learned knowledge. Specifically, the algorithm is mainly formulated based on the recurrent neural network (RNN) and the Transformer architecture. The backbone of our network is a hierarchical RNN module composed of two long short‐term memory (LSTM) units, in which the first LSTM is utilized to embed the posture information of the historical frames into a latent space, and the second one is employed to predict the human posture for the next frame. Moreover, our framework contains two Transformer‐based controllers, which are used to model the constraints of the root trajectory and the velocity factor respectively, so as to better utilize the temporal context of the frames and achieve fine‐grained motion control. We verify the proposed approach on a dance dataset containing a wide range of contemporary dance. The results of three quantitative analyses validate the superiority of our algorithm. The video and qualitative experimental results demonstrate that the complex motion sequences generated by our algorithm can achieve diverse and smooth motion transitions between keyframes, even for long‐term synthesis. JunJun Pan, Junxuan Bai, Ju Dai |
Comput. Graph. Forum | 1 |
| 2021 | EmoDescriptor: A hybrid feature for emotional classification in dance movementsabstractAbstract Similar to language and music, dance performances provide an effective way to express human emotions. With the abundance of the motion capture data, content‐based motion retrieval and classification have been fiercely investigated. Although researchers attempt to interpret body language in terms of human emotions, the progress is limited by the scarce 3D motion database annotated with emotion labels. This article proposes a hybrid feature for emotional classification in dance performances. The hybrid feature is composed of an explicit feature and a deep feature. The explicit feature is calculated based on the Laban movement analysis, which considers the body, effort, shape, and space properties. The deep feature is obtained from latent representation through a 1D convolutional autoencoder. Eventually, we present an elaborate feature fusion network to attain the hybrid feature that is almost linearly separable. The abundant experiments demonstrate that our hybrid feature is superior to the separate features for the emotional classification in dance performances. Junxuan Bai, Rong Dai, Ju Dai, JunJun Pan |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Generalized Separable Nonnegative Matrix FactorizationabstractNonnegative matrix factorization (NMF) is a linear dimensionality technique for nonnegative data with applications such as image analysis, text mining, audio source separation, and hyperspectral unmixing. Given a data matrix M and a factorization rank r, NMF looks for a nonnegative matrix W with r columns and a nonnegative matrix H with r rows such that M ≈ WH. NMF is NP-hard to solve in general. However, it can be computed efficiently under the separability assumption which requires that the basis vectors appear as data points, that is, that there exists an index set K such that W = M(:,K). In this article, we generalize the separability assumption. We only require that for each rank-one factor W(:,k)H(k,:) for k=1,2,…,r, either W(:,k) = M(:,j) for some j or H(k,:) = M(i,:) for some i. We refer to the corresponding problem as generalized separable NMF (GS-NMF). We discuss some properties of GS-NMF and propose a convex optimization model which we solve using a fast gradient method. We also propose a heuristic algorithm inspired by the successive projection algorithm. To verify the effectiveness of our methods, we compare them with several state-of-the-art separable NMF and standard NMF algorithms on synthetic, document and image data sets. JunJun Pan, Nicolas Gillis |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Augmented reality-based visual-haptic modeling for thoracoscopic surgery training systemsabstractCompared with traditional thoracotomy, video-assisted thoracoscopic surgery (VATS) has less minor trauma, faster recovery, higher patient compliance, but higher requirements for surgeons. Virtual surgery training simulation systems are important and have been widely used in Europe and America. Augmented reality (AR) in surgical training simulation systems significantly improve the training effect of virtual surgical training, although AR technology is still in its initial stage. Mixed reality has gained increased attention in technology-driven modern medicine but has yet to be used in everyday practice. This study proposed an immersive AR lobectomy within a thoracoscope surgery training system, using visual and haptic modeling to study the potential benefits of this critical technology. The content included immersive AR visual rendering, based on the cluster-based extended position-based dynamics algorithm of soft tissue physical modeling. Furthermore, we designed an AR haptic rendering systems, whose model architecture consisted of multi-touch interaction points, including kinesthetic and pressure-sensitive points. Finally, based on the above theoretical research, we developed an AR interactive VATS surgical training platform. Twenty-four volunteers were recruited from the First People's Hospital of Yunnan Province to evaluate the VATS training system. Face, content, and construct validation methods were used to assess the tactile sense, visual sense, scene authenticity, and simulator performance. The results of our construction validation demonstrate that the simulator is useful in improving novice and surgical skills that can be retained after a certain period of time. The video-assisted thoracoscopic system based on AR developed in this study is effective and can be used as a training device to assist in the development of thoracoscopic skills for novices. Yonghang Tai, Junsheng Shi, JunJun Pan, Aimin Hao, Victor Chang 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Flower Factory: A Component-based Approach for Rapid Flower ModelingabstractThe rapid 3D objects modeling provides an effective way to enrich digital content, which is one of the essential tasks in VR/AR research. Flowers are frequently utilized in real-time applications, such as video games and VR/AR scenes. Technically, a realistic flower generation using the existing 3D modeling software is complicated and time-consuming for designers. Moreover, it is difficult to create imaginary and surreal flowers, which might be more interesting and attractive for the artists and game players. In this paper, we propose a component-based framework for rapid flower modeling, called Flower Factory. The flowers are assembled by different components, e.g., petals, stamens, receptacles and leaves. The shape of these components are created using simple primitives such as points and splines. After the shape of models are determined, the textures are synthesized automatically based on a predefine mask, according to a number of rules from real flowers. The whole modeling process can be controlled by several parameters, which describe the physical attributes of the flowers. Our technique is capable of producing a variety of flowers rapidly. Even novices without any modeling skills are able to control and model the 3D flowers. Furthermore, the developed system will be integrated in a lightweight application of smartphone due to its low computational cost. JunJun Pan, Junxuan Bai, Jinglei Wang |
ISMAR | 2 |
| 2020 | Real-time VR Simulation of Laparoscopic Cholecystectomy based on Parallel Position-based Dynamics in GPUabstractIn recent years, virtual reality (VR) based training has greatly changed surgeons learning mode. It can simulate the surgery from the visual, auditory, and tactile aspects. VR medical simulator can greatly reduce the risk of the real patient and the cost of hospitals. Laparoscopic cholecystectomy is one of the typical representatives in minimal invasive surgery (MIS). Due to the large incidence of cholecystectomy, the application of its VR-based simulation is vital and necessary for the residents' surgical training. In this paper, we present a VR simulation framework based on position-based dynamics (PBD) for cholecystectomy. To further accelerate the deformation of organs, PBD constraints are solved in parallel by a graph coloring algorithm. We introduce a bio-thermal conduction model to improve the realism of the fat tissue electrocautery. Finally, we design a hybrid multi-model connection method to handle the interaction and simulation of the liver-gallbladder separation. This simulation system has been applied to laparoscopic cholecystectomy training in several hospitals. From the experimental results, users can operate in real-time with high stability and fidelity. The simulator is also evaluated by a number of digestive surgeons through preliminary studies. They believed that the system can offer great help to the improvement of surgical skills. JunJun Pan, Leiyu Zhang, Yang Shen 0009, Haimin Hao, Hong Qin 0001 |
VR | 1 |
| 2020 | Hybrid features for skeleton-based action recognition based on network fusionabstractAbstract In recent years, the topic of skeleton‐based human action recognition has attracted significant attention from researchers and practitioners in graphics, vision, animation, and virtual environments. The most fundamental issue is how to learn an effective and accurate representation from spatiotemporal action sequences towards improved performance, and this article aims to address the aforementioned challenge. In particular, we design a novel method of hybrid features' extraction based on the construction of multistream networks and their organic fusion. First, we train a convolution neural networks (CNN) model to learn CNN‐based features with the raw skeleton coordinates and their temporal differences serving as input signals. The attention mechanism is injected into the CNN model to weigh more effective and important information. Then, we employ long short‐term memory (LSTM) to obtain long‐term temporal features from action sequences. Finally, we generate the hybrid features by fusing the CNN and LSTM networks, and we classify action types with the hybrid features. The extensive experiments are performed on several large‐scale publically available databases, and promising results demonstrate the efficacy and effectiveness of our proposed framework. Zhangmeng Chen, JunJun Pan, Xiaosong Yang, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2020 | Real-time suturing simulation for virtual reality medical trainingabstractAbstract At present, virtual reality (VR) ‐based medical simulators provide an efficient and cost‐effective alternative without exposing risk to the traditional training approaches. As an essential and indispensable task in fundamental surgical skills training, the research of suturing simulation still remains insufficient in the field of virtual surgery. In this paper, we present a real‐time suturing simulation framework which can handle the complex interactions between surgical instruments and soft tissue. The simulation consists of two stages: external interaction and internal coupling. External interaction involves the interplay between needle/suture and the soft tissue, which are both deformed by position‐based dynamics (PBD) with different constraints. At the internal coupling stage, once the force exceeds a threshold, the needle tip will puncture and penetrate into the soft tissue and generate a path. To guarantee the needle/suture accurately following the path inside the soft tissue, we propose a novel coupling method by matching and generating the constraints among needle, suture, and penetration path. We have applied this suturing simulation into a VR laparoscopic surgery simulator with haptic force. Our experimental results demonstrate that our approach can achieve real‐time performance with a high degree of visual realism and haptic fidelity. JunJun Pan, Hong Qin 0001, Aimin Hao |
Comput. Animat. Virtual Worlds | 2 |
| 2020 | Fast character modeling with sketch-based PDE surfacesabstractAbstract Virtual characters are 3D geometric models of characters. They have a lot of applications in multimedia. In this paper, we propose a new physics-based deformation method and efficient character modelling framework for creation of detailed 3D virtual character models. Our proposed physics-based deformation method uses PDE surfaces. Here PDE is the abbreviation of Partial Differential Equation, and PDE surfaces are defined as sculpting force-driven shape representations of interpolation surfaces. Interpolation surfaces are obtained by interpolating key cross-section profile curves and the sculpting force-driven shape representation uses an analytical solution to a vector-valued partial differential equation involving sculpting forces to quickly obtain deformed shapes. Our proposed character modelling framework consists of global modeling and local modeling. The global modeling is also called model building, which is a process of creating a whole character model quickly with sketch-guided and template-based modeling techniques. The local modeling produces local details efficiently to improve the realism of the created character model with four shape manipulation techniques. The sketch-guided global modeling generates a character model from three different levels of sketched profile curves called primary, secondary and key cross-section curves in three orthographic views. The template-based global modeling obtains a new character model by deforming a template model to match the three different levels of profile curves. Four shape manipulation techniques for local modeling are investigated and integrated into the new modelling framework. They include: partial differential equation-based shape manipulation, generalized elliptic curve-driven shape manipulation, sketch assisted shape manipulation, and template-based shape manipulation. These new local modeling techniques have both global and local shape control functions and are efficient in local shape manipulation. The final character models are represented with a collection of surfaces, which are modeled with two types of geometric entities: generalized elliptic curves (GECs) and partial differential equation-based surfaces. Our experiments indicate that the proposed modeling approach can build detailed and realistic character models easily and quickly. Lihua You, Xiaosong Yang, JunJun Pan, Tong-Yee Lee, Shaojun Bian, Kun Qian 0009, Zulfiqar Habib, Allah Bux Sargano, Ismail Khalid Kazmi, Jian J. Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Specular Reflections Removal for Endoscopic Image Sequences With Adaptive-RPCA DecompositionabstractSpecular reflections (i.e., highlight) always exist in endoscopic images, and they can severely disturb surgeons' observation and judgment. In an augmented reality (AR)-based surgery navigation system, the highlight may also lead to the failure of feature extraction or registration. In this paper, we propose an adaptive robust principal component analysis (Adaptive-RPCA) method to remove the specular reflections in endoscopic image sequences. It can iteratively optimize the sparse part parameter during RPCA decomposition. In this new approach, we first adaptively detect the highlight image based on pixels. With the proposed distance metric algorithm, it then automatically measures the similarity distance between the sparse result image and the detected highlight image. Finally, the low-rank and sparse results are obtained by enforcing the similarity distance between the two types of images to fall within a certain range. Our method has been verified by multiple different types of endoscopic image sequences in minimally invasive surgery (MIS). The experiments and clinical blind tests demonstrate that the new Adaptive-RPCA method can obtain the optimal sparse decomposition parameters directly and can generate robust highlight removal results. Compared with the state-of-the-art approaches, the proposed method not only achieves the better highlight removal results but also can adaptively process image sequences. Ranyang Li, JunJun Pan, Yaqing Si, Hong Qin 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Real-Time Tracking of Corneal Contour in Dalk Surgical Navigation Using Deep Neural NetworksabstractCorneal disease is one of the most common causes of blindness for human beings in the world. Deep anterior lamellar k-eratoplasty (DALK) is a widely-used corneal transplantation technique, which requires precise control of surgical tools. This paper proposes a deep learning framework of augmented reality (AR) based surgical navigation to guide the suturing process in DALK. It aims to track the cutting corneal contour robustly through semantic segmentation and occlusion reconstruction. We devise a novel optical flow inpainting network to restore the missing motion caused by occlusion. The occluded regions are obtained using weakly-supervised segmentation of surgical tools and reconstructed by the key-frame warping along the completed optical flow. We introduce two kinds of loss functions to adapt the inpainting network to the optical flow space. The performance of our techniques is evaluated using real surgery videos from Shandong Eye Hospital. All experimental results show that our approach can achieve accurate corneal contour tracking subject to complex disturbance of tools in real-time surgical scenarios. Pu Ge, JunJun Pan, Fanghong Li, Weiyun Shi, Hong Qin 0001 |
ICIP | 2 |
| 2019 | Accurate and Fast Classification of Foot Gestures for Virtual LocomotionabstractThis work explores the use of foot gestures for locomotion in virtual environments. Foot gestures are represented as the distribution of plantar pressure and detected by three sparsely-located sensors on each insole. The Long Short-Term Memory model is chosen as the classifier to recognize the performer's foot gesture based on the captured signals of pressure information. The trained classifier directly takes the noisy and sparse input of sensor data, and handles seven categories of foot gestures (stand, walk forward/backward, run, jump, slide left and right) without manual definition of signal features for classifying these gestures. This classifier is capable of recognizing the foot gestures, even with the existence of large sensor-specific, inter-person and intra-person variations. Results show that an accuracy of ~80% can be achieved across different users with different shoe sizes and ~85% for users with the same shoe size. A novel method, Dual-Check Till Consensus, is proposed to reduce the latency of gesture recognition from 2 seconds to 0.5 seconds and increase the accuracy to over 97%. This method offers a promising solution to achieve lower latency and higher accuracy at a minor cost of computation workload. The characteristics of high accuracy and fast classification of our method could lead to wider applications of using foot patterns for human-computer interaction. JunJun Pan, Zeyong Hu, Juncong Lin, Shihui Guo, Minghong Liao |
ISMAR | 2 |
| 2019 | Real-time Animation and Motion Retargeting of Virtual Characters Based on Single RGB-D CameraabstractThe rapid generation and flexible reuse of characters animation by commodity devices are of significant importance to rich digital content production in virtual reality. This paper aims to handle the challenges of current motion imitation for human body in several indoor scenes (e.g., fitness training). We develop a real-time system based on single Kinect device, which is able to capture stable human motions and retarget to virtual characters. A large variety of motions and characters are tested to validate the efficiency and effectiveness of our system. Ning Kang 0006, Junxuan Bai, JunJun Pan, Hong Qin 0001 |
VR | 3 |
| 2019 | Interactive animation generation of virtual characters using single RGB-D camera
Ning Kang 0006, Junxuan Bai, JunJun Pan, Hong Qin 0001 |
Vis. Comput. | 3 |
| 2019 | Real-time simulation of electrocautery procedure using meshfree methods in laparoscopic cholecystectomy
JunJun Pan, Yang Gao 0032, Hong Qin 0001, Yaqing Si |
Vis. Comput. | 1 |
| 2018 | Structured Convex Optimization Method for Orthogonal Nonnegative Matrix FactorizationabstractOrthogonal nonnegative matrix factorization plays an important role for data clustering and machine learning. In this paper, we propose a new optimization model for orthogonal nonnegative matrix factorization based on the structural properties of orthogonal nonnegative matrix. The new model can be solved by a novel convex relaxation technique which can be employed quite efficiently. Numerical examples in document clustering, image segmentation and hyperspectral unmixing are used to test the performance of the proposed model. The performance of our method is better than the other testing methods in terms of clustering accuracy. JunJun Pan, Michael Kwok-Po Ng, Xiongjun Zhang |
ICPR | 1 |
| 2018 | Real-time fish animation generation by monocular camera
Xiangfei Meng, JunJun Pan, Hong Qin 0001, Pu Ge |
Comput. Graph. | 2 |
| 2018 | Novel metaballs-driven approach with dynamic constraints for character articulation
Junxuan Bai, JunJun Pan, Hong Qin 0001 |
Sci. China Inf. Sci. | 2 |
| 2018 | Automatic skinning and weight retargeting of articulated characters using extended position-based dynamics
JunJun Pan, Hong Qin 0001 |
Vis. Comput. | 1 |
| 2018 | Real-time dissection of organs via hybrid coupling of geometric metaballs and physics-centric mesh-free method
JunJun Pan, Shizeng Yan, Hong Qin 0001, Aimin Hao |
Vis. Comput. | 1 |
| 2017 | Motion Capture and Retargeting of Fish by Monocular CameraabstractAccurate motion capture and flexible retargeting of underwater creatures such as fish remain to be difficult due to the long-lasting challenges of marker attachment and feature description for soft bodies in the underwater environment. Despite limited new research progresses appeared in recent years, the fish motion retargeting with a desirable motion pattern in real-time remains elusive. Strongly motivated by our ambitious goal of achieving high-quality data-driven fish animation with a light-weight, mobile device, this paper develops a novel framework of motion capturing and retargeting for a fish. We capture the motion of actual fish by a monocular camera without the utility of any marker. The elliptical Fourier coefficients are then integrated into the contour-based feature extraction process to analyze the fish swimming patterns. This novel approach can obtain the motion information in a robust way, with smooth medial axis as the descriptor for a soft fish body. For motion retargeting, we propose a two-level scheme to properly transfer the captured motion into new models, such as 2D meshes (with texture) generated from pictures or 3D models designed by artists, regardless of different body geometry and fin proportions among various species. Both motion capture and retargeting processes are functioning in real time. Hence, the system can simultaneously create fish animation with variation, while obtaining video sequences of real fish by a monocular camera. Xiangfei Meng, JunJun Pan, Hong Qin 0001 |
CW | 2 |
| 2017 | Essential techniques for laparoscopic surgery simulationabstractAbstract Laparoscopic surgery is a complex minimum invasive operation that requires long learning curve for the new trainees to have adequate experience to become a qualified surgeon. With the development of virtual reality technology, virtual reality‐based surgery simulation is playing an increasingly important role in the surgery training. The simulation of laparoscopic surgery is challenging because it involves large non‐linear soft tissue deformation, frequent surgical tool interaction and complex anatomical environment. Current researches mostly focus on very specific topics (such as deformation and collision detection) rather than a consistent and efficient framework. The direct use of the existing methods cannot achieve high visual/haptic quality and a satisfactory refreshing rate at the same time, especially for complex surgery simulation. In this paper, we proposed a set of tailored key technologies for laparoscopic surgery simulation, ranging from the simulation of soft tissues with different properties, to the interactions between surgical tools and soft tissues to the rendering of complex anatomical environment. Compared with the current methods, our tailored algorithms aimed at improving the performance from accuracy, stability and efficiency perspectives. We also abstract and design a set of intuitive parameters that can provide developers with high flexibility to develop their own simulators. Copyright © 2016 John Wiley & Sons, Ltd. Kun Qian 0009, Junxuan Bai, Xiaosong Yang, JunJun Pan, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2015 | Virtual reality based laparoscopic surgery simulationabstractWith the development of computer graphic and haptic devices, training surgeons with virtual reality technology has proven to be very effective in surgery simulation. Many successful simulators have been deployed for training medical students. However, due to the various unsolved technical issues, the laparoscopic surgery simulation has not been widely used. Such issues include modeling of complex anatomy structure, large soft tissue deformation, frequent surgical tools interactions, and the rendering of complex material under the illumination of headlight. A successful laparoscopic surgery simulator should integrate all these required components in a balanced and efficient manner to achieve both visual/haptic quality and a satisfactory refreshing rate. In this paper, we propose an efficient framework integrating a set of specially tailored and designed techniques, ranging from deformation simulation, collision detection, soft tissue dissection and rendering. We optimize all the components based on the actual requirement of laparoscopic surgery in order to achieve an improved overall performance of fidelity and responding speed. Kun Qian 0009, Junxuan Bai, Xiaosong Yang, JunJun Pan, Jian J. Zhang 0001 |
VRST | 4 |
| 2015 | Real-time haptic manipulation and cutting of hybrid soft tissue models by extended position-based dynamicsabstractAbstract This paper systematically describes an interactive dissection approach for hybrid soft tissue models governed by extended position‐based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh with high‐precision geometric structure and texture at the detailed level is employed to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral mesh, which is also employed as physical model participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost during simulation. For physical deformation, we design and implement an extended position‐based dynamics approach that supports topology modification and material heterogeneities of soft tissue. Besides stretching and volume conservation constraints, it enforces the energy preserving constraints, which take the different spring stiffness of material into account and improve the visual performance of soft tissue deformation. Furthermore, we develop mechanical modeling of dissection behavior and analyze the system stability. The experimental results have shown that our approach affords real‐time and robust cutting without sacrificing realistic visual performance. Our novel dissection technique has already been integrated into a virtual reality‐based laparoscopic surgery simulator. Copyright © 2015 John Wiley & Sons, Ltd. JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Metaballs-based physical modeling and deformation of organs for virtual surgery
JunJun Pan, Chengkai Zhao, Xin Zhao 0025, Aimin Hao, Hong Qin 0001 |
Vis. Comput. | 1 |
| 2014 | Dissection of hybrid soft tissue models using position-based dynamicsabstractThis paper describes an interactive dissection approach for hybrid soft tissue models governed by position-based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh is used to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral meshes, which are also employed as physical models participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost of deformation and geometric subdivision during dissection. For physical deformation, we design and implement a position-based dynamics approach that supports topology modification and enforces the volume-preserving constraint. Experimental results have shown that, this hybrid dissection method affords real-time and robust cutting simulation without sacrificing realistic visual performance. JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001 |
VRST | 1 |
| 2009 | Automatic rigging for animation characters with 3D silhouetteabstractAbstract Animating an articulated 3D character requires the specification of its interior skeleton structure which defines how the skin surface is deformed during animation. Currently this task is to a large extent accomplished manually, which consumes a large amount of animators' time. This paper presents an automatic rigging method making use of a new geometry entity called the 3D silhouette. The first step is to extract a coarse 3D curve skeleton and some skeletal joints of a character. This curve skeleton is then refined with a perpendicular silhouette. According to the connectivity of the skeletal joints, the hierarchical animation skeleton is finally constructed. By avoiding complicated computation such as voxelization and pruning, this method is simple and efficient, much faster than existing methods. It proves very useful for quick animation production, with applications including games design and prototype graphical systems. Copyright © 2009 John Wiley & Sons, Ltd. JunJun Pan, Xiaosong Yang, Philip J. Willis, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 1 |