VLDB 2026 Research / reviewers in the wild / expert
Xiaoming Deng 0001
dblp:96/2411-1
· DBLP profile ↗
50ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-8576-1201ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 27 · 4 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCoRE: Standardized Human Evaluation Provides a Reliable Measure for Semantic Consistency of Text-to-Image Generation
Zejian Li, Qi Liu 0076, Jiaman Pan, Lefan Hou, Xiangfei Hu, Jiarui Ma, Shengyuan Zhang, Jiesi Zhang, Xuetao Tian, Xiaoming Deng 0001 |
Int. J. Comput. Vis. | 14 |
| 2025 | HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data AugmentationabstractUnderstanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, which prevents further improvement of bimanual hand-object interaction-related baselines. In this work, we propose a new 3D Gaussian Splatting based data augmentation framework for bimanual hand-object interaction, which is capable of augmenting existing dataset to large-scale photorealistic data with various hand-object pose and viewpoints. First, we use mesh-based 3DGS to model objects and hands, and to deal with the rendering blur problem due to multi-resolution input images used, we design a super-resolution module. Second, we extend the single hand grasping pose optimization module for the bimanual hand object to generate various poses of bimanual hand-object interaction, which can significantly expand the pose distribution of the dataset. Third, we conduct an analysis for the impact of different aspects of the proposed data augmentation on the understanding of the bimanual hand-object interaction. We perform our data augmentation on two benchmarks, H2O and Arctic, and verify that our method can improve the performance of the baselines. Wentian Qu, Jiahe Li 0006, Jian Cheng 0006, Chenyu Meng, CuiXia Ma, Hongan Wang, Xiaoming Deng 0001, Yinda Zhang 0001 |
AAAI | 8 |
| 2025 | Universal Features Guided Zero-Shot Category-Level Object Pose EstimationabstractObject pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to establish semantic similarity-based correspondences and can be extended to unseen categories without additional model fine-tuning. Our method begins with combining efficient 2D universal features to find sparse correspondences between intra-category objects and gets initial coarse pose. To handle the correspondence degradation of 2D universal features if the pose deviates much from the target pose, we use an iterative strategy to optimize the pose. Subsequently, to resolve pose ambiguities due to shape differences between intra-category objects, the coarse pose is refined by optimizing with dense alignment constraint of 3D universal features. Our method outperforms previous methods on the REAL275 and Wild6D benchmarks for unseen categories. Wentian Qu, Chenyu Meng, Heng Li 0009, Jian Cheng 0006, CuiXia Ma, Hongan Wang, Xiao Zhou 0023, Xiaoming Deng 0001, Ping Tan 0002 |
AAAI | 8 |
| 2025 | DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelabstractGenerating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in different sizes and their different motion trajectories, while ensuring strong grasping realism and guaranteeing the coordination of movement in all body parts. Contrasting with existing work, which either generates human interaction motion sequences without detailed hand grasping poses or only models a static grasping pose, we propose a simple yet effective framework that jointly models the relationship between the body, hands, and the given object motion sequences within a single diffusion model. To guide our network in perceiving the object's spatial position and learning more natural grasping poses, we introduce novel contact-aware losses and incorporate a data-driven, carefully designed guidance. Experimental results demonstrate that our approach outperforms the state-of-the-art method and generates plausible results. Yonghao Zhang 0002, Yanguang Wan, Yinda Zhang 0001, Xiaoming Deng 0001, CuiXia Ma, Hongan Wang |
AAAI | 5 |
| 2025 | Sketch-Guided Scene-Level Image Editing with Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Yaokun Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
CVM (2) | 3 |
| 2025 | SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Cangjun Gao, Yaxian Shan, Haoxiang Hu, Qingkun Li, Xiaoming Deng 0001, CuiXia Ma, Yukun Lai, Yong-Jin Liu 0001, Feng Tian 0001, Guozhong Dai, Hongan Wang |
UIST | 6 |
| 2024 | SpaceGTN: A Time-Agnostic Graph Transformer Network for Handwritten Diagram Recognition and SegmentationabstractOnline handwriting recognition is pivotal in domains like note-taking, education, healthcare, and office tasks. Existing diagram recognition algorithms mainly rely on the temporal information of strokes, resulting in a decline in recognition performance when dealing with notes that have been modified or have no temporal information. The current datasets are drawn based on templates and cannot reflect the real free-drawing situation. To address these challenges, we present SpaceGTN, a time-agnostic Graph Transformer Network, leveraging spatial integration and removing the need for temporal data. Extensive experiments on multiple datasets have demonstrated that our method consistently outperforms existing methods and achieves state-of-the-art performance. We also propose a pipeline that seamlessly connects offline and online handwritten diagrams. By integrating a stroke restoration technique with SpaceGTN, it enables intelligent editing of previously uneditable offline diagrams at the stroke level. In addition, we have also launched the first online handwritten diagram dataset, OHSD, which is collected using a free-drawing method and comes with modification annotations. Haoxiang Hu, Cangjun Gao, Yaokun Li, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
AAAI | 4 |
| 2024 | Complementing Event Streams and RGB Frames for Hand Mesh ReconstructionabstractReliable hand mesh reconstruction (HMR) from commonly-used color and depth sensors is challenging es-pecially under scenarios with varied illuminations and fast motions. Event camera is a highly promising alternative for its high dynamic range and dense temporal resolution properties, but it lacks salient texture appearance for hand mesh reconstruction. In this paper, we propose EvRGBHand - the first approach for 3D hand mesh reconstruction with an event camera and an RGB camera compensating for each other. By fusing two modalities of data across time, space, and information dimensions, EvRGBHand can tackle overexposure and motion blur issues in RGB-based HMR and foreground scarcity as well as background overflow issues in event-based HMR. We further propose EvRGBDegrader, which allows our model to generalize effectively in challenging scenes, even when trained solely on standard scenes, thus reducing data acquisition costs. Experiments on real-world data demonstrate that EvRGBHand can effectively solve the challenging issues when using either type of camera alone via retaining the merits of both, and shows the potential of generalization to outdoor scenes and another type of event camera. For code, models, and dataset, please refer to https://alanjiang98.github.io/.evrgbhand.github.io/. Bingxuan Wang, Xiaoming Deng 0001, Boxin Shi |
CVPR | 4 |
| 2024 | SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Cangjun Gao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IJCAI | 3 |
| 2024 | EvHandPose: Event-Based 3D Hand Pose Estimation With Sparse SupervisionabstractEvent camera shows great potential in 3D hand pose estimation, especially addressing the challenges of fast motion and high dynamic range in a low-power way. However, due to the asynchronous differential imaging mechanism, it is challenging to design event representation to encode hand motion information especially when the hands are not moving (causing motion ambiguity), and it is infeasible to fully annotate the temporally dense event stream. In this paper, we propose EvHandPose with novel hand flow representations in Event-to-Pose module for accurate hand pose estimation and alleviating the motion ambiguity issue. To solve the problem under sparse annotation, we design contrast maximization and hand-edge constraints in Pose-to-IWE (Image with Warped Events) module and formulate EvHandPose in a weakly-supervision framework. We further build EvRealHands, the first large-scale real-world event-based hand pose dataset on several challenging scenes to bridge the real-synthetic domain gap. Experiments on EvRealHands demonstrate that EvHandPose outperforms previous event-based methods under all evaluation scenes, achieves accurate and stable hand pose estimation with high temporal resolution in fast motion and strong light scenes compared with RGB-based methods, generalizes well to outdoor scenes and another type of event camera, and shows the potential for the hand gesture recognition task. Jiahe Li 0006, Baowen Zhang, Xiaoming Deng 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking EffectivenessabstractAs communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack of evidence-based analytical systems for people to comprehensively evaluate online speeches and further discover possibilities for improvement. This paper introduces SpeechMirror, a visual analytics system facilitating reflection on a speech based on insights from a collection of online speeches. The system estimates the impact of different speech techniques on effectiveness and applies them to a speech to give users awareness of the performance of speech techniques. A similarity recommendation approach based on speech factors or script content supports guided exploration to expand knowledge of presentation evidence and accelerate the discovery of speech delivery possibilities. SpeechMirror provides intuitive visualizations and interactions for users to understand speech factors. Among them, SpeechTwin, a novel multimodal visual summary of speech, supports rapid understanding of critical speech factors and comparison of different speech samples, and SpeechPlayer augments the speech video by integrating visualization of the speaker's body language with interaction, for focused analysis. The system utilizes visualizations suited to the distinct nature of different speech factors for user comprehension. The proposed system and visualization techniques were evaluated with domain experts and amateurs, demonstrating usability for users with low visualization literacy and its efficacy in assisting users to develop insights for potential improvement. Kevin T. Maher, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Sheng Feng Qin, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Novel-view Synthesis and Pose Estimation for Hand-Object Interaction from Sparse ViewsabstractHand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose estimation system for hand-object interaction from sparse views, which can also enable 3D hand-object interaction editing. We share the inspiration from recent scene understanding work that shows a scene specific model built beforehand can significantly improve and unblock vision tasks especially when inputs are sparse, and extend it to the dynamic hand-object interaction scenario and propose to solve the problem in two stages. We first learn the shape and appearance prior knowledge of hands and objects separately with the neural representation at the offline stage. During the online stage, we design a rendering-based joint model fitting framework to understand the dynamic hand-object interaction with the pre-built hand and object models as well as interaction priors, which thereby overcomes penetration and separation issues between hand and object and also enables novel view synthesis. In order to get stable contact during the hand-object interaction process in a sequence, we propose a stable contact loss to make the contact region to be consistent. Experiments demonstrate that our method outperforms the state-of-the-art methods. Code and dataset are available in project web-page https://iscas3dv.github.io/HO-NeRF. Wentian Qu, Zhaopeng Cui, Yinda Zhang 0001, Chenyu Meng, CuiXia Ma, Xiaoming Deng 0001, Hongan Wang |
ICCV | 6 |
| 2023 | Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable ObjectsabstractLearning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton poses for human bodies or animals, which require extra annotation effort and suffer from error accumulation, and they are limited to specific domain. In this paper, we propose a novel self-supervised approach to learn neural implicit shape representation for deformable objects, which can represent shapes with a template shape and dense correspondence in 3D. Our method does not require the priors of skeleton and skinning weight, and only requires a collection of shapes represented in signed distance fields. To handle the large deformation, we constrain the learned template shape in the same latent space with the training shapes, design a new formulation of local rigid constraint that enforces rigid transformation in local region and addresses local reflection issue, and present a new hierarchical rigid constraint to reduce the ambiguity due to the joint learning of template shape and correspondences. Extensive experiments show that our model can represent shapes with large deformations. We also show that our shape representation can support two typical applications, such as texture transfer and shape editing, with competitive performance. The code and models are available at https://iscas3dv.github.io/deformshape. Baowen Zhang, Jiahe Li 0006, Xiaoming Deng 0001, Yinda Zhang 0001, CuiXia Ma, Hongan Wang |
ICCV | 3 |
| 2023 | Recurrent 3D Hand Pose Estimation Using Cascaded Pose-Guided 3D Alignmentsabstract3D hand pose estimation is a challenging problem in computer vision due to the high degrees-of-freedom of hand articulated motion space and large viewpoint variation. As a consequence, similar poses observed from multiple views can be dramatically different. In order to deal with this issue, view-independent features are required to achieve state-of-the-art performance. In this paper, we investigate the impact of view-independent features on 3D hand pose estimation from a single depth image, and propose a novel recurrent neural network for 3D hand pose estimation, in which a cascaded 3D pose-guided alignment strategy is designed for view-independent feature extraction and a recurrent hand pose module is designed for modeling the dependencies among sequential aligned features for 3D hand pose estimation. In particular, our cascaded pose-guided 3D alignments are performed in 3D space in a coarse-to-fine fashion. First, hand joints are predicted and globally transformed into a canonical reference frame; Second, the palm of the hand is detected and aligned; Third, local transformations are applied to the fingers to refine the final predictions. The proposed recurrent hand pose module for aligned 3D representation can extract recurrent pose-aware features and iteratively refines the estimated hand pose. Our recurrent module could be utilized for both single-view estimation and sequence-based estimation with 3D hand pose tracking. Experiments show that our method improves the state-of-the-art by a large margin on popular benchmarks with the simple yet efficient alignment and network architectures. Xiaoming Deng 0001, Dexin Zuo, Yinda Zhang 0001, Zhaopeng Cui, Jian Cheng 0006, Ping Tan 0002, Liang Chang 0001, Marc Pollefeys, Sean Ryan Fanello, Hongan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | MMPosE: Movie-Induced Multi-Label Positive Emotion Classification Through EEG SignalsabstractEmotional information plays an important role in various multimedia applications. Movies, as a widely available form of multimedia content, can induce multiple positive emotions and stimulate people's pursuit of a better life. Different from negative emotions, positive emotions are highly correlated and difficult to distinguish in the emotional space. Since different positive emotions are often induced simultaneously by movies, traditional single-target or multi-class methods are not suitable for the classification of movie-induced positive emotions. In this paper, we proposeTransEEG, a model for multi-label positive emotion classification from a viewer's brain activities when watching emotional movies. The key features ofTransEEGinclude (1) explicitly modeling the spatial correlation and temporal dependencies of multi-channel EEG signals using the Transformer structure based model, which effectively addresses long-distance dependencies, (2) exploiting the label-label correlations to guide the discriminative EEG representation learning, for that we design an Inter-Emotion Mask for guiding the Multi-Head Attention to learn the inter-emotion correlations, and (3) constructing an attention score vector from the representation-label correlation matrix to refine emotion-relevant EEG features. To evaluate the ability of our model for multi-label positive emotion classification, we demonstrate our model on a state-of-the-art positive emotion database CPED. Extensive experimental results show that our proposed method achieves superior performance over the competitive approaches. Xiaobing Du, Xiaoming Deng 0001, Hangyu Qin, Yezhi Shu, Fang Liu 0035, Guozhen Zhao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Fine-Grained Video Retrieval With Scene SketchesabstractBenefiting from the intuitiveness and naturalness of sketch interaction, sketch-based video retrieval (SBVR) has received considerable attention in the video retrieval research area. However, most existing SBVR research still lacks the capability of accurate video retrieval with fine-grained scene content. To address this problem, in this paper we investigate a new task, which focuses on retrieving the target video by utilizing a fine-grained storyboard sketch depicting the scene layout and major foreground instances' visual characteristics (e.g., appearance, size, pose, etc.) of video; we call such a task "fine-grained scene-level SBVR". The most challenging issue in this task is how to perform scene-level cross-modal alignment between sketch and video. Our solution consists of two parts. First, we construct a scene-level sketch-video dataset called SketchVideo, in which sketch-video pairs are provided and each pair contains a clip-level storyboard sketch and several keyframe sketches (corresponding to video frames). Second, we propose a novel deep learning architecture called Sketch Query Graph Convolutional Network (SQ-GCN). In SQ-GCN, we first adaptively sample the video frames to improve video encoding efficiency, and then construct appearance and category graphs to jointly model visual and semantic alignment between sketch and video. Experiments show that our fine-grained scene-level SBVR framework with SQ-GCN architecture outperforms the state-of-the-art fine-grained retrieval methods. The SketchVideo dataset and SQ-GCN code are available in the project webpage https://iscas-mmsketch.github.io/FG-SL-SBVR/. Ran Zuo, Xiaoming Deng 0001, Yukun Lai, Fang Liu 0035, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 2 |
| 2023 | Stroke-based semantic segmentation for scene-level free-hand sketches
Xiaoming Deng 0001, Jinyao Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
Vis. Comput. | 2 |
| 2022 | Efficient Virtual View Selection for 3D Hand Pose Estimationabstract3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual view selection and fusion module for 3D hand pose estimation from single depth. We propose to automatically select multiple virtual viewpoints for pose estimation and fuse the results of all and find this empirically delivers accurate and robust pose estimation. In order to select most effective virtual views for pose fusion, we evaluate the virtual views based on the confidence of virtual views using a light-weight network via network distillation. Experiments on three main benchmark datasets including NYU, ICVL and Hands2019 demonstrate that our method outperforms the state-of-the-arts on NYU and ICVL, and achieves very competitive performance on Hands2019-Task1, and our proposed virtual view selection and fusion module is both effective for 3D hand pose estimation. Jian Cheng 0006, Yanguang Wan, Dexin Zuo, CuiXia Ma, Ping Tan 0002, Hongan Wang, Xiaoming Deng 0001, Yinda Zhang 0001 |
AAAI | 8 |
| 2022 | An Efficient LSTM Network for Emotion Recognition From Multichannel EEG SignalsabstractMost previous EEG-based emotion recognition methods studied hand-crafted EEG features extracted from different electrodes. In this article, we study the relation among different EEG electrodes and propose a deep learning method to automatically extract the spatial features that characterize the functional relation between EEG signals at different electrodes. Our proposed deep model is calledATtention-basedLSTMwithDomainDiscriminator (ATDD-LSTM), a model based on Long Short-Term Memory (LSTM) for emotion recognition that can characterize nonlinear relations among EEG signals of different electrodes. To achieve state-of-the-art emotion recognition performance, the architecture of ATDD-LSTM has two distinguishing characteristics: (1) By applying the attention mechanism to the feature vectors produced by LSTM, ATDD-LSTM automatically selects suitable EEG channels for emotion recognition, which makes the learned model concentrate on the emotion related channels in response to a given emotion; (2) To minimize the significant feature distribution shift between different sessions and/or subjects, ATDD-LSTM uses a domain discriminator to modify the data representation space and generate domain-invariant features. We evaluate the proposed ATDD-LSTM model on three public EEG emotional databases (DEAP, SEED and CMEED) for emotion recognition. The experimental results demonstrate that our ATDD-LSTM model achieves superior performance on subject-dependent (for the same subject), subject-independent (for different subjects) and cross-session (for the same subject) evaluation. Xiaobing Du, CuiXia Ma, Jinyao Li, Yukun Lai, Guozhen Zhao, Xiaoming Deng 0001, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 7 |
| 2022 | SketchMaker: Sketch Extraction and Reuse for Interactive Scene Sketch CompositionabstractSketching is an intuitive and simple way to depict sciences with various object form and appearance characteristics. In the past few years, widely available touchscreen devices have increasingly made sketch-based human-AI co-creation applications popular. One key issue of sketch-oriented interaction is to prepare input sketches efficiently by non-professionals because it is usually difficult and time-consuming to draw an ideal sketch with appropriate outlines and rich details, especially for novice users with no sketching skills. Thus, sketching brings great obstacles for sketch applications in daily life. On the other hand, hand-drawn sketches are scarce and hard to collect. Given the fact that there are several large-scale sketch datasets providing sketch data resources, but they usually have a limited number of objects and categories in sketch, and do not support users to collect new sketch materials according to their personal preferences. In addition, few sketch-related applications support the reuse of existing sketch elements. Thus, knowing how to extract sketches from existing drawings and effectively re-use them in interactive scene sketch composition will provide an elegant way for sketch-based image retrieval (SBIR) applications, which are widely used in various touch screen devices. In this study, we first conduct a study on current SBIR to better understand the main requirements and challenges in sketch-oriented applications. Then we develop the SketchMaker as an interactive sketch extraction and composition system to help users generate scene sketches via reusing object sketches in existing scene sketches with minimal manual intervention. Moreover, we demonstrate how SBIR improves from composited scene sketches to verify the performance of our interactive sketch processing system. We also include a sketch-based video localization task as an alternative application of our sketch composition scheme. Our pilot study shows that our system is effective and efficient, and provides a way to promote practical applications of sketches. Fang Liu 0035, Xiaoming Deng 0001, Jian-Cheng Song, Yukun Lai, Yong-Jin Liu 0001, Hao Wang 0005, CuiXia Ma, Sheng Feng Qin, Hongan Wang |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2022 | SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNsabstractSketch-based image retrieval (SBIR) is a long-standing research topic in computer vision. Existing methods mainly focus on category-level or instance-level image retrieval. This paper investigates the fine-grained scene-level SBIR problem where a free-hand sketch depicting a scene is used to retrieve desired images. This problem is useful yet challenging mainly because of two entangled facts: 1) achieving an effective representation of the input query data and scene-level images is difficult as it requires to model the information across multiple modalities such as object layout, relative size and visual appearances, and 2) there is a great domain gap between the query sketch input and target images. We present SceneSketcher-v2, a Graph Convolutional Network (GCN) based architecture to address these challenges. SceneSketcher-v2 employs a carefully designed graph convolution network to fuse the multi-modality information in the query sketch and target images and uses a triplet training process and end-to-end training manner to alleviate the domain gap. Extensive experiments demonstrate SceneSketcher-v2 outperforms state-of-the-art scene-level SBIR models with a significant margin. Fang Liu 0035, Xiaoming Deng 0001, Changqing Zou, Yukun Lai, Ran Zuo, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 2 |
| 2022 | E-ffective: A Visual Analytic System for Exploring the Emotion and Effectiveness of Inspirational SpeechesabstractWhat makes speeches effective has long been a subject for debate, and until today there is broad controversy among public speaking experts about what factors make a speech effective as well as the roles of these factors in speeches. Moreover, there is a lack of quantitative analysis methods to help understand effective speaking strategies. In this paper, we propose E-ffective, a visual analytic system allowing speaking experts and novices to analyze both the role of speech factors and their contribution in effective speeches. From interviews with domain experts and investigating existing literature, we identified important factors to consider in inspirational speeches. We obtained the generated factors from multi-modal data that were then related to effectiveness data. Our system supports rapid understanding of critical factors in inspirational speeches, including the influence of emotions by means of novel visualization methods and interaction. Two novel visualizations include E-spiral (that shows the emotional shifts in speeches in a visually compact way) and E-script (that connects speech content with key speech delivery information). In our evaluation we studied the influence of our system on experts' domain knowledge about speech factors. We further studied the usability of the system by speaking novices and experts on assisting analysis of inspirational speech effectiveness. Kevin T. Maher, Jian-Cheng Song, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageabstractIn this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solution space. In order to address the occlusion and similar appearance between hands that may confuse the network, we design a hand pose-aware attention module to extract features associated to each individual hand respectively. We then leverage the two hand context presented in interaction to propose a context-aware cascaded refinement that improves the hand pose and shape accuracy of each hand conditioned on the context between interacting hands. Extensive experiments on the main benchmark datasets demonstrate that our method predicts accurate 3D hand pose and shape from single color image, and achieves the state-of-the-art performance. Code is available in project webpage https://baowenz.github.io/Intershape/. Baowen Zhang, Yangang Wang 0001, Xiaoming Deng 0001, Yinda Zhang 0001, Ping Tan 0002, CuiXia Ma, Hongan Wang |
ICCV | 3 |
| 2021 | Sequential 3D Human Pose Estimation Using Adaptive Point Cloud Sampling Strategyabstract3D human pose estimation is a fundamental problem in artificial intelligence, and it has wide applications in AR/VR, HCI and robotics. However, human pose estimation from point clouds still suffers from noisy points and estimated jittery artifacts because of handcrafted-based point cloud sampling and single-frame-based estimation strategies. In this paper, we present a new perspective on the 3D human pose estimation method from point cloud sequences. To sample effective point clouds from input, we design a differentiable point cloud sampling method built on density-guided attention mechanism. To avoid the jitter caused by previous 3D human pose estimation problems, we adopt temporal information to obtain more stable results. Experiments on the ITOP dataset and the NTU-RGBD dataset demonstrate that all of our contributed components are effective, and our method can achieve state-of-the-art performance. Lei Hu 0008, Xiaoming Deng 0001, Shihong Xia |
IJCAI | 3 |
| 2021 | Hand Pose Understanding With Large-Scale Photo-Realistic Rendering DatasetabstractHand pose understanding is essential to applications such as human computer interaction and augmented reality. Recently, deep learning based methods achieve great progress in this problem. However, the lack of high-quality and large-scale dataset prevents the further improvement of hand pose related tasks such as 2D/3D hand pose from color and depth from color. In this paper, we develop a large-scale and high-quality synthetic dataset, PBRHand. The dataset contains millions of photo-realistic rendered hand images and various ground truths including pose, semantic segmentation, and depth. Based on the dataset, we firstly investigate the effect of rendering methods and used databases on the performance of three hand pose related tasks: 2D/3D hand pose from color, depth from color and 3D hand pose from depth. This study provides insights that photo-realistic rendering dataset is worthy of synthesizing and shows that our new dataset can improve the performance of the state-of-the-art on these tasks. This synthetic data also enables us to explore multi-task learning, while it is expensive to have all the ground truth available on real data. Evaluations show that our approach can achieve state-of-the-art or competitive performance on several public datasets. Xiaoming Deng 0001, Yinda Zhang 0001, Yuying Zhu 0002, Dachuan Cheng, Dexin Zuo, Zhaopeng Cui, Ping Tan 0002, Liang Chang 0001, Hongan Wang |
IEEE Trans. Image Process. | 1 |
| 2021 | Weakly Supervised Learning for Single Depth-Based Hand Shape RecoveryabstractRecent emerging technologies such AR/VR and HCI are drawing high demand on more comprehensive hand shape understanding, requiring not only 3D hand skeleton pose but also hand shape geometry. In this paper, we propose a deep learning framework to produce 3D hand shape from a single depth image. To address the challenge that capturing ground truth 3D hand shape in the training dataset is non-trivial, we leverage synthetic data to construct a statistical hand shape model and adopt weak supervision from widely accessible hand skeleton pose annotation. To bridge the gap due to the different hand skeleton definitions in the existing public datasets, we propose a joint regression network for hand pose adaptation. To reconstruct the hand shape, we use Chamfer loss between the predicted hand shape and the point cloud from the input depth to learn the shape reconstruction model in a weakly-supervised manner. Experiments demonstrate that our model adapts well to the real data and produces accurate hand shapes that outperform the state-of-the-art methods both qualitatively and quantitatively. Xiaoming Deng 0001, Yuying Zhu 0002, Yinda Zhang 0001, Zhaopeng Cui, Ping Tan 0002, Wentian Qu, CuiXia Ma, Hongan Wang |
IEEE Trans. Image Process. | 1 |
| 2020 | SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches
Fang Liu 0035, Changqing Zou, Xiaoming Deng 0001, Ran Zuo, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
ECCV (19) | 3 |
| 2020 | Image-based Pose Representation for Action Recognition and Hand Gesture RecognitionabstractIn this paper, we propose an effective and compact image-based pose representation named Poseimage Pyramid, which encodes the spatial and temporal information of human pose or hand pose as an image pyramid. Poseimage is constructed by the normalized distance between pairwise joints, and it has the advantage of its invariant to similarity transformations. With our Poseimage representation we can design the pose based action recognition or hand gesture recognition model using existing image or video classification models. In order to adapt to different actions with a variety of movement speed, we design Poseimage Pyramid to encode the multi-scale temporal information of human pose or hand pose. Experiments demonstrate that our pose representation is effective, and we achieve state-of-the-art performance on the action recognition datasets and the hand gesture recognition datasets. Our pose presentation is also complementary to video and optical flow streams in the seminal action recognition network I3D, and we achieve the state-of-the-art performance on the JHMDB, HMDB and UCF101 datasets by integrating our pose representation with I3D. Zeyi Lin, Xiaoming Deng 0001, CuiXia Ma, Hongan Wang |
FG | 3 |
| 2020 | Face-sketch learning with human sketch-drawing order enforcement
Liang Chang 0001, Lihua Jin, Lifen Weng, Wentao Chao, Xiaoming Deng 0001, Qiulei Dong |
Sci. China Inf. Sci. | 6 |
| 2020 | Leveraging 3D blendshape for facial expression recognition using CNN
Sa Wang, Zhengxin Cheng, Xiaoming Deng 0001, Liang Chang 0001, Fuqing Duan, Ke Lu 0002 |
Sci. China Inf. Sci. | 3 |
| 2020 | Edge-guided single facial depth map super-resolution using CNNabstractIn recent years, consumer depth cameras have been widely used in digital entertainment and human‐machine interaction due to the advantages of real‐time performance and low cost. Facial depth maps have shown great potential in 3D‐face‐related studies. However, disadvantages of low resolution and precision limit its further applications. In this work, the authors propose an edge‐guided convolutional neural network for single facial depth map super‐resolution. It consists of two parts: an edge prediction sub‐network and a depth reconstruction sub‐network. The edge prediction sub‐network generates an edge guidance map to guide the depth reconstruction sub‐network to recover sharp edges and fine structures. Effective data augmentation methods are proposed as well. The network is patch‐based and able to cope with any size of the input depth maps. In addition, it is insensitive to the face pose since the synthetic training dataset they generated covers a wide range of face poses. The proposed method is validated with three datasets including a synthetic facial depth data set, a real Kinect V2 facial depth data set and Middlebury Stereo Data set. Experimental results show that it outperforms the state‐of‐the‐art methods on all the three data sets. Fan Zhang 0062, Liang Chang 0001, Fuqing Duan, Xiaoming Deng 0001 |
IET Image Process. | 5 |
| 2020 | Weakly Supervised Adversarial Learning for 3D Human Pose Estimation from Point CloudsabstractPoint clouds-based 3D human pose estimation that aims to recover the 3D locations of human skeleton joints plays an important role in many AR/VR applications. The success of existing methods is generally built upon large scale data annotated with 3D human joints. However, it is a labor-intensive and error-prone process to annotate 3D human joints from input depth images or point clouds, due to the self-occlusion between body parts as well as the tedious annotation process on 3D point clouds. Meanwhile, it is easier to construct human pose datasets with 2D human joint annotations on depth images. To address this problem, we present a weakly supervised adversarial learning framework for 3D human pose estimation from point clouds. Compared to existing 3D human pose estimation methods from depth images or point clouds, we exploit both the weakly supervised data with only annotations of 2D human joints and fully supervised data with annotations of 3D human joints. In order to relieve the human pose ambiguity due to weak supervision, we adopt adversarial learning to ensure the recovered human pose is valid. Instead of using either 2D or 3D representations of depth images in previous methods, we exploit both point clouds and the input depth image. We adopt 2D CNN to extract 2D human joints from the input depth image, 2D human joints aid us in obtaining the initial 3D human joints and selecting effective sampling points that could reduce the computation cost of 3D human pose regression using point clouds network. The used point clouds network can narrow down the domain gap between the network input i.e. point clouds and 3D joints. Thanks to weakly supervised adversarial learning framework, our method can achieve accurate 3D human pose from point clouds. Experiments on the ITOP dataset and EVAL dataset demonstrate that our method can achieve state-of-the-art performance efficiently. Lei Hu 0008, Xiaoming Deng 0001, Shihong Xia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | STA-GCN: two-stream graph convolutional network with spatial-temporal attention for hand gesture recognition
Zeyi Lin, Jian Cheng 0006, CuiXia Ma, Xiaoming Deng 0001, Hongan Wang |
Vis. Comput. | 5 |
| 2019 | SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial NetworkabstractHand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often incomplete, which makes sketch recognition a challenging problem. In this paper, we propose SketchGAN, a new generative adversarial network (GAN) based approach that jointly completes and recognizes a sketch, boosting the performance of both tasks. Specifically, we use a cascade Encode-Decoder network to complete the input sketch in an iterative manner, and employ an auxiliary sketch recognition task to recognize the completed sketch. Experiments on the Sketchy database benchmark demonstrate that our joint learning approach achieves competitive sketch completion and recognition performance compared with the state-of-the-art methods. Further experiments using several sketch-based applications also validate the performance of our method. Fang Liu 0035, Xiaoming Deng 0001, Yukun Lai, Yong-Jin Liu 0001, CuiXia Ma, Hongan Wang |
CVPR | 2 |
| 2019 | Cascaded Point Network for 3D Hand Pose EstimationabstractRecent PointNet-family hand pose methods have the advantages of high pose estimation performance and small model size, and it is a key problem to get effective sample points for PointNet-family methods. In this paper, we propose a two-stage coarse to fine hand pose estimation method, which belongs to PointNet-family methods and explores a new sample point strategy. In the first stage, we use 3D coordinate and surface normal of normalized point cloud as input to regress coarse hand joints. In the second stage, we use the hand joints in the first stage as the initial sample points to refine the hand joints. Experiments on widely used datasets demonstrate that using joints as sample points is more effective and our method achieves top-rank performance. Yikun Dou, Yuying Zhu 0002, Xiaoming Deng 0001, CuiXia Ma, Liang Chang 0001, Hongan Wang |
ICASSP | 4 |
| 2019 | High-Fidelity Face Sketch-To-Photo Synthesis Using Generative Adversarial NetworkabstractFace sketch-photo synthesis has important usage in law enforcement and human authentication. Due to the sparse information (no color or texture), the abstraction level, the diversity of sketches, and the domain gap between sketch and photo, it is challenging to synthesize a photo-realistic photo from an input sketch. Moreover, the deficiency of data also restricts the synthesis performance. In this paper, we present a high-fidelity face sketch-photo synthesis method using Generation Adversarial Network (GAN). Our network adopts a deep residual U-Net as generator and a Patch-GAN with residual blocks as discriminator. We design effective loss functions by enforcing pixels, edges and high-level features of the produced face photos. Moreover, we augment the CUHK sketch dataset using an effective sampling method. With the improved GAN and augmented dataset, we achieve high-fidelity face photos. Qualitative and quantitative experiments demonstrate the approach outperforms other method. Further experiments with a sketch-based photo editing application also validate the performance of our method. Wentao Chao, Liang Chang 0001, Jian Cheng 0006, Xiaoming Deng 0001, Fuqing Duan |
ICIP | 5 |
| 2018 | Text2Sketch: Learning Face Sketch from Facial Attribute TextabstractFace sketch is the main approach to find suspect in law enforcement, especially in many cases when facial attribute descriptions of suspects by witnesses are available. Face sketch synthesized from facial attribute text can also be used in sketch based face recognition. While most previous work focus on face photo to sketch synthesis, the problem of sketch synthesis with facial attribute text has not been explored yet. The problem is challenging due to two facts: firstly, no database of face attribute text to sketch is available; secondly, it is hard to synthesize high-quality face sketches due to the ambiguity and complexity of text description. In this paper, we propose a face sketch synthesis approach with text using Stagewise-GAN. Our contributions lie in two aspects: 1) we construct the first text to face sketch database. The database, namely Text2Sketch dataset, is annotated with CUFSF dataset of 1194 sketches. For each sketch, an attribute description is labelled; 2) we synthesize vivid face sketches using Stagewise-GAN. We use user study, face retrieval performance with synthesized sketch, and quantitative results for evaluation. Experimental results show the effectiveness of our approach. Liang Chang 0001, Lihua Jin, Zhengxin Cheng, Xiaoming Deng 0001, Fuqing Duan |
ICIP | 6 |
| 2018 | Learning Scene Illumination by Pairwise Photos from Rear and Front Mobile CamerasabstractAbstract Illumination estimation is an essential problem in computer vision, graphics and augmented reality. In this paper, we propose a learning based method to recover low‐frequency scene illumination represented as spherical harmonic (SH) functions by pairwise photos from rear and front cameras on mobile devices. An end‐to‐end deep convolutional neural network (CNN) structure is designed to process images on symmetric views and predict SH coefficients. We introduce a novel Render Loss to improve the rendering quality of the predicted illumination. A high quality high dynamic range (HDR) panoramic image dataset was developed for training and evaluation. Experiments show that our model produces visually and quantitatively superior results compared to the state‐of‐the‐arts. Moreover, our method is practical for mobile‐based applications. Dachuan Cheng, Yanyun Chen, Xiaoming Deng 0001, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 4 |
| 2018 | Joint Hand Detection and Rotation Estimation Using CNNabstractHand detection is essential for many hand related tasks, e.g., recovering hand pose and understanding gesture. However, hand detection in uncontrolled environments is challenging due to the flexibility of wrist joint and cluttered background. We propose a convolutional neural network (CNN), which formulates in-plane rotation explicitly to solve hand detection and rotation estimation jointly. Our network architecture adopts the backbone of faster R-CNN to generate rectangular region proposals and extract local features. The rotation network takes the feature as input and estimates an in-plane rotation which manages to align the hand, if any in the proposal, to the upward direction. A derotation layer is then designed to explicitly rotate the local spatial feature map according to the rotation network and feed aligned feature map for detection. Experiments show that our method outperforms the state-of-the-art detection models on widely-used benchmarks, such as Oxford and Egohands database. Further analysis show that rotation estimation and classification can mutually benefit each other. Xiaoming Deng 0001, Yinda Zhang 0001, Shuo Yang 0002, Ping Tan 0002, Liang Chang 0001, Hongan Wang |
IEEE Trans. Image Process. | 1 |
| 2015 | Face sketch synthesis using non-local means and patch-based seamingabstractThis paper proposed a face sketch synthesis method by using non-local means (NL-Means), which takes the advantage of the non-local self-similarity of face photo and sketch patches. With a learning database of individuals described by one face photo and one face sketch, we assume that, for a given individual, the NL-Means coefficient of a given face photo patch is the same as its corresponding sketch patch. In order to handle the visible seam due to intensity difference of neighbor overlapping patches, we use patch based optimal seam to enforce the consistency of synthesized overlapping sketch patches. Experimental results on CUHK Face Sketch Database illustrate that our method has the advantage of easy implementation and much less required training samples, meanwhile our method can achieve fairly competitive synthesis results. Liang Chang 0001, Yves Rozenholc, Xiaoming Deng 0001, Fuqing Duan |
ICIP | 3 |
| 2014 | Motion estimation of multiple depth cameras using spheresabstractAutomatic motion estimation of multiple depth cameras has remained a challenging topic in computer vision due to its reliance on the image correspondence problem. In this paper, spherical objects are employed to estimate motion parameters between multiple depth cameras. We move a sphere several times in the common view of depth cameras. We fit the spherical point clouds to get the sphere centers in each depth camera system, and then introduce a factorization based approach to estimate motions between the depth cameras. Both simulated and real experiments show the robustness and effectiveness of our method. Xiaoming Deng 0001, Jie Liu 0029, Feng Tian 0001, Liang Chang 0001, Hongan Wang |
ICIP | 1 |
| 2014 | Automatic Gait Motion Capture with Missing-Marker FillingsabstractAlthough marker-based optical motion capture has been a useful method for computer animation during the past decades, automatic and robust motion tracking from multiple video sequences is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to track and identify the reconstructed 3D points after image matching process? How to handle the heavy occlusion problem? This paper gives a careful investigation of the above issues. In particular, we propose a novel way to track and identify proper markers, and a new method of filling missing markers by taking account of the human model constraints. Experiments are presented to show its accuracy and robustness. Xiaoming Deng 0001, Shihong Xia, Wenzhong Wang, Liang Chang 0001, Hongan Wang |
ICPR | 1 |
| 2012 | Smoothness-constrained face photo-sketch synthesis using sparse representation
Liang Chang 0001, Xiaoming Deng 0001, Fuqing Duan, Zhongke Wu |
ICPR | 2 |
| 2012 | Self-calibration of hybrid central catadioptric and perspective cameras
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Fuqing Duan, Liang Chang 0001, Hongan Wang |
Comput. Vis. Image Underst. | 1 |
| 2012 | Calibrating effective focal length for central catadioptric cameras using one space line
Fuqing Duan, Fuchao Wu, Xiaoming Deng 0001, Yun Tian 0002 |
Pattern Recognit. Lett. | 4 |
| 2011 | Calibration of central catadioptric camera with one-dimensional object undertaking general motionsabstractAID object is a segment with several known-distance markers, and calibration methods with ID objects are more flexible than those with 2D/3D objects. Under the pinhole camera model, it is proved that the calibration with free-moving ID objects is not possible. For a central catadioptric camera setup, can the camera be calibrated by a ID object under general motions? In this paper, we prove that a central catadioptric camera can indeed be calibrated, and propose a catadioptric camera calibration method using ID objects undertaking general motions. The proposed method consists of two steps. Firstly, the principal point is calculated with geometric invariants under catadioptric camera model; Secondly, we use images of ID object to calibrate the focal lengths, skew factor and mirror parameter. The method needs neither prior knowledge of catadioptric parameters nor conic fitting, and it is linear, which makes it easy to implement. Experiments demonstrate its usefulness and stability. Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Liang Chang 0001, Wei Liu 0023, Hongan Wang |
ICIP | 1 |
| 2010 | Face Sketch Synthesis via Sparse RepresentationabstractFace sketch synthesis with a photo is challenging due to that the psychological mechanism of sketch generation is difficult to be expressed precisely by rules. Current learning-based sketch synthesis methods concentrate on learning the rules by optimizing cost functions with low-level image features. In this paper, a new face sketch synthesis method is presented, which is inspired by recent advances in sparse signal representation and neuroscience that human brain probably perceives images using high-level features which are sparse. Sparse representations are desired in sketch synthesis due to that sparseness can adaptively selects the most relevant samples which give best representations of the input photo. We assume that the face photo patch and its corresponding sketch patch follow the same sparse representation. In the feature extraction, we select succinct high-level features by using the sparse coding technique, and in the sketch synthesis process each sketch patch is synthesized with respect to high-level features by solving an l1-norm optimization. Experiments have been given on CUHK database to show that our method can resemble the true sketch fairly well. Liang Chang 0001, Yanjun Han, Xiaoming Deng 0001 |
ICPR | 4 |
| 2009 | Learning local models for 2D human motion trackingabstractWe present a novel approach to tracking 2D human motion in uncalibrated monocular videos. Human motion usually exhibits time-varying patterns, and we propose to use locally learnt prior models to capture this characteristics. For each input image, our method automatically learns a local probability density model and a local dynamical model from a set of training examples that are close matches to the input. We evaluate the image likelihood by matching a deformable 2D human body model to the input images. The local models and the image likelihood are integrated to optimize the pose for the current input. Experiments on both synthetic and real videos demonstrate the effectiveness of our method. Wenzhong Wang, Xiaoming Deng 0001, Xianjie Qiu, Shihong Xia |
ICIP | 2 |
| 2008 | Visual metrology with uncalibrated radial distorted imagesabstractVisual metrology methods with radial distorted images usually require a radial distortion model and a pre-calibration. In this paper, we propose a novel 3D metrology algorithm with at least three uncalibrated radial distorted images, and also derive a 2D metrology algorithm with at least two uncalibrated radial distorted images. The algorithm does not require a radial distortion model or calibrating camera intrinsic parameters except for radial distortion center, which can be usually known as a prior or computed easily, and correspondences of control points with known coordinates. The algorithm is of high accuracy, and robust to noise due to no requirement for specific radial distortion models. Experimental results show the feasibility and accuracy of the algorithm. Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Fuqing Duan |
ICPR | 1 |
| 2008 | A new normalized method on line-based homography estimation
Xiaoming Deng 0001, Zhanyi Hu |
Pattern Recognit. Lett. | 2 |