EDBT 2026 Demo / reviewers in the wild / expert
CuiXia Ma
dblp:55/7026 · also Cui-Xia Ma, Cuixia Ma
· DBLP profile ↗
44ranked-venue papers
7as first author
22since 2021 · last 2025
0000-0003-3999-7429ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 12 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data AugmentationabstractUnderstanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, which prevents further improvement of bimanual hand-object interaction-related baselines. In this work, we propose a new 3D Gaussian Splatting based data augmentation framework for bimanual hand-object interaction, which is capable of augmenting existing dataset to large-scale photorealistic data with various hand-object pose and viewpoints. First, we use mesh-based 3DGS to model objects and hands, and to deal with the rendering blur problem due to multi-resolution input images used, we design a super-resolution module. Second, we extend the single hand grasping pose optimization module for the bimanual hand object to generate various poses of bimanual hand-object interaction, which can significantly expand the pose distribution of the dataset. Third, we conduct an analysis for the impact of different aspects of the proposed data augmentation on the understanding of the bimanual hand-object interaction. We perform our data augmentation on two benchmarks, H2O and Arctic, and verify that our method can improve the performance of the baselines. Wentian Qu, Jiahe Li 0006, Jian Cheng 0006, Chenyu Meng, CuiXia Ma, Hongan Wang, Xiaoming Deng 0001, Yinda Zhang 0001 |
AAAI | 6 |
| 2025 | Universal Features Guided Zero-Shot Category-Level Object Pose EstimationabstractObject pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to establish semantic similarity-based correspondences and can be extended to unseen categories without additional model fine-tuning. Our method begins with combining efficient 2D universal features to find sparse correspondences between intra-category objects and gets initial coarse pose. To handle the correspondence degradation of 2D universal features if the pose deviates much from the target pose, we use an iterative strategy to optimize the pose. Subsequently, to resolve pose ambiguities due to shape differences between intra-category objects, the coarse pose is refined by optimizing with dense alignment constraint of 3D universal features. Our method outperforms previous methods on the REAL275 and Wild6D benchmarks for unseen categories. Wentian Qu, Chenyu Meng, Heng Li 0009, Jian Cheng 0006, CuiXia Ma, Hongan Wang, Xiao Zhou 0023, Xiaoming Deng 0001, Ping Tan 0002 |
AAAI | 5 |
| 2025 | DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelabstractGenerating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in different sizes and their different motion trajectories, while ensuring strong grasping realism and guaranteeing the coordination of movement in all body parts. Contrasting with existing work, which either generates human interaction motion sequences without detailed hand grasping poses or only models a static grasping pose, we propose a simple yet effective framework that jointly models the relationship between the body, hands, and the given object motion sequences within a single diffusion model. To guide our network in perceiving the object's spatial position and learning more natural grasping poses, we introduce novel contact-aware losses and incorporate a data-driven, carefully designed guidance. Experimental results demonstrate that our approach outperforms the state-of-the-art method and generates plausible results. Yonghao Zhang 0002, Yanguang Wan, Yinda Zhang 0001, Xiaoming Deng 0001, CuiXia Ma, Hongan Wang |
AAAI | 6 |
| 2025 | Sketch-Guided Scene-Level Image Editing with Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Yaokun Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
CVM (2) | 6 |
| 2025 | SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Cangjun Gao, Yaxian Shan, Haoxiang Hu, Qingkun Li, Xiaoming Deng 0001, CuiXia Ma, Yukun Lai, Yong-Jin Liu 0001, Feng Tian 0001, Guozhong Dai, Hongan Wang |
UIST | 7 |
| 2024 | SpaceGTN: A Time-Agnostic Graph Transformer Network for Handwritten Diagram Recognition and SegmentationabstractOnline handwriting recognition is pivotal in domains like note-taking, education, healthcare, and office tasks. Existing diagram recognition algorithms mainly rely on the temporal information of strokes, resulting in a decline in recognition performance when dealing with notes that have been modified or have no temporal information. The current datasets are drawn based on templates and cannot reflect the real free-drawing situation. To address these challenges, we present SpaceGTN, a time-agnostic Graph Transformer Network, leveraging spatial integration and removing the need for temporal data. Extensive experiments on multiple datasets have demonstrated that our method consistently outperforms existing methods and achieves state-of-the-art performance. We also propose a pipeline that seamlessly connects offline and online handwritten diagrams. By integrating a stroke restoration technique with SpaceGTN, it enables intelligent editing of previously uneditable offline diagrams at the stroke level. In addition, we have also launched the first online handwritten diagram dataset, OHSD, which is collected using a free-drawing method and comes with modification annotations. Haoxiang Hu, Cangjun Gao, Yaokun Li, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
AAAI | 6 |
| 2024 | SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Cangjun Gao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IJCAI | 7 |
| 2024 | EEG-Based Evaluation of Aesthetic Experience Using BiLSTM NetworkabstractEvaluation of aesthetic design fulfills a pivotal function in product development, which urges for an efficacious objective method to measure customers’ experience. The stability and effectiveness of electroencephalography (EEG) make it a suitable tool for aesthetic experience measurement. Nevertheless, existing studies have several limitations, especially regarding the stimuli and the algorithm. The potential of an EEG-based deep learning model has not been verified in pinpointing subtle differences in physical product aesthetics. To fill the research gap in this issue, we recorded EEG signals in real-life scenarios when participants were presented with different types of physical smartphones, and asked participants to rate them from four dimensions of aesthetic experience (arousal, valence, likeness, and aesthetic evaluation). Then, the time–frequency data were fed into a spatial feature extraction network and an attention-based bidirectional long short-term memory (BiLSTM) optimized by the cross-entropy loss function. The result showed that at 16s window size, the four outcome models yielded the best joint recognition performance of aesthetic experience with an average accuracy of over 85% (arousal: 88.10%, valence: 87.97%, likeness: 85.99%, and aesthetic evaluation: 87.23%). It provides an objective cross-subject recognition method with multi-faceted evaluation results of aesthetic experience. Additionally, we verified the ability of EEG as a reliable and informative resource in terms of aesthetic experience evaluation, even with subtle differences. More practically, a future direction of incorporating EEG signals into subjective product aesthetics measurement could be given more credit. Peishan Wang, Haibei Feng, Xiaobing Du, Yudi Lin, CuiXia Ma |
Int. J. Hum. Comput. Interact. | 6 |
| 2024 | SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking EffectivenessabstractAs communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack of evidence-based analytical systems for people to comprehensively evaluate online speeches and further discover possibilities for improvement. This paper introduces SpeechMirror, a visual analytics system facilitating reflection on a speech based on insights from a collection of online speeches. The system estimates the impact of different speech techniques on effectiveness and applies them to a speech to give users awareness of the performance of speech techniques. A similarity recommendation approach based on speech factors or script content supports guided exploration to expand knowledge of presentation evidence and accelerate the discovery of speech delivery possibilities. SpeechMirror provides intuitive visualizations and interactions for users to understand speech factors. Among them, SpeechTwin, a novel multimodal visual summary of speech, supports rapid understanding of critical speech factors and comparison of different speech samples, and SpeechPlayer augments the speech video by integrating visualization of the speaker's body language with interaction, for focused analysis. The system utilizes visualizations suited to the distinct nature of different speech factors for user comprehension. The proposed system and visualization techniques were evaluated with domain experts and amateurs, demonstrating usability for users with low visualization literacy and its efficacy in assisting users to develop insights for potential improvement. Kevin T. Maher, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Sheng Feng Qin, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Novel-view Synthesis and Pose Estimation for Hand-Object Interaction from Sparse ViewsabstractHand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose estimation system for hand-object interaction from sparse views, which can also enable 3D hand-object interaction editing. We share the inspiration from recent scene understanding work that shows a scene specific model built beforehand can significantly improve and unblock vision tasks especially when inputs are sparse, and extend it to the dynamic hand-object interaction scenario and propose to solve the problem in two stages. We first learn the shape and appearance prior knowledge of hands and objects separately with the neural representation at the offline stage. During the online stage, we design a rendering-based joint model fitting framework to understand the dynamic hand-object interaction with the pre-built hand and object models as well as interaction priors, which thereby overcomes penetration and separation issues between hand and object and also enables novel view synthesis. In order to get stable contact during the hand-object interaction process in a sequence, we propose a stable contact loss to make the contact region to be consistent. Experiments demonstrate that our method outperforms the state-of-the-art methods. Code and dataset are available in project web-page https://iscas3dv.github.io/HO-NeRF. Wentian Qu, Zhaopeng Cui, Yinda Zhang 0001, Chenyu Meng, CuiXia Ma, Xiaoming Deng 0001, Hongan Wang |
ICCV | 5 |
| 2023 | Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable ObjectsabstractLearning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton poses for human bodies or animals, which require extra annotation effort and suffer from error accumulation, and they are limited to specific domain. In this paper, we propose a novel self-supervised approach to learn neural implicit shape representation for deformable objects, which can represent shapes with a template shape and dense correspondence in 3D. Our method does not require the priors of skeleton and skinning weight, and only requires a collection of shapes represented in signed distance fields. To handle the large deformation, we constrain the learned template shape in the same latent space with the training shapes, design a new formulation of local rigid constraint that enforces rigid transformation in local region and addresses local reflection issue, and present a new hierarchical rigid constraint to reduce the ambiguity due to the joint learning of template shape and correspondences. Extensive experiments show that our model can represent shapes with large deformations. We also show that our shape representation can support two typical applications, such as texture transfer and shape editing, with competitive performance. The code and models are available at https://iscas3dv.github.io/deformshape. Baowen Zhang, Jiahe Li 0006, Xiaoming Deng 0001, Yinda Zhang 0001, CuiXia Ma, Hongan Wang |
ICCV | 5 |
| 2023 | MMPosE: Movie-Induced Multi-Label Positive Emotion Classification Through EEG SignalsabstractEmotional information plays an important role in various multimedia applications. Movies, as a widely available form of multimedia content, can induce multiple positive emotions and stimulate people's pursuit of a better life. Different from negative emotions, positive emotions are highly correlated and difficult to distinguish in the emotional space. Since different positive emotions are often induced simultaneously by movies, traditional single-target or multi-class methods are not suitable for the classification of movie-induced positive emotions. In this paper, we proposeTransEEG, a model for multi-label positive emotion classification from a viewer's brain activities when watching emotional movies. The key features ofTransEEGinclude (1) explicitly modeling the spatial correlation and temporal dependencies of multi-channel EEG signals using the Transformer structure based model, which effectively addresses long-distance dependencies, (2) exploiting the label-label correlations to guide the discriminative EEG representation learning, for that we design an Inter-Emotion Mask for guiding the Multi-Head Attention to learn the inter-emotion correlations, and (3) constructing an attention score vector from the representation-label correlation matrix to refine emotion-relevant EEG features. To evaluate the ability of our model for multi-label positive emotion classification, we demonstrate our model on a state-of-the-art positive emotion database CPED. Extensive experimental results show that our proposed method achieves superior performance over the competitive approaches. Xiaobing Du, Xiaoming Deng 0001, Hangyu Qin, Yezhi Shu, Fang Liu 0035, Guozhen Zhao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 8 |
| 2023 | Fine-Grained Video Retrieval With Scene SketchesabstractBenefiting from the intuitiveness and naturalness of sketch interaction, sketch-based video retrieval (SBVR) has received considerable attention in the video retrieval research area. However, most existing SBVR research still lacks the capability of accurate video retrieval with fine-grained scene content. To address this problem, in this paper we investigate a new task, which focuses on retrieving the target video by utilizing a fine-grained storyboard sketch depicting the scene layout and major foreground instances' visual characteristics (e.g., appearance, size, pose, etc.) of video; we call such a task "fine-grained scene-level SBVR". The most challenging issue in this task is how to perform scene-level cross-modal alignment between sketch and video. Our solution consists of two parts. First, we construct a scene-level sketch-video dataset called SketchVideo, in which sketch-video pairs are provided and each pair contains a clip-level storyboard sketch and several keyframe sketches (corresponding to video frames). Second, we propose a novel deep learning architecture called Sketch Query Graph Convolutional Network (SQ-GCN). In SQ-GCN, we first adaptively sample the video frames to improve video encoding efficiency, and then construct appearance and category graphs to jointly model visual and semantic alignment between sketch and video. Experiments show that our fine-grained scene-level SBVR framework with SQ-GCN architecture outperforms the state-of-the-art fine-grained retrieval methods. The SketchVideo dataset and SQ-GCN code are available in the project webpage https://iscas-mmsketch.github.io/FG-SL-SBVR/. Ran Zuo, Xiaoming Deng 0001, Yukun Lai, Fang Liu 0035, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 7 |
| 2023 | Stroke-based semantic segmentation for scene-level free-hand sketches
Xiaoming Deng 0001, Jinyao Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
Vis. Comput. | 5 |
| 2022 | Efficient Virtual View Selection for 3D Hand Pose Estimationabstract3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual view selection and fusion module for 3D hand pose estimation from single depth. We propose to automatically select multiple virtual viewpoints for pose estimation and fuse the results of all and find this empirically delivers accurate and robust pose estimation. In order to select most effective virtual views for pose fusion, we evaluate the virtual views based on the confidence of virtual views using a light-weight network via network distillation. Experiments on three main benchmark datasets including NYU, ICVL and Hands2019 demonstrate that our method outperforms the state-of-the-arts on NYU and ICVL, and achieves very competitive performance on Hands2019-Task1, and our proposed virtual view selection and fusion module is both effective for 3D hand pose estimation. Jian Cheng 0006, Yanguang Wan, Dexin Zuo, CuiXia Ma, Ping Tan 0002, Hongan Wang, Xiaoming Deng 0001, Yinda Zhang 0001 |
AAAI | 4 |
| 2022 | An Efficient LSTM Network for Emotion Recognition From Multichannel EEG SignalsabstractMost previous EEG-based emotion recognition methods studied hand-crafted EEG features extracted from different electrodes. In this article, we study the relation among different EEG electrodes and propose a deep learning method to automatically extract the spatial features that characterize the functional relation between EEG signals at different electrodes. Our proposed deep model is calledATtention-basedLSTMwithDomainDiscriminator (ATDD-LSTM), a model based on Long Short-Term Memory (LSTM) for emotion recognition that can characterize nonlinear relations among EEG signals of different electrodes. To achieve state-of-the-art emotion recognition performance, the architecture of ATDD-LSTM has two distinguishing characteristics: (1) By applying the attention mechanism to the feature vectors produced by LSTM, ATDD-LSTM automatically selects suitable EEG channels for emotion recognition, which makes the learned model concentrate on the emotion related channels in response to a given emotion; (2) To minimize the significant feature distribution shift between different sessions and/or subjects, ATDD-LSTM uses a domain discriminator to modify the data representation space and generate domain-invariant features. We evaluate the proposed ATDD-LSTM model on three public EEG emotional databases (DEAP, SEED and CMEED) for emotion recognition. The experimental results demonstrate that our ATDD-LSTM model achieves superior performance on subject-dependent (for the same subject), subject-independent (for different subjects) and cross-session (for the same subject) evaluation. Xiaobing Du, CuiXia Ma, Jinyao Li, Yukun Lai, Guozhen Zhao, Xiaoming Deng 0001, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | SketchMaker: Sketch Extraction and Reuse for Interactive Scene Sketch CompositionabstractSketching is an intuitive and simple way to depict sciences with various object form and appearance characteristics. In the past few years, widely available touchscreen devices have increasingly made sketch-based human-AI co-creation applications popular. One key issue of sketch-oriented interaction is to prepare input sketches efficiently by non-professionals because it is usually difficult and time-consuming to draw an ideal sketch with appropriate outlines and rich details, especially for novice users with no sketching skills. Thus, sketching brings great obstacles for sketch applications in daily life. On the other hand, hand-drawn sketches are scarce and hard to collect. Given the fact that there are several large-scale sketch datasets providing sketch data resources, but they usually have a limited number of objects and categories in sketch, and do not support users to collect new sketch materials according to their personal preferences. In addition, few sketch-related applications support the reuse of existing sketch elements. Thus, knowing how to extract sketches from existing drawings and effectively re-use them in interactive scene sketch composition will provide an elegant way for sketch-based image retrieval (SBIR) applications, which are widely used in various touch screen devices. In this study, we first conduct a study on current SBIR to better understand the main requirements and challenges in sketch-oriented applications. Then we develop the SketchMaker as an interactive sketch extraction and composition system to help users generate scene sketches via reusing object sketches in existing scene sketches with minimal manual intervention. Moreover, we demonstrate how SBIR improves from composited scene sketches to verify the performance of our interactive sketch processing system. We also include a sketch-based video localization task as an alternative application of our sketch composition scheme. Our pilot study shows that our system is effective and efficient, and provides a way to promote practical applications of sketches. Fang Liu 0035, Xiaoming Deng 0001, Jian-Cheng Song, Yukun Lai, Yong-Jin Liu 0001, Hao Wang 0005, CuiXia Ma, Sheng Feng Qin, Hongan Wang |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2022 | SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNsabstractSketch-based image retrieval (SBIR) is a long-standing research topic in computer vision. Existing methods mainly focus on category-level or instance-level image retrieval. This paper investigates the fine-grained scene-level SBIR problem where a free-hand sketch depicting a scene is used to retrieve desired images. This problem is useful yet challenging mainly because of two entangled facts: 1) achieving an effective representation of the input query data and scene-level images is difficult as it requires to model the information across multiple modalities such as object layout, relative size and visual appearances, and 2) there is a great domain gap between the query sketch input and target images. We present SceneSketcher-v2, a Graph Convolutional Network (GCN) based architecture to address these challenges. SceneSketcher-v2 employs a carefully designed graph convolution network to fuse the multi-modality information in the query sketch and target images and uses a triplet training process and end-to-end training manner to alleviate the domain gap. Extensive experiments demonstrate SceneSketcher-v2 outperforms state-of-the-art scene-level SBIR models with a significant margin. Fang Liu 0035, Xiaoming Deng 0001, Changqing Zou, Yukun Lai, Ran Zuo, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 7 |
| 2022 | E-ffective: A Visual Analytic System for Exploring the Emotion and Effectiveness of Inspirational SpeechesabstractWhat makes speeches effective has long been a subject for debate, and until today there is broad controversy among public speaking experts about what factors make a speech effective as well as the roles of these factors in speeches. Moreover, there is a lack of quantitative analysis methods to help understand effective speaking strategies. In this paper, we propose E-ffective, a visual analytic system allowing speaking experts and novices to analyze both the role of speech factors and their contribution in effective speeches. From interviews with domain experts and investigating existing literature, we identified important factors to consider in inspirational speeches. We obtained the generated factors from multi-modal data that were then related to effectiveness data. Our system supports rapid understanding of critical factors in inspirational speeches, including the influence of emotions by means of novel visualization methods and interaction. Two novel visualizations include E-spiral (that shows the emotional shifts in speeches in a visually compact way) and E-script (that connects speech content with key speech delivery information). In our evaluation we studied the influence of our system on experts' domain knowledge about speech factors. We further studied the usability of the system by speaking novices and experts on assisting analysis of inspirational speech effectiveness. Kevin T. Maher, Jian-Cheng Song, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageabstractIn this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solution space. In order to address the occlusion and similar appearance between hands that may confuse the network, we design a hand pose-aware attention module to extract features associated to each individual hand respectively. We then leverage the two hand context presented in interaction to propose a context-aware cascaded refinement that improves the hand pose and shape accuracy of each hand conditioned on the context between interacting hands. Extensive experiments on the main benchmark datasets demonstrate that our method predicts accurate 3D hand pose and shape from single color image, and achieves the state-of-the-art performance. Code is available in project webpage https://baowenz.github.io/Intershape/. Baowen Zhang, Yangang Wang 0001, Xiaoming Deng 0001, Yinda Zhang 0001, Ping Tan 0002, CuiXia Ma, Hongan Wang |
ICCV | 6 |
| 2021 | Multi-scale visualization based on sketch interaction for massive surveillance video data
Ran Zuo, Ti Zhou, CuiXia Ma, Hongan Wang |
Pers. Ubiquitous Comput. | 7 |
| 2021 | Weakly Supervised Learning for Single Depth-Based Hand Shape RecoveryabstractRecent emerging technologies such AR/VR and HCI are drawing high demand on more comprehensive hand shape understanding, requiring not only 3D hand skeleton pose but also hand shape geometry. In this paper, we propose a deep learning framework to produce 3D hand shape from a single depth image. To address the challenge that capturing ground truth 3D hand shape in the training dataset is non-trivial, we leverage synthetic data to construct a statistical hand shape model and adopt weak supervision from widely accessible hand skeleton pose annotation. To bridge the gap due to the different hand skeleton definitions in the existing public datasets, we propose a joint regression network for hand pose adaptation. To reconstruct the hand shape, we use Chamfer loss between the predicted hand shape and the point cloud from the input depth to learn the shape reconstruction model in a weakly-supervised manner. Experiments demonstrate that our model adapts well to the real data and produces accurate hand shapes that outperform the state-of-the-art methods both qualitatively and quantitatively. Xiaoming Deng 0001, Yuying Zhu 0002, Yinda Zhang 0001, Zhaopeng Cui, Ping Tan 0002, Wentian Qu, CuiXia Ma, Hongan Wang |
IEEE Trans. Image Process. | 7 |
| 2020 | SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches
Fang Liu 0035, Changqing Zou, Xiaoming Deng 0001, Ran Zuo, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
ECCV (19) | 6 |
| 2020 | Image-based Pose Representation for Action Recognition and Hand Gesture RecognitionabstractIn this paper, we propose an effective and compact image-based pose representation named Poseimage Pyramid, which encodes the spatial and temporal information of human pose or hand pose as an image pyramid. Poseimage is constructed by the normalized distance between pairwise joints, and it has the advantage of its invariant to similarity transformations. With our Poseimage representation we can design the pose based action recognition or hand gesture recognition model using existing image or video classification models. In order to adapt to different actions with a variety of movement speed, we design Poseimage Pyramid to encode the multi-scale temporal information of human pose or hand pose. Experiments demonstrate that our pose representation is effective, and we achieve state-of-the-art performance on the action recognition datasets and the hand gesture recognition datasets. Our pose presentation is also complementary to video and optical flow streams in the seminal action recognition network I3D, and we achieve the state-of-the-art performance on the JHMDB, HMDB and UCF101 datasets by integrating our pose representation with I3D. Zeyi Lin, Xiaoming Deng 0001, CuiXia Ma, Hongan Wang |
FG | 4 |
| 2020 | EmotionMap: Visual Analysis of Video Emotional Content on a Map
CuiXia Ma, Jian-Cheng Song, Qian Zhu 0010, Kevin T. Maher, Hongan Wang |
J. Comput. Sci. Technol. | 1 |
| 2020 | STA-GCN: two-stream graph convolutional network with spatial-temporal attention for hand gesture recognition
Zeyi Lin, Jian Cheng 0006, CuiXia Ma, Xiaoming Deng 0001, Hongan Wang |
Vis. Comput. | 4 |
| 2019 | SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial NetworkabstractHand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often incomplete, which makes sketch recognition a challenging problem. In this paper, we propose SketchGAN, a new generative adversarial network (GAN) based approach that jointly completes and recognizes a sketch, boosting the performance of both tasks. Specifically, we use a cascade Encode-Decoder network to complete the input sketch in an iterative manner, and employ an auxiliary sketch recognition task to recognize the completed sketch. Experiments on the Sketchy database benchmark demonstrate that our joint learning approach achieves competitive sketch completion and recognition performance compared with the state-of-the-art methods. Further experiments using several sketch-based applications also validate the performance of our method. Fang Liu 0035, Xiaoming Deng 0001, Yukun Lai, Yong-Jin Liu 0001, CuiXia Ma, Hongan Wang |
CVPR | 5 |
| 2019 | Cascaded Point Network for 3D Hand Pose EstimationabstractRecent PointNet-family hand pose methods have the advantages of high pose estimation performance and small model size, and it is a key problem to get effective sample points for PointNet-family methods. In this paper, we propose a two-stage coarse to fine hand pose estimation method, which belongs to PointNet-family methods and explores a new sample point strategy. In the first stage, we use 3D coordinate and surface normal of normalized point cloud as input to regress coarse hand joints. In the second stage, we use the hand joints in the first stage as the initial sample points to refine the hand joints. Experiments on widely used datasets demonstrate that using joints as sample points is more effective and our method achieves top-rank performance. Yikun Dou, Yuying Zhu 0002, Xiaoming Deng 0001, CuiXia Ma, Liang Chang 0001, Hongan Wang |
ICASSP | 5 |
| 2017 | Leveraging weak segmentation for multi-object tracking systemabstractObject Tracking is an important task in Computer Vision, which has gained increasing attention from academia to industry. In this paper, we propose a real-time tracking system based on weak segmentation. Different from general tracking by detection systems, we do not classify objects into car, cat or bike, instead we just classify the image into object area and non-object area. Many tracking systems simply combine weak segmentation results with Kalman filter, which cannot track objects well when the target moves irregularly. In this paper, we use state-of-art background subtraction methods to accomplish weak segmentation and incorporates weak segmentation results with the local invariant feature to create robust object association, finally we proposal a status machine to convert temporary object association into long-term object tracking results. The experiment shows our algorithm outperforms the state-of-art tracking by weak segmentation algorithms. CuiXia Ma, Hao Wang 0005, Hongan Wang |
SMC | 2 |
| 2016 | VideoMap: An interactive and scalable visualization for exploring video contentabstractLarge-scale dynamic relational data visualization has attracted considerable research attention recently. We introduce dynamic data visualization into the multimedia domain, and present an interactive and scalable system, VideoMap, for exploring large-scale video content. A long video or movie has much content; the associations between the content are complicated. VideoMap uses new visual representations to extract meaningful information from video content. Map-based visualization naturally and easily summarizes and reveals important features and events in video. Multi-scale descriptions are used to describe the layout and distribution of temporal information, spatial information, and associations between video content. Firstly, semantic associations are used in which map elements correspond to video contents. Secondly, video contents are visualized hierarchically from a large scale to a fine-detailed scale. VideoMap uses a small set of sketch gestures to invoke analysis, and automatically completes charts by synthesizing visual representations from the map and binding them to the underlying data. Furthermore, VideoMap allows users to use gestures to move and resize the view, as when using a map, facilitating interactive exploration. Our experimental evaluation of VideoMap demonstrates how the system can assist in exploring video content as well as significantly reducing browsing time when trying to understand and find events of interest. CuiXia Ma, Hongan Wang |
Comput. Vis. Media | 1 |
| 2016 | An Interactive SpiralTape Video SummarizationabstractA majority of video summarization systems use linear representations, such as rectangular storyboards and timelines at linear scales. In this paper, we propose a novel nonlinear dynamic representation called SpiralTape that summarizes a video in a smooth spiral pattern. SpiralTape provides an unusual and fresh activity suitable for stimulating environments such as science and technology museums, in which children or young individuals can have enjoyable experiences that create meaningful learning outcomes. In addition, SpiralTape provides an uninterrupted overall structure of video content and takes design principles including compactness, continuity, efficient overview, and interactivity into consideration. A working SpiralTape system was developed and deployed in pilot applications and exhibitions. Elaborate user studies with evaluation benchmarks on multiple metrics were conducted to compare SpiralTape with two representative linear video summarization methods and a state-of-the-art radial video visualization. The evaluation results demonstrate the effectiveness and natural interaction performance of SpiralTape. Yong-Jin Liu 0001, CuiXia Ma, Guozhen Zhao, Xiaolan Fu, Hongan Wang, Guozhong Dai, Lexing Xie |
IEEE Trans. Multim. | 2 |
| 2016 | Visualizing and Analyzing Video Content With Interactive Scalable MapsabstractVisualizing and communicating insights through maps offers an intuitive and familiar way to explore large-scale dynamic relational data. In this paper, we present VideoMap, which is a novel approach for presenting and interacting with relational video content by taking advantage of the map metaphor. VideoMap employs a metaphor to visualize video content by elements of a map with the aim of enabling exploration of video content as if reading a map. Video content is visualized in a hierarchal structure from a very large scale to a small scale of finely detailed representation. VideoMap recognizes a small set of sketch gestures for semantic zooming in and out, annotating the map, and automatically completing path navigation. To achieve this, VideoMap synthesizes map-derived visuals and binds them to the underlying data by operating the map with sketch interaction to facilitate interactive exploration. Extensive user studies were conducted to evaluate VideoMap, and the results demonstrated the effectiveness of VideoMap for facilitating the exploration and understanding of large video content. CuiXia Ma, Yong-Jin Liu 0001, Guozhen Zhao, Hongan Wang |
IEEE Trans. Multim. | 1 |
| 2015 | Exploring the Benefits of Text and Sketch in Video Retrieval of Complex QueriesabstractThe booming of mobile devices and networks leads to an explosive growth in video resources. Efficient video research styles are appealing for facilely exploring video content appropriate to user's intention with a low cognitive load. The style of input becomes particularly relevant during the process of finding a target video clip in a large-scale database on tablets or other mobile devices. Some users have strong allegiance to input text, while others only input sketches. In this paper, we present the first systematic comparison of these two input styles and analyze the responses and feedbacks of users. An elaborated user study was conducted to test two different styles of inputting the semantics. Some users preferred to text input because it could describe their objective easily in a short time, yet some users also liked sketch because it helped illustrate the action or orientation clearly and immediately. Combining text with sketch ("sketch-text") is efficient for searching video of complex queries. The evaluation results show users' enjoying "sketch-text" and its higher performance than other input styles. CuiXia Ma, Hongan Wang |
VINCI | 2 |
| 2014 | A Sketch-Based Approach for Interactive Organization of Video ClipsabstractWith the rapid growth of video resources, techniques for efficient organization of video clips are becoming appealing in the multimedia domain. In this article, a sketch-based approach is proposed to intuitively organize video clips by: (1) enhancing their narrations using sketch annotations and (2) structurizing the organization process by gesture-based free-form sketching on touch devices. There are two main contributions of this work. The first is a sketch graph, a novel representation for the narrative structure of video clips to facilitate content organization. The second is a method to perform context-aware sketch recommendation scalable to large video collections, enabling common users to easily organize sketch annotations. A prototype system integrating the proposed approach was evaluated on the basis of five different aspects concerning its performance and usability. Two sketch searching experiments showed that the proposed context-aware sketch recommendation outperforms, in terms of accuracy and scalability, two state-of-the-art sketch searching methods. Moreover, a user study showed that the sketch graph is consistently preferred over traditional representations such as keywords and keyframes. The second user study showed that the proposed approach is applicable in those scenarios where the video annotator and organizer were the same person. The third user study showed that, for video content organization, using sketch graph users took on average 1/3 less time than using a mass-market tool Movie Maker and took on average 1/4 less time than using a state-of-the-art sketch alternative. These results demonstrated that the proposed sketch graph approach is a promising video organization tool. Yong-Jin Liu 0001, CuiXia Ma, Qiu-Fang Fu, Xiaolan Fu, Sheng Feng Qin, Lexing Xie |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | An approach to visual analysis for task flow management
Dongxing Teng, Shaolei Song, CuiXia Ma, Hongan Wang, Guozhong Dai |
Sci. China Inf. Sci. | 4 |
| 2013 | Interactive multi-scale structures for summarizing video content
Hongan Wang, CuiXia Ma |
Sci. China Inf. Sci. | 2 |
| 2013 | Collaborative Interaction for Videos on Mobile Devices Based on Sketch Gestures
Jinkai Zhang, CuiXia Ma, Yong-Jin Liu 0001, Qiu-Fang Fu, Xiaolan Fu |
J. Comput. Sci. Technol. | 2 |
| 2013 | User-Adaptive Sketch-Based 3-D CAD Model Retrievalabstract3-D CAD models are an important digital resource in the manufacturing industry. 3-D CAD model retrieval has become a key technology in product lifecycle management enabling the reuse of existing design data. In this paper, we propose a new method to retrieve 3-D CAD models based on 2-D pen-based sketch inputs. Sketching is a common and convenient method for communicating design intent during early stages of product design, e.g., conceptual design. However, converting sketched information into precise 3-D engineering models is cumbersome, and much of this effort can be avoided by reuse of existing data. To achieve this purpose, we present a user-adaptive sketch-based retrieval method in this paper. The contributions of this work are twofold. First, we propose a statistical measure for CAD model retrieval: the measure is based on sketch similarity and accounts for users' drawing habits. Second, for 3-D CAD models in the database, we propose a sketch generation pipeline that represents each 3-D CAD model by a small yet sufficient set of sketches that are perceptually similar to human drawings. User studies and experiments that demonstrate the effectiveness of the proposed method in the design process are presented. Yong-Jin Liu 0001, Ajay Joneja, CuiXia Ma, Xiaolan Fu, Dawei Song 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2012 | Sketch-Based Annotation and Visualization in Video AuthoringabstractAuthoring context-aware, interactive video representation is usually a complex process. A user-friendly multimedia authoring environment is thus solicited to explore and express users' design ideas efficiently and naturally. In this paper we present a sketch-based two-layer representation, called scene structure graph (SSG), to facilitate the video authoring process. One layer in SSG uses sketches as a concise form with which the visualization of scene information is easily understood and the other layer uses a graph to represent and edit the narrative structure in the authoring process. With SSG, the authoring process works in two stages. In the first stage, various sketch forms such as symbols and hand-drawing illustrations are used as basic primitives to annotate the video clips and the hyperlinks encoding spatio-temporal relations are established in SSG. In the second stage, sketches in SSGs are modified and new SSG is composed for any particular authoring purpose. Three user studies are elaborated, showing that the SSG is user-friendly and can achieve a good balance between expressiveness of users' intent and ease of use for authoring of interactive video. CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang, Dongxing Teng, Guozhong Dai |
IEEE Trans. Multim. | 1 |
| 2011 | KnitSketch: A Sketch Pad for Conceptual Design of 2D Garment PatternsabstractIn this paper, we present a new sketch-based system - KnitSketch, to improve the efficiency of process planning for knitting garments at an early design stage. The KnitSketch system utilizes sketching interface with the pen-paper metaphor and users only need to draw outlines of different parts of the garment. Based on sketching understanding, the system automatically makes reasonable geometric inferences about the process-planning data of the garment. The system is designed for nonprofessional users and can design diverse garment styles by freehand drawings. The contributions of this work include contextual extraction of reusable data from sketches, a MDG structure for sketch beautification, and an integrated system with natural expression and effective communication that reduces the cognitive load of human beings. User experience shows that the proposed system helps designers focus on the task instead of the designing tools, and thus improves the efficiency and productivity of human beings. CuiXia Ma, Yong-Jin Liu 0001, Dongxing Teng, Hongan Wang, Guozhong Dai |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2010 | Cooperative concept map based on cognitive model for visual analysisabstractThe ability to recognize users and their intention during the visual analysis process is important to interact intelligently and effortlessly in a cooperative environment. In this paper, we present the framework for creating and modifying a concept map that allows analysts to collaborate and utilize live collection of enterprise data in a visual interface. Based on modeling users and tasks in visual analysis, a user model supporting collaborative analysis is given. Furthermore, we propose the algorithm of concept map layout to facilitate the visual analysis. The interface presents the analyst with the dashboard configuration of certain data that users deems relevant, provides tools to facilitate the decision making process. Finally, we apply it to a collaborative visual decision support system. Experimental results show that it enhances the accessibility and the individualization of collaborative concept map for configuring enterprise data and helps to enhance user experience during the visual analysis process. Yi Du 0010, CuiXia Ma, Dongxing Teng, Guozhong Dai |
VINCI | 2 |
| 2008 | Online Personalised Non-photorealistic Rendering Technique for 3D Geometry from Incremental SketchingabstractAbstract This paper presents an online personalised non‐photorealistic rendering (NPR) technique for 3D models generated from interactively sketched input. This technique has been integrated into a sketch‐based modelling system. It lets users interact with computers by drawing naturally, without specifying the number, order, or direction of strokes. After sketches are interpreted as 3D objects, they can be rendered with personalised drawing styles so that the reconstructed 3D model can be presented in a sketchy style similar in appearance to what have been drawn for the 3D model. This technique captures the user's drawing style without using template or prior knowledge of the sketching style. The personalised rendering style can be applied to both visible and initially invisible geometry. The rendering strokes are intelligently selected from the input sketches and mapped to edges of the 3D object. In addition, non‐geometric information such as surface textures can be added to the recognised object in different sketching modes. This will integrate sketch‐based incremental 3D modelling and NPR into conceptual design. Day Chyi Ku, Sheng Feng Qin, David K. Wright 0001, CuiXia Ma |
Comput. Graph. Forum | 4 |
| 2002 | A Gesture Language for Collaborative Conceptual DesignabstractUbiquitous computing provides the new research challenges for the next generation interactions. Natural User Interface is one of the important aspects. Pen and paper are the natural and simple tools to use in our daily life, which can fit for the imprecise and flexible design mode in conceptual design. Gestures are the useful and natural mode in the pen-based interaction which can be introduced and used through the similar mode to the existing resources in interaction. Inherent problems exist with collaborative conceptual design because the simultaneous group interaction required for users to smoothly and effectively work together in the same virtual space. In this paper, we provide solutions to these problems by devising a fast and direct gesture-based user interface, a set of visual effects that better enable a user's awareness of the operations done by other participants, and a set of tools for enhancing visual communication between participants. A gesture description language (GDL) is given to provide the useful gestures to the applications conveniently. CuiXia Ma, Hongan Wang, Guozhong Dai |
CSCWD | 1 |
| 2001 | Research on Network Based Conceptual DesignabstractConceptual design has played a very important role in the design process. Knowledge of all the design requirements and constraints during this early phase of a product's life cycle is usually approximate, imprecise, and unknown. Conceptual design is now recognized as a highly important process when considering both the concurrent engineering aspects and the commitment of resources necessary to design and develop a product. With the development of the Internet, collaborative design is imperative and effective. The technology of CSCW is adopted widely and necessarily in the domain of conceptual design. It can fit the nature of the design. It shortens the distance of designers in different places and collects and shares all the information in time. In conceptual design, collaborative design mainly includes function, principle, layout and initial structure, etc. We developed a system of conceptual design based on the current CAD system, which realized the seamless link with later detail design. For individual design, multi-modal technology, which integrates many interactions such as gesture and speech, supports an advanced interface to assist humans in developing applications. CuiXia Ma, Hongan Wang, Guozhong Dai |
CSCWD | 1 |