VLDB 2026 Research / reviewers in the wild / expert
Lichun Wang 0002
dblp:80/5367-2
· DBLP profile ↗
34ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0002-4977-0183ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Motion-VLN: Learning ego-motion from optical flow for VLN-CE
Lichun Wang 0002, Zijin Liu |
Neurocomputing | 2 |
| 2026 | Deep Residual Discriminative Dictionary Learning for image classification
Lichun Wang 0002, Jianjia Xin, Kai Xu 0012, Huiyong Zhang, Shaofan Wang 0001, Dehui Kong |
Knowl. Based Syst. | 2 |
| 2026 | VRGraspNet: Toward Viewpoint Robust 6-DoF Grasp Pose Estimationabstract6-DoF grasp pose estimation is crucial for achieving robust robot manipulation. Despite significant progress in data-driven methods, the cross-view adaptability of 6-DoF grasp pose estimation remains insufficient. Under the condition of observing some certain viewpoints, the performance of pose estimation is relatively low. In response, this paper introduces VRGraspNet, a novel 6-DoF grasp pose estimation model designed to enhance the robustness and performance of robotic grasping across diverse viewpoints. The key to the VRGraspNet lies in filling in the holes in the scene point cloud and learning multimodal features for seed points. The former provides dense neighborhood points for the seed point, while the latter provides richer information for extracting geometric features of the local area to which the seed point belongs. Furthermore, this paper proposes a Performance-Viewpoint jointly Weighted Loss (PVWL), with the key being two weight factors: a static viewpoint position dependent weight factor that focuses on challenging samples, and a dynamic performance related weight factor that focuses on samples difficult to learn. Extensive experiments on the GraspNet-1Billion dataset demonstrate that VRGraspNet achieves SOTA performance and strong cross-view robustness. Real-world robot experiments further validate its practicality in robotic manipulation tasks. Our source code is available at https://github.com/huamo555/VRGraspNet. Yuming Gao, Lichun Wang 0002, Jiaqi Zheng 0020, Kai Xu 0012, Huayang Yao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Self-Knowledge Distillation and Its Application: A SurveyabstractAs an efficient model compression technique, knowledge distillation has become an important research topic in the field of deep learning. However, the requirement of pre-trained teacher networks makes the process cumbersome and inefficient, which prompted researchers to propose a more efficient mechanism. Therefore, self-knowledge distillation is proposed, which does not require assistance from additional teacher networks. In previous surveys, self-knowledge distillation has usually been considered a special case of knowledge distillation. In recent years, significant progress has been made in self-knowledge distillation, which has evolved beyond the functions or roles of traditional knowledge distillation. However, there is no individual and intensive survey of self-knowledge distillation methods up to now. Therefore, this paper reviews and investigates existing self-knowledge distillation methods from a comprehensive perspective. Specifically, first, according to different sources of knowledge, this paper categorizes self-knowledge distillation methods into three types, label knowledge-based, feature knowledge-based and data knowledge-based. Then, this paper introduces the evaluation protocol and performance of SKD. In particular, the commonly used experimental datasets and evaluation networks are summarized, aiming to encourage researchers to choose common network architectures and evaluation datasets for promoting the standardization of self-knowledge distillation's comparison. Finally, this paper introduces the applications of self-knowledge distillation in different task scenarios, enabling researchers to quickly locate the relevant task fields. Furthermore, this paper introduces the technical purposes of applying self-knowledge distillation and explores the core motivations for using SKD across different tasks, promoting researchers' extensive attempts at self-distillation technology in their own research fields that have not yet applied self-knowledge distillation. Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Ensemble predicate decoding for unbiased scene graph generation
Jiasong Feng, Lichun Wang 0002, Kai Xu 0012 |
Neurocomputing | 2 |
| 2025 | Bcgn: BLIP-based cross-modal grasping network for language-conditioned robotic grasping
Kai Xu 0012, Lichun Wang 0002, Jianjia Xin |
Multim. Syst. | 2 |
| 2025 | Memory Transmission Based Referring Video Object Segmentation
Zijin Liu, Lichun Wang 0002, Yongli Hu |
Neural Networks | 2 |
| 2025 | Scene Adaptive Context Modeling and Balanced Relation Prediction for Scene Graph GenerationabstractScene graph generation (SGG) aims to perceive objects and their relations in images, which can bridge the gap between upstream detection tasks and downstream high-level visual understanding tasks. For SGG models, over-fitting head predicates can lead to bias in the generated scene graph, which has become a consensus. A series of debiasing methods have been proposed to solve the problem. However, some existing debiasing SGG methods have a tendency to over-fit tail predicates, which is another type of bias. In order to eliminate the one-way over-fitting of head or tail predicates, this article proposes a balanced relation prediction (BRP) module which is model-agnostic and compatible with existing re-balancing methods. Moreover, because the relation prediction is based on object feature representation, this article proposes a scene adaptive context fusion (SACF) module to refine the object feature representation. Specifically, SACF models the context based on a chain structure, where the order of objects in the chain structure is adaptively arranged according to the scene content, achieving visual information fusion that adapts to the scene where the objects are located. Experiments on VG and GQA datasets show that the proposed method achieves competitive results on the comprehensive metric of R@K and mR@K. Kai Xu 0012, Lichun Wang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Self-Knowledge Distillation with Learning from Role-Model SamplesabstractSelf-knowledge distillation does not require a pre-trained teacher network like traditional knowledge distillation. Existing methods either require additional parameters or require additional memory consumption. To alleviate this problem, this paper proposes a more efficient self-knowledge distillation method, named LRMS (learning from role-model samples). In every mini-batch, LRMS selects out a role-model sample for each sampled category, and takes its prediction as the proxy semantic for the corresponding category. Then, predictions of the other samples are constrained to be consistent with the proxy semantics, which makes the distribution of predictions for samples within the same category more compact. Meanwhile, the regularization targets corresponding to proxy semantics are set with a higher distillation temperature to better utilize the classificatory information about the categories. Experimental results show that diverse architectures achieve improvements on four image classification datasets by using LRMS. Code is acaliable: https://github.com/KAI1179/LRMS Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang |
ICASSP | 2 |
| 2024 | Area-keywords cross-modal alignment for referring image segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012 |
Neurocomputing | 2 |
| 2024 | OASNet: Object Affordance State Recognition Network With Joint Visual Features and Relational Semantic EmbeddingsabstractTraditional affordance learning tasks aim to understand object’s interactive functions in an image, such as affordance recognition and affordance detection. However, these tasks cannot determine whether the object is currently interacting, which is crucial for many follow-up tasks, including robotic manipulation and planning task. To fill this gap, this paper proposes a novel object affrodance state (OAS) recognition task, i.e., simultaneously recognizing an object’s affordances and the partner objects that are interacting with it. Accordingly, to facilitate the application of deep learning technology, an OAS recognition task related dataset OAS10k is constructed by collecting and labeling over 10k images. In the dataset, a sample is defined as a set of an image and its OAS labels, each label is represented as$\left \langle{ \rm {\textit {subject, subject's affrodance, interacted object}} }\right \rangle $. These triplet labels have rich relational semantic information, which can improve OAS recognition performance. We hence construct a directed OAS knowledge graph of affordance states, and extract an OAS matrix from it for modelling the semantic relationships of the triplets. Based on the matrix, we propose an OAS recognition network (OASNet), which utilizes GCN to capture the relational semantic embeddings, and uses a transformer to fuse them with the visual features from an image to recognize the affordance states of objects in the image. Experimental results on OAS10k dataset and other triplet label recognition datasets demonstrate that the proposed OASNet achieves the best performance compared to the state-of-the-art methods. The dataset and codes will be released onhttps://github.com/mxmdpc/OAS. Dongpan Chen, Dehui Kong, Lichun Wang 0002, Junna Gao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A New Training Data Organization Form and Training Mode for Unbiased Scene Graph GenerationabstractThe current mainstream studies on Scene Graph Generation (SGG) devote to the long-tailed predicate distribution problem to generate unbiased scene graph. The long-tailed predicate distribution exists in VG dataset and is more severe during the SGG network training process. Most existing de-biasing methods solve the problem by applying re-sampling or re-weighting in a mini-batch, with the main idea being to provide unbiased attention to different predicate categories based on prior predicate distributions. During the training process of SGG models, existing training mode samples several images into a mini-batch to obtain training data, thus providing sparse and scattered predicate instances for training. However, sampling predicate instances from a limited set of predicate samples in terms of quantity and category poses difficulties in training unbiased SGG models. In order to provide a wider range for sampling predicate instances, this paper reorganizes the images in VG training set with a new form, i.e. object-pairs, and constructs VG-OP (VG Object-Pair) training set to save object-pairs. Meanwhile, this paper introduces a new SGG network training mode, which can realize unbiased SGG without resampling or re-weighting. In particular, a Predicate-balanced Sampling Network (PS-Net) is designed to validate the new training mode. Extensive experiments on VG test set demonstrate that our method achieves competitive or state-of-the-art unbiased SGG performance. Lichun Wang 0002, Kai Xu 0012, Fangyu Fu, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Self-Distillation With Augmentation in Feature SpaceabstractCompared with traditional knowledge distillation, self-distillation does not require a pre-trained teacher network, which is more concise. Among them, data augmentation-based methods provide an elegant solution without modifying the network structure or additional memory consumption. However, when employing data augmentation in the input space, the forward propagations for augmented data bring additional computation costs and the augmentation methods need be adaptive to the modality of input data. Meanwhile, we note that from a generalization perspective, under the condition of being able to distinguish from other classes, a dispersed intra-class feature distribution is superior to compact intra-class feature distribution, especially for categories with larger sample differences. Based on the above considerations, this paper proposes a feature augmentation based self-distillation method (FASD) based on the idea of feature extrapolation. For each source feature, two augmentations are generated by subtraction between features. The one is subtracting the temporary class center computed with samples belonging to the same category, and another one is subtracting a sample feature belonging to other categories with the closest distance. Then, the predicted outputs of the augmented features are constrained to be consistent with that of the source feature. The consistent constraint on the previous augmented feature expands the learned class feature distribution, leading to greater overlap with the unknown feature distribution of test samples, thereby improving the generalization performance of the network. The consistent constraint on the latter augmented feature increases the distance between samples from different categories, which enhances the distinguishability between categories. Experimental results on image classification task demonstrate the effectiveness and efficiency of the proposed method. Meanwhile, experiments on text and audio tasks prove the universality of the method for classification tasks with different modalities. Kai Xu 0012, Lichun Wang 0002, Jianjia Xin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning From Teacher's Failure: A Reflective Learning Paradigm for Knowledge DistillationabstractKnowledge Distillation transfers knowledge learned by a teacher network to a student network. A common mode of knowledge transfer is directly using the teacher network’s experience for all samples without differentiating whether the experience of teacher is successful or not. According to common sense, experience varies with its nature. Successful experience is used for guidance, and failed experience is used for correction. Inspired by that, this paper analyzes the failure of teacher and proposes a reflective learning paradigm, which additionally uses heuristic knowledge extracted from the teacher’s failure besides following the authority of teacher. Specifically, this paper defines Mutual Error Distance (MED) based on the teacher’s wrong predictions. MED measures the adequacy of the decision boundary learned by teacher, which concretizes the failure of teacher. Then, this paper proposes DCGD (divide-and-conquer grouping distillation) to critically transfer the teacher’s knowledge by grouping the target task into small-scale subtasks and designing multi-branch networks on the basis of MED. Finally, a switchable training mechanism is designed to integrate a regular student which provides an option of student network without parameter addition compared with the multi-branch student network. Extensive experiments on three image classification benchmarks (CIFAR-10, CIFAR-100 and TinyImageNet) show the effectiveness of the proposed paradigm. Especially on CIFAR-100 dataset, the average error of students using DCGD+DKD decreased by 4.28%. In addition, the experiment results show that the paradigm is also applicable to self-distillation. Kai Xu 0012, Lichun Wang 0002, Jianjia Xin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A Balanced Relation Prediction Framework for Scene Graph Generation
Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang |
ICANN (4) | 2 |
| 2023 | A Novel Encoder and Label Assignment for Instance Segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012 |
ICANN (6) | 2 |
| 2023 | Distance-Aware Vector-Field and Vector Screening Strategy for 6D Object Pose Estimation
Lichun Wang 0002, Chao Yang 0040, Jianjia Xin |
ICIG (2) | 1 |
| 2023 | Scale-Adaptive Multi-area Representation for Instance Segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012 |
ICIG (4) | 2 |
| 2023 | Augmented Spatial Context Fusion Network for Scene Graph GenerationabstractScene graph generation provides high-order semantic information by understanding the objects and their relations in images. In order to improve the performance of scene graph generation, context fusion has been widely used in scene graph generation tasks, LSTM and Vision-Transformer are commonly used fusion modules. Both LSTM and Vision-Transformer realize context fusion by stacking multiple basic units, which needs to learn a large number of parameters of the units. However, the model computational efficiency of scene graph generation as a mid-level semantic understanding task to support downstream tasks is crucial. To simplify the context fusion computation, this paper proposes ASCF -Net (Augmented Spatial Context Fusion Network) which computes the spatial context of designated object by searching the nearest neighbor objects with high relevance and strengthens the context with random noise. Without learning parameters, the above computational process essentially simulates the attention mechanism. Experiments on VG dataset show that ASCF -Net uses 15.26% of the parameters of Bi-LSTM and 13.34% of the parameters of Vision-Transformer for context fusion based on the same baseline and achieves higher performance than using the two fusion modules. At the same time, ASCF -Net uses simple fusion module to obtain competitive results on VG dataset compared with the mainstream scene generation models. Lichun Wang 0002, Kai Xu 0012, Fangyu Fu, Qingming Huang |
IJCNN | 2 |
| 2023 | PASIFTNet: Scale-and-Directional-Aware Semantic Segmentation of Point Clouds
Shaofan Wang 0001, Lichun Wang 0002 |
Comput. Aided Des. | 3 |
| 2023 | Hierarchical Coupled Discriminative Dictionary Learning for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize images of novel classes, but does not use any images belonging to the novel classes during model training, which is realized by exploiting the auxiliary semantic information. Recently, most ZSL methods focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the novel classes. Visual-semantic embedding is usually established based on the visual features of images and the semantic information of classes, i.e., class attributes. However, image features are extracted at the individual level, while class attributes are obtained at the group level, so the granularity of these features is different, which makes it difficult to match the two kinds of features. To tackle such problem, we propose hierarchical coupled discriminative dictionary learning (HCDDL) method to hierarchically establish visual-semantic embedding at class-level and image-level with a coarse-to-fine way. Firstly, a class-level coupled dictionary is trained to build basic and coarse-grained connection between visual space and semantic space. Using the class-level coupled dictionary, image attributes are generated. Based on the fine-grained image attributes and images features, an image-level coupled dictionary is learned. In addition, during the learning of hierarchical coupled dictionaries, the discriminative losses are adopted to ensure dictionaries learn more accurate representation, which is beneficial to the recognition task. Recognition of unseen images is performed through searching the class nearest to the unseen image in multiple spaces. Experiments on four widely used benchmark datasets show the effectiveness of the proposed method, and sufficient ablation experiments demonstrate that the coarse-to-fine way leads to good performances. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Real-Time Human Action Recognition Using Locally Aggregated Kinematic-Guided Skeletonlet and Supervised Hashing-by-Analysis Modelabstract3-D action recognition is referred to as the classification of action sequences which consist of 3-D skeleton joints. While many research works are devoted to 3-D action recognition, it mainly suffers from three problems: 1) highly complicated articulation; 2) a great amount of noise; and 3) low implementation efficiency. To tackle all these problems, we propose a real-time 3-D action-recognition framework by integrating the locally aggregated kinematic-guided skeletonlet (LAKS) with a supervised hashing-by-analysis (SHA) model. We first define the skeletonlet as a few combinations of joint offsets grouped in terms of the kinematic principle and then represent an action sequence using LAKS, which consists of a denoising phase and a locally aggregating phase. The denoising phase detects the noisy action data and adjusts it by replacing all the features within it with the features of the corresponding previous frame, while the locally aggregating phase sums the difference between an offset feature of the skeletonlet and its cluster center together over all the offset features of the sequence. Finally, the SHA model combines sparse representation with a hashing model, aiming at promoting the recognition accuracy while maintaining high efficiency. Experimental results on MSRAction3D, UTKinectAction3D, and Florence3DAction datasets demonstrate that the proposed method outperforms state-of-the-art methods in both recognition accuracy and implementation efficiency. Shaofan Wang 0001, Dehui Kong, Lichun Wang 0002 |
IEEE Trans. Cybern. | 4 |
| 2021 | Zero-shot Recognition with Image Attributes Generation using Hierarchical Coupled Dictionary LearningabstractZero-shot learning (ZSL) aims to recognize images from unseen (novel) classes with the training images from seen classes. The attributes of each class is exploited as auxiliary semantic information. Recently most ZSL approaches focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the unseen classes. However, few works study whether the auxiliary semantic information in the class-level is extensive enough or not for the ZSL task. To tackle such problem, we propose a hierarchical coupled dictionary learning (HCDL) approach to hierarchically align the visual-semantic structures in both the class-level and the image-level. Firstly, the class-level coupled dictionary is trained to establish a basic connection between visual space and semantic space. Then, the image attributes are generated based on the basic connection. Finally, the fine-grained information can be embedded by training the image-level coupled dictionary. Zero-shot recognition is performed in multiple spaces by searching the nearest neighbor class of the unseen image. Experiments on two widely used benchmark datasets show the effectiveness of the proposed approach. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
MMAsia | 2 |
| 2021 | Joint Transferable Dictionary Learning and View Adaptation for Multi-view Human Action RecognitionabstractMulti-view human action recognition remains a challenging problem due to large view changes. In this article, we propose a transfer learning-based framework called transferable dictionary learning and view adaptation (TDVA) model for multi-view human action recognition. In the transferable dictionary learning phase, TDVA learns a set of view-specific transferable dictionaries enabling the same actions from different views to share the same sparse representations, which can transfer features of actions from different views to an intermediate domain. In the view adaptation phase, TDVA comprehensively analyzes global, local, and individual characteristics of samples, and jointly learns balanced distribution adaptation, locality preservation, and discrimination preservation, aiming at transferring sparse features of actions of different views from the intermediate domain to a common domain. In other words, TDVA progressively bridges the distribution gap among actions from various views by these two phases. Experimental results on IXMAS, ACT4 2 , and NUCLA action datasets demonstrate that TDVA outperforms state-of-the-art methods. Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | Hardness-Aware Dictionary Learning: Boosting Dictionary for RecognitionabstractSparse representation is a powerful tool in many visual applications since images can be represented effectively and efficiently with a dictionary. Conventional dictionary learning methods usually treat each training sample equally, which would lead to the degradation of recognition performance when the samples from same category distribute dispersedly. This is because the dictionary focuses more on easy samples (known as highly clustered samples), and those hard samples (known as widely distributed samples) are easily ignored. As a result, the test samples which exhibit high dissimilarities to most of intra-category samples tend to be misclassified. To circumvent this issue, this paper proposes a simple and effective hardness-aware dictionary learning (HADL) method, which considers training samples discriminatively based on the AdaBoost mechanism. Different from learning one optimal dictionary, HADL learns a set of dictionaries and corresponding sub-classifiers jointly in an iterative fashion. In each iteration, HADL learns a dictionary and a sub-classifier, and updates the weights based on the classification errors given by current sub-classifier. Those correctly classified samples are assigned with small weights while those incorrectly classified samples are assigned with large weights. Through the iterated learning procedure, the hard samples are associated with different dictionaries. Finally, HADL combines the learned sub-classifiers linearly to form a strong classifier, which improves the overall recognition accuracy effectively. Experiments on well-known benchmarks show that HADL achieves promising classification results. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Multim. | 1 |
| 2021 | Discriminative matrix-variate restricted Boltzmann machine classification model
Pengyu Tian, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
Wirel. Networks | 4 |
| 2020 | Matrix-variate variational auto-encoder with applications to image process
Huixia Yan, Junbin Gao, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Effective human action recognition using global and local offsets of skeleton joints
Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
Multim. Tools Appl. | 4 |
| 2015 | Synthesis of sign language co-articulation based on key frames
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 2 |
| 2015 | High-Resolution Light Field Capture With Coded ApertureabstractAcquiring light field with larger angular resolution and higher spatial resolution in low cost is the goal of light field capture. Combining or modifying traditional optical cameras is a usual method for designing light field capture equipment, among which most models should deliberate trade-off between angular and spatial resolution, but augmenting coded aperture avoids this consideration by multiplexing information from different views. On the basis of coded aperture, this paper suggests an improved light field camera model that has double measurements and one mask. The two compressive measurements are respectively realized by a coded aperture and a random convolution CMOS imager, the latter is used as imaging sensor of the camera. The single mask design permits high light efficiency, which enables the sampling images to have high clarity. The double measurement design keeps more correlation information, which is conductive to enhancing the reconstructed light field. The higher clarity and more correlation of samplings mean higher quality of rebuilt light field, which also means higher resolution under condition of a lower PSNR requirement for rebuilt light field. Experimental results have verified advantage of the proposed design: compared with the representative mask-based light field camera models, the proposed model has the highest reconstruction quality and a higher light efficiency. Lichun Wang 0002, Dehui Kong |
IEEE Trans. Image Process. | 2 |
| 2014 | Chinese Sign Language animation generation considering context
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 3 |
| 2014 | Similarity Assessment Model for Chinese Sign Language VideosabstractThis paper proposes a model for measuring similarity between videos which content is Chinese Sign Language (CSL), vision and sign language semantic are considered for the model. Vision component of the model is distance based on Volume Local Binary Patterns (VLBP), which is robust for motion and illumination. Semantic component of the model computes semantic distance based on definition of sign language semantic, which is defined as hand shape, location, orientation and movements. While quantizing the sign language semantic, contour is used to measure shape and orientation; trajectory is used for measuring location and movement. Experiment results show that proposed assessment model is effective and assessing result given by the model is close to subjective scoring. Lichun Wang 0002, Dehui Kong |
IEEE Trans. Multim. | 1 |
| 2012 | Making smooth transitions based on a multi-dimensional transition database for joining Chinese sign-language videos
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 2 |
| 2009 | CSLML: a markup language for expressive Chinese sign language synthesisabstractAbstract This paper presents a Chinese Sign Language Markup Language (CSLML), which is developed for expressive Chinese sign language synthesis by introducing features and structure of sign language prosody. The tags of CSLML are divided into two levels: function level and phonetic level. Function level provides abstract information about signed content and prosody, so it facilitates text annotating for text‐driven automatic synthetic system and adapts to diversified synthetic methods, such as motion capture animation or image‐based synthesis, which may not be good at processing lower‐level information. Phonetic level provides detailed behavioral manners based on phonetics and phonology to interpret the meaning in function level. It facilitates creation and edit of any motions. The two levels co‐exist in CSLML documents and high‐level description can be mapped into corresponding low‐level behavior to provide one‐to‐many variability and expression of synthesis. Therefore, we also introduce a framework for this mapping processing and exhibit results of our animated prototype system based on this framework. Copyright © 2009 John Wiley & Sons, Ltd. Kejia Ye, Lichun Wang 0002 |
Comput. Animat. Virtual Worlds | 3 |