EDBT 2026 Demo / reviewers in the wild / expert
Changsong Liu
dblp:77/6356
· DBLP profile ↗
65ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 9 first-author · 13 since 2021Databases, data management, data science and information retrieval · 19 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image-plane geometric decoding for view-invariant indoor scene reconstruction
Yimeng Fan, Changsong Liu, Lixue Xu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Spiking Depth: Depth estimation from sparse events with spiking neural networks
Dongze Liu, Yimeng Fan, Wenrui Lu, Changsong Liu, Wei Zhang 0055 |
Expert Syst. Appl. | 4 |
| 2026 | Dual-stage decoupling network for RGB-thermal salient object detection
Rudong Jing, Yimeng Fan, Dongze Liu, Changsong Liu |
Neurocomputing | 8 |
| 2026 | MS2Edge: Towards energy-efficient and crisp edge detection with multi-scale residual learning in SNNs
Yimeng Fan, Changsong Liu, Yuzhou Dai, Yanyan Liu 0001, Wei Zhang 0055 |
Pattern Recognit. | 2 |
| 2025 | Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental DisordersabstractMental disorders, such as anxiety and depression, have become a global concern that affects people of all ages. Early detection and treatment are crucial to mitigate the negative effects these disorders can have on daily life. Although AI-based detection methods show promise, progress is hindered by the lack of publicly available large-scale datasets. To address this, we introduce the Multi-Modal Psychological assessment corpus (MMPsy), a large-scale dataset containing audio recordings and transcripts from Mandarin-speaking adolescents undergoing automated anxiety/depression assessment interviews. MMPsy also includes self-reported anxiety/depression evaluations using standardized psychological questionnaires. Leveraging this dataset, we propose Mental-Perceiver, a deep learning model for estimating mental disorders from audio and textual data. Extensive experiments on MMPsy and the DAIC-WOZ dataset demonstrate the effectiveness of Mental-Perceiver in anxiety and depression detection. Jinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu, Qianying Huang, Rumin Zhang |
AAAI | 2 |
| 2025 | A hybrid architecture of sparse convolutional neural network-transformer for enhanced spatial-geometric feature learning in surface reconstruction
Changsong Liu, Yimeng Fan, Lixue Xu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Cycle pixel difference network for crisp edge detection
Changsong Liu, Wei Zhang 0055, Yanyan Liu 0001, Yimeng Fan, Xiangnan Bai |
Neurocomputing | 1 |
| 2025 | Learning to utilize image second-order derivative information for crisp edge detection
Changsong Liu, Yimeng Fan, Wei Zhang 0055, Yanyan Liu 0001 |
Knowl. Based Syst. | 1 |
| 2024 | SFOD: Spiking Fusion Object DetectorabstractEvent cameras, characterized by high temporal resolution, high dynamic range, low power consumption, and high pixel bandwidth, offer unique capabilities for object detection in specialized contexts. Despite these advantages, the inherent sparsity and asynchrony of event data pose challenges to existing object detection algorithms. Spiking Neural Networks (SNNs), inspired by the way the human brain codes and processes information, offer a potential solution to these difficulties. However, their performance in object detection using event cameras is limited in current imple-mentations. In this paper, we propose the Spiking Fusion Object Detector (SFOD), a simple and efficient approach to SNN-based object detection. Specifically, we design a Spiking Fusion Module, achieving the first-time fusion of feature maps from different scales in SNNs applied to event cameras. Additionally, through integrating our analysis and experiments conducted during the pretraining of the back-bone network on the NCAR dataset, we delve deeply into the impact of spiking decoding strategies and loss functions on model performance. Thereby, we establish state-of-the-art classification results based on SNNs, achieving 93.7% accuracy on the NCAR dataset. Experimental results on the GEN1 detection dataset demonstrate that the SFOD achieves a state-of-the-art mAP of 32.1%, outperforming existing SNN-based approaches. Our research not only underscores the potential of SNNs in object detection with event cameras but also propels the advancement of SNNs. Code is available at https://github.com/yimeng-fan/SFOD. Yimeng Fan, Changsong Liu, Wenrui Lu |
CVPR | 3 |
| 2024 | Generating crisp boundaries using multi-scale features and mixed loss function
Changsong Liu, Wei Zhang 0055, Yanyan Liu 0001, Rudong Jing |
Appl. Intell. | 1 |
| 2024 | An effective method for small object detection in low-resolution images
Rudong Jing, Wei Zhang 0055, Yanyan Liu 0001, Changsong Liu |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | An Empirical Study on Punctuation Restoration for English, Mandarin, and Code-Switching Speech
Changsong Liu, Thi-Nga Ho, Chng Eng Siong |
ACIIDS (2) | 1 |
| 2022 | Domain Adaptation via Mutual Information Maximization for Handwriting RecognitionabstractDeep learning models for handwriting recognition have been developed in recent years. To improve the model’s generalization ability for sequence modeling task, this paper proposes to use domain adaptation with statistical distribution alignment and entropy regularization. For statistical distribution alignment, a domain adaptation loss function is proposed by using both the first and second order statistical information of deep feature representations, which is equivalent to maximizing the mutual information in feature spaces of the source domain and target domain. For entropy regularization, the entropy of the predicted text symbols of unlabeled samples in the target domain is also utilized as an additional loss function, which maximizes the mutual information between the feature space and pattern space in the target domain. Experimental results on the IAM handwriting dataset have demonstrated the effectiveness of the proposed domain adaptation method for sequence modeling task. Pei Tang, Liangrui Peng, Ruijie Yan, Haodong Shi, Changsong Liu |
ICASSP | 6 |
| 2022 | An efficient fire and smoke detection algorithm based on an end-to-end structured network
Wei Zhang 0055, Yanyan Liu 0001, Rudong Jing, Changsong Liu |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Occlusion Robust Face Recognition Based on Mask Learning With Pairwise Differential Siamese NetworkabstractDeep Convolutional Neural Networks (CNNs) have been pushing the frontier of face recognition over past years. However, existing CNN models are far less accurate when handling partially occluded faces. These general face models generalize poorly for occlusions on variable facial areas. Inspired by the fact that human visual system explicitly ignores the occlusion and only focuses on the non-occluded facial areas, we propose a mask learning strategy to find and discard corrupted feature elements from recognition. A mask dictionary is firstly established by exploiting the differences between the top conv features of occluded and occlusion-free face pairs using innovatively designed pairwise differential siamese network (PDSN). Each item of this dictionary captures the correspondence between occluded facial areas and corrupted feature elements, which is named Feature Discarding Mask (FDM). When dealing with a face image with random partial occlusions, we generate its FDM by combining relevant dictionary items and then multiply it with the original features to eliminate those corrupted feature elements from recognition. Comprehensive experiments on both synthesized and realistic occluded face datasets show that the proposed algorithm significantly outperforms the state-of-the-art systems. Lingxue Song, Dihong Gong, Zhifeng Li 0001, Changsong Liu, Wei Liu 0005 |
ICCV | 4 |
| 2019 | TH-GAN: Generative Adversarial Network Based Transfer Learning for Historical Chinese Character RecognitionabstractHistorical Chinese character recognition faces problems including low image quality and lack of labeled training samples. We propose a generative adversarial network (GAN) based transfer learning method to ease these problems. The proposed TH-GAN architecture includes a discriminator and a generator. The network structure of the discriminator is based on a convolutional neural network (CNN). Inspired by Wasserstein GAN, the loss function of the discriminator aims to measure the probabilistic distribution distance of the generated images and the target images. The network structure of the generator is a CNN based encoder-decoder. The loss function of the generator aims to minimize the distribution distance between the real samples and the generated samples. In order to preserve the complex glyph structure of a historical Chinese character, a weighted mean squared error (MSE) criterion by incorporating both the edge and the skeleton information in the ground truth image is proposed as the weighted pixel loss in the generator. These loss functions are used for joint training of the discriminator and the generator. Experiments are conducted on two tasks to evaluate the performance of the proposed TH-GAN. The first task is carried out on style transfer mapping for multi-font printed traditional Chinese character samples. The second task is carried out on transfer learning for historical Chinese character samples by adding samples generated by TH-GAN. Experimental results show that the proposed TH-GAN is effective. Junyang Cai, Liangrui Peng, Yejun Tang, Changsong Liu, Pengchao Li |
ICDAR | 4 |
| 2019 | A Modified Inception-ResNet Network with Discriminant Weighting Loss for Handwritten Chinese Character RecognitionabstractHandwritten Chinese character recognition (HCCR) is a representative large character set pattern classification task. Recently, convolutional neural networks have provided promising solutions for this challenging task. This paper adopts the modified Inception-ResNet network for handwritten Chinese character recognition, and proposes a discriminant weighting method for cross-entropy loss calculation which focuses on recognition errors in the training stage. Sparse training technique is also incorporated. Under the specific condition of utilizing the testing mini-batch mean and variance for batch normalization, the proposed method achieves improved performance on the ICDAR-2013 offline handwritten Chinese character competition dataset. Linhui Chen, Liangrui Peng, Changsong Liu, Xudong Zhang 0001 |
ICDAR | 4 |
| 2018 | Learning with rethinking: Recurrently improving convolutional neural networks through feedback
Zequn Jie, Jiashi Feng, Changsong Liu, Shuicheng Yan |
Pattern Recognit. | 4 |
| 2018 | Heteroscedastic Max-Min Distance Analysis for Dimensionality ReductionabstractMax-min distance analysis (MMDA) performs dimensionality reduction by maximizing the minimum pairwise distance between classes in the latent subspace under the homoscedastic assumption, which can address the class separation problem caused by the Fisher criterion, but is incapable of tackling heteroscedastic data properly. In this paper, we propose two heteroscedastic MMDA (HMMDA) methods to employ the differences of class covariances. Whitened HMMDA (WHMMDA) extends MMDA by utilizing the Chernoff distance as the separability measure between classes in the whitened space. Orthogonal HMMDA (OHMMDA) incorporates the maximization of the minimal pairwise Chernoff distance and the minimization of class compactness into a trace quotient formulation with an orthogonal constraint of the transformation, which can be solved by bisection search. Two variants of OHMMDA further encode the margin information by using only neighboring samples to construct the intra-class and inter-class scatters. Experiments on several UCI datasets and two face databases demonstrate the effectiveness of the HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | FoveaNet: Perspective-Aware Urban Scene ParsingabstractParsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consider the geometry property of car-captured urban scene images. Thus, they suffer from heterogeneous object scales caused by perspective projection of cameras on actual scenes and inevitably encounter parsing failures on distant objects as well as other boundary and recognition errors. In this work, we propose a new FoveaNet model to fully exploit the perspective geometry of scene images and address the common failures of generic parsing models. FoveaNet estimates the perspective geometry of a scene image through a convolutional network which integrates supportive evidence from contextual objects within the image. Based on the perspective geometry information, FoveaNet “undoes” the camera perspective projection - analyzing regions in the space of the actual scene, and thus provides much more reliable parsing results. Furthermore, to effectively address the recognition errors, FoveaNet introduces a new dense CRFs model that takes the perspective geometry as a prior potential. We evaluate FoveaNet on two urban scene parsing datasets, Cityspaces and CamVid, which demonstrates that FoveaNet can outperform all the well-established baselines and provide new state-of-the-art performance. Zequn Jie, Wei Wang 0108, Changsong Liu, Jimei Yang, Xiaohui Shen, Zhe Lin 0001, Qiang Chen 0007, Shuicheng Yan, Jiashi Feng |
ICCV | 4 |
| 2017 | Semi-Supervised Transfer Learning for Convolutional Neural Network Based Chinese Character RecognitionabstractAlthough transfer learning has aroused researchers' great interest, how to utilize the unlabeled data is still an open and important problem in this area. We propose a novel semi-supervised transfer learning (STL) method by incorporating Multi-Kernel Maximum Mean Discrepancy (MK-MMD) loss into the traditional fine-tuned Convolutional Neural Network (CNN) transfer learning framework for Chinese character recognition. The proposed method includes three steps. First, a CNN model is trained by massive labeled samples in the source domain. Then the CNN model is fine-tuned by a few labeled samples in the target domain. Finally, the CNN model is trained with both a large number of unlabeled samples and the limited labeled samples in the target domain to minimize the MK-MMD loss. Experiments investigate detailed configurations and parameters of the proposed STL method with several frequently used CNN structures including AlexNet, GoogLeNet, and ResNet. Experimental results on practical Chinese character transfer learning tasks, such as Dunhuang historical Chinese character recognition, indicate that the proposed method can significantly improve recognition accuracy in the target domain. Yejun Tang, Liangrui Peng, Changsong Liu |
ICDAR | 4 |
| 2017 | Layout and Perspective Distortion Independent Recognition of Captured Chinese Document ImageabstractThis paper introduced a layout and perspective distortion independent recognition framework for captured Chinese document image. Under the framework, 1) Conditional random field (CRF) is employed for text line extraction from a global point of view. As the text line extraction is layout independent it could be widely used in different type of document images 2) A text line image based perspective distortion correction method is detailed and used in three different ways. 3) The text line extraction and perspective distortion correction are combined with character recognition to construct a recognition system. On three captured document image datasets, the proposed framework improves the accuracies from 94.03% to 95.20%, 13.01% to 93.71% and 10.63% to 92.68% respectively for different distortion degrees. The experimental results demonstrate that the introduced recognition framework is promising for solving layout and perspective distortion problems in captured document image recognition. Yuefang Sun, Changsong Liu |
ICDAR | 3 |
| 2017 | Local Discriminant Training and Global Optimization for Convolutional Neural Network Based Handwritten Chinese Character RecognitionabstractThis paper investigates local discriminant training and global optimization methods for Convolutional Neural Network (CNN) to improve its discriminant ability and recognition accuracy. For local discriminant training, we propose to combine triplet loss and softmax with cross-entropy loss as the loss function. The triplet loss is incorporated into an additional fully-connected layer before the final fully-connected layer of a CNN model. For global optimization, we use Conditional Random Field (CRF) to further utilize the pairwise distance of the CNN feature vectors trained with triplet loss. Experiments with different CNN models on handwritten Chinese character samples show that the combined local discriminant training and global optimization scheme achieves better character recognition accuracy and confidence analysis performance. Xiangsheng Zeng, Donglai Xiang, Liangrui Peng, Changsong Liu, Xiaoqing Ding |
ICDAR | 4 |
| 2017 | Discriminative Transformation for Multi-Dimensional Temporal SequencesabstractFeature space transformation techniques have been widely studied for dimensionality reduction in vector-based feature space. However, these techniques are inapplicable to sequence data because the features in the same sequence are not independent. In this paper, we propose a method called max-min inter-sequence distance analysis (MMSDA) to transform features in sequences into a low-dimensional subspace such that different sequence classes are holistically separated. To utilize the temporal dependencies, MMSDA first aligns features in sequences from the same class to an adapted number of temporal states, and then, constructs the sequence class separability based on the statistics of these ordered states. To learn the transformation, MMSDA formulates the objective of maximizing the minimal pairwise separability in the latent subspace as a semi-definite programming problem and provides a new tractable and effective solution with theoretical proofs by constraints unfolding and pruning, convex relaxation, and within-class scatter compression. Extensive experiments on different tasks have demonstrated the effectiveness of MMSDA. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Hao Wang 0005, Ying Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Jointly Learning Grounded Task Structures from Language Instruction and Visual DemonstrationabstractTo enable language-based communication and collaboration with cognitive robots, this paper presents an approach where an agent can learn task models jointly from language instruction and visual demonstration using an And-Or Graph (AoG) representation.The learned AoG captures a hierarchical task structure where linguistic labels (for language communication) are grounded to corresponding state changes from the physical environment (for perception and action).Our empirical results on a cloth-folding domain have shown that, although state detection through visual processing is full of uncertainties and error prone, by a tight integration with language the agent is able to learn an effective AoG for task representation.The learned AoG can be further applied to infer and interpret on-going actions from new visual demonstration using linguistic labels at different levels of granularity. Changsong Liu, Sari Saba-Sadiya, Nishant Shukla, Yunzhong He, Song-Chun Zhu, Joyce Y. Chai |
EMNLP | 1 |
| 2016 | Grounded Semantic Role LabelingabstractShaohua Yang, Qiaozi Gao, Changsong Liu, Caiming Xiong, Song-Chun Zhu, Joyce Y. Chai. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Qiaozi Gao, Changsong Liu, Caiming Xiong, Song-Chun Zhu, Joyce Y. Chai |
HLT-NAACL | 3 |
| 2016 | Combining multiple biometric traits with an order-preserving score fusion algorithm
Yicong Liang, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue |
Neurocomputing | 3 |
| 2016 | An empirical investigation into the effect of slice types on slice-based cohesion metrics
Yibiao Yang, Changsong Liu, Hongmin Lu, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 3 |
| 2016 | Sufficient Canonical Correlation AnalysisabstractCanonical correlation analysis (CCA) is an effective way to find two appropriate subspaces in which Pearson's correlation coefficients are maximized between projected random vectors. Due to its well-established theoretical support and relatively efficient computation, CCA is widely used as a joint dimension reduction tool and has been successfully applied to many image processing and computer vision tasks. However, as reported, the traditional CCA suffers from overfitting in many practical cases. In this paper, we propose sufficient CCA (S-CCA) to relieve CCA's overfitting problem, which is inspired by the theory of sufficient dimension reduction. The effectiveness of S-CCA is verified both theoretically and experimentally. Experimental results also demonstrate that our S-CCA outperforms some of CCA's popular extensions during the prediction phase, especially when severe overfitting occurs. Yiwen Guo, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue |
IEEE Trans. Image Process. | 3 |
| 2015 | Learning to Mediate Perceptual Differences in Situated Human-Robot DialogueabstractIn human-robot dialogue, although a robot and its human partner are co-present in a shared environment, they have significantly mismatched perceptual capabilities (e.g., recognizing objects in the surroundings). When a shared perceptual basis is missing, it becomes difficult for the robot to identify referents in the physical world that are referred to by the human (i.e., a problem of referential grounding). To overcome this problem, we have developed an optimization based approach that allows the robot to detect and adapt to perceptual differences. Through online interaction with the human, the robot can learn a set of weights indicating how reliably/unreliably each dimension (e.g., object type, object color, etc.) of its perception of the environment maps to the human's linguistic descriptors and thus adjust its word models accordingly. Our empirical evaluation has shown that this weight-learning approach can successfully adjust the weights to reflect the robot's perceptual limitations. The learned weights, together with updated word models, can lead to a significant improvement for referential grounding in future dialogues. Changsong Liu, Joyce Y. Chai |
AAAI | 1 |
| 2015 | Heteroscedastic max-min distance analysisabstractMany discriminant analysis methods such as LDA and HLDA actually maximize the average pairwise distances between classes, which often causes the class separation problem. Max-min distance analysis (MMDA) addresses this problem by maximizing the minimum pairwise distance in the latent subspace, but it is developed under the homoscedastic assumption. This paper proposes Heteroscedastic MMDA (HMMDA) methods that explore the discriminative information in the difference of intra-class scatters for dimensionality reduction. WHMMDA maximizes the minimal pairwise Chenoff distance in the whitened space. OHMMDA incorporates this objective and the minimization of class compactness into a trace quotient formulation and imposes an orthogonal constraint to the final transformation, which can be solved by a bisection search algorithm. Two variants of OHMMDA are further proposed to encode the margin information. Experiments on several UCI Machine Learning datasets and the Yale Face database demonstrate the effectiveness of the proposed HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
CVPR | 3 |
| 2015 | Application of Face Verification in Automated Passenger Clearance System
Wentao Shen, Yicong Liang, Xiaoqing Ding, Changsong Liu |
ICPRAM (2) | 4 |
| 2015 | Restoring camera-captured distorted document images
Changsong Liu, Baokang Wang, Xiaoqing Ding |
Int. J. Document Anal. Recognit. | 1 |
| 2015 | Exploring More Representative States of Hidden Markov Model in Optical Character Recognition: A Clustering-Based Model Pre-Training ApproachabstractHidden Markov Model (HMM) is an effective method to describe sequential signals in many applications. As to model estimation issue, common training algorithm only focuses on the optimization of model parameters. However, model structure influences system performance as well. Although some structure optimization methods are proposed, they are usually implemented as an independent module before parameter optimization. In this paper, the clustering feature of states in HMM is discussed through comparing the mechanism of Quadratic Discriminant Function (QDF) classifier and HMM. Then, through the clustering effect of Viterbi training and Baum–Welch training, a novel clustering-based model pre-training approach is proposed. It can optimize model parameters and model structure by turns, until the representative states of all models are explored. Finally, the proposed approach is evaluated on two typical OCR applications, printed and handwritten Arabic text line recognition. And it is compared with some other optimization methods. The improvement of character recognition performance proves the proposed approach can make more precise state allocation. And the representative states are benefit to HMM decoding. Xiaoqing Ding, Liangrui Peng, Changsong Liu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2015 | Importance sampling based discriminative learning for large scale offline handwritten Chinese character recognition
Xiaoqing Ding, Changsong Liu |
Pattern Recognit. | 4 |
| 2014 | Collaborative effort towards common ground in situated human-robot dialogueabstractIn situated human-robot dialogue, although humans and robots are co-present in a shared environment, they have significantly mismatched capabilities in perceiving the shared environment. Their representations of the shared world are misaligned. In order for humans and robots to communicate with each other successfully using language, it is important for them to mediate such differences and to establish common ground. To address this issue, this paper describes a dialogue system that aims to mediate a shared perceptual basis during human-robot dialogue. In particular, we present an empirical study that examines the role of the robot's collaborative effort and the performance of natural language processing modules in dialogue grounding. Our empirical results indicate that in situated human-robot dialogue, a low collaborative effort from the robot may lead its human partner to believe a common ground is established. However, such beliefs may not reflect true mutual understanding. To support truly grounded dialogues, the robot should make an extra effort by making its partner aware of its internal representation of the shared world. Joyce Y. Chai, Lanbo She, Spencer Ottarson, Cody Littley, Changsong Liu, Kenneth Hanson |
HRI | 6 |
| 2014 | An MQDF-CNN Hybrid Model for Offline Handwritten Chinese Character RecognitionabstractAn MQDF-CNN hybrid model is presented for offline handwritten Chinese character recognition. The main idea behind MQDF-CNN hybrid model is that the significant difference on features and classification mechanisms between MQDF and CNN can complement each other. Linear confidence accumulation and multiplication confidence criteria are used for fusion outputs of MQDF and CNN. Experiments have been conducted on CASIA-HWDB1.1 and ICDAR2013 offline handwritten Chinese character recognition competition dataset. On both datasets, CNN beats MQDF by more than 1% of the accuracy, and the MQDF-CNN hybrid model has achieved the test accuracies of 92.03% and 94.44% respectively. The result on competition dataset is comparable to the state-of-the-art result though less training samples and only one CNN is used. Xin Li 0144, Changsong Liu, Xiaoqing Ding, Youxin Chen |
ICFHR | 3 |
| 2014 | Topic Language Model Adaption for Recognition of Homologous Offline Handwritten Chinese Text ImageabstractAs the content of a full text page usually focuses on a specific topic, a topic language model adaption method is proposed to improve the recognition performance of homologous offline handwritten Chinese text image. Firstly, the text images are recognized with a character based bi-gram language model. Secondly, the topic of the text image is matched adaptively. Finally, the text image is recognized again with the best matched topic language model. To obtain a tradeoff between the recognition performance and computational complexity, a restricted topic language model adaption method is further presented. The methods have been evaluated on 100 offline Chinese text images. Compared to the general language model, the topic language model adaption has reduced the relative error rate by 11.94%. The restricted topic language model has lessened the running time by 19.22% at the cost of losing 0.35% of the accuracy. Xiaoqing Ding, Changsong Liu |
IEEE Signal Process. Lett. | 3 |
| 2013 | Towards Situated Dialogue: Revisiting Referring Expression GenerationabstractIn situated dialogue, humans and agents have mismatched capabilities of perceiving the shared environment.Their representations of the shared world are misaligned.Thus referring expression generation (REG) will need to take this discrepancy into consideration.To address this issue, we developed a hypergraph-based approach to account for group-based spatial relations and uncertainties in perceiving the environment.Our empirical results have shown that this approach outperforms a previous graph-based approach with an absolute gain of 9%.However, while these graph-based approaches perform effectively when the agent has perfect knowledge or perception of the environment (e.g., 84%), they perform rather poorly when the agent has imperfect perception of the environment (e.g., 45%).This big performance gap calls for new solutions to REG that can mediate a shared perceptual basis in situated dialogue. Changsong Liu, Lanbo She, Joyce Y. Chai |
EMNLP | 2 |
| 2013 | Cross-Language Sensitive Words Distribution Map: A Novel Recognition-Based Document Understanding Method for Uighur and TibetanabstractCross-language document recognition and understanding have urgent realistic needs and extensive application prospects. In this paper, we propose a novel recognition-based Uighur and Tibetan document understanding method, termed "cross-language sensitive words distribution map" (CSWDM). In our unified recognition-understanding framework, digital Uighur/Tibetan document images are first recognized using OCR technology, and then CSWDM labels the Chinese information of sensitive words on the recognized transcriptions or directly on the original digital images, thus the space location and occurrence frequency of these sensitive words can be intuitively represented. With such information, readers can roughly understand the theme and meaning of the cross-language documents. Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu |
ICDAR | 4 |
| 2013 | A Novel Baseline-independent Feature Set for Arabic Handwriting RecognitionabstractHMM-based analytical methods have been widely used for Arabic handwriting recognition. A key factor influencing the performance of HMM-based systems is the features extracted from a sliding window. In this paper, we propose a novel baseline-independent feature set extracted from a wider sliding window to directly capture the contextual information. This feature set is a combination of center of mass based log-space distribution features and inverse percentile features. Center of mass based log-space distribution features use a normalized histogram to describe the distribution of foreground pixels in different direction and distances with respect to the center of mass. Experiments on the IFN/ENIT database demonstrate the effectiveness of the proposed feature set. Further, this feature set can be combined with some popular baseline-independent features to form a large feature set, which achieves comparable results with several state-of-the-art systems using a simple HMM-based architecture. Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu |
ICDAR | 4 |
| 2013 | Similar Pattern Discriminant Analysis for Improving Chinese Character Recognition AccuracyabstractIn this paper, a similar pattern discriminant analysis method is proposed. It optimizes the feature projection matrix based on similar pattern pairs and aims to extract targeted features for similar pattern discrimination. For improving Chinese character recognition accuracy, we introduce a cascade modified quadratic discriminant function (MQDF) model to combine linear discriminant analysis (LDA) and similar pattern discriminant analysis. The proposed method is investigated and compared with compound Mahalanobis function (CMF) on two data sets. The results indicate that the cascade MQDF achieves a better improvement and higher recognition accuracies than CMF. The relative recognition errors have been decreased up to 19.73% and 15.59% respectively on HCL2000 and THU-HCD datasets with respect to single MQDF. Changsong Liu, Xiaoqing Ding |
ICDAR | 2 |
| 2013 | Modeling Collaborative Referring for Situated Referential Grounding
Changsong Liu, Lanbo She, Joyce Y. Chai |
SIGDIAL Conference | 1 |
| 2012 | Integrating word acquisition and referential grounding towards physical world interactionabstractIn language-based interaction between a human and an artificial agent (e.g., robot) in a physical world, because the human and the agent have different knowledge and capabilities in perceiving the shared environment, referential grounding is very difficult. To facilitate such interaction, it is important for the agent to continuously learn and acquire knowledge about the environment through interactions with humans and incorporate the learned knowledge in grounding references from human utterances. To address this issue, this paper presents a graph-based approach for referential grounding and examines how referential grounding and word acquisition influence each other in physical world interaction. Our empirical results have shown that for most words, automated word acquisition through interaction improves referential grounding performance. However, this is not the case for words describing object types, where human supervision is important. Nevertheless, better referential grounding enables more accurate acquisition of word meanings, which in turn further improves grounding performance for references in subsequent utterances. Changsong Liu, Joyce Y. Chai |
ICMI | 2 |
| 2012 | Analyzing the information entropy of states to optimize the number of states in an HMM-based off-line handwritten Arabic word recognizer
Xiaoqing Ding, Liangrui Peng, Changsong Liu |
ICPR | 4 |
| 2012 | Towards Mediating Shared Perceptual Basis in Situated Dialogue
Changsong Liu, Joyce Y. Chai |
SIGDIAL Conference | 1 |
| 2011 | A Multi-scale Text Line Segmentation Method in Freestyle Handwritten DocumentsabstractText lines in free-style handwritten documents are often curved, touch or overlap with each other, which presents a challenge for text line segmentation. In this paper, we proposed a novel text line segmentation method that utilizes the advantages of algorithms in both the small scale and large scale. A path is dynamically detected between each pair of neighboring text lines to separate them. During the process, the line-separating path's coordinate in each step is determined by a three-stage multi-scale method that combines (1) a simple local minima search algorithm, (2) the technique based on following the contour of the foreground component and (3) the piecewise projection profile. Without training, our method has achieved a high segmentation accuracy on plenty of samples, which proves its strong adaptability to various line conditions. Experimental results show that the proposed method outperforms traditional methods. Yangdong Gao, Xiaoqing Ding, Changsong Liu |
ICDAR | 3 |
| 2011 | A Novel Short Merged Off-line Handwritten Chinese Character String Segmentation Algorithm Using Hidden Markov ModelabstractHidden Markov model (called "HMM" for short) has been a widespread method to segment sequential data in speech recognition and DNA sequence analysis. According to the same principle, it can be also used in segmenting short merged off-line handwritten Chinese character strings, which is a tough issue but often met in practice. Because HMM is still not a common method in this field nowadays, in this paper, we will introduce a novel algorithm using HMM for the segmentation issue above. Eventually, this segmentation algorithm can achieve an applicable performance even when 3755 character classes are compressed into similar characters classes with only 1% amount of original ones, and it also shows an enormous potential of segmenting long text lines. Xiaoqing Ding, Changsong Liu |
ICDAR | 3 |
| 2011 | On-line Chinese Character Recognition System for Overlapping SamplesabstractWe proposed a new process strategy for on-line handwriting Chinese Character recognition and applied it to overlapping samples. On one hand, those samples are evaluated on stroke level by support vector machine, on the other hand, we do character level evaluation basing on a character pair search model. Then a merging strategy was proposed to filter out correct segmentation positions. We test our strategy on samples from real context, verifying that our strategy performs better than traditional over-segmentation and merging method. Changsong Liu, Yanming Zou |
ICDAR | 2 |
| 2011 | MQDF Discriminative Learning Based Offline Handwritten Chinese Character RecognitionabstractThis paper has proposed a discriminative learning method of modified quadratic discriminant function (MQDF) based on sample importance weights. Firstly, sample importance function is derived from distance based recognition results under bayes decision rule. It weights samples according to extended recognition confidence. On these weighted samples, parameters of MQDF are modulated indirectly by re-estimating the mean vector and covariance matrix. The proposed method is investigated and compared with other discriminative learning methods about MQDF on THU-HCD offline Chinese handwriting sets. The results show that the proposed method has improved the basic MQDF drastically and outperforms other methods compared. Xiaoqing Ding, Changsong Liu |
ICDAR | 3 |
| 2011 | An Improved Scene Text Extraction Method Using Conditional Random Field and Optical Character RecognitionabstractOver the past few years, research on scene text extraction has developed rapidly. Recently, condition random field (CRF) has been used to give connected components (CCs) 'text' or 'non-text' labels. However, a burning issue in CRF model comes from multiple text lines extraction. In this paper, we propose a two-step iterative CRF algorithm with a Belief Propagation inference and an OCR filtering stage. Two kinds of neighborhood relationship graph are used in the respective iterations for extracting multiple text lines. Furthermore, OCR confidence is used as an indicator for identifying the text regions, while a traditional OCR filter module only considered the recognition results. The first CRF iteration aims at finding certain text CCs, especially in multiple text lines, and sending uncertain CCs to the second iteration. The second iteration gives second chance for the uncertain CCs and filter false alarm CCs with the help of OCR. Experiments based on the public dataset of ICDAR 2005 prove that the proposed method is comparative with the existing algorithms. Changsong Liu, Xiaoqing Ding, Kongqiao Wang |
ICDAR | 2 |
| 2010 | Multi-font printed Mongolian document recognition system
Liangrui Peng, Changsong Liu, Xiaoqing Ding, Jianming Jin, Youshou Wu, Yanhua Bao |
Int. J. Document Anal. Recognit. | 2 |
| 2008 | Arbitrary warped document image restoration based on segmentation and Thin-Plate SplinesabstractWarping is a common appearance in camera captured document images. It is the primary factor that makes such kind of document images hard to be recognized. Therefore it is necessary to restore warped document image before recognition. In this paper, a novel restore method is presented. The method takes a rough line segmentation and character segmentation firstly in order to estimate the warping direction. Then several pairs of key points mapping between the original image and the restored image are determined and thin-plate splines (TPS), which is an interpolation algorithm, is introduced to restore the image. Such process can effectively describe the warping direction of the document and successfully restore the image. Some experimental results show the effect of the image restoration and compare the recognition rate before and after the restoration based on a same OCR application. Changsong Liu, Xiaoqing Ding, Yanming Zou |
ICPR | 2 |
| 2005 | Gabor filters-based feature extraction for character recognition
Xiaoqing Ding, Changsong Liu |
Pattern Recognit. | 3 |
| 2004 | Contextual post-processing based on the confusion matrix in offline handwritten Chinese script recognition
Chew Lim Tan, Xiaoqing Ding, Changsong Liu |
Pattern Recognit. | 4 |
| 2003 | A Cylindrical Surface Model to Rectify the Bound Document ImageabstractWe propose a novel approach on how to rectify the photo image of the bound document. The surface of the document is modeled by a cylindrical surface. By the geometry of camera image formation, the equations using the cue of directrixes to map the points on the surface in the 3D scene to the points on the image plane are achieved. Baselines of the horizontal text line are extracted as projections of directrixes to estimate the bending extent of the surface, and then the images are rectified. The proposed method needs no auxiliary device. Experimental results are presented to demonstrate the feasibility and the application of the method. Huaigu Cao, Xiaoqing Ding, Changsong Liu |
ICCV | 3 |
| 2003 | Rectifying the Bound Document Image Captured by the Camera: A Model Based ApproachabstractA model based approach for rectifying the camera image of the bound document has been developed, i.e., the surface of the document is represented by a general cylindrical surface. The principle of using the model to unwrap the image is discussed. Practically, the skeleton of each horizontal text line is extracted to help estimate the parameter of the model, and rectify the images. To use the model, only a few priori is required, and no more auxiliary device is necessary. Experiment results are given to demonstrate the feasibility and the stability of the method. Huaigu Cao, Xiaoqing Ding, Changsong Liu |
ICDAR | 3 |
| 2002 | Automatic performance evaluation of printed Chinese character recognition systems
Chi Fang, Changsong Liu, Liangrui Peng, Xiaoqing Ding |
Int. J. Document Anal. Recognit. | 2 |
| 2001 | An Automatic Performance Evaluation Method for Document Page SegmentationabstractAutomatic performance evaluation for a document page segmentation module is necessary, as OCR products are used to manipulate large scale of documents with complex layout, especially for newspapers. The paper presents a region-based method to evaluate the performance of a page segmentation module by analyzing geometric region relationships between the segmentation results and the preset ground-truth. The ground-truth is not only the correct answer to page segmentation, but also the comparison benchmark of the automatic evaluation, so it has more restricted geometric constraints. The region-matching algorithm is realized by searching the equal region in the segmentation results for each region in the ground-truth. The performance parameters are calculated based on the matching results. An experiment is given to test two page segmentation modules in a popular Chinese OCR product-THOCR2000, and the results show this method is effective. Liangrui Peng, Changsong Liu, Xiaoqing Ding, Jirong Zheng |
ICDAR | 3 |
| 2001 | Character Extraction and Recognition in Natural Scene ImagesabstractWith the proposal of the concept of a "smart camera", character recognition in natural scene images has become an interesting but difficult task nowadays. In this paper, we propose an algorithm for extracting characters from text regions of natural scene images with complex backgrounds. Our method first clusters the color feature vectors of the text regions into a number of color classes by applying a modified coarse-fine fuzzy c-means algorithm. Then, different slices are constructed according to these color classes. Characters are eventually extracted from the images using the information of segmentation and recognition. Some experiments have shown that this method is a promising starting point for such applications. Xiaoqing Ding, Changsong Liu |
ICDAR | 3 |
| 2001 | Form Frame Line Detection with Directional Single-Connected ChainabstractIn this paper, a novel form frame line detection algorithm is proposed based on the directional single-connected chain (DSCC). Defined as an array of black pixel run-lengths, DSCC works very well as an image structure element or vector in our vectorization algorithm. By merging multiple DSCCs under some constraints, we are able to extract the form frame lines automatically yet fast. The speed of our algorithm is comparable with some well-known projection methods. Experiments show that our algorithm is fast, resistant to moderate serious line breaks and can detect diagonal lines with any angle. Yefeng Zheng 0001, Changsong Liu, Xiaoqing Ding, Shiyan Pan |
ICDAR | 2 |
| 2001 | Location and interpretation of destination addresses on handwritten Chinese envelopes
Junliang Xue, Xiaoqing Ding, Changsong Liu, Weiwei Qian |
Pattern Recognit. Lett. | 3 |
| 2000 | Gray-Scale Character Image Recognition Based on Fuzzy DCT Transform FeaturesabstractWe propose a method for recognizing gray-scale text images with noise and of extraordinary low resolution. First, a fuzzy classification of pixels in the gray-scale character images is applied to form the fuzzy attributed pixel graph using integral ratio techniques. Second, a DCT transform is applied on this fuzzy graph and parts of the frequency domain components are selected and compressed to produce the features for recognition. By using the modified quadratic discriminant function in the classifier, we achieve satisfactory recognition results with the average recognition rate of 96.82% on 15 pixel/spl times/15 pixel character images and at the speed of over 30 characters/s. Xiaoqing Ding, Changsong Liu |
ICPR | 3 |
| 2000 | Multi-Scale Feature Extraction and Nested-Subset Classifier Design for High Accuracy Handwritten Character RecognitionabstractBoth efficient representation and robust classification are essential to high-performance cursive offline handwritten Chinese character recognition. A novel multi-scale feature extraction method is presented based on the information entropy theory. Feature detection and compression are thus combined into an integrated optimization process. A series of optimal feature-spaces are constructed at varying values of the scale parameter and the best one is obtained with the maximum LDA criterion over the scale interval. For more robust classification, we introduce a structure into the Mahalanobis distance classifier and strike the balance between machine capacity and the performance on the training data in light of the ideas of structural risk minimization. A high accuracy recognition system is developed based on the new methods and for the first time, 4 widely different databases ranging from regular to completely unconstrained with several structural distortions and stroke connections are fully tested. The accuracies of 99.S% on regular database and 88.4% on cursive one at the speed of over 40 characters/s are achieved. Jiayong Zhang, Xiaoqing Ding, Changsong Liu |
ICPR | 3 |
| 1999 | Destination Address Block Location on Handwritten Chinese EnvelopeabstractHandwritten Chinese envelopes have three obvious characteristics: handwritten text lines are the main part of the envelope image; there is no distinct border between the destination address block (DAB) and sender address block (SAB); and handwritten Chinese characters are composed of complicated strokes with more freedom. A method for segmenting the envelope image into several candidate DABs and selecting the best one as the DAB is not yet suitable for handwritten Chinese envelopes. This paper presents a new method of DAB location on envelopes based on extraction of text lines, according to the characteristics of handwritten Chinese envelopes. Our method divides the DAB location problem into several simpler sub-problems and solves them respectively. Experimental results of extraction rate for DAB and average processing time on 500 samples directly sampled from mail sorting machines shows that this method is effective and takes less CPU time. Junliang Xue, Xiaoqing Ding, Changsong Liu, Shiyan Pan, Hongwei Kong |
ICDAR | 3 |