VLDB 2026 Research / reviewers in the wild / expert
Heng Zhang 0028
dblp:55/826-28
· DBLP profile ↗
18ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-9448-4031ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CHSAM: Efficient Scene Text Segmentation via SAM with Convolutional Adapters and Hierarchical Decoding
Jing-Yao Zhang, Heng Zhang 0028 |
ICDAR (3) | 2 |
| 2024 | Adaptive Scaling and Refined Pyramid Feature Fusion Network for Scene Text Segmentation
Tian-Zuo Li, Heng Zhang 0028, Xiao-Hui Li 0012 |
ICDAR (5) | 2 |
| 2024 | Deep Metric Learning with Cross-Writer Attention for Offline Signature Verification
Lu-Rong Ling, Heng Zhang 0028, Cheng-Lin Liu 0001 |
ICDAR (2) | 2 |
| 2024 | An approach for handwritten Chinese text recognition unifying character segmentation and recognition
Mingming Yu, Heng Zhang 0028, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2022 | An Efficient Prototype-Based Model for Handwritten Text Recognition with Multi-loss Fusion
Mingming Yu, Heng Zhang 0028, Cheng-Lin Liu 0001 |
ICFHR | 2 |
| 2021 | Regularizing CTC in Expectation-Maximization Framework with Application to Handwritten Text RecognitionabstractConnectionist Temporal Classification (CTC) is an objective function for sequence learning and has shown promising results in speech and text recognition tasks. However, its inherent mechanism has not been investigated thoroughly. In this paper, we propose a theoretical explanation of CTC from the perspective of the Expectation-Maximization (EM) algorithm. Based on the EM analysis, we propose a pseudo-label-based L1 regularization and voting decoding algorithm to improve the performance of text recognition. The L1 regularization can reduce the pseudo-label estimation error, while the voting decoding algorithm modifies the built-in decoding logic of CTC and introduces a voting mechanism to the inference process. Experiments of handwritten text recognition show that the proposed method consistently improves over the CTC baseline and yields state-of-the-art results on three benchmark datasets. Likun Gao, Heng Zhang 0028, Cheng-Lin Liu 0001 |
IJCNN | 2 |
| 2020 | Table detection and cell segmentation in online handwritten documents with graph attention networksabstractIn this paper, we propose a multi-task learning approach for table detection and cell segmentation with densely connected graph attention networks in free form online documents. Each online document is regarded as a graph, where nodes represent strokes and edges represent the relationships between strokes. Then we propose a graph attention network model to classify nodes and edges simultaneously. According to node classification results, tables can be detected in each document. By combining node and edge classification resutls, cells in each table can be segmented. To improve information flow in the network and enable efficient reuse of features among layers, dense connectivity among layers is used. Our proposed model has been experimentally validated on an online handwritten document dataset IAMOnDo and achieved encouraging results. Heng Zhang 0028, Xiao-Long Yun, Jun-Yu Ye, Cheng-Lin Liu 0001 |
MMAsia | 2 |
| 2019 | Oracle Character Recognition by Nearest Neighbor Classification with Deep Metric LearningabstractOracle character is one kind of the earliest hieroglyphics, which can be dated back to Shang Dynasty in China. Oracle character recognition is important for modern archaeology, ancient text understanding, and historical chronology, etc. To overcome the limitation and class imbalance of training data in oracle character recognition, we propose a classification method based on deep metric learning. We use a convolutional neural network (CNN) to map the character images to an Euclidean space where the distance between different samples can measure their similarities such that classification can be performed by the Nearest Neighbor (NN) rule. Because new categories are still being discovered in reality, our model enables the rejection of unseen categories and the configuration of new categories. To accelerate NN classification, we also propose a prototype pruning method with little loss of accuracy. The proposed method exceeds the state of the art on the public dataset Oracle-20K and outperforms CNN with softmax layer on a new dataset Oracle-AYNU. Heng Zhang 0028, Yong-Ge Liu, Qing Yang 0002, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2017 | Keyword spotting in handwritten chinese documents using semi-markov conditional random fields
Heng Zhang 0028, Cheng-Lin Liu 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2016 | Improving short text classification by learning vector representations of both words and hidden topics
Heng Zhang 0028, Guoqiang Zhong 0001 |
Knowl. Based Syst. | 1 |
| 2015 | Towards end-to-end speech recognition for Chinese Mandarin using long short-term memory recurrent neural networksabstractEnd-to-end speech recognition systems have been successfully designed for English. Taking into account the distinctive characteristics between Chinese Mandarin and English, it is worthy to do some additional work to transfer these approaches to Chinese. In this paper, we attempt to build a Chinese speech recognition system using end-to-end learning method. The system is based on a combination of deep Long Short-Term Memory Projected (LSTMP) network architecture and the Connectionist Temporal Classification objective function (CTC). The Chinese characters (the number is about 6,000) are used as the output labels directly. To integrate language model information during decoding, the CTC Beam Search method is adopted and optimized to make it more effective and more efficient. We present the first-pass decoding results which are obtained by decoding from scratch using CTC-trained network and language model. Although these results are not as good as the performance of DNN-HMMs hybrid system, they indicate that it is feasible to choose Chinese characters as the output alphabet in the end-toend speech recognition system. Jie Li 0032, Heng Zhang 0028, Xinyuan Cai, Bo Xu 0002 |
INTERSPEECH | 2 |
| 2014 | Short Text Hashing Improved by Integrating Topic Features and Tags
Jiaming Xu 0001, Bo Xu 0002, Jun Zhao 0001, Guanhua Tian, Heng Zhang 0028, Hongwei Hao |
ICONIP (2) | 5 |
| 2014 | A robust framework for short text categorization based on topic model and integrated classifierabstractIn this paper, we propose a method for short text categorization using topic model and integrated classifier. To enrich the representation of short text, the Latent Dirichlet Allocation (LDA) model is used to extract latent topic information. While for classification, we combine two classifiers for achieving high reliability. Particularly, we train LDA models with variable number of topics using the Wikipedia corpus as external knowledge base, and extend labeled Web snippets by potential topics extracted by LDA. Then, the enriched representation of snippets are used to learn Maximum Entropy (MaxEnt) and support vector machine (SVM) classifiers separately. Finally, viewing that the most possible predicted result will appear in the top two candidates selected by MaxEnt classifier, we develop a novel scheme that if the gap between these candidates is large enough, the predicted result is considered to be reliable; otherwise, the SVM classifier will be integrated with MaxEnt classifier to make a comprehensive prediction. Experimental results show that our framework is effective and can outperform the state-of-the-art techniques. Peng Wang 0079, Heng Zhang 0028, Yu-Fang Wu, Bo Xu 0002, Hongwei Hao |
IJCNN | 2 |
| 2014 | Short Text Feature Enrichment Using Link Analysis on Topic-Keyword Graph
Peng Wang 0079, Heng Zhang 0028, Bo Xu 0002, Hongwei Hao |
NLPCC | 2 |
| 2014 | Character confidence based on N-best list for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2013 | Keyword Spotting from Online Chinese Handwritten Documents using One-versus-All Character Classification ModelabstractIn this paper, we propose a method for text-query-based keyword spotting from online Chinese handwritten documents using character classification model. The similarity between the query word and handwriting is obtained by combining the character classification scores. The classifier is trained by one-versus-all strategy so that it gives high similarity to the target class and low scores to the others. Using character classification-based word similarity also helps overcome the out-of-vocabulary (OOV) problem. We use a character-synchronous dynamic search algorithm to efficiently spot the query word in large database. The retrieval performance is further improved by using competing character confusion and writer-adaptive thresholds. Our experimental results on a large handwriting database CASIA-OLHWDB justify the superiority of one-versus-all trained classifiers and the benefits of confidence transformation, character confusion and adaptive thresholds. Particularly, a one-versus-all trained prototype classifier performs as well as a linear support vector machine (SVM) classifier, but consumes much less storage of index file. The experimental comparison with keyword spotting based on handwritten text recognition also demonstrates the effectiveness of the proposed method. Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001, Horst Bunke |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | A confidence-based method for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
ICPR | 1 |
| 2010 | Keyword Spotting from Online Chinese Handwritten Documents Using One-vs-All Trained Character ClassifierabstractThis paper presents a text query-based method for keyword spotting from online Chinese handwritten documents. The similarity between a text word and handwriting is obtained by combining the character similiarity scores given by a character classifier. To overcome the ambiguity of character segmentation, multiple candidates of character patterns are generated by over-segmentation, and sequences of candidate characters are matched with the query word in beam search. The character classifier is trained by one-vs-all strategy so that it gives high similarity to the target class and low scores to the others. Particularly, we use a one-vs-all trained prototype classifier and a support vector machine (SVM) classifier for similarity scoring. The method yielded promising performance in experiments on a database containing 550 pages of 110 writers. For words of four characters, the recall, precision and F measure are 87.25%, 94.84% and 90.88%, respectively. Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
ICFHR | 1 |