VLDB 2026 Research / reviewers in the wild / expert
Hanlin Tang 0001
dblp:179/3388-1
· DBLP profile ↗
7ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | POS Tagging on Code Identifiers: How Far Are We?abstractPart-of-Speech (POS) tags are natural attributes of words in natural languages, and they are fundamental for natural language analysis. Many automated approaches have been proposed to tag natural language texts. Identifiers in source code have POS tags as well, which are useful for various source code analysis tasks, like code search, code comment generation, and code completion. Currently, state-of-the-art POS taggers originally designed for natural languages are often employed to tag source code identifiers. However, identifiers in source code are significantly different from natural languages. Consequently, POS taggers designed for natural languages could be less accurate in source code identifiers. Recently, several identifier-specific taggers have been proposed within the field of software engineering, but their adoption in practical software engineering tasks remains limited. This raises the question of why these taggers have not been more widely utilized in such tasks. In this article, we investigate the performance of natural language POS taggers on source code identifiers, specifically method names, parameter names, and class names. To do so, we manually annotated identifiers from open source projects in Java, C, and Python, creating a large dataset IDData for evaluation. We then evaluated six widely used natural language POS taggers: NLTK, CoreNLP, OpenNLP, spaCy, Flair, and Stanza, alongside three identifier-specific taggers: SWUM, POSSE, and Ensemble Tagger. Our evaluation reveals that while natural language-oriented POS taggers outperform identifier-specific taggers, their performance on identifiers is still significantly lower compared to their performance on natural language sentences. To understand the underlying reasons for this, we conducted an in-depth analysis, examining factors such as identifier length, POS distribution, syntactic structures, and special tags, which differentiate identifiers from natural language sentences. To further improve POS tagging performance on identifiers, we created a large-scale method name dataset MNTrain with manually labeled tags and retrained the natural language taggers on this new dataset. The results show substantial improvements in method name POS tagging performance, with taggers achieving performance comparable to their results on natural language sentences. Finally, we discuss the significance and practical implications of our findings, offering insights for future research. Hanlin Tang 0001, Yanjie Jiang, Yuxia Zhang, Nan Niu, Hui Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | A Flat Tree-based Transformer for Nested Named Entity Recognition
Hongli Mao, Xianling Mao, Hanlin Tang 0001, Xiaoyan Gao 0001, Heyan Huang |
Knowl. Based Syst. | 3 |
| 2024 | Span Graph Transformer for Document-Level Named Entity RecognitionabstractNamed Entity Recognition (NER), which aims to identify the span and category of entities within text, is a fundamental task in natural language processing. Recent NER approaches have featured pre-trained transformer-based models (e.g., BERT) as a crucial encoding component to achieve state-of-the-art performance. However, due to the length limit for input text, these models typically consider text at the sentence-level and cannot capture the long-range contextual dependency within a document. To address this issue, we propose a novel Span Graph Transformer (SGT) method for document-level NER, which constructs long-range contextual dependencies at both the token and span levels. Specifically, we first retrieve relevant contextual sentences in the document for each target sentence, and jointly encode them by BERT to capture token-level dependencies. Then, our proposed model extracts candidate spans from each sentence and integrates these spans into a document-level span graph, where nested spans within sentences and identical spans across sentences are connected. By leveraging the power of Graph Transformer and well-designed position encoding, our span graph can fully exploit span-level dependencies within the document. Extensive experiments on both resource-rich nested and flat NER datasets, as well as low-resource distantly supervised NER datasets, demonstrate that proposed SGT model achieves better performance than previous state-of-the-art models. Hongli Mao, Xianling Mao, Hanlin Tang 0001, Yuming Shang, Heyan Huang |
AAAI | 3 |
| 2024 | Span-based Unified Named Entity Recognition Framework via Contrastive Learning
Hongli Mao, Xianling Mao, Hanlin Tang 0001, Yuming Shang, Xiaoyan Gao 0001, Ao-Jie Ma, Heyan Huang |
IJCAI | 3 |
| 2023 | SQL#: A Language for Maintainable and Debuggable Database QueriesabstractStructured Query Language (SQL) is the dominant language for managing relational databases. However, complex SQL queries are hard to write and maintain because of the intricate inter-table and inter-column relations. To this end, we propose a novel query language called SQL#, which allows programmers to construct complex queries module by module and explicitly specify the relations between different modules according to the logical steps of constructing queries. Besides, we design a SQL#-based system, aiming to facilitate the maintenance of SQL# queries. Specifically, the system renders a SQL# program into a hierarchical graph, which could help programmers understand the high-level structures of SQL# programs and the intricate relations between different components within SQL# programs. In addition, the system can ease the generation of the intermediate tables that correspond to the logical steps of constructing queries, which could help programmers debug complex SQL# queries. Notably, the design of SQL# makes it easy for the system to generate the hierarchical graph and the intermediate tables. Controlled experiments suggest that the SQL#-based system reduces the durations of writing and understanding database queries by 79% and 39%, respectively, compared to raw SQL code. Yamin Hu, Hanlin Tang 0001, Zongyao Hu |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2020 | A Span-Based Distantly Supervised NER with Self-learning
Hongli Mao, Hanlin Tang 0001, Heyan Huang, Xianling Mao |
NLPCC (1) | 2 |
| 2020 | LSTM-based argument recommendation for non-API methods
Guangjie Li, Hui Liu 0003, Ge Li 0001, Sijie Shen, Hanlin Tang 0001 |
Sci. China Inf. Sci. | 5 |