Jun Wang 0100

dblp:125/8189-100 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Language model collaboration for relation extraction from classical Chinese historical documents
Xuemei Tang, Linxu Wang, Jun Wang 0100
Inf. Process. Manag.3
2024 Curating the Chinese ancient book catalogs: Leveraging the dual roles of humanities scholars as experts and users in collaborative practice
abstract
Abstract Chinese ancient book catalogs are important cultural heritage and academic resources for the study of ancient Chinese history and culture. These catalogs need to be curated so that their value can be fully exploited in today's digital environment. This study is based on a collaborative curation project where eight representative ancient catalogs were curated into a diachronic dataset and tools to discover and analyze the data were developed. We reviewed literature and consulted humanities scholars to derive the characteristics and curation requirements of the ancient catalogs. A collaborative model was proposed based on the requirements to guide the curation process. This model reveals the duality of humanities scholars' role in collaborative curation and depicts main curation activities including metadata and description, appraisal and selection, data processing, developing tools, access and use, and evolution. Lessons learned from the curation practice include two main issues—project personnel and humanities scholars' acceptance of visualization. The study also yields a dataset and a set of tools that can be directly used by scholars interested in knowledge organization and ancient catalog related topics.
Wenqi Li 0002, Jun Wang 0100
J. Assoc. Inf. Sci. Technol.2
2023 Diachronic Named Entity Disambiguation for Ancient Chinese Historical Records
Zekun Deng, Jun Wang 0100
ICONIP (11)3
2023 Modeling Chinese Ancient Book Catalog
Linxu Wang, Jun Wang 0100
KSEM (2)2
2020 Exploiting Hybrid Subword Information for Chinese Historical Named Entity Recognition
abstract
Historical named entity recognition plays a very important role on the historical studies, especially for Chinese historical domain. This paper introduces a novel Deep learning (DL)-based approach with rich subword information of Chinese characters, which shows a powerful entity recognition of Chinese classical literatures. Our experiments on a manually constructed corpus of CMAG indicates that our model can achieve the best performance compared to other state-of-the-art approaches. The ablation studies on entity categories shows that the morphological information of Chinese characters can significantly improve the performance of NER models, especially combined with a gate-based neural unit GLU and the multi-head attention mechanism. This research provides some meaningful suggestions for automatic extraction of named entities in Chinese history, even extended to other low-resource domains.
Chengxi Yan, Jun Wang 0100
IEEE BigData2
2009 An extensive study on automated Dewey Decimal Classification
abstract
Abstract In this paper, we present a theoretical analysis and extensive experiments on the automated assignment of Dewey Decimal Classification (DDC) classes to bibliographic data with a supervised machine‐learning approach. Library classification systems, such as the DDC, impose great obstacles on state‐of‐art text categorization (TC) technologies, including deep hierarchy, data sparseness, and skewed distribution. We first analyze statistically the document and category distributions over the DDC, and discuss the obstacles imposed by bibliographic corpora and library classification schemes on TC technology. To overcome these obstacles, we propose an innovative algorithm to reshape the DDC structure into a balanced virtual tree by balancing the category distribution and flattening the hierarchy. To improve the classification effectiveness to a level acceptable to real‐world applications, we propose an interactive classification model that is able to predict a class of any depth within a limited number of user interactions. The experiments are conducted on a large bibliographic collection created by the Library of Congress within the science and technology domains over 10 years. With no more than three interactions, a classification accuracy of nearly 90% is achieved, thus providing a practical solution to the automatic bibliographic classification problem.
Jun Wang 0100
J. Assoc. Inf. Sci. Technol.1
2006 Automatic thesaurus development: Term extraction from title metadata
abstract
Abstract The application of thesauri in networked environments is seriously hampered by the challenges of introducing new concepts and terminology into the formal controlled vocabulary, which is critical for enhancing its retrieval capability. The author describes an automated process of adding new terms to thesauri as entry vocabulary by analyzing the association between words/phrases extracted from bibliographic titles and subject descriptors in the metadata record (subject descriptors are terms assigned from controlled vocabularies of thesauri to describe the subjects of the objects [e.g., books, articles] represented by the metadata records). The investigated approach uses a corpus of metadata for scientific and technical (S&T) publications in which the titles contain substantive words for key topics. The three steps of the method are (a) extracting words and phrases from the title field of the metadata; (b) applying a method to identify and select the specific and meaningful keywords based on the associated controlled vocabulary terms from the thesaurus used to catalog the objects; and (c) inserting selected keywords into the thesaurus as new terms (most of them are in hierarchical relationships with the existing concepts), thereby updating the thesaurus with new terminology that is being used in the literature. The effectiveness of the method was demonstrated by an experiment with the Chinese Classification Thesaurus (CCT) and bibliographic data in China Machine‐Readable Cataloging Record (MARC) format (CNMARC) provided by Peking University Library. This approach is equally effective in large‐scale collections and in other languages.
Jun Wang 0100
J. Assoc. Inf. Sci. Technol.1