Wenbo Li 0011

dblp:51/3185-11 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Class-Specific Word Sense Aware Topic Modeling via Soft Orthogonalized Topics
abstract
We propose a word sense aware topic model for document classification based on soft orthogonalized topics. An essential problem for this task is to capture word senses related to classes, i.e., class-specific word senses. Traditional models mainly introduce semantic information of knowledge libraries for word sense discovery. However, this information may not align with the classification targets, because these targets are often subjective and task-related. We aim to model the class-specific word senses in topic space. The challenge is to optimize the class separability of the senses, i.e., obtaining sense vectors with (a) high intra-class and (b) low inter-class similarities. Most existing models predefine specific topics for each class to specify the class-specific sense vectors. We call them hard orthogonalization based methods. These methods can hardly achieve both (a) and (b) since they assume the conditional independence of topics to classes and inevitably lose topic information. To this problem, we propose soft orthogonalization for topics. Specifically, we reserve all the topics and introduce a group of class-specific weights for each word to handle the importance of topic dimensions to class separability. Besides, we detect and use highly class-specific words in each document to guide sense estimation. Our experiments on two standard datasets show that our proposal outperforms other state-of-the-art models in terms of accuracy of sense estimation, document classification, and topic modeling. In addition, our joint learning experiments with the pre-trained language model BERT showcased the best complementarity of our model in most cases compared to other topic models.
Wenbo Li 0011, Einoshin Suzuki
CIKM1
2021 Adaptive and hybrid context-aware fine-grained word sense disambiguation in topic modeling based document representation
Wenbo Li 0011, Einoshin Suzuki
Inf. Process. Manag.1
2021 Topic modeling for sequential documents based on hybrid inter-document topic dependency
Wenbo Li 0011, Hiroto Saigo, Bin Tong, Einoshin Suzuki
J. Intell. Inf. Syst.1
2020 Hybrid Context-Aware Word Sense Disambiguation in Topic Modeling based Document Representation
abstract
We propose a hybrid context based topic model for word sense disambiguation in document representation. Document representation is an essential part of various document based tasks, and word sense disambiguation is to capture the distinctions of word senses in the representation. Traditional methods mainly rely on knowledge libraries for data enrichment; however, semantics division for a word may vary from different domain-specific datasets. We aim to discover more particular word semantic differences for each input dataset and handle the disambiguation problem without data enrichment. The challenge for this disambiguation is to (1) divide various senses for each polysemous word while (2) preserve the differences between synonyms. Most of the existing models are either based on separate context clusters or integrating an auxiliary module to specify word senses. They can hardly achieve both (1) and (2) since different senses of a word are assumed to be independent and their intrinsic relationships are ignored. To solve this problem, we estimate a word sense by both the context in which it occurs and the contexts of its other occurrences. Besides, we introduce the “Bag-of-Senses” (BoS) assumption: a document is a multiset of word senses, and the senses are generated instead of the words. Our experiments on three standard datasets show that our proposal outperforms other state-of-the-art methods in terms of accuracy of word sense estimation, topic modeling, and document classification.
Wenbo Li 0011, Einoshin Suzuki
ICDM1
2020 Context-Aware Latent Dirichlet Allocation for Topic Segmentation
Wenbo Li 0011, Tetsu Matsukawa, Hiroto Saigo, Einoshin Suzuki
PAKDD (1)1