VLDB 2026 Research / reviewers in the wild / expert
Wenbo Li 0011
dblp:51/3185-11
· DBLP profile ↗
5ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Class-Specific Word Sense Aware Topic Modeling via Soft Orthogonalized TopicsabstractWe propose a word sense aware topic model for document classification based on soft orthogonalized topics. An essential problem for this task is to capture word senses related to classes, i.e., class-specific word senses. Traditional models mainly introduce semantic information of knowledge libraries for word sense discovery. However, this information may not align with the classification targets, because these targets are often subjective and task-related. We aim to model the class-specific word senses in topic space. The challenge is to optimize the class separability of the senses, i.e., obtaining sense vectors with (a) high intra-class and (b) low inter-class similarities. Most existing models predefine specific topics for each class to specify the class-specific sense vectors. We call them hard orthogonalization based methods. These methods can hardly achieve both (a) and (b) since they assume the conditional independence of topics to classes and inevitably lose topic information. To this problem, we propose soft orthogonalization for topics. Specifically, we reserve all the topics and introduce a group of class-specific weights for each word to handle the importance of topic dimensions to class separability. Besides, we detect and use highly class-specific words in each document to guide sense estimation. Our experiments on two standard datasets show that our proposal outperforms other state-of-the-art models in terms of accuracy of sense estimation, document classification, and topic modeling. In addition, our joint learning experiments with the pre-trained language model BERT showcased the best complementarity of our model in most cases compared to other topic models. Wenbo Li 0011, Einoshin Suzuki |
CIKM | 1 |
| 2021 | Adaptive and hybrid context-aware fine-grained word sense disambiguation in topic modeling based document representation
Wenbo Li 0011, Einoshin Suzuki |
Inf. Process. Manag. | 1 |
| 2021 | Topic modeling for sequential documents based on hybrid inter-document topic dependency
Wenbo Li 0011, Hiroto Saigo, Bin Tong, Einoshin Suzuki |
J. Intell. Inf. Syst. | 1 |
| 2020 | Hybrid Context-Aware Word Sense Disambiguation in Topic Modeling based Document RepresentationabstractWe propose a hybrid context based topic model for word sense disambiguation in document representation. Document representation is an essential part of various document based tasks, and word sense disambiguation is to capture the distinctions of word senses in the representation. Traditional methods mainly rely on knowledge libraries for data enrichment; however, semantics division for a word may vary from different domain-specific datasets. We aim to discover more particular word semantic differences for each input dataset and handle the disambiguation problem without data enrichment. The challenge for this disambiguation is to (1) divide various senses for each polysemous word while (2) preserve the differences between synonyms. Most of the existing models are either based on separate context clusters or integrating an auxiliary module to specify word senses. They can hardly achieve both (1) and (2) since different senses of a word are assumed to be independent and their intrinsic relationships are ignored. To solve this problem, we estimate a word sense by both the context in which it occurs and the contexts of its other occurrences. Besides, we introduce the “Bag-of-Senses” (BoS) assumption: a document is a multiset of word senses, and the senses are generated instead of the words. Our experiments on three standard datasets show that our proposal outperforms other state-of-the-art methods in terms of accuracy of word sense estimation, topic modeling, and document classification. Wenbo Li 0011, Einoshin Suzuki |
ICDM | 1 |
| 2020 | Context-Aware Latent Dirichlet Allocation for Topic Segmentation
Wenbo Li 0011, Tetsu Matsukawa, Hiroto Saigo, Einoshin Suzuki |
PAKDD (1) | 1 |