Dang Pham

dblp:207/5218 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Pre-Trained Language Models with Topic Attention for Supervised Document Structure Learning
abstract
The discourse-level structure of a document can be captured through learning the rhetorical functions of sentences in that document. Existing supervised methods based on pre-trained language models for classifying rhetorical functions of sentences usually focus on utilizing rhetorical words but ignore the topics of sentences. Since topic words can provide additional information for enhancing the learning of the document structure, we present a neural topic model that is integrated with a BERT-based language model through a unified probabilistic generative process for learning both the rhetorical structure and topic structure of documents. For inference, we design a topic attention mechanism to utilize the learned topic words from previous sentences to improve the prediction of the current sentence’s rhetorical label. The extensive experiments on four real-world datasets of different domains show that the proposed model improves the detection of rhetorical functions of sentences and is effective in document modeling and extracting coherent topics.
Dang Pham, Tuan M. V. Le
IEEE Big Data1
2023 Utilizing Textual Reviews for Visualizing and Understanding User Preferences
abstract
Latent factor models are widely used in recommender systems. In these models, users and items are represented as vectors in a joint latent factor space. The inner products of user vectors and item vectors are used to model the user-item interactions (e.g., ratings). A review is often posted by the user to explain the given rating. Therefore, reviews can be used to understand how users rate the items and to interpret the latent dimensions of user and item vectors. In this paper, we propose a probabilistic model that learns latent vectors of users and items in a two- or three-dimensional space for visualization. Our proposed model also extracts review topics and visualizes them in the same visualization space for interpreting the ratings. We model the user-item interactions by using the distances between users and items in the visualization space. Extensive experiments using several real-world datasets demonstrate the effectiveness of our proposed model in recommendation and visualization tasks.
Dang Pham, Tuan M. V. Le
ASONAM1
2021 Neural Topic Models for Hierarchical Topic Detection and Visualization
Dang Pham, Tuan M. V. Le
ECML/PKDD (3)1
2020 Auto-Encoding Variational Bayes for Inferring Topics and Visualization
abstract
Visualization and topic modeling are widely used approaches for text analysis. Traditional visualization methods find low-dimensional representations of documents in the visualization space (typically 2D or 3D) that can be displayed using a scatterplot. In contrast, topic modeling aims to discover topics from text, but for visualization, one needs to perform a post-hoc embedding using dimensionality reduction methods. Recent approaches propose using a generative model to jointly find topics and visualization, allowing the semantics to be infused in the visualization space for a meaningful interpretation. A major challenge that prevents these methods from being used practically is the scalability of their inference algorithms. We present, to the best of our knowledge, the first fast Auto-Encoding Variational Bayes based inference method for jointly inferring topics and visualization. Since our method is black box, it can handle model changes efficiently with little mathematical rederivation effort. We demonstrate the efficiency and effectiveness of our method on real-world large datasets and compare it with existing baselines.
Dang Pham, Tuan M. V. Le
COLING1