Nhu-Thuat Tran

dblp:253/9132 · DBLP profile ↗
← Back
8ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0001-5496-6749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Bridging LLM Embeddings and VAE Parameters for Disentangled Recommendation
abstract
Disentangled recommendation within the Variational Autoencoder (VAE) framework aims to capture multiple user interests. While effective, these VAEs are fundamentally constrained by their reliance on interaction data alone, lacking the rich external semantic knowledge needed to properly structure and separate latent interests. Meanwhile, Large Language Models (LLMs) excel at deriving profound user preference signals from textual data. Prevailing methods for integrating LLMs into recommendation, however, either focus on single-interest modeling or perform a shallow fusion by aligning LLM and VAE representation spaces. Thus, they fail to fundamentally shape the VAE's latent space for multi-interest learning, hindering recommendation performance. To bridge this gap, we propose to fundamentally shift the integration point: instead of aligning representation spaces, we bridge the LLM-generated embedding space directly to the VAE's parameter space. Our framework designs hypernetworks to generate parameters governing interest discovery of a disentangled VAE conditioned on LLM-derived user embeddings. This direct bridge injects rich semantic knowledge into the model's learning foundation, preserving the VAE's power for disentangled representation while dramatically enhancing its modeling capability with informative, LLM-structured priors. Extensive experiments on benchmark datasets show that our method yields significant performance gains, unveiling the untapped potential of using LLMs for disentangled recommendation.
Nhu-Thuat Tran, Hady Wirawan Lauw
SIGIR1
2025 Parameter-Efficient Variational AutoEncoder for Multimodal Multi-Interest Recommendation
abstract
Learning user preferences in recommendation systems is enriched by multimodal features, such as textual and visual content, and amplified by multi-interest modeling with Variational AutoEncoders (VAEs). However, prior efforts are limited by single modality focus and cumbersome, parameter-heavy architecture designs. To address these limitations, we introduce an innovative solution that blends the semantic richness of multimodal data with the representational power of multi-representation VAEs. Drawing inspiration from Mixture of Experts (MoE), we cast each VAE as an expert tailored to a specific modality, then fuse them via a novel parameter-merging function into a lean, unified model. This approach efficiently captures diverse user preferences behind multimodal data with minimal complexity. Rigorous experiments on real-world benchmarks show our method outshines state-of-the-art baselines while slashing parameter counts. Our work sets a new, streamlined standard for multimodal, multi-interest recommendation systems.
Nhu-Thuat Tran, Hady Wirawan Lauw
ACM Multimedia1
2025 Optimal Transport Alignment of User Preferences from Ratings and Texts
abstract
Modeling hidden factors driving user preferences is crucial for recommendation yet challenging due to sparse rating data. While aligning preference factors from ratings and texts, as a solution, shows improvements, existing methods impose restrictive one-to-one factor correspondences and underutilize cross-modal interest signals. We propose an optimal transport (OT) approach to address these gaps. By modeling rating- and text-based preference factors as distributions, we compute an OT plan that captures their probabilistic relationships. This plan serves dual roles: 1) to regularize cross-modal preference factors without rigid correspondence assumptions, and 2) to blend preference signals across modalities through barycentric mapping. Experiments on real-world datasets validate our method’s effectiveness over competitive baselines, highlighting its novel use of OT for adaptive preference factor alignment, an underexplored direction in recommender system research.
Nhu-Thuat Tran, Hady Wirawan Lauw
UAI1
2025 VARIUM: Variational Autoencoder for Multi-Interest Representation with Inter-User Memory
abstract
Frameworks for discovering multiple user interest factors based on Variational AutoEncoder (VAE) has demonstrated competitive recommendation performance. However, as VAE only considers one user as input at a time, sharing across like-minded users may not be adequately facilitated. Moreover, interest sharing between users is not always available and thus, poses a challenge for VAE to explicitly model this information. To resolve this, we introduce an inter-user memory-based mechanism to unsupervisedly discover latent interest sharing between users under VAE framework. Concretely, we design a memory including an array of prototypes, each hypothetically representing a group of users sharing a particular interest. These memory prototypes are jointly trained with the backbone VAE-based recommendation model. For each user, we first discover multiple intra-user interest factors behind their item adoptions. Next, intra-user interest factors query to memory to retrieve the inter-user interest clues from like-minded users. This query-retrieve process is performed sequentially via a series of attention-transformation steps. Then, interest clues retrieved from memory are incorporated into interest factor representations of each user to increase their expressiveness. Thorough experiments on real-world datasets verify the strength of our method over an array of baselines. We further conduct qualitative analysis to understand the inner working of our memory-based refinement approach.
Nhu-Thuat Tran, Hady Wirawan Lauw
WSDM1
2024 Learning Multi-Faceted Prototypical User Interests
abstract
We seek to uncover the latent interest units from behavioral data to better learn user preferences under the VAE framework. Existing practices tend to ignore the multiple facets of item characteristics, which may not capture it at appropriate granularity. Moreover, current studies equate the granularity of item space to that of user interests, which we postulate is not ideal as user interests would likely map to a small subset of item space. In addition, the compositionality of user interests has received inadequate attention, preventing the modeling of interactions between explanatory factors driving a user's decision. To resolve this, we propose to align user interests with multi-faceted item characteristics. First, we involve prototype-based representation learning to discover item characteristics along multiple facets. Second, we compose user interests from uncovered item characteristics via binding mechanism, separating the granularity of user preferences from that of item space. Third, we design a dedicated bi-directional binding block, aiding the derivation of compositional user interests. On real-world datasets, the experimental results demonstrate the strong performance of our proposed method compared to a series of baselines.
Nhu-Thuat Tran, Hady Wirawan Lauw
ICLR1
2023 Multi-Representation Variational Autoencoder via Iterative Latent Attention and Implicit Differentiation
abstract
Variational Autoencoder (VAE) offers a non-linear probabilistic modeling of user's preferences. While it has achieved remarkable performance at collaborative filtering, it typically samples a single vector for representing user's preferences, which may be insufficient to capture the user's diverse interests. Existing solutions extend VAE to model multiple interests of users by resorting a variant of self-attentive method, i.e., employing prototypes to group items into clusters, each capturing one topic of user's interests. Despite showing improvements, the current design could be more effective since prototypes are randomly initialized and shared across users, resulting in uninformative and non-personalized clusters.
Nhu-Thuat Tran, Hady Wirawan Lauw
CIKM1
2023 Memory Network-Based Interpreter of User Preferences in Content-Aware Recommender Systems
abstract
This article introduces a novel architecture for two objectives recommendation and interpretability in a unified model. We leverage textual content as a source of interpretability in content-aware recommender systems. The goal is to characterize user preferences with a set of human-understandable attributes, each is described by a single word, enabling comprehension of user interests behind item adoptions. This is achieved via a dedicated architecture, which is interpretable by design, involving two components for recommendation and interpretation. In particular, we seek an interpreter , which accepts holistic user’s representation from a recommender to output a set of activated attributes describing user preferences. Besides encoding interpretability properties such as fidelity, conciseness and diversity, the proposed memory network-based interpreter enables the generalization of user representation by discovering relevant attributes that go beyond her adopted items’ textual content. We design experiments involving both human- and functionally-grounded evaluations of interpretability. Results on four real-world datasets show that our proposed model not only discovers highly relevant attributes for interpreting user preferences, but also enjoys comparable or better recommendation accuracy than a series of baselines.
Nhu-Thuat Tran, Hady Wirawan Lauw
ACM Trans. Intell. Syst. Technol.1
2022 Aligning Dual Disentangled User Representations from Ratings and Textual Content
abstract
Classical recommendation methods typically render user representation as a single vector in latent space. Oftentimes, a user's interactions with items are influenced by several hidden factors. To better uncover these hidden factors, we seek disentangled representations. Existing disentanglement methods for recommendations are mainly concerned with user-item interactions alone. To further improve not only the effectiveness of recommendations but also the interpretability of the representations, we propose to learn a second set of disentangled user representations from textual content and to align the two sets of representations with one another. The purpose of this coupling is two-fold. For one benefit, we leverage textual content to resolve sparsity of user-item interactions, leading to higher recommendation accuracy. For another benefit, by regularizing factors learned from user-item interactions with factors learned from textual content, we map uninterpretable dimensions from user representation into words. An attention-based alignment is introduced to align and enrich hidden factors representations. A series of experiments conducted on four real-world datasets show the efficacy of our methods in improving recommendation quality.
Nhu-Thuat Tran, Hady Wirawan Lauw
KDD1