Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wenzhe Du

dblp:348/5487 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
0009-0001-7897-0344ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Recommender systems · 82% Data mining · 18%
Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Vision and language · 50%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › interactive recommendation
conversational recommendation
1.422024
Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data Setting · ACM Multimedia 2024
Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation · ACM Multimedia 2023
Recommender systems › interactive recommendation › conversational recommendation
multimodal conversational recommendation
1.422024
Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data Setting · ACM Multimedia 2024
Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation · ACM Multimedia 2023
Data mining
semi-supervised learning
0.812024
Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data Setting · ACM Multimedia 2024
Recommender systems › representation learning for recommendation
product representation learning
0.712023
Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation · ACM Multimedia 2023
Computer vision › Vision and language
cross-modal interaction
0.212023
Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation · ACM Multimedia 2023
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.212023
Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

gated multi-view encoder · 1.3attention mechanism · 1.3semi-supervised learning · 0.8dialogue state encoder · 0.8correlation regularization · 0.8
YearPublicationVenuePosition
2024 Sample Efficiency Matters: Training Multimodal Conversational Recommendation Systems in a Small Data Setting
abstract
With the increasing prevalence of virtual assistants, multimodal conversational recommendation systems (multimodal CRS) becomes essential for boosting customer engagement, improving conversion rates, and enhancing user satisfaction. Yet conversational samples, as training data for such a system, are difficult to obtain in large quantities, particularly in new platforms. To effectively train multimodal CRS in a small data setting, we enhance data quality to make up for the small data quantity by augmenting conversations with dialogue states. We then devise an effective dialogue state encoder to bridge the semantic gap between conversation and product representations for recommendation. To further reduce the cost of dialogue state annotation, a semi-supervised learning method is developed to effectively train the dialogue state encoder with a small set of labeled conversations. In addition, we design a correlation regularisation that leverages knowledge in the multimodal product database to help align textual and visual modalities. Experiments on the dataset MMD demonstrate the effectiveness of our method. Particularly, with only 5% of the MMD training set, our method (namely SeMANTIC) obtains better NDCG scores than those of baseline models trained on the full MMD training set.
Wenzhe Du, Xiaoliang Wang 0001, Cam-Tu Nguyen
ACM Multimedia2
2023 Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation
abstract
Multimodal Conversational Recommendation aims to find appropriate products based on a multi-turn dialogue, where user requests and products can be presented in both visual and textual modalities. While previous studies have focused on understanding user preferences from conversational contexts, the task of product modeling has been relatively unexplored. This study targets to fill this gap and demonstrates that information from multiple product views and cross-view interactions are essential for recommendation, along with dialog information. To this end, a product image is first encoded using a gated multi-view image encoder, and representations for the global and local views are obtained. On the textual side, two views are considered: the structure view (product attributes) and the sequence view (product description/reviews). Two forms of inter-modal interactions for product representation are then modeled: interactions between the global image view and the textual structure view, and interactions between the local image view and the textual sequence view. Furthermore, the representation is enhanced to attend to the latest user request in the dialog context, resulting in query-aware product representation. The experimental results indicate that our method, named Enteract, achieves state-of-the-art performance on two well-known datasets (MMD and SIMMC).
Wenzhe Du, Cam-Tu Nguyen
ACM Multimedia1