EDBT 2026 Demo / reviewers in the wild / expert
Yan Fan 0004
dblp:35/8310-4
· DBLP profile ↗
10ranked-venue papers
5as first author
3since 2021 · last 2023
0009-0004-9070-9351ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning What to Ask: Mining Product Attributes for E-commerce Sales from Massive Dialogue CorporaabstractConversational Recommender Systems (CRSs) are extensively applied in e-commercial platforms that recommend items to users. To ensure accurate recommendation, agents usually ask for users' preferences towards specific product attributes which are pre-defined by humans. In e-commercial platforms, however, the number of products easily reaches to billions, making it prohibitive to pre-define decisive attributes for efficient recommendation due to the lack of substantial human resources and the scarce domain expertise. In this work, we present AliMeMOSAIC, a novel knowledge mining and conversational assistance framework that extracts core product attributes from massive dialogue corpora for better conversational recommendation experience. It first extracts user-agent interaction utterances from massive corpora that contain product attributes. A Joint Attribute and Value Extraction (JAVE) network is designed to extract product attributes from user-agent interaction utterances. Finally, AliMeMOSAIC generates attribute sets that frequently appear in dialogues as the target attributes for agents to request, and serve as an assistant to guide the dialogue flow. To prove the effectiveness of AliMeMOSAIC, we show that it consistently improves the overall recommendation performance of our CRS system. An industrial demonstration scenario is further presented to show how it benefits online shopping experiences. Yan Fan 0004, Chengyu Wang 0001, Hengbin Cui, Yuchuan Wu, Yongbin Li 0001 |
CIKM | 1 |
| 2023 | U-NEED: A Fine-grained Dataset for User Needs-Centric E-commerce Conversational RecommendationabstractConversational recommender systems ( CRS s) aim to understand the information needs and preferences expressed in a dialogue to recommend suitable items to the user. Most of the existing conversational recommendation datasets are synthesized or simulated with crowdsourcing, which has a large gap with real-world scenarios. To bridge the gap, previous work contributes a dataset E-ConvRec, based on pre-sales dialogues between users and customer service staff in E-commerce scenarios. However, E-ConvRec only supplies coarse-grained annotations and general tasks for making recommendations in pre-sales dialogues. Different from it, we use real user needs as a clue to explore the E-commerce conversational recommendation in complex pre-sales dialogues, namely user needs-centric E-commerce conversational recommendation (UNECR). Yuanxing Liu 0001, Weinan Zhang 0003, Baohua Dong, Yan Fan 0004, Ziyu Zhuang, Hengbin Cui, Yongbin Li 0001, Wanxiang Che |
SIGIR | 4 |
| 2022 | Building Multi-turn Query Interpreters for E-commercial Chatbots with Sparse-to-dense Attentive ModelingabstractPredicting query intents is crucial for understanding user demands in chatbots. In real-world applications, accurate query intent classification can be highly challenging as human-machine interactions are often conducted in multiple turns, which requires the models to capture related information from the entire contexts. In addition, query intents tend to be fine-grained (up to hundreds of classes), containing lots of casual chats without clear intents. Hence, it is difficult for standard transformer-based models to capture complicated language characteristics of dialogues to support these applications. In this demo, we present AliMeTerp, a multi-turn query interpretation system, which can be seamlessly integrated into e-commercial chatbots in order to generate appropriate responses. Specifically, in AliMeTerp, we introduce SAM-BERT, a pre-trained language model for fine-grained query intent understanding, based on Sparse-to-dense Attentive Modeling. For model pre-training, a stack of Sparse-to-dense Attentive Encoders are employed to model the complicated dialogue structures from different levels. We further design Hierarchical Multi-grained Classification tasks for model fine-tuning. Experiments show SAM-BERT consistently outperforms strong baselines over multiple multi-turn chatbot datasets. We further show how AliMeTerp is deployed in real-world e-commercial chatbots to support real-time customer service. Yan Fan 0004, Chengyu Wang 0001, Yunhua Hu |
WSDM | 1 |
| 2019 | SPMM: A Soft Piecewise Mapping Model for Bilingual Lexicon InductionabstractBilingual Lexicon Induction (BLI) aims at inducing word translations in two distinct languages. The generated bilingual dictionaries via BLI are essential for cross-lingual NLP applications. Most existing methods assume that a mapping matrix can be learned to project the embedding of a word in the source language to that of a word in the target language which shares the same meaning. However, a single matrix may not be able to provide sufficiently large parameter space and to tailor to the semantics of words across different domains and topics due to the complicated nature of linguistic regularities. In this paper, we propose a Soft Piecewise Mapping Model (SPMM). It generates word alignments in two languages by learning multiple mapping matrices with orthogonal constraint. Each matrix encodes the embedding translation knowledge over a distribution of latent topics in the embedding spaces. Such learning problem can be formulated as an extended version of the Wahba's problem, with a closed-form solution derived. To address the limited size of training data for low-resourced languages and emerging domains, an iterative boosting method based on SPMM is used to augment training dictionaries. Experiments conducted on both general and domain-specific corpora show that SPMM is effective and outperforms previous methods. Yan Fan 0004, Chengyu Wang 0001, Boxing Chen, Zhongkai Hu |
SDM | 1 |
| 2019 | A Family of Fuzzy Orthogonal Projection Models for Monolingual and Cross-lingual Hypernymy PredictionabstractHypernymy is a semantic relation, expressing the “is-a” relation between a concept and its instances. Such relations are building blocks for large-scale taxonomies, ontologies and knowledge graphs. Recently, much progress has been made for hypernymy prediction in English using textual patterns and/or distributional representations. However, applying such techniques to other languages is challenging due to the high language dependency of these methods and the lack of large training datasets of lower-resourced languages. Chengyu Wang 0001, Yan Fan 0004, Aoying Zhou |
WWW | 2 |
| 2019 | Predicting hypernym-hyponym relations for Chinese taxonomy learning
Chengyu Wang 0001, Yan Fan 0004, Aoying Zhou |
Knowl. Inf. Syst. | 2 |
| 2019 | Decoding Chinese User Generated Categories for Fine-Grained Knowledge HarvestingabstractUser Generated Categories (UGCs) are short but informative phrases that reflect how people describe and organize entities. UGCs express semantic relations among entities implicitly hence serve as a rich data source for knowledge harvesting. However, most UGC relation extraction methods focus on English and heavily rely on lexical and syntactic patterns. Applying them directly to Chinese UGCs poses significant challenges because Chinese is an analytic language with flexible language expressions. In this paper, we aim at harvesting fine-grained relations from Chinese UGCs automatically. Based on neural networks and negative sampling, we introduce two word embedding projection models to identify is-a relations. The accuracy of prediction results is improved via a collective refinement algorithm and a hypernym expansion method. We further propose a graph clique mining algorithm to harvest non-taxonomic relations from UGCs, together with their textual patterns. Two experiments are conducted to validate our approach based on Chinese Wikipedia. The first experiment verifies the is-a relation extraction approach achieves high accuracy, outperforming state-of-the-art methods. The second experiment shows that the proposed method can harvest non-taxonomic relations of large quantity and high accuracy, with minimal human intervention. Chengyu Wang 0001, Yan Fan 0004, Aoying Zhou |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Exploratory Neural Relation Classification for Domain Knowledge AcquisitionabstractThe state-of-the-art methods for relation classification are primarily based on deep neural net- works. This kind of supervised learning method suffers from not only limited training data, but also the large number of low-frequency relations in specific domains. In this paper, we propose the task of exploratory relation classification for domain knowledge harvesting. The goal is to learn a classifier on pre-defined relations and discover new relations expressed in texts. A dynamically structured neural network is introduced to classify entity pairs to a continuously expanded relation set. We further propose the similarity sensitive Chinese restaurant process to discover new relations. Experiments conducted on a large corpus show the effectiveness of our neural network, while new relations are discovered with high precision and recall. Yan Fan 0004, Chengyu Wang 0001 |
COLING | 1 |
| 2017 | DKGBuilder: An Architecture for Building a Domain Knowledge Graph from Scratch
Yan Fan 0004, Chengyu Wang 0001, Guomin Zhou |
DASFAA (2) | 1 |
| 2017 | Learning Fine-grained Relations from Chinese User Generated CategoriesabstractUser generated categories (UGCs) are short texts that reflect how people describe and organize entities, expressing rich semantic relations implicitly.While most methods on UGC relation extraction are based on pattern matching in English circumstances, learning relations from Chinese UGCs poses different challenges due to the flexibility of expressions.In this paper, we present a weakly supervised learning framework to harvest relations from Chinese UGCs.We identify is-a relations via word embedding based projection and inference, extract non-taxonomic relations and their category patterns by graph mining.We conduct experiments on Chinese Wikipedia and achieve high accuracy, outperforming state-of-the-art methods. Chengyu Wang 0001, Yan Fan 0004, Aoying Zhou |
EMNLP | 2 |