Yan Liang 0004

dblp:08/2359-4 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
5since 2021 · last 2024
0009-0004-3176-7703ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Explainable and Coherent Complement Recommendation Based on Large Language Models
abstract
A complementary item is an item that pairs well with another item when consumed together. In the context of e-commerce, providing recommendations for complementary items is essential for both customers and stores. Current models for suggesting complementary items often rely heavily on user behavior data, such as co-purchase relationships. However, just because two items are frequently bought together does not necessarily mean they are truly complementary. Relying solely on co-purchase data may not align perfectly with the goal of making meaningful complementary recommendations. In this paper, we introduce the concept of "coherent complement recommendation", where "coherent" implies that recommended item pairs are compatible and relevant. Our approach builds upon complementary item pairs, with a focus on ensuring that recommended items are well used together and contextually relevant. To enhance the explainability and coherence of our complement recommendations, we fine-tune the Large Language Model (LLM) with coherent complement recommendation and explanation generation tasks since LLM has strong natural language explanation generation ability and multi-task fine-tuning enhances task understanding. Experimental results indicate that our model can provide more coherent complementary recommendations than existing state-of-the-art methods, and human evaluation validates that our approach achieves up to a 48% increase in the coherent rate of complement recommendations.
Zelong Li 0001, Yan Liang 0004, Sungro Yoon, Jiaying Shi, Xiang He 0007, Wenyi Wu, Hanbo Wang, Jin Li 0003, Jim Chan, Yongfeng Zhang 0003
CIKM2
2022 TAED: Topic-Aware Event Detection
abstract
Identifying event trigger words and classifying event types known as the event detection task is a fundamental step for extracting event-related knowledge from textual sources. Examples of the topics within documents include "military conflict," "earthquake," "concert tour," "wrestling," and others. Topical information embedded within documents where the events are extracted from is rarely explored. Rich topic information could be a helpful indicator of the event’s type. Semantically similar topics share similar event types, while event types are quite different between distinguishable document topics. In this paper, we explored a novel method of integrating document topic information to complete the event detection task. We summarized our contribution as the following: we used the topic information of the documents to generate topic comprehensive sentence representations. We adopted a multi-task deep neural network, trained with event detection and topic classification t asks. We evaluated our method with two datasets that are designed for more diverse and general event types event detection MAVEN [1] and RAMS [2]. We demonstrated that the topic-aware model outperformed the baseline model F1score on both MAVEN and RAMS datasets. An analysis in the few-shot event types scenario showed that topic-aware model can beat the baseline by up to 13.34% on the F1score for the rare event types.
Yan Liang 0004, Christan Grant
IEEE Big Data1
2021 AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding
abstract
Jun Yan, Nasser Zalmout, Yan Liang, Christan Grant, Xiang Ren, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jun Yan 0012, Nasser Zalmout, Yan Liang 0004, Christan Grant, Xiang Ren 0001, Xin Dong 0001
ACL/IJCNLP (1)3
2021 PAM: Understanding Product Images in Cross Product Category Attribute Extraction
abstract
Understanding product attributes plays an important role in improving online shopping experience for customers and serves asan integral part for constructing a product knowledge graph. Most existing methods focus on attribute extraction from text description or utilize visual information from product images such as shape and color. Compared to the inputs considered in prior works, a product image in fact contains more information, represented by a rich mixture of words and visual clues with a layout carefully designed to impress customers. This work proposes a more inclusive framework that fully utilizes these different modalities for attribute extraction.Inspired by recent works in visual question answering, we use a transformer based sequence to sequence model to fuse representations of product text, Optical Character Recognition (OCR) tokens and visual objects detected in the product image. The framework is further extended with the capability to extract attribute value across multiple product categories with a single model, by training the decoder to predict both product category and attribute value and conditioning its output on product category. The model provides a unified attribute extraction solution desirable at an e-commerce platform that offers numerous product categories with a diverse body of product attributes. We evaluated the model on two product attributes, one with many possible values and one with a small set of possible values, over 14 product categories and found the model could achieve 15% gain on the Recall and 10% gain on the F1 score compared to existing methods using text-only features.
Rongmei Lin, Xiang He 0007, Nasser Zalmout, Yan Liang 0004, Li Xiong 0001, Xin Dong 0001
KDD5
2021 All You Need to Know to Build a Product Knowledge Graph
abstract
Knowledge graphs have been pivotal in supporting downstream applications like search, recommendation, and question answering, among others. Therefore, knowledge graphs have naturally become key enabling technologies in e-Commerce platforms. Developing a high coverage product knowledge graph is more challenging than generic knowledge graphs. The highly specific and complex domain, the sparsity of training data, along with the dynamic taxonomies and product types, can constrain the resulting knowledge graphs. In this tutorial we present best practices and ML innovations in industry towards building a scalable product knowledge graph. Contributions in this domain benefit from the general literature in areas including information extraction and data mining, tailored to address the specific characteristics of e-Commerce platforms.
Nasser Zalmout, Yan Liang 0004, Xin Dong 0001
KDD4
2020 AutoKnow: Self-Driving Knowledge Collection for Products of Thousands of Types
abstract
Can one build a knowledge graph (KG) for all products in the world? Knowledge graphs have firmly established themselves as valuable sources of information for search and question answering, and it is natural to wonder if a KG can contain information about products offered at online retail sites. There have been several successful examples of generic KGs, but organizing information about products poses many additional challenges, including sparsity and noise of structured data for products, complexity of the domain with millions of product types and thousands of attributes, heterogeneity across large number of categories, as well as large and constantly growing number of products.
Xin Dong 0001, Xiang He 0007, Andrey Kan, Yan Liang 0004, Jun Ma 0029, Yifan Ethan Xu, Tong Zhao 0002, Gabriel Blanco Saldana, Saurabh Deshpande, Alexandre Michetti Manduca, Jay Ren, Surender Pal Singh, Fan Xiao 0001, Haw-Shiuan Chang, Giannis Karamanolakis, Yuning Mao, Yaqing Wang 0001, Christos Faloutsos, Andrew McCallum, Jiawei Han 0001
KDD5
2017 [Research paper] formalizing interruptible algorithms for human over-the-loop analytics
abstract
Traditional data mining algorithms are exceptional at seeing patterns in data that humans cannot, but are often confused by details that are obvious to the organic eye. Algorithms that include humans “in-the-loop” have proved beneficial for accuracy by allowing a user to provide direction in these situations, but the slowness of human interactions causes execution times to increase exponentially. Thus, we seek to formalize frameworks that include humans “over-the-loop”, giving the user an option to intervene when they deem it necessary while not having user feedback be an execution requirement. With this strategy, we hope to increase the accuracy of solutions with minimal losses in execution time. This paper describes our vision of this strategy and associated problems.
Austin Graham, Yan Liang 0004, Le Gruenwald, Christan Grant
IEEE BigData2
2017 Adaptive scalable pipelines for political event data generation
abstract
Political event data has been increasingly important for researchers to study and predict global events. Until recently the majority of political events were hand-coded from text, limiting the timeliness and coverage of event data sets. Recent systems have successfully employed big data systems for extracting events from text. These automated event systems have been limited by either the slow performance or high infrastructure demands. In this work, we present a new approach to big data systems that allow for faster extractions when compared to existing systems. We describe a modular system, Biryani, that adaptively extracts events from batches of documents. We use distributed containers to process streams of incoming documents. The number of containers processing documents can be increased or reduced depending on the number of available resources. The optimal configuration for event extraction is learned, and the system adapts to maximize the throughput of coded documents. We show the adaptability through experiments running on laptops and multiple commodity machines. We use this system to extract a new political event data set from several terabytes of text data.
Andrew Halterman, Jill Irvine, Manar Landis, Phanindra Jalla, Yan Liang 0004, Christan Grant, Mohiuddin Solaimani
IEEE BigData5