EDBT 2026 Demo / reviewers in the wild / expert
Yan Liang 0004
dblp:08/2359-4
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2024
0009-0004-3176-7703ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Explainable and Coherent Complement Recommendation Based on Large Language ModelsabstractA complementary item is an item that pairs well with another item when consumed together. In the context of e-commerce, providing recommendations for complementary items is essential for both customers and stores. Current models for suggesting complementary items often rely heavily on user behavior data, such as co-purchase relationships. However, just because two items are frequently bought together does not necessarily mean they are truly complementary. Relying solely on co-purchase data may not align perfectly with the goal of making meaningful complementary recommendations. In this paper, we introduce the concept of "coherent complement recommendation", where "coherent" implies that recommended item pairs are compatible and relevant. Our approach builds upon complementary item pairs, with a focus on ensuring that recommended items are well used together and contextually relevant. To enhance the explainability and coherence of our complement recommendations, we fine-tune the Large Language Model (LLM) with coherent complement recommendation and explanation generation tasks since LLM has strong natural language explanation generation ability and multi-task fine-tuning enhances task understanding. Experimental results indicate that our model can provide more coherent complementary recommendations than existing state-of-the-art methods, and human evaluation validates that our approach achieves up to a 48% increase in the coherent rate of complement recommendations. Zelong Li 0001, Yan Liang 0004, Sungro Yoon, Jiaying Shi, Xiang He 0007, Wenyi Wu, Hanbo Wang, Jin Li 0003, Jim Chan, Yongfeng Zhang 0003 |
CIKM | 2 |
| 2022 | TAED: Topic-Aware Event DetectionabstractIdentifying event trigger words and classifying event types known as the event detection task is a fundamental step for extracting event-related knowledge from textual sources. Examples of the topics within documents include "military conflict," "earthquake," "concert tour," "wrestling," and others. Topical information embedded within documents where the events are extracted from is rarely explored. Rich topic information could be a helpful indicator of the event’s type. Semantically similar topics share similar event types, while event types are quite different between distinguishable document topics. In this paper, we explored a novel method of integrating document topic information to complete the event detection task. We summarized our contribution as the following: we used the topic information of the documents to generate topic comprehensive sentence representations. We adopted a multi-task deep neural network, trained with event detection and topic classification t asks. We evaluated our method with two datasets that are designed for more diverse and general event types event detection MAVEN [1] and RAMS [2]. We demonstrated that the topic-aware model outperformed the baseline model F1score on both MAVEN and RAMS datasets. An analysis in the few-shot event types scenario showed that topic-aware model can beat the baseline by up to 13.34% on the F1score for the rare event types. Yan Liang 0004, Christan Grant |
IEEE Big Data | 1 |
| 2021 | AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive DecodingabstractJun Yan, Nasser Zalmout, Yan Liang, Christan Grant, Xiang Ren, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Yan 0012, Nasser Zalmout, Yan Liang 0004, Christan Grant, Xiang Ren 0001, Xin Dong 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | PAM: Understanding Product Images in Cross Product Category Attribute ExtractionabstractUnderstanding product attributes plays an important role in improving online shopping experience for customers and serves asan integral part for constructing a product knowledge graph. Most existing methods focus on attribute extraction from text description or utilize visual information from product images such as shape and color. Compared to the inputs considered in prior works, a product image in fact contains more information, represented by a rich mixture of words and visual clues with a layout carefully designed to impress customers. This work proposes a more inclusive framework that fully utilizes these different modalities for attribute extraction.Inspired by recent works in visual question answering, we use a transformer based sequence to sequence model to fuse representations of product text, Optical Character Recognition (OCR) tokens and visual objects detected in the product image. The framework is further extended with the capability to extract attribute value across multiple product categories with a single model, by training the decoder to predict both product category and attribute value and conditioning its output on product category. The model provides a unified attribute extraction solution desirable at an e-commerce platform that offers numerous product categories with a diverse body of product attributes. We evaluated the model on two product attributes, one with many possible values and one with a small set of possible values, over 14 product categories and found the model could achieve 15% gain on the Recall and 10% gain on the F1 score compared to existing methods using text-only features. Rongmei Lin, Xiang He 0007, Nasser Zalmout, Yan Liang 0004, Li Xiong 0001, Xin Dong 0001 |
KDD | 5 |
| 2021 | All You Need to Know to Build a Product Knowledge GraphabstractKnowledge graphs have been pivotal in supporting downstream applications like search, recommendation, and question answering, among others. Therefore, knowledge graphs have naturally become key enabling technologies in e-Commerce platforms. Developing a high coverage product knowledge graph is more challenging than generic knowledge graphs. The highly specific and complex domain, the sparsity of training data, along with the dynamic taxonomies and product types, can constrain the resulting knowledge graphs. In this tutorial we present best practices and ML innovations in industry towards building a scalable product knowledge graph. Contributions in this domain benefit from the general literature in areas including information extraction and data mining, tailored to address the specific characteristics of e-Commerce platforms. Nasser Zalmout, Yan Liang 0004, Xin Dong 0001 |
KDD | 4 |
| 2020 | AutoKnow: Self-Driving Knowledge Collection for Products of Thousands of TypesabstractCan one build a knowledge graph (KG) for all products in the world? Knowledge graphs have firmly established themselves as valuable sources of information for search and question answering, and it is natural to wonder if a KG can contain information about products offered at online retail sites. There have been several successful examples of generic KGs, but organizing information about products poses many additional challenges, including sparsity and noise of structured data for products, complexity of the domain with millions of product types and thousands of attributes, heterogeneity across large number of categories, as well as large and constantly growing number of products. Xin Dong 0001, Xiang He 0007, Andrey Kan, Yan Liang 0004, Jun Ma 0029, Yifan Ethan Xu, Tong Zhao 0002, Gabriel Blanco Saldana, Saurabh Deshpande, Alexandre Michetti Manduca, Jay Ren, Surender Pal Singh, Fan Xiao 0001, Haw-Shiuan Chang, Giannis Karamanolakis, Yuning Mao, Yaqing Wang 0001, Christos Faloutsos, Andrew McCallum, Jiawei Han 0001 |
KDD | 5 |
| 2017 | [Research paper] formalizing interruptible algorithms for human over-the-loop analyticsabstractTraditional data mining algorithms are exceptional at seeing patterns in data that humans cannot, but are often confused by details that are obvious to the organic eye. Algorithms that include humans “in-the-loop” have proved beneficial for accuracy by allowing a user to provide direction in these situations, but the slowness of human interactions causes execution times to increase exponentially. Thus, we seek to formalize frameworks that include humans “over-the-loop”, giving the user an option to intervene when they deem it necessary while not having user feedback be an execution requirement. With this strategy, we hope to increase the accuracy of solutions with minimal losses in execution time. This paper describes our vision of this strategy and associated problems. Austin Graham, Yan Liang 0004, Le Gruenwald, Christan Grant |
IEEE BigData | 2 |
| 2017 | Adaptive scalable pipelines for political event data generationabstractPolitical event data has been increasingly important for researchers to study and predict global events. Until recently the majority of political events were hand-coded from text, limiting the timeliness and coverage of event data sets. Recent systems have successfully employed big data systems for extracting events from text. These automated event systems have been limited by either the slow performance or high infrastructure demands. In this work, we present a new approach to big data systems that allow for faster extractions when compared to existing systems. We describe a modular system, Biryani, that adaptively extracts events from batches of documents. We use distributed containers to process streams of incoming documents. The number of containers processing documents can be increased or reduced depending on the number of available resources. The optimal configuration for event extraction is learned, and the system adapts to maximize the throughput of coded documents. We show the adaptability through experiments running on laptops and multiple commodity machines. We use this system to extract a new political event data set from several terabytes of text data. Andrew Halterman, Jill Irvine, Manar Landis, Phanindra Jalla, Yan Liang 0004, Christan Grant, Mohiuddin Solaimani |
IEEE BigData | 5 |