Chunyuan Yuan

dblp:248/8262 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0001-9794-5032ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 5 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2025 ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce
abstract
Relevance modeling in e-commerce search remains challenged by semantic gaps in term-matching methods (e.g., BM25) and neural models' reliance on the scarcity of domain-specific hard samples. We propose ADORE, a self-sustaining framework that synergizes three innovations: (1) A Rule-aware Relevance Discrimination module, where a Chain-of-Thought LLM generates intent-aligned training data, refined via Kahneman-Tversky Optimization (KTO) to align with user behavior; (2) An Error-type-aware Data Synthesis module that auto-generates adversarial examples to harden robustness; and (3) A Key-attribute-enhanced Knowledge Distillation module that injects domain-specific attribute hierarchies into a deployable student model. ADORE automates annotation, adversarial generation, and distillation, overcoming data scarcity while enhancing reasoning. Large-scale experiments and online A/B testing verify the effectiveness of ADORE. The framework establishes a new paradigm for resource-efficient, cognitively aligned relevance modeling in industrial applications.
Donghao Xie, Ming Pang, Chunyuan Yuan, Changping Peng, Zhangang Lin
SIGIR4
2025 Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising
abstract
The retrieval system is a crucial module in e-commerce search advertising that matches user queries with ads. The diverse expressions of users often produce massive tail queries that cannot match merchant bidwords, leading to poor retrieval efficiency. Existing methods, such as query log mining and vector matching, fail to optimize relevance, authenticity, and ad revenue of the rewrite.
Zhenhui Liu, Chunyuan Yuan, Ming Pang, Li Yuan 0007, Changping Peng, Zhangang Lin, Jingping Shao
SIGIR2
2023 Learning Query-aware Embedding Index for Improving E-commerce Dense Retrieval
abstract
The embedding index has become an essential part of the dense retrieval (DR) system, which enables a fast search for billion of items in online E-commerce applications. To accelerate the retrieval process in industrial scenarios, most of the previous studies only utilize item embeddings. However, the product quantization process without query embeddings will lead to inconsistency between queries and items. A straightforward solution is to put query embedding into the product quantization process. But we found that the distance of the positive query and item embedding pairs is too large, which means the query and item embeddings learned by the two-tower are not fully aligned. This problem would lead to performance decay when directly putting query embeddings into the product quantization.
Chunyuan Yuan, Jingwei Zhuo, Songlin Wang, Sulong Xu
SIGIR2
2022 Cascade-Enhanced Graph Convolutional Network for Information Diffusion Prediction
Lingwei Wei, Chunyuan Yuan, Yinan Bao, Wei Zhou 0019, Xian Zhu, Songlin Hu 0001
DASFAA (1)3
2021 Label-Specific Dual Graph Neural Network for Multi-Label Text Classification
abstract
Qianwen Ma, Chunyuan Yuan, Wei Zhou, Songlin Hu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Songlin Hu 0001
ACL/IJCNLP (1)2
2021 GEDIT: Geographic-Enhanced and Dependency-Guided Tagging for Joint POI and Accessibility Extraction at Baidu Maps
abstract
Providing timely accessibility reminders (such as closed and relocated) of a point-of-interest (POI) plays a vital role in improving user satisfaction of finding places and making visiting decisions. However, it is difficult to keep the POI database in sync with the real-world counterparts due to the dynamic nature of business changes and innovations. To alleviate this problem, we formulate and present a practical solution that jointly extracts POI mentions and identifies their coupled accessibility labels from unstructured text (hereafter referred to as joint POI and accessibility extraction). We approach this task as a sequence tagging problem, where the goal is to produce (POI name, accessibility label) pairs from unstructured text. This task is challenging because of two main issues: (1) POI names are often newly-coined words so as to successfully register new entities or brands and (2) there may exist multiple pairs in the text, which necessitates dealing with one-to-many or many-to-one mapping to make each POI coupled with its matching accessibility label. To this end, we propose a Geographic-Enhanced and Dependency-guIded sequence Tagging (GEDIT) model to concurrently address the two challenges. First, to alleviate challenge #1, we develop a geographic-enhanced pre-trained model to learn the text representations, which is able to significantly relieve the problem of newly-coined words. Second, to mitigate challenge #2, we apply a relational graph convolutional network to learn the tree node representations from the parsed dependency tree, which enables us to establish a correlation between a POI and its accessibility label. Finally, we construct a neural sequence tagging model by integrating and feeding the previously pre-learned representations into a CRF layer. Extensive experiments conducted on a real-world dataset demonstrate the superiority and effectiveness of GEDIT. In addition, it has already been deployed in production at Baidu Maps, and it successfully keeps processing hundreds of thousands of Web documents every week. Statistics show that the proposed solution can save significant human effort and labor costs to deal with the same amount of documents, which confirms that it is a practical way for POI accessibility maintenance.
Jizhou Huang, Chunyuan Yuan, Haifeng Wang 0001, Ming Liu 0004, Bing Qin 0001
CIKM3
2021 SRLF: A Stance-aware Reinforcement Learning Framework for Content-based Rumor Detection on Social Media
abstract
The rapid development of social media changes the lifestyle of people and simultaneously provides an ideal place for publishing and disseminating rumors, which severely exacerbates social panic and triggers a crisis of social trust. Early content-based methods focused on finding clues from the text and user profiles for rumor detection. Recent studies combine the stances of users' comments with news content to capture the difference between true and false rumors. Although the user's stance is effective for rumor detection, the manual labeling process is time-consuming and labor-intensive, which limits the application of utilizing it to facilitate rumor detection. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework.
Chunyuan Yuan, Wanhui Qian, Qianwen Ma, Wei Zhou 0019, Songlin Hu 0001
IJCNN1
2021 SRLF: A Stance-aware Reinforcement Learning Framework for Content-based Rumor Detection on Social Media
abstract
The rapid development of social media changes the lifestyle of people and simultaneously provides an ideal place for publishing and disseminating rumors, which severely exacerbates social panic and triggers a crisis of social trust. Early content-based methods focused on finding clues from the text and user profiles for rumor detection. Recent studies combine the stances of users' comments with news content to capture the difference between true and false rumors. Although the user's stance is effective for rumor detection, the manual labeling process is time-consuming and labor-intensive, which limits the application of utilizing it to facilitate rumor detection. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework. In this paper, we first finetune a pre-trained BERT model on a small labeled dataset and leverage this model to annotate weak stance labels for users' comment data to overcome the problem mentioned above. Then, we propose a novel Stance-aware Reinforcement Learning Framework (SRLF) to select high-quality labeled stance data for model training and rumor detection. Both the stance selection and rumor detection tasks are optimized simultaneously to promote both tasks mutually. We conduct experiments on two commonly used real-world datasets. The experimental results demonstrate that our framework outperforms the state-of-the-art models significantly, which confirms the effectiveness of the proposed framework.
Chunyuan Yuan, Wanhui Qian, Qianwen Ma, Wei Zhou 0019, Songlin Hu 0001
IJCNN1
2021 HGAMN: Heterogeneous Graph Attention Matching Network for Multilingual POI Retrieval at Baidu Maps
abstract
The increasing interest in international travel has raised the demand of retrieving point of interests (POIs) in multiple languages. This is even superior to find local venues such as restaurants and scenic spots in unfamiliar languages when traveling abroad. Multilingual POI retrieval, enabling users to find desired POIs in a demanded language using queries in numerous languages, has become an indispensable feature of today's global map applications such as Baidu Maps. This task is non-trivial because of two key challenges: (1) visiting sparsity and (2) multilingual query-POI matching. To this end, we propose a Heterogeneous Graph Attention Matching Network (HGAMN) to concurrently address both challenges. Specifically, we construct a heterogeneous graph that contains two types of nodes: POI node and query node using the search logs of Baidu Maps. First, to alleviate challenge #1, we construct edges between different POI nodes to link the low-frequency POIs with the high-frequency ones, which enables the transfer of knowledge from the latter to the former. Second, to mitigate challenge #2, we construct edges between POI and query nodes based on the co-occurrences between queries and POIs, where queries in different languages and formulations can be aggregated for individual POIs. Moreover, we develop an attention-based network to jointly learn node representations of the heterogeneous graph and further design a cross-attention module to fuse the representations of both types of nodes for query-POI relevance scoring. In this way, the relevance ranking between multilingual queries and POIs with different popularity can be better handled. Extensive experiments conducted on large-scale real-world datasets from Baidu Maps demonstrate the superiority and effectiveness of HGAMN. In addition, HGAMN has already been deployed in production at Baidu Maps, and it successfully keeps serving hundreds of millions of requests every day. Compared with the previously deployed model, HGAMN achieves significant performance improvement, which confirms that HGAMN is a practical and robust solution for large-scale real-world multilingual POI retrieval service.
Jizhou Huang, Haifeng Wang 0001, Zhengjie Huang, Chunyuan Yuan, Yawen Li 0001
KDD6
2020 Who Are Controlled by The Same User? Multiple Identities Deception Detection via Social Interaction Activity (Student Abstract)
abstract
Social media has become a preferential place for sharing information. However, some users may create multiple accounts and manipulate them to deceive legitimate users. Most previous studies utilize verbal or behavior features based methods to solve this problem, but they are only designed for some particular platforms, leading to low universalness.In this paper, to support multiple platforms, we construct interaction tree for each account based on their social interactions which is common characteristic of social platforms. Then we propose a new method to calculate the social interaction entropy of each account and detect the accounts which are controlled by the same user. Experimental results on two real-world datasets show that the method has robust superiority over state-of-the-art methods.
Chunyuan Yuan, Wei Zhou 0019, Jingli Wang, Songlin Hu 0001
AAAI2
2020 Early Detection of Fake News by Utilizing the Credibility of News, Publishers, and Users based on Weakly Supervised Learning
abstract
The dissemination of fake news significantly affects personal reputation and public trust.Recently, fake news detection has attracted tremendous attention, and previous studies mainly focused on finding clues from news content or diffusion path.However, the required features of previous models are often unavailable or insufficient in early detection scenarios, resulting in poor performance.Thus, early fake news detection remains a tough challenge.Intuitively, the news from trusted and authoritative sources or shared by many users with a good reputation is more reliable than other news.Using the credibility of publishers and users as prior weakly supervised information, we can quickly locate fake news in massive news and detect them in the early stages of dissemination.In this paper, we propose a novel Structure-aware Multi-head Attention Network (SMAN), which combines the news content, publishing, and reposting relations of publishers and users, to jointly optimize the fake news detection and credibility prediction tasks.In this way, we can explicitly exploit the credibility of publishers and users for early fake news detection.We conducted experiments on three real-world datasets, and the results show that SMAN can detect fake news in 4 hours with an accuracy of over 91%, which is much faster than the state-of-the-art models.
Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
COLING1
2020 Adversarial Learning for Overlapping Community Detection and Network Embedding
abstract
Network Embedding (NE) aims at modeling network graph by encoding vertices and edges into a low-dimensional space. These learned vectors which preserve proximities can be used for subsequent applications, such as vertex classification and link prediction. Skip-gram with negative sampling is the most widely used method for existing NE models to approximate their objective functions. However, this method only focuses on learning representation from the local connectivity of vertices (i.e., neighbors). In real-world scenarios, a vertex may have multifaceted aspects and should belong to overlapping communities. For example, in a social network, a user may subscribe to political, economic and sports channels simultaneously, but the politics share more common attributes with the economy and less with the sports. In this paper, we propose an adversarial learning approach for modeling overlapping communities of vertices. Each community and vertex are mapped into an embedding space, while we also learn the association between each pair of community and vertex. The experimental results show that our proposed model not only can outperform the state-of-the-art (including GANs-based) models on vertex classification tasks but also can achieve superior performances on overlapping community detection.
Junyang Chen 0001, Zhiguo Gong, Quanyu Dai, Chunyuan Yuan, Weiwen Liu
ECAI4
2020 Exploiting Heterogeneous Artist and Listener Preference Graph for Music Genre Classification
abstract
Music genres are useful for indexing, organizing, searching, and recommending songs and albums. Therefore, the automatic classification of music genres is an essential part of almost all kinds of music applications. Recent works focus on exploiting text, audio, or multi-modal information for genre classification, without considering the influence of the artists' and listeners' preference. However, intuitively, artists have their composing preferences, and listeners also have their music tastes. Both of them provide helpful hints to the music genre from different views, which are crucial to improve classification performance.
Chunyuan Yuan, Qianwen Ma, Junyang Chen 0001, Wei Zhou 0019, Xiaodan Zhang 0004, Xuehai Tang, Jizhong Han, Songlin Hu 0001
ACM Multimedia1
2020 DyHGCN: A Dynamic Heterogeneous Graph Convolutional Network to Learn Users' Dynamic Preferences for Information Diffusion Prediction
Chunyuan Yuan, Wei Zhou 0019, Xiaodan Zhang 0004, Songlin Hu 0001
ECML/PKDD (3)1
2020 Beyond Statistical Relations: Integrating Knowledge Relations into Style Correlations for Multi-Label Music Style Classification
abstract
Automatically labeling multiple styles for every song is a comprehensive application in all kinds of music websites. Recently, some researches explore review-driven multi-label music style classification and exploit style correlations for this task. However, their methods focus on mining the statistical relations between different music styles and only consider shallow style relations. Moreover, these statistical relations suffer from the underfitting problem because some music styles have little training data. To tackle these problems, we propose a novel knowledge relations integrated framework (KRF) to capture the complete style correlations, which jointly exploits the inherent relations between music styles according to external knowledge and their statistical relations. Based on the two types of relations, we use graph convolutional network to learn the deep correlations between styles automatically. Experimental results show that our framework significantly outperforms the state-of-the-art methods. Further studies demonstrate that our framework can effectively alleviate the underfitting problem and learn meaningful style correlations.
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
WSDM2
2019 Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based Chatbots
abstract
Chunyuan Yuan, Wei Zhou, Mingming Li, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chunyuan Yuan, Wei Zhou 0019, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu 0001
EMNLP/IJCNLP (1)1
2019 Jointly Embedding the Local and Global Relations of Heterogeneous Graph for Rumor Detection
abstract
The development of social media has revolutionized the way people communicate, share information and make decisions, but it also provides an ideal platform for publishing and spreading rumors. Existing rumor detection methods focus on finding clues from text content, user profiles, and propagation patterns. However, the local semantic relation and global structural information in the message propagation graph have not been well utilized by previous works. In this paper, we present a novel global-local attention network (GLAN) for rumor detection, which jointly encodes the local semantic and global structural information. We first generate a better integrated representation for each source tweet by fusing the semantic information of related retweets with the attention mechanism. Then, we model the global relationships among all source tweets, retweets, and users as a heterogeneous graph to capture the rich structural information for rumor detection. We conduct experiments on three real-world datasets, and the results demonstrate that GLAN significantly outperforms the state-of-the-art models in both rumor detection and early detection scenarios.
Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
ICDM1
2019 Learning Review Representations from user and Product Level Information for Spam Detection
abstract
Opinion spam has become a widespread problem in social media, where hired spammers write deceptive reviews to promote or demote products to mislead the consumers for profit or fame. Existing works mainly focus on manually designing discrete textual or behavior features, which cannot capture complex global semantics of reviews. Although recent works apply deep learning methods to learn review-level semantic features, their models ignore the impact of the user-level and product-level information on learning review semantics and the inherent user-review-product relationship information. In this paper, we propose a Hierarchical Fusion Attention Network (HFAN) to automatically learn the semantics of reviews from user and product level. Specifically, we design a multiattention unit to extract user(product)-related review information. Then, we use orthogonal decomposition and fusion attention to learn a user, review, and product representation from the review information. Finally, we take the review as a relation between user and product entity and apply TransH to jointly encode this relationship into review representation. Experimental results obtained more than 10% absolute precision improvement over the state-of-the-art performances on four real-world datasets, which show the effectiveness and versatility of the model.
Chunyuan Yuan, Wei Zhou 0019, Qianwen Ma, Shangwen Lv, Jizhong Han, Songlin Hu 0001
ICDM1
2019 Fusion Convolutional Attention Network for Opinion Spam Detection
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
ICONIP (1)3