Yanjiang Chen

dblp:320/9733 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0008-6951-4225ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generative Query Augmentation with Dual-view Contrastive Learning for Dense Retrieval in Conversational Search
Wenyu Yan, Aoran Gan, Xukai Liu, Yanjiang Chen, Kai Zhang 0038, Qi Liu 0003
DASFAA (6)4
2025 Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document Relationship
abstract
Dense retrieval has emerged as the leading approach in information retrieval, aiming to find semantically relevant documents based on natural language queries. Given that a single document can be retrieved by multiple distinct queries, existing methods aim to represent a document with multiple vectors. Each vector is aligned with a different query to model the many-to-one relationship between queries and documents. However, these multiple vector-based approaches encounter challenges such as Increased Storage, Vector Collapse, and Search Efficiency. To address these issues, we introduce the Distribution-Driven Dense Retrieval framework (DDR). Specifically, we use vectors to represent queries and distributions to represent documents. This approach not only captures the relationships between multiple queries corresponding to the same document but also avoids the need to use multiple vectors to represent the document. Furthermore, to ensure search efficiency for DDR, we propose a dot product-based computation method to calculate the similarity between documents represented by distributions and queries represented by vectors. This allows for seamless integration with existing approximate nearest neighbor (ANN) search algorithms for efficient search. Finally, we conduct extensive experiments on real-world datasets, which demonstrate that our method significantly outperforms traditional dense retrieval methods.
Junfeng Kang, Rui Li 0093, Qi Liu 0003, Zhenya Huang, Zheng Zhang 0048, Yanjiang Chen, Linbo Zhu, Yu Su 0002
AAAI6
2025 PQR: Improving Dense Retrieval via Potential Query Modeling
abstract
Dense retrieval has now become the mainstream paradigm in information retrieval. The core idea of dense retrieval is to align document embeddings with their corresponding query embeddings by maximizing their dot product. The current training data is quite sparse, with each document typically associated with only one or a few labeled queries. However, a single document can be retrieved by multiple different queries. Aligning a document with just one or a limited number of labeled queries results in a loss of its semantic information. In this paper, we propose a training-free Potential Query Retrieval (PQR) framework to address this issue. Specifically, we use a Gaussian mixture distribution to model all potential queries for a document, aiming to capture its comprehensive semantic information. To obtain this distribution, we introduce three sampling strategies to sample a large number of potential queries for each document and encode them into a semantic space. Using these sampled queries, we employ the Expectation-Maximization algorithm to estimate parameters of the distribution. Finally, we also propose a method to calculate similarity scores between user queries and documents under the PQR framework. Extensive experiments demonstrate the effectiveness of the proposed method.
Junfeng Kang, Rui Li 0093, Qi Liu 0003, Yanjiang Chen, Zheng Zhang 0048, Junzhe Jiang 0001, Yu Su 0002
ACL (1)4
2025 MASS: Mitigating Aspect-Oriented Semantic Sparsity for Fine-Grained Sentiment Analysis
Yanjiang Chen, Kai Zhang 0038, Linan Yue, Kun Zhang 0015, Qi Liu 0003
DASFAA (1)1
2025 Learnable Relational Knowledge Distillation For Language Model Compression
Feng Hu 0005, Kai Zhang 0038, Ye Liu 0011, Meikai Bao, Xukai Liu, Yanjiang Chen, Qi Liu 0003
DASFAA (6)6
2025 Enhancing Protein-Ligand Binding Affinity Prediction via Parameter-Efficient Fine-Tuning of Protein and Chemical Language Models
Ruikang Li, Jiaxian Yan, Kai Zhang 0038, Yanjiang Chen, Qi Liu 0003, Min Gao 0017, Enhong Chen
DASFAA (2)4
2024 Dynamic Multi-granularity Attribution Network for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity of a specific aspect within a given sentence.Most existing methods predominantly leverage semantic or syntactic information based on attention scores, which are susceptible to interference caused by irrelevant contexts and often lack sentiment knowledge at a data-specific level.In this paper, we propose a novel Dynamic Multigranularity Attribution Network (DMAN) from the perspective of attribution.Initially, we leverage Integrated Gradients to dynamically extract attribution scores for each token, which contain underlying reasoning knowledge for sentiment analysis.Subsequently, we aggregate attribution representations from multiple semantic granularities in natural language, enhancing a profound understanding of the semantics.Finally, we integrate attribution scores with syntactic information to capture the relationships between aspects and their relevant contexts more accurately during the sentence understanding process.Extensive experiments on five benchmark datasets demonstrate the effectiveness of our proposed method.
Yanjiang Chen, Kai Zhang 0038, Feng Hu 0005, Xianquan Wang, Ruikang Li, Qi Liu 0003
EMNLP1
2024 I-AM-G: Interest Augmented Multimodal Generator for Item Personalization
abstract
The emergence of personalized generation has made it possible to create texts or images that meet the unique needs of users.Recent advances mainly focus on style or scene transfer based on given keywords.However, in e-commerce and recommender systems, it is almost an untouched area to explore user historical interactions, automatically mine user interests with semantic associations, and create item representations that closely align with user individual interests.In this paper, we propose a brand new framework called Interest Augmented Multimodal Generator (I-AM-G).The framework first extracts tags from the multimodal information of items that the user has interacted with, and the most frequently occurred ones are extracted to rewrite the text description of the item.Then, the framework uses a decoupled text-to-text and image-to-image retriever to search for the top-K similar item text and image embeddings from the item pool.Finally, the Attention module for user interests fuses the retrieved information in a cross-modal manner and further guides the personalized generation process collaborating with the rewritten text.We conducted extensive and comprehensive experiments to demonstrate that our framework can effectively generate results aligned with user preferences, which potentially provides a new paradigm of Rewrite and Retrieve for personalized generation.
Xianquan Wang, Likang Wu, Shukang Yin, Zhi Li 0057, Yanjiang Chen, Hufeng Hufeng, Yu Su 0002, Qi Liu 0003
EMNLP5
2024 Get Rid of Isolation: A Continuous Multi-task Spatio-Temporal Learning Framework
abstract
Spatiotemporal learning has become a pivotal technique to enable urban intelligence. Traditional spatiotemporal models mostly focus on a specific task by assuming a same distribution between training and testing sets. However, given that urban systems are usually dynamic, multi-sourced with imbalanced data distributions, current specific task-specific models fail to generalize to new urban conditions and adapt to new domains without explicitly modeling interdependencies across various dimensions and types of urban data. To this end, we argue that there is an essential to propose a Continuous Multi-task Spatio-Temporal learning framework (CMuST) to empower collective urban intelligence, which reforms the urban spatiotemporal learning from single-domain to cooperatively multi-dimensional and multi-task learning. Specifically, CMuST proposes a new multi-dimensional spatiotemporal interaction network (MSTI) to allow cross-interactions between context and main observations as well as self-interactions within spatial and temporal aspects to be exposed, which is also the core for capturing task-level commonality and personalization. To ensure continuous task learning, a novel Rolling Adaptation training scheme (RoAda) is devised, which not only preserves task uniqueness by constructing data summarization-driven task prompts, but also harnesses correlated patterns among tasks by iterative model behavior modeling. We further establish a benchmark of three cities for multi-task spatiotemporal learning, and empirically demonstrate the superiority of CMuST via extensive evaluations on these datasets. The impressive improvements on both few-shot streaming data and new domain tasks against existing SOAT methods are achieved. Code is available at https://github.com/DILab-USTCSZ/CMuST.
Zhongchao Yi, Zhengyang Zhou, Qihe Huang, Yanjiang Chen, Liheng Yu, Yang Wang 0015
NeurIPS4
2024 Harnessing domain insights: A prompt knowledge tuning method for aspect-based sentiment analysis
Kai Zhang 0038, Qi Liu 0003, Meikai Bao, Yanjiang Chen
Knowl. Based Syst.5