Yujie Lin 0003

dblp:126/0783-3 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 35% Language models and text generation · 25% Trustworthy machine learning · 20%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
mathematical reasoning
1.012026
Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration · AAAI 2026
Machine learning › Reinforcement learning
reward design
1.012026
Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration · AAAI 2026
Natural language and speech › Information extraction and text analysis › relation extraction
open relation extraction
0.912025
LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models · EMNLP 2025
Natural language and speech › Information extraction and text analysis
relation extraction
0.912025
LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
in-context learning
0.312025
LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.0heuristic reward function · 1.0self-correcting inference · 0.9large language model prompting · 0.9cross-validation · 0.9
YearPublicationVenuePosition
2026 Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration
abstract
Critical thinking is essential for building robust AI systems, preventing them from blindly accepting flawed data or biased reasoning. However, prior work has primarily focused on passive critical thinking, where models simply reject problematic queries without taking constructive steps to address user requests. In this work, we introduce proactive critical thinking, a paradigm where models actively seek missing or clarifying information from users to resolve their queries better. To evaluate this capability, we present GSM-MC and GSM-MCE, two novel benchmarks based on GSM8K for assessing mathematical reasoning under incomplete or misleading conditions. Experiments on Qwen3 and Llama series models show that, while these models excel in traditional reasoning tasks, they struggle with proactive critical thinking, especially smaller ones. However, we demonstrate that reinforcement learning (RL) can significantly improve this ability. By incorporating heuristic information into the reward function, we achieve substantial gains, boosting the Qwen3-1.7B's accuracy from 0.15% to 73.98% on GSM-MC. We hope this work advances models that collaborate more effectively with users in problem-solving through proactive critical thinking.
Ante Wang, Yujie Lin 0003, Suhang Wu, Xinyan Xiao, Jinsong Su
AAAI2
2026 DocTER: Evaluating document-based knowledge editing
Suhang Wu, Ante Wang, Minlong Peng, Yujie Lin 0003, Mingming Sun 0001, Jinsong Su
Inf. Process. Manag.4
2025 LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models
abstract
The goal of open relation extraction (OpenRE) is to develop an RE model that can generalize to new relations not encountered during training. Existing studies primarily formulate OpenRE as a clustering task. They first cluster all test instances based on the similarity between the instances, and then manually assign a new relation to each cluster. However, their reliance on human annotation limits their practicality. In this paper, we propose an OpenRE framework based on large language models (LLMs), which directly predicts new relations for test instances by leveraging their strong language understanding and generation abilities, without human intervention. Specifically, our framework consists of two core components: (1) a relation discoverer (RD), designed to predict new relations for test instances based on demonstrations formed by training instances with known relations; and (2) a relation predictor (RP), used to select the most likely relation for a test instance from n candidate relations, guided by demonstrations composed of their instances. To enhance the ability of our framework to predict new relations, we design a self-correcting inference strategy composed of three stages: relation discovery, relation denoising, and relation prediction. In the first stage, we use RD to preliminarily predict new relations for all test instances. Next, we apply RP to select some high-reliability test instances for each new relation from the prediction results of RD through a cross-validation method. During the third stage, we employ RP to re-predict the relations of all test instances based on the demonstrations constructed from these reliable test instances. Extensive experiments on three OpenRE datasets demonstrate the effectiveness of our framework. We release our code at https://github.com/XMUDeepLIT/LLM-OREF.git.
Hongyao Tu, Yujie Lin 0003, Haibo Zhang 0013, Long Zhang 0012, Jinsong Su
EMNLP3
2025 A Dual-Perspective Metaphor Detection Framework Using Large Language Models
abstract
Metaphor detection, a critical task in natural language processing, involves identifying whether a particular word in a sentence is used metaphorically. Traditional approaches often rely on supervised learning models that implicitly encode semantic relationships based on metaphor theories. However, these methods often suffer from a lack of transparency in their decision-making processes, which undermines the reliability of their predictions. Recent research indicates that LLMs (large language models) exhibit significant potential in metaphor detection. Nevertheless, their reasoning capabilities are constrained by predefined knowledge graphs. To overcome these limitations, we propose DMD, a novel dual-perspective framework that harnesses both implicit and explicit applications of metaphor theories to guide LLMs in metaphor detection and adopts a self-judgment mechanism to validate the responses from the aforementioned forms of guidance. In comparison to previous methods, our framework offers more transparent reasoning processes and delivers more reliable predictions. Experimental results prove the effectiveness of DMD, demonstrating state-of-the-art performance across widely-used datasets.
Yujie Lin 0003, Ante Wang, Jinsong Su
ICASSP1
2024 Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment
abstract
Cross-modal text-molecule retrieval model aims to learn a shared feature space of the text and molecule modalities for accurate similarity calculation, which facilitates the rapid screening of molecules with specific properties and activities in drug design. However, previous works have two main defects. First, they are inadequate in capturing modality-shared features considering the significant gap between text sequences and molecule graphs. Second, they mainly rely on contrastive learning and adversarial training for cross-modality alignment, both of which mainly focus on the first-order similarity, ignoring the second-order similarity that can capture more structural information in the embedding space. To address these issues, we propose a novel cross-modal text-molecule retrieval model with two-fold improvements. Specifically, on the top of two modality-specific encoders, we stack a memory bank based feature projector that contain learnable memory vectors to extract modality-shared features better. More importantly, during the model training, we calculate four kinds of similarity distributions (text-to-text, text-to-molecule, molecule-to-molecule, and molecule-to-text similarity distributions) for each instance, and then minimize the distance between these similarity distributions (namely second-order similarity losses) to enhance cross-modal alignment. Experimental results and analysis strongly demonstrate the effectiveness of our model. Particularly, our model achieves SOTA performance, outperforming the previously-reported best result by 6.4%.
Wanru Zhuang, Yujie Lin 0003, Chunyan Li 0002, Jinsong Su, Xiaochen Bo
BIBM3