Wei Li 0076

dblp:64/6025-76 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-8077-7025ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
abstract
Large Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing need for models capable of integrating multimodal information in the chemical domain. In this paper, we introduce ChemVLM, an open-source chemical multimodal large language model specifically designed for chemical applications. ChemVLM is trained on a carefully curated bilingual multimodal dataset that enhances its ability to understand both textual and visual chemical information, including molecular structures, reactions, and chemistry examination questions. We develop three datasets for comprehensive evaluation, tailored to Chemical Optical Character Recognition (OCR), Multimodal Chemical Reasoning (MMCR), and Multimodal Molecule Understanding tasks. We benchmark ChemVLM against a range of open-source and proprietary multimodal large language models on various tasks. Experimental results demonstrate that ChemVLM achieves competitive performance across all evaluated tasks.
Junxian Li 0001, Di Zhang 0026, Xunzhi Wang, Zeying Hao, Jingdi Lei, Cai Zhou, Wei Liu 0123, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen 0013, Wenhai Wang, Wei Li 0076, Mao Su, Shufei Zhang, Wanli Ouyang, Dongzhan Zhou
AAAI14
2025 Can a Large Language Model be a Gaslighter?
abstract
Large language models (LLMs) have gained human trust due to their capabilities and helpfulness. However, this in turn may allow LLMs to affect users' mindsets by manipulating language. It is termed as gaslighting, a psychological effect. In this work, we aim to investigate the vulnerability of LLMs under prompt-based and fine-tuning-based gaslighting attacks. Therefore, we propose a two-stage framework DeepCoG designed to: 1) elicit gaslighting plans from LLMs with the proposed DeepGaslighting prompting template, and 2) acquire gaslighting conversations from LLMs through our Chain-of-Gaslighting method. The gaslighting conversation dataset along with a corresponding safe dataset is applied to fine-tuning-based attacks on open-source LLMs and anti-gaslighting safety alignment on these LLMs. Experiments demonstrate that both prompt-based and fine-tuning-based attacks transform three open-source LLMs into gaslighters. In contrast, we advanced three safety alignment strategies to strengthen~(by $12.05\%$) the safety guardrail of LLMs. Our safety alignment strategies have minimal impacts on the utility of LLMs. Empirical studies indicate that an LLM may be a potential gaslighter, even if it passed the harmfulness test on general dangerous queries.
Wei Li 0076, Ruixi Lin, Rui Mao 0010
ICLR1
2025 MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
abstract
Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothesis discovery, which entails generating detailed, experimentally actionable hypotheses from coarse initial research directions. We frame this as a combinatorial optimization problem and investigate the upper limits of LLMs' capacity to solve it when maximally leveraged. Specifically, we explore four foundational questions: (1) how to best harness an LLM's internal heuristics to formulate the fine-grained hypothesis it itself would judge as the most promising among all the possible hypotheses it might generate, based on its own internal scoring-thus defining a latent reward landscape over the hypothesis space; (2) whether such LLM-judged better hypotheses exhibit stronger alignment with ground-truth hypotheses; (3) whether shaping the reward landscape using an ensemble of diverse LLMs of similar capacity yields better outcomes than defining it with repeated instances of the strongest LLM among them; and (4) whether an ensemble of identical LLMs provides a more reliable reward landscape than a single LLM. To address these questions, we propose a hierarchical search method that incrementally proposes and integrates details into the hypothesis, progressing from general concepts to specific experimental configurations. We show that this hierarchical process smooths the reward landscape and enables more effective optimization. Empirical evaluations on a new benchmark of expert-annotated fine-grained hypotheses from recent literature show that our method consistently outperforms strong baselines.
Zonglin Yang 0001, Wanhao Liu, Ben Gao, Wei Li 0076, Tong Xie, Lidong Bing, Wanli Ouyang, Erik Cambria, Dongzhan Zhou
NeurIPS5
2025 A Discourse Structure- and Interlocutor-Guided Network for Dialogue Act Recognition and Sentiment Classification
abstract
Dialogue act recognition and sentiment classification (DAR and DSC) are closely related tasks in dialogue systems, both benefiting from joint modeling of their interdependencies. Recent advancements have improved performance on these tasks by integrating them, yet many approaches oversimplify dialogues by treating them as monologues and assuming uniform influence across all utterances. This neglects the inherent structural and interactive nature of dialogues. To address these issues, we propose a novelDiscourseStructure- andInterlocutor-Guided (DSIG) network that fuses dialogue act recognition and sentiment classification. Our network synergizes structural dialogue relationships with interlocutor identity information, enabling effective modeling of utterance flow and cross-task interactions. Specifically, we utilize a shared encoder that functions as a dialogue discourse parser to dynamically construct utterance connections, thereby integrating dialogue structural relationships. In addition, we embed interlocutor affiliations into the fusion process during encoding, semantic modeling, and decoding, enhancing dialogue understanding and dual-task reasoning. The core component is a collaborative updating graph interaction layer, which filters redundant connections and introduces interlocutor nodes for exclusive information exchange. Experimental results on two benchmark datasets demonstrate DSIG achieves state-of-the-art performance by effectively fusing dialogue structure and interlocutor information, improving F1 scores by 6.9% and 2.1% for DSC and DAR, respectively, on Mastodon, and by 13.7% and 4.7% on DailyDialog. Model variants trained independently show promise for extension to other dialogue-related tasks.
Yunhe Xie, Rui Mao 0010, Wei Li 0076, Atika Qazi, Erik Cambria
IEEE Trans. Affect. Comput.3
2023 SKIER: A Symbolic Knowledge Integrated Model for Conversational Emotion Recognition
abstract
Emotion recognition in conversation (ERC) has received increasing attention from the research community. However, the ERC task is challenging, largely due to the complex and unstructured properties of multi-party conversations. Besides, the majority of daily dialogues take place in a specific context or circumstance, which requires rich external knowledge to understand the background of a certain dialogue. In this paper, we address these challenges by explicitly modeling the discourse relations between utterances and incorporating symbolic knowledge into multi-party conversations. We first introduce a dialogue parsing algorithm into ERC and further improve the algorithm through a transfer learning method. Moreover, we leverage different symbolic knowledge graph relations to learn knowledge-enhanced features for the ERC task. Extensive experiments on three benchmarks demonstrate that both dialogue structure graphs and symbolic knowledge are beneficial to the model performance on the task. Additionally, experimental results indicate that the proposed model surpasses baseline models on several indices.
Wei Li 0076, Rui Mao 0010, Erik Cambria
AAAI1
2023 PAED: Zero-Shot Persona Attribute Extraction in Dialogues
abstract
Persona attribute extraction is critical for personalized human-computer interaction.Dialogue is an important medium that communicates and delivers persona information.Although there is a public dataset for triplet-based persona attribute extraction from conversations, its automatically generated labels present many issues, including unspecific relations and inconsistent annotations.We fix such issues by leveraging more reliable text-label matching criteria to generate high-quality data for persona attribute extraction.We also propose a contrastive learning-and generation-based model with a novel hard negative sampling strategy for generalized zero-shot persona attribute extraction.We benchmark our model with stateof-the-art baselines on our dataset and a public dataset, showing outstanding accuracy gains.Our sampling strategy also exceeds others by a large margin in persona attribute extraction.
Wei Li 0076, Rui Mao 0010, Vlad Pandelea, Erik Cambria
ACL (1)2
2023 ECPEC: Emotion-Cause Pair Extraction in Conversations
abstract
Conversational sentiment analysis (CSA) and emotion-cause pair extraction (ECPE) tasks have attracted increasing attention in recent years. The former aims to predict the sentiment states of speakers in a conversation, and the latter is about extracting emotion-cause clauses in a document. However, one drawback of CSA is that it cannot model the causal reasoning among emotion and neutral utterances from different speakers. In this work, we propose a new task: emotion-cause pair extraction in conversations (ECPEC), which aims to extract pairs of emotional utterances and corresponding cause utterances in conversations. The utterance-level ECPEC task is more challenging since the distance between emotion and cause utterances is larger than that of the clause-level ECPE task. To this end, we build a novel dataset ConvECPE and propose a specifically designed two-step framework for the new ECPEC task. Experimental results on ConvECPE dataset demonstrate the feasibility of the ECPEC task as well as the effectiveness of our framework.
Wei Li 0076, Yang Li 0055, Vlad Pandelea, Mengshi Ge, Erik Cambria
IEEE Trans. Affect. Comput.1
2023 The Biases of Pre-Trained Language Models: An Empirical Study on Prompt-Based Sentiment Analysis and Emotion Detection
abstract
Thanks to the breakthrough of large-scale pre-trained language model (PLM) technology, prompt-based classification tasks, e.g., sentiment analysis and emotion detection, have raised increasing attention. Such tasks are formalized as masked language prediction tasks which are in line with the pre-training objects of most language models. Thus, one can use a PLM to infer the masked words in a downstream task, then obtaining label predictions with manually defined label-word mapping templates. Prompt-based affective computing takes the advantages of both neural network modeling and explainable symbolic representations. However, there still remain many unclear issues related to the mechanisms of PLMs and prompt-based classification. We conduct a systematic empirical study on prompt-based sentiment analysis and emotion detection to study the biases of PLMs towards affective computing. We find that PLMs are biased in sentiment analysis and emotion detection tasks with respect to the number of label classes, emotional label-word selections, prompt templates and positions, and the word forms of emotion lexicons.
Rui Mao 0010, Qian Liu 0012, Kai He 0001, Wei Li 0076, Erik Cambria
IEEE Trans. Affect. Comput.4
2022 BiERU: Bidirectional emotional recurrent unit for conversational sentiment analysis
Wei Li 0076, Wei Shao 0009, Shaoxiong Ji, Erik Cambria
Neurocomputing1
2021 Taylor's theorem: A new perspective for neural tensor networks
Wei Li 0076, Erik Cambria
Knowl. Based Syst.1
2021 SentiVec: Learning Sentiment-Context Vector via Kernel Optimization Function for Sentiment Analysis
abstract
Deep learning-based sentiment analysis (SA) methods have drawn more attention in recent years, which calls for more precise word embedding methods. This article proposes SentiVec, a kernel optimization function system for sentiment word embedding, which is based on two phases. The first phase is a supervised learning method, and the second phase consists of two unsupervised updating models, object-word-to-surrounding-words reward model (O2SR) and context-to-object-word reward model (C2OR). SentiVec is aimed at: 1) integrating the statistical information and sentiment orientation into sentiment word vectors and 2) propagating and updating the semantic information to all the word representations in a corpus. Extensive experimental results show that the optimal sentiment vectors successfully extract the features in terms of semantic and sentiment information, which makes it outperform the baseline methods on word similarity, word analogy, and SA tasks.
Wei Li 0076, Yong Shi 0001, Kun Guo 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 DWWP: Domain-specific new words detection and word propagation system for sentiment analysis in the tourism domain
Wei Li 0076, Kun Guo 0001, Yong Shi 0001, Yuanchun Zheng
Knowl. Based Syst.1