Yuyan Chen

dblp:96/11155 · DBLP profile ↗
← Back
33ranked-venue papers
19as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 13 first-author · 19 since 2021Databases, data management, data science and information retrieval · 11 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 MMIFEvol: Towards Evolutionary Multimodal Instruction Following
abstract
Multimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model's ability to ``solve tasks" and ``follow constraints" into a single metric, which fails to accurately reflect its instruction-following capability. (c) Lack of Multi-Task Instructions, they overlook the fact that real-world user instructions often consist of multiple combined tasks. This paper proposes MMIFEvol, a framework for multimodal instruction evolving and benchmarking. First, we define the essential components of a carefully curated multimodal instruction set and establish corresponding difficulty levels, based on which we synthesize diverse instruction data. Next, we decouple the evaluation criteria for the instruction following into three different metrics to construct a high-quality benchmark and assess existing models. Experimental results demonstrate that current models still struggle with following complex instructions, while fine-tuning using MMIFEvol data effectively improves models' responsiveness to multimodal instructions.
Sihang Jiang 0001, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yanghua Xiao
AAAI4
2026 S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection
abstract
Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizability. In this paper, we introduce a new large vision–language model (LVLM) dubbed S³-MSD for explainable and generalizable MSD through three key components. For explainability, we develop (1) a self-training paradigm that automatically bootstraps answers with explanations, and (2) a self-calibrating mechanism that rectifies flawed explanations. For generalizability, we design (3) a self-focusing module that amplifies visual semantic entities through preference optimization, thereby mitigating textual over-reliance. Experimental results on both in-distribution and out-of-distribution (OOD) benchmarks demonstrate that S³-MSD consistently outperforms state-of-the-art methods in detection performance. Furthermore, the proposed S³-MSD provides persuasive explanations, as verified by both quantitative metrics and human evaluations.
Zhihong Zhu 0001, Fan Zhang 0111, Yunyan Zhang, Jinghan Sun, Guimin Hu, Hao Wu 0094, Yuyan Chen, Xian Wu 0001
AAAI7
2026 Efficient Multimodal Serving via Module Multiplexing
abstract
Multimodal learning enables models to process and reason over diverse information sources, unlocking human-like perceptual and cognitive capabilities. As such models gain adoption, efficiently serving them on GPUs has become increasingly important. However, the modular architecture of multimodal models poses significant challenges to existing unimodal serving systems, which treat models as monolithic and overlook inter-module heterogeneity. This results in severe GPU underutilization. To address this, we propose Eevee, a multimodal serving system based on a new scheduling paradigm we call module multiplexing. Unlike prior approaches that execute all modules sequentially with uniform batch sizes, Eevee schedules modality-specific modules concurrently on the same GPU with independently tuned batching and resource allocation. This design enables fine-grained GPU sharing, boosting intra-GPU parallelism and improving request-level throughput. We implement a prototype of Eevee and evaluate it on several representative multimodal models (e.g., CLIP, BLIP, LLaVA, InternVL). Our results show that Eevee significantly outperforms state-of-the-art serving systems in both throughput and GPU utilization.
Zicong Hong, Yuyan Chen, Peng Li 0017, Wuhui Chen, Song Guo 0001
EuroSys2
2026 FlashServe: Adaptive Kernel Provision for Quantized LLM Serving
Yuyan Chen, Junyuan Liang, Wuhui Chen, Zicong Hong, Song Guo 0001, Ruiyan Zhuang, Yi Quan
ICDCS1
2026 Constructing Commonsense Knowledge Graph for Persona Consistency
abstract
Ensuring consistent persona in interactive AI systems presents a significant challenge, especially in diverse application scenarios ranging from virtual assistants to customer service bots. Such capability is often constrained by the system's understanding of direct and explicit persona conflicts. Traditional approaches primarily focus on detecting discrepancies between machine responses and its predefined profile, or the contextual inconsistencies between the responses at the semantic level rather than the persona level. Due to the lack of a comprehensive persona-specific Commonsense Knowledge Graph, some indirect and implicit persona inconsistencies between machine responses can hardly be identified. In this paper, we build the first persona commonsense knowledge graph (PersonaKG), based on which we then construct a large-scale persona consistency dialogue dataset (PersonaCOM) containing both explicit and implicit persona conflicts between machine responses. With the guidance of the persona commonsense knowledge, we propose a Recognize-Rewrite framework (R2) which first recognizes the responses that are inconsistent in persona with the previous responses, and then rewrites them into consistent ones. The empirical study demonstrates that utilizing R2 method on PersonaCOM with PersonaKG results in a significant improvement of 12.20% in automatic metrics and 10.09% in manual evaluation compared to not using the R2 method and PersonaKG.
Lei Xia 0003, Yuyan Chen, Xiangqin Chen, Jixiang Fan, Weinan Dai, Zhixu Li
WSDM2
2025 Attributive Reasoning for Hallucination Diagnosis of Large Language Models
abstract
In recent years, large language models (LLMs) have demonstrated outstanding capabilities in various tasks. However, LLMs also have various drawbacks, especially hallucination. Hallucination refers to the generation of content that does not align with the user input, contradicts previously generated content or world knowledge. Current research on hallucination mainly include knowledge retrieval, prompt engineering, training data improvement, reinforcement learning, etc. However, these methods do not involve different categories of hallucinations which is important on hallucination analysis, and make detailed investigation for the internal state of LLMs which indicates the direction on hallucination occurrence. Therefore, in our research, we introduce an attribution framework to trace the origins of hallucinations based on the internal signals of LLMs. To support this framework, we develop a new benchmark named RelQA-Cate, which includes eight categories of hallucinations for the answers generated by LLMs. After that, we present a novel Differential Penalty Decoding (DPD) strategy for reducing hallucinations through adjusting post-probabilities of each answer. We conduct a series of experiments and the performance on answer reliability has significant improvement, achieving 28.25% at most, which demonstrates the effectiveness of our proposed DPD and its generalization in mitigating hallucination in LLMs.
Yuyan Chen, Shuangjie You, Jingwen Chang, Weinan Dai, Qingpei Guo, Yanghua Xiao
AAAI1
2025 Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
abstract
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
Yi Zhang 0109, Chun-Wun Cheng, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I. Avilés-Rivero
AAAI6
2025 VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions
abstract
Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language models (MLLMs) may accomplish this task by grasping the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. To tackle this task, we first collect a new dedicated Complex VQA dataset named CVQA and then propose VQAGuider, an innovative framework planning a few atomic visual recognition tools by video-related API matching. VQAGuider facilitates a deep engagement with video content and precise responses to complex video-related questions by MLLMs, which is beyond aligning visual and language features for simple VQA tasks. Our experiments demonstrate VQAGuider is capable of navigating the complex VQA tasks by MLLMs and improves the accuracy by 29.6% and 17.2% on CVQA and the existing VQA datasets, respectively, highlighting its potential in advancing MLLMs’s capabilities in video understanding.
Yuyan Chen, Jiyuan Jia, Yu Guan 0001, Ming Yang 0007, Qingpei Guo
ACL (1)1
2025 High-Context Empathy in Conversations for Large Language Models
abstract
Large Language Models (LLMs) exhibit remarkable capabilities across various downstream tasks, including empathetic dialogues. However, a non-trivial question arises: Do they possess high-context empathy and can they generate emotional interactions with humans? High-context empathy, which tends to be more indirect and concise like Chinese-style empathy, differs from the current empathy capabilities of LLMs. These capabilities are predominantly low-context empathy, which is often direct and lengthy, resembling English-style empathy. In this paper, We first construct a comprehensive Chinese High-context Empathy Dialogue dataset (HED), which consists of emotional, role-based emotional, personality-based emotional, and role-personality-based emotional dialogues. Next, we explore whether LLMs have high-context empathy in conversations. After that, we propose an innovative High-context Empathy Network (HEN) to improve LLMs' capabilities in generating high-context empathetic responses. Our empirical study demonstrates that there is much room for LLMs in generating high-context empathetic responses, and the proposed HEN can not only significantly improve LLMs' capabilities in generating high-context empathetic responses, but also has positive effects for LLMs in solving similar sentiment-related tasks.
Yuyan Chen, Lei Xia 0003, Jinghan Cao, Zhendong Hou, Weinan Dai, Zhixu Li
CIKM1
2025 Engage for All: Making Ordinary Image Descriptions Appealing Again!
Yuyan Chen, Jinghan Cao, Yu Guan 0001, Ming-Hsuan Yang 0001, Qing Guo 0005
ICCV1
2025 Open-Insect: Benchmarking Open-Set Recognition of Novel Species in Biodiversity Monitoring
abstract
Global biodiversity is declining at an unprecedented rate, yet little information isknown about most species and how their populations are changing. Indeed, some90% Earth’s species are estimated to be completely unknown. Machine learning hasrecently emerged as a promising tool to facilitate long-term, large-scale biodiversitymonitoring, including algorithms for fine-grained classification of species fromimages. However, such algorithms typically are not designed to detect examplesfrom categories unseen during training – the problem of open-set recognition(OSR) – limiting their applicability for highly diverse, poorly studied taxa such asinsects. To address this gap, we introduce Open-Insect, a large-scale, fine-graineddataset to evaluate unknown species detection across different geographic regionswith varying difficulty. We benchmark 38 OSR algorithms across three categories:post-hoc, training-time regularization, and training with auxiliary data, finding thatsimple post-hoc approaches remain a strong baseline. We also demonstrate how toleverage auxiliary data to improve species discovery in regions with limited data.Our results provide timely insights to guide the development of computer visionmethods for biodiversity monitoring and species discovery.
Yuyan Chen, Nico Lang, B. Christian Schmidt, Yves Basset, Sara Beery, Maxim Larrivée, David Rolnick
NeurIPS1
2025 MedTransTab: Advancing Medical Cross-Table Tabular Data Generation
abstract
In medical research, clinical trials are pivotal. While prospective clinical research provides a systematic approach to collecting patient data, it grapples with challenges like long durations, increased costs, and most crucially, data scarcity. To address above-mentioned challenge, this paper introduces a novel approach: using cross-table generation to create relevant data. Unlike existing work focused on single-table operations, our method leverages data from multiple sources across various tables, integrating diverse data types and ensuring data consistency across multiple tables. We develop a new framework, MedTransTab, tailored for cross-table tabular data generation in the medical context. This framework extends our previous efforts and is built upon the newly constructed PMC-Struct, derived from an unstructured PMC-patient dataset. Our MedTransTab can generate high-quality patient records, synthesizing detailed biomedical information to align with real or simulated tables from multiple sources. The experiments show that the proposed method significantly improves performance in cross-table tasks. On the PMC-Struct-Plus dataset, we observe an average improvement of 28.85% in data generation and prediction. Similarly, on the Out-Of-Domain (OOD) dataset, there's an average improvement of 22.56%, indicating substantial progress in medical data analysis.
Yuyan Chen, Qingpei Guo, Shuangjie You, Zhixu Li
WSDM1
2024 Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor Interpretation
abstract
Humor is a crucial part of human communication. Understanding humor and generating humorous responses in dialogue can provide natural and empathic human-computer interactions. However, most existing pre-trained language models (PLMs) perform unsatisfactorily in humor generation. On one hand, the serious shortage of humor corpus and datasets pose challenges for constructing models that can understand and generate humorous expressions. On the other hand, humor generation relies on rich knowledge and commonsense, which is often tacit and unspoken. In this paper, we construct the largest Chinese Explainable Humor Response Dataset to date with chain-of-humor and humor mind map annotations, which can be used to comprehensively evaluate as well as improve the humorous response ability of PLMs. We further design humor-related auxiliary tasks to further enhance PLMs' humorous response performance. Extensive evaluations demonstrate that our proposed dataset and auxiliary tasks effectively help PLMs to generate humorous responses, laying the groundwork for future humor research.
Yuyan Chen, Panjun Liu, Dayiheng Liu, Qinghao Guan, Mengfei Guo, Haiming Peng, Bang Liu 0003, Zhixu Li, Yanghua Xiao
AAAI1
2024 Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
abstract
Teachers are important to imparting knowledge and guiding learners, and the role of large language models (LLMs) as potential educators is emerging as an important area of study.Recognizing LLMs' capability to generate educational content can lead to advances in automated and personalized learning.While LLMs have been tested for their comprehension and problem-solving skills, their capability in teaching remains largely unexplored.In teaching, questioning is a key skill that guides students to analyze, evaluate, and synthesize core concepts and principles.Therefore, our research introduces a benchmark to evaluate the questioning capability in education as a teacher of LLMs through evaluating their generated educational questions, utilizing Anderson and Krathwohl's taxonomy across general, monodisciplinary, and interdisciplinary domains.We shift the focus from LLMs as learners to LLMs as educators, assessing their teaching capability through guiding them to generate questions.We apply four metrics, including relevance, coverage, representativeness, and consistency, to evaluate the educational quality of LLMs' outputs.Our results indicate that GPT-4 demonstrates significant potential in teaching general, humanities, and science courses; Claude2 appears more apt as an interdisciplinary teacher.Furthermore, the automatic scores align with human perspectives.
Yuyan Chen, Songzhou Yan, Panjun Liu, Yanghua Xiao
ACL (1)1
2024 DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm Detection
abstract
Zhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu, Xian Wu, Yefeng Zheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhihong Zhu 0001, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu 0027, Xian Wu 0001, Yefeng Zheng 0001
EMNLP5
2024 Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator
abstract
In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all simply learn discriminative patterns by modelling low-level relationships between adjacent points of trajectories, while completely ignoring the inherent geometric structures of characters. Instead, we propose a novel Graph-guided Cross-modality Translator for IAHTR, which further explicitly exploits the geometric structures of characters for guiding the decoding of trajectories via graph-guided cross-modality attention mechanism without introducing extra annotation costs. Experiments on benchmarks IAHEW-UCAS2016 & IAM-OnDB show that our method has achieved state-of-the-art performance for handwritten text recognition.
Yuyan Chen, Ji Gan, Jiaxu Leng, Yan Zhang 0108, Xinbo Gao 0001
ICASSP1
2024 Kenet: Knowledge-Enhanced DOC-Label Attention Network for Multi-Label Text Classification
abstract
Multi-Label Text Classification (MLTC) is a fundamental task in the field of Natural Language Processing (NLP) that involves the assignment of multiple labels to a given text. MLTC has gained significant importance and has been widely applied in various domains such as topic recognition, recommendation systems, sentiment analysis, and information retrieval. However, traditional machine learning and Deep neural network have not yet addressed certain issues, such as the fact that some documents are brief but have a large number of labels and how to establish relationships between the labels. It is imperative to additionally acknowledge that the significance of knowledge is substantiated in the realm of MLTC. To address this issue, we provide a novel approach known as Knowledge-enhanced Doc-Label Attention Network (KeNet). Specifically, we design an Attention Network that incorporates external knowledge, label embedding, and a comprehensive attention mechanism. In contrast to conventional methods, we use comprehensive representation of documents, knowledge and labels to predict all labels for each single text. Our approach has been validated by comprehensive research conducted on three multi-label datasets. Experimental results demonstrate that our method outperforms state-of-the-art MLTC method. Additionally, a case study is undertaken to illustrate the practical implementation of KeNet.
Bo Li 0131, Yuyan Chen
ICASSP2
2024 XMeCap: Meme Caption Generation with Sub-Image Adaptability
abstract
Humor, deeply rooted in societal meanings and cultural details, poses a unique challenge for machines. While advances have been made in natural language processing, real-world humor often thrives in a multi-modal context, encapsulated distinctively by memes. This paper poses a particular emphasis on the impact of multi-images on meme captioning. After that, we introduce the XMeCap framework, a novel approach that adopts supervised fine-tuning and reinforcement learning based on an innovative reward model, which factors in both global and local similarities between visuals and text. Our results, benchmarked against contemporary models, manifest a marked improvement in caption generation for both single-image and multi-image memes, as well as different meme categories. XMeCap achieves an average evaluation score of 75.85 for single-image memes and 66.32 for multi-image memes, outperforming the best baseline by 3.71% and 4.82%, respectively. This research not only establishes a new frontier in meme-related studies but also underscores the potential of machines in understanding and generating humor in a multi-modal setting.
Yuyan Chen, Songzhou Yan, Zhihong Zhu 0001, Zhixu Li, Yanghua Xiao
ACM Multimedia1
2024 InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory Processing
abstract
Multi-modal intent detection (MID) aims to comprehend users' intentions through diverse modalities, which has received widespread attention in dialogue systems. Despite the promising advancements in complex fusion mechanisms or architecture designs, challenges remain due to: (1) various noise and redundancy in both visual and audio modalities and (2) long-tailed distributions of intent categories. In this paper, to tackle the above two issues, we propose InMu-Net, a simple yet effective framework for MID from the Information bottleneck and Multi-sensory processing perspective. Our contributions lie in three aspects. First, we devise a denoising bottleneck module to filter out the intent-irrelevant information in the fused feature; Second, we introduce a saliency preservation loss to prevent the dropping of intent-relevant information; Ultimately, kurtosis regulation is introduced to maintain representation smoothness during the filtering process, mitigating the adverse impact of the long tail distribution. Comprehensive experiments on two MID benchmark datasets demonstrate the effectiveness of InMu-Net and its vital components. Impressively, a series of analyses reveal our denoising potential and robustness in low-resource, modality corruption, cross-architecture and cross-task scenarios.
Zhihong Zhu 0001, Xuxin Cheng, Zhaorun Chen, Yuyan Chen, Yunyan Zhang, Xian Wu 0001, Yefeng Zheng 0001
ACM Multimedia4
2024 TemporalMed: Advancing Medical Dialogues with Time-Aware Responses in Large Language Models
abstract
Medical dialogue models predominantly emphasize generating coherent and clinically accurate responses. However, in many clinical scenarios, time plays a pivotal role, often dictating subsequent patient management and interventions. Recognizing the latent importance of temporal dynamics, this paper introduces a novel dimension to medical dialogues: timestamps. We advocate that the integration of time-sensitive directives can profoundly impact medical advice, using an illustrative example of post-surgery care with and without timestamps. Our contributions are three-fold: Firstly, we highlight the intrinsic significance of timestamps in medical conversations, marking a paradigm shift in dialogue modeling. Secondly, we present an innovative dataset and framework explicitly tailored for time-stamped medical dialogues, facilitating the model to not only provide medical counsel but also chronologically outline care regimens. Lastly, empirical evaluations indicate our method's proficiency in time-stamped tasks and reveal an uptick in performance in broader medical Q&A domains. Through our endeavors, we aspire to set new benchmarks in patient-centric and time-sensitive medical dialogue systems.
Yuyan Chen, Jin Zhao 0004, Zhihao Wen, Zhixu Li, Yanghua Xiao
WSDM1
2024 A multi-granularity facial extreme makeup transfer and removal model with local-global collaboration
Yuyan Chen, Jing Chi, Tianshu Shen, Bingyi You, Caiming Zhang 0001
Appl. Intell.1
2024 XMQAs: Constructing Complex-Modified Question-Answering Dataset for Robust Question Understanding
abstract
Question understanding is an important issue to the success of a Knowledge-based Question Answering (KBQA) system.However, the existing study does not pay enough attention to this issue given that the questions in the existing KBQA datasets are usually expressed in simple and straightforward way. This is not in line with the actual linguistic conventions, which often use a lot of modifiers. To facilitate the study on evaluating and enhancing the question understanding ability of the KBQA systems, this paper proposes to construct a complex-modified question-answering (XMQAs) dataset based on existing KBQA datasets. With the help of knowledge bases and dictionaries, three kinds of modifiers are defined and applied to original simple-expressed questions. These modifiers could make the expression of these questions complex without changing their semantics. Based on XMQAs, we then propose a novel question understanding algorithm upon existing KBQA models, which greatly improves the robustness of their question understanding abilities. We conduct extensive experiments on XMQAs and two widely acknowledged KBQA datasets. The empirical results demonstrate that our proposed algorithm can improve the performance of KBQA models on not only the complex-modified questions, but also simple-expressed questions.
Yuyan Chen, Yanghua Xiao, Zhixu Li, Bang Liu 0003
IEEE Trans. Knowl. Data Eng.1
2023 Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
abstract
Recent years, Pre-trained Language models (PLMs) have swept into various fields of artificial intelligence and achieved great success. However, most PLMs, such as T5 and GPT3, have a huge amount of parameters, fine-tuning them is often expensive and time consuming, and storing them takes up a lot of space. Therefore, it is necessary to adopt a parameter-efficient approach to reduce parameters of PLMs in fine-tuning without compromising their performance in downstream tasks. In this paper, we design a novel adapter which only acts on self-attention outputs in PLMs. This adapter adopts element-wise linear transformation using Hadamard product, hence named as Hadamard adapter, requires the fewest parameters compared to previous parameter-efficient adapters. In addition, we also summarize some tuning patterns for Hadamard adapter shared by various downstream tasks, expecting to provide some guidance for further parameter reduction with shared adapters in future studies. The experiments conducted on the widely-used GLUE benchmark with several SOTA PLMs prove that the Hadamard adapter achieves competitive performance with only 0.033% parameters compared with full fine-tuning, and it has the fewest parameters compared with other adapters. Moreover, we further find that there is also some redundant layers in the Hadamard adapter which can be removed to achieve more parameter efficiency with only 0.022% parameters.
Yuyan Chen, Qiang Fu 0015, Ge Fan, Lun Du, Jian-Guang Lou, Shi Han, Dongmei Zhang 0001, Zhixu Li, Yanghua Xiao
CIKM1
2023 Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
abstract
Large language models (LLMs) have gained widespread adoption in various natural language processing tasks, including question answering and dialogue systems. However, a major drawback of LLMs is the issue of hallucination, where they generate unfaithful or inconsistent content that deviates from the input source, leading to severe consequences. In this paper, we propose a robust discriminator named RelD to effectively detect hallucination in LLMs' generated answers. RelD is trained on the constructed RelQA, a bilingual question-answering dialogue dataset along with answers generated by LLMs and a comprehensive set of metrics. Our experimental results demonstrate that the proposed RelD successfully detects hallucination in the answers generated by diverse LLMs. Moreover, it performs well in distinguishing hallucination in LLMs' generated answers from both in-distribution and out-of-distribution datasets. Additionally, we also conduct a thorough analysis of the types of hallucinations that occur and present valuable insights. This research significantly contributes to the detection of reliable answers generated by LLMs and holds noteworthy implications for mitigating hallucination in the future work.
Yuyan Chen, Qiang Fu 0015, Zhihao Wen, Ge Fan, Dayiheng Liu, Dongmei Zhang 0001, Zhixu Li, Yanghua Xiao
CIKM1
2023 Can Pre-trained Language Models Understand Chinese Humor?
abstract
Humor understanding is an important and challenging research in natural language processing. As the popularity of pre-trained language models (PLMs), some recent work makes preliminary attempts to adopt PLMs for humor recognition and generation. However, these simple attempts do not substantially answer the question: whether PLMs are capable of humor understanding? This paper is the first work that systematically investigates the humor understanding ability of PLMs. For this purpose, a comprehensive framework with three evaluation steps and four evaluation tasks is designed. We also construct a comprehensive Chinese humor dataset, which can fully meet all the data requirements of the proposed evaluation framework. Our empirical study on the Chinese humor dataset yields some valuable observations, which are of great guiding value for future optimization of PLMs in humor understanding and generation.
Yuyan Chen, Zhixu Li, Jiaqing Liang, Yanghua Xiao, Bang Liu 0003, Yunwen Chen
WSDM1
2023 Characters as graphs: Interpretable handwritten Chinese character recognition via Pyramid Graph Transformer
Ji Gan, Yuyan Chen, Bo Hu 0008, Jiaxu Leng, Weiqiang Wang 0001, Xinbo Gao 0001
Pattern Recognit.2
2022 Grow-and-Clip: Informative-yet-Concise Evidence Distillation for Answer Explanation
abstract
Interpreting the predictions of existing Question Answering (QA) models is critical to many real-world intelligent applications, such as QA systems for healthcare, education, and finance. However, existing QA models lack interpretability and provide no feedback or explanation for end-users to help them understand why a specific prediction is the answer to a question. In this research, we argue that the evidences of an answer is critical to enhancing the interpretability of QA models. Unlike previous research that simply extracts several sentence(s) in the context as evidence, we are the first to explicitly define the concept of evidence as the supporting facts in a context which are informative, concise, and readable. Besides, we provide effective strategies to quantitatively measure the informativeness, conciseness and readability of evidence. Furthermore, we propose Grow-and-Clip Evidence Distillation (GCED) algorithm to extract evidences from the contexts by trade-off informativeness, conciseness, and readability. We conduct extensive experiments on the SQuAD and TriviaQA datasets with several baseline models to evaluate the effect of GCED on interpreting answers to questions. Human evaluation are also carried out to check the quality of distilled evidences. Experimental results show that automatic distilled evidences have human-like informativeness, conciseness and readability, which can enhance the interpretability of the answers to questions.
Yuyan Chen, Yanghua Xiao, Bang Liu 0003
ICDE1
2021 Brain Connectivity: Exploring from a High-Level Topological Perspective
Wei Sheng, Shaoqiang Han, Yun-Shuang Fan, Yuyan Chen, Huafu Chen
ICIG (2)7
2020 AMQAN: Adaptive Multi-Attention Question-Answer Networks for Answer Selection
Haitian Yang, Weiqing Huang, Xuan Zhao 0011, Yan Wang 0081, Yuyan Chen, Rui Mao 0004
ECML/PKDD (3)5
2019 Gated Relational Graph Neural Network for Semi-supervised Learning on Knowledge Graphs
Yuyan Chen, Lei Zou 0001, Zongyue Qin
WISE1
2019 Energy management strategy for HEV based on KFCM and neural network
abstract
Summary Aiming at the deficiency of optimal control energy management strategy, a model of energy management controller for hybrid electric vehicle (HEV) is constructed based on Kernel Fuzzy C‐means Clustering (KFCM) and multi‐neural network. Using energy management control strategy based on PMP, the operational parameters of the four driving modes for HEV is extracted; the data cluster corresponding to the driving mode is generated by clustering through the KFCM method and is used as the training samples for the feedforward neural network. Taking the battery SOC, needed power and speed as the inputs of neural network, and taking engine power as the output of neural network, four sub‐neural network models are established. Taking the vehicle driving needed power at the current moment and the engine output power at the previous moment as characteristic parameters, the corresponding sub‐neural network model is selected for output prediction according to the proportional relationship between the driving demand torque and the engine output power. The simulation results show that, compared with the energy management strategy based on PMP, the calculation time is greatly shortened using the proposed control strategy, and the real‐time performance is better. The fuel economy is a little decreased under the condition of meeting the requirements, but better dynamic performance can be obtained.
Yeqin Wang, Aoyun Xia, Yuyan Chen, Zhongyi Tang
Concurr. Comput. Pract. Exp.5
2019 What are the factors affecting the handover process in open source development?
Bohan Liu 0003, Guoping Rong, Liming Dong 0001, He Zhang 0001, Danni Chen, Tiange Chen, Yuyan Chen
J. Syst. Softw.7
2012 Computing of the contribution rate of scientific and technological progress to economic growth in Chinese regions
Haixiang Guo, Jinglu Hu, Shiwei Yu, Yuyan Chen
Expert Syst. Appl.5