Kaili Sun

dblp:211/3333 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Decoding visual neural representations by multimodal with dynamic balancing
Kaili Sun, Xingyu Miao, Bing Zhai, Haoran Duan 0001, Yang Long 0001
Expert Syst. Appl.1
2026 Cross-modal progressive modeling for neuro-visual representation learning
abstract
Neural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG dataset, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our methods narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines.
Jiyao Pu, Kaili Sun, Zeyu Fu, Haoran Duan 0001, Yang Long 0001
Neurocomputing3
2025 RAIDEN Benchmark: Evaluating Role-playing Conversational Agents with Measurement-Driven Custom Dialogues
abstract
As Large-scale Language Models (LLMs) advance, the development of engaging Role-Playing Conversational Agents (RPCAs) has gained prominence. Despite this progress, there is a notable absence of benchmarks designed around dialogues, rather than question-answering formats, to assess the effectiveness of RPCA interactions. This paper introduces the RAIDEN benchmark, containing a comprehensive dataset specifically developed for RPCA evaluation, comprising over 40,000 multi-turn utterances across 135 characters. The benchmark focuses on assessing particular dimensions at different stages of a conversation, facilitated through interactions conducted by annotators. This approach allows the evaluation phase to concentrate on specific response dimensions, and thus subjectivity in dialogue evaluation is reduced. To further enhance objectivity, evaluators compare responses from two different models rather than assessing a single response in isolation. Besides, we introduce RPCAJudger, a specialized judging LLM tailored for automatic RPCA evaluation. The evaluations conducted by RPCAJudger closely mirror human judgments, and its API-free methodology serves to prevent potential data leakage. All the models and all non-private leaderboard data will be made publicly available.
Bowen Wu 0001, Kaili Sun, Ziwei Bai, Ying Li 0012, Baoxun Wang
COLING2
2025 Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agents
abstract
Large Language Models (LLMs) have demonstrated significant advancements in various fields, notably in Role-Playing Conversational Agents (RPCAs).However, when confronted with role-specific professional inquiries, LLMsbased RPCAs tend to underperform due to their excessive emphasis on the conversational abilities of characters rather than effectively invoking and integrating relevant expert knowledge.This often results in inaccurate responses.We refer to this phenomenon as the "Knowledge Misalignment" which underscores the limitations of RPCAs in integrating expert knowledge.To mitigate this issue, we have introduced an Anchoring-Guidance Fine-Tuning (AnGFT) Framework into the RPCAs' training process.This involves initially linking the Anchoring-Based System Prompt (ASP) with the LLM's relevant expert domains through diverse prompt construction strategies and supervised fine-tuning (SFT).Following the roleplay enriched SFT, the integration of ASP enables LLMs to better associate with relevant expert knowledge, thus enhancing their response capabilities in role-specific expert domains.Moreover, we have developed four comprehensive metrics-helpfulness, thoroughness, credibility, and feasibility-to evaluate the proficiency of RPCAs in responding to professional questions.Our method was tested across four professional fields, and the experimental outcomes suggest that the proposed AnGFT Framework substantially improves the RPCAs' performance in handling role-specific professional queries, while preserving their robust role-playing abilities.
Qibin Li, Shengyuan Bai, Nianmin Yao, Kaili Sun, Baoxun Wang
EMNLP5
2024 Contextual Augmented Global Contrast for Multimodal Intent Recognition
abstract
Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language, visual, and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Existing MIR methods tend to model the individual video independently, ignoring global contextual information across videos. This learning manner inevitably introduces perception biases, exacerbated by the inconsistencies of the multimodal representation, amplifying the intent uncertainty. This challenge motivates us to explore effective global context modeling. Thus, we propose a context-augmented global contrast (CAGC) method to capture rich global context features by mining both intra-and cross-video context interactions for MIR. Concretely, we design a context-augmented transformer module to extract global context dependencies across videos. To further alleviate error accumulation and interference, we develop a cross-video bank that retrieves effective video sources by considering both intentional tendency and video similarity. Furthermore, we introduce a global context-guided contrastive learning scheme, designed to mitigate inconsistencies arising from global context and individual modalities in different feature spaces. This scheme incorporates global cues as the supervision to capture robust the multimodal intent representation. Experiments demonstrate CAGC obtains superior performance than state-of-the-art MIR methods. We also generalize our approach to a closely related task, multimodal sentiment analysis, achieving the comparable performance.
Kaili Sun, Zhiwen Xie, Mang Ye, Huyin Zhang
CVPR1
2024 Fuzzy Concession Strategy for Emotional Human-Computer Negotiation*
abstract
This paper proposes an emotion-based human-computer negotiation system. Existing systems handle basic dialogues but struggle with complex negotiations involving human emotions. To address this, we design a fuzzy concession strategy that adjusts tactics through sentiment analysis, managing states like anger, happiness, and anxiety. Experiments show significant improvement in negotiation success and overall gains when considering emotional factors. Results demonstrate that this model enhances customer satisfaction and user experience by achieving agreements that satisfy users. This research highlights the potential of integrating fuzzy reasoning, sentiment analysis, and natural language processing into AI systems, laying the foundation for more intelligent, human-like dialogue agents and improving human-computer interaction and e-commerce engagement.
Kaili Sun
ICTAI3
2024 An Emotion-Aware Human-Computer Negotiation Model Powered by Pretrained Language Model
Zhiqi Deng, Kaili Sun, Pingping Lin
KSEM (4)3
2024 SDGIN: Structure-aware dual-level graph interactive network with semantic roles for visual dialog
Kaili Sun, Zhiwen Xie, Chi Guo, Huyin Zhang
Knowl. Based Syst.1
2023 Sentiment Analysis Based on Pretrained Language Models: Recent Progress
Binxia Yang, Xudong Luo 0001, Kaili Sun, Michael Y. Luo
ICONIP (12)3
2023 Recent Progress on Text Summarisation Based on BERT and GPT
Binxia Yang, Xudong Luo 0001, Kaili Sun, Michael Y. Luo
KSEM (4)3
2022 A Survey of Sentiment Analysis Based on Pretrained Language Models
abstract
Pretrained Language Models (PLMs) can be applied to downstream tasks with only fine-tuning, without learning the model from scratch. In particular, PLMs have been applied to Sentiment Analysis (SA), which detects, analyses, and extracts the polarity of the sentiment expressed in texts. To help researchers quickly grasp the state-of-art PLM-based SA, we survey PLM-based methods for mono-lingual and cross-lingual SA in this paper. Specifically, we brief these methods, compare their per-formance and point out the challenges for future research.
Kaili Sun, Xudong Luo 0001, Michael Y. Luo
ICTAI1
2022 A Survey of Pretrained Language Models
Kaili Sun, Xudong Luo 0001, Michael Y. Luo
KSEM (2)1
2022 HVLM: Exploring Human-Like Visual Cognition and Language-Memory Network for Visual Dialog
Kaili Sun, Chi Guo, Huyin Zhang
Inf. Process. Manag.1
2022 Syntax-Aware graph convolutional network for the recognition of chinese implicit inter-sentence relations
Kaili Sun, Huyin Zhang, Chi Guo, Linfei Yuan, Quan Hu
J. Supercomput.1
2019 The express decay effect of time delays for globally exponentially stable nonlinear stochastic systems
Kaili Sun, Song Zhu
Peer-to-Peer Netw. Appl.1
2019 Global Anti-Synchronization of Complex-Valued Memristive Neural Networks With Time Delays
abstract
This paper formulates a class of complex-valued memristive neural networks as well as investigates the problem of anti-synchronization for complex-valued memristive neural networks. Under the concept of drive-response, several sufficient conditions for guaranteeing the anti-synchronization are given by employing suitable Lyapunov functional and some inequality techniques. The proposed results of this paper are less conservative than existing literatures due to the characteristics of memristive complex-valued neural networks. Moreover, the proposed results are easy to be validated with the parameters of system itself. Finally, two examples with numerical simulations are showed to demonstrate the efficiency of our theoretical results.
Dan Liu 0005, Song Zhu, Kaili Sun
IEEE Trans. Cybern.3
2018 New results for exponential stability of complex-valued memristive neural networks with variable delays
Dan Liu 0005, Song Zhu, Kaili Sun
Neurocomputing3
2018 Anti-synchronization of complex-valued memristor-based delayed neural networks
Dan Liu 0005, Song Zhu, Kaili Sun
Neural Networks3