Lida Shi

dblp:262/7358 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-5011-6931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
abstract
Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical texts. First, the scarcity of historical language samples renders unsupervised learning approaches based on large text corpora highly inefficient, hindering effective pre-training. Moreover, due to the considerable temporal gap and complex evolution of ancient scripts, the absence of comprehensive character encoding schemes limits the digitization and computational processing of ancient texts, particularly in early Chinese writing. To address these challenges, we introduce InteChar, a unified and extensible character list that integrates unencoded oracle bone characters with traditional and modern Chinese. InteChar enables consistent digitization and representation of historical texts, providing a foundation for robust modeling of ancient scripts. To evaluate the effectiveness of InteChar, we construct the Oracle Corpus Set (OracleCS), an ancient Chinese corpus that combines expert-annotated samples with LLM-assisted data augmentation, centered on Chinese oracle bone inscriptions. Extensive experiments show that models trained with InteChar on OracleCS achieve substantial improvements across various historical language understanding tasks, confirming the effectiveness of our approach and establishing a solid foundation for future research in ancient Chinese NLP.
Xiaolei Diao, Zhihan Zhou 0003, Lida Shi, Ting Wang 0019, Ruihua Qi, Daqian Shi, Hao Xu 0012
AAAI3
2026 AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora
abstract
Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient characters. Existing Chinese benchmarks are mostly targeted at modern Chinese and transmitted documents in ancient Chinese, but the part of excavated documents in ancient Chinese is not covered. To meet this need, we propose the AncientBench, which aims to evaluate the comprehension of ancient characters, especially in the scenario of excavated documents. The AncientBench is divided into four dimensions, which correspond to the four competencies of ancient character comprehension: glyph comprehension, pronunciation comprehension, meaning comprehension, and contextual comprehension. The benchmark also contains ten tasks, including radical, phonetic radical, homophone, cloze, translation, and more, providing a comprehensive framework for evaluation. We convened archaeological researchers to conduct experimental evaluations, proposed an ancient model as baseline, and conducted extensive experiments on the currently best-performing large language models. The experimental results reveal the great potential of large language models in ancient textual scenarios as well as the gap with humans. Our research aims to promote the development and application of large language models in the field of archaeology and ancient Chinese language.
Zhihan Zhou 0003, Daqian Shi, Rui Song 0008, Lida Shi, Xiaolei Diao, Hao Xu 0012
AAAI4
2026 Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
abstract
In recent years, rapid advances in Multimodal Large Language Models (MLLMs) have increasingly stimulated research on ancient Chinese scripts.As the evolution of written characters constitutes a fundamental pathway for understanding cultural transformation and historical continuity, how MLLMs can be systematically leveraged to support and advance text evolution analysis remains an open and largely underexplored problem.To bridge this gap, we construct a comprehensive benchmark comprising 11 tasks and over 130,000 instances, specifically designed to evaluate the capability of MLLMs in analyzing the evolution of ancient Chinese scripts.We conduct extensive evaluations across multiple widely used MLLMs and observe that, while existing models demonstrate a limited ability in glyph-level comparison, their performance on core tasks-such as character recognition and evolutionary reasoning-remains substantially constrained.Motivated by these findings, we propose a glyph-driven fine-tuning framework (GEVO) that explicitly encourages models to capture evolutionary consistency in glyph transformations and enhances their understanding of text evolution.Experimental results show that even models at the 2B scale achieve consistent and comprehensive performance improvements across all evaluated tasks.To facilitate future research, we publicly release both the benchmark and the trained models 1 .
Rui Song 0008, Lida Shi, Ruihua Qi, Yingji Li, Hao Xu 0012
ACL (1)2
2026 Learn from the best: A universal self-distillation approach with historical logits
Lida Shi, Fausto Giunchiglia, Hongda Zhang, Daqian Shi, Rui Song 0008, Jian Li 0080, Xiaolei Diao, Alan Zhao, Hao Xu 0012
Expert Syst. Appl.1
2026 An empirical study of LLMs via in-context learning for stance classification
Lida Shi, Fausto Giunchiglia, Ran Luo 0005, Daqian Shi, Rui Song 0008, Xiaolei Diao, Hao Xu 0012
Inf. Process. Manag.1
2026 From text mining to intelligent debate: Task frameworks and technological evolution in computational argumentation
Lida Shi, Fausto Giunchiglia, Yongqi Cheng, Rui Song 0008, Daqian Shi, Xiaolei Diao, Hao Xu 0012
Inf. Process. Manag.1
2025 A Dual-Mind Framework for Strategic and Expressive Negotiation Agent
abstract
Negotiation agents need to influence the attitudes or intentions of users to reach a consensus. Strategy planning and expressive optimization are crucial aspects of effective negotiations. However, previous studies have typically focused on only one of these aspects, neglecting the fact that their combined synergistic effect can lead to better performance. Inspired by the dual-process theory in human cognition, we propose a Dual-Mind Negotiation Agent (DMNA) framework. This framework integrates an intuitive module for rapid, experience-based response and a deliberative module for slow, expression optimization. The intuitive module is trained using Monte Carlo Tree Search (MCTS) and Direct Preference Optimization (DPO), enabling it to make suitable strategic planning and expression. The deliberative module employs a multifaceted reflexion mechanism to enhance the quality of expression. Experiments conducted on negotiation datasets confirm that DMNA achieves state-of-the-art results, demonstrating an enhancement in the negotiation ability of agents.
Lida Shi, Rui Song 0008, Hao Xu 0012
ACL (1)2
2025 DGJA: Dependency Graph-enhanced Joint Attention Structure for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection (MSD) leverages multimodal data, including both images and text, to detect whether the input content contains sarcastic information. Despite recent advances, existing MSD approaches often overlook the imbalance in sarcastic content between text and image modalities, where text typically carries more sarcastic cues conveyed through complex semantic relationships. To address this issue, we propose a dependency graph-enhanced joint attention network that integrates the flat representations extracted from the joint attention mechanisms and the graph-based representations learned from the dependency graph. Specifically, we design a joint attention module to capture the rich sarcastic cues and semantic relationships within the text, and use the graph-based representations to enhance the flat representations. Evaluated on the HFM dataset, our method achieved 0.54% improvement in F1 score and 0.81% in accuracy, demonstrating its effectiveness.
Rui Song 0008, Lida Shi, Hao Xu 0012
ICASSP3
2025 Counterfactual contrastive learning for robust text classification based on word group search
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Lida Shi, Hao Xu 0012
Inf. Sci.4
2024 The SES framework and Frequency domain information fusion strategy for Human activity recognition
abstract
Human activity recognition (HAR) is a task designed to identify and classify physical activities or behaviors in people’s daily lives. This field relies on data collected from various sensors. Due to differences in environment, user posture and habits, as well as equipment placement, the data exhibits severe heterogeneity. The channel information from different sensors exhibits more complex and uncertain characteristics in terms of shape, noise level, and data quality. These properties can lead to poor generalization of deep learning models, posing challenges for the effectiveness of deep learning algorithms and the widespread use of specific embedded devices in this field. Therefore, this article proposes an adaptive channel signal scaling method to calibrate channel characteristics from the perspective of sensor channel information. Additionally, we propose a stable feature completion strategy to enrich the data feature information using different fusion strategies at the data level. Extensive experiments were conducted on three publicly available HAR datasets, and the experimental results demonstrate that our proposed method significantly improves the performance of state-of-the-art deep learning methods.
Haotian Feng, Lida Shi, Hongda Zhang, Hao Xu 0012
IJCNN3
2024 ATFA: Adversarial Time-Frequency Attention network for sensor-based multimodal human activity recognition
Haotian Feng, Qiang Shen 0005, Rui Song 0008, Lida Shi, Hao Xu 0012
Expert Syst. Appl.4
2024 Learning neighbor-enhanced region representations and question-guided visual representations for visual question answering
Hongda Zhang, Nan Sheng, Lida Shi, Hao Xu 0012
Expert Syst. Appl.4
2023 Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level Annotations
abstract
Optical character recognition (OCR) methods have been applied to diverse tasks, e.g., street view text recognition and document analysis. Recently, zero-shot OCR has piqued the interest of the research community because it considers a practical OCR scenario with unbalanced data distribution. However, there is a lack of benchmarks for evaluating such zero-shot methods that apply a divide-and-conquer recognition strategy by decomposing characters into radicals. Meanwhile, radical recognition, as another important OCR task, also lacks radical-level annotation for model training. In this paper, we construct an ancient Chinese character image dataset that contains both radical-level and character-level annotations to satisfy the requirements of the above-mentioned methods, namely, ACCID, where radical-level annotations include radical categories, radical locations, and structural relations. To increase the adaptability of ACCID, we propose a splicing-based synthetic character algorithm to augment the training samples and apply an image denoising method to improve the image quality. By introducing character decomposition and recombination, we propose a baseline method for zero-shot OCR. The experimental results demonstrate the validity of ACCID and the baseline model quantitatively and qualitatively.
Xiaolei Diao, Daqian Shi, Jian Li 0080, Lida Shi, Mingzhe Yue, Ruihua Qi, Hao Xu 0012
ACM Multimedia4
2023 SUNET: Speaker-utterance interaction Graph Neural Network for Emotion Recognition in Conversations
Rui Song 0008, Fausto Giunchiglia, Lida Shi, Qiang Shen 0005, Hao Xu 0012
Eng. Appl. Artif. Intell.3
2023 Fine-grained classification of intracranial haemorrhage subtypes in head CT scans
abstract
Abstract Intracranial haemorrhage (ICH) is a haemorrhagic disease that occurs in the ventricle or brain tissue and has a high probability of mortality and disability. For ICH, it is important to obtain a correct diagnosis in the early stages. Currently, ICH classification mainly depends on professional radiologists for manual diagnosis. Therefore, it is necessary to develop a method that can efficiently and rapidly diagnose ICH. In the field of ICH subtype classification, most studies directly use the existing convolutional neural network (CNN) to extract CT slice features. However, these existing networks have the following shortcomings: (1) insufficient discrimination of CT slice features leads to an inability to achieve satisfactory classification performance. (2) Most CT slice data sets of ICH have the serious problem of sample imbalance. (3) There is a correlation between subtypes; however, in previous studies, this correlation has been ignored. To solve these problems, the authors propose a classification algorithm for ICH subtypes applied to CT images. The CNN–RNN architecture was adopted to classify ICH subtypes. In the CNN module, the problem is viewed from a fine‐grained perspective, which solves the problem of insufficient feature discrimination in existing methods. A new loss function is also proposed to solve the problems of unbalanced data distribution and neglected dependencies among the labels. These parts are integrated into the proposed fine‐grained network architecture. The image embeddings were obtained by the CNN module and then input to the RNN module. The authors’ method was evaluated on the Radiological Society of North America 2019 Brain CT Haemorrhage (RSNA‐2019) benchmark. The experimental results demonstrated that the performance of the proposed method is state‐of‐the‐art.
Pingping Liu, Gangjun Ning, Lida Shi, Qiuzhan Zhou
IET Comput. Vis.3
2023 Measuring and mitigating language model biases in abusive language detection
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Lida Shi, Hao Xu 0012
Inf. Process. Manag.4
2022 A Simple Contrastive Learning Framework for Interactive Argument Pair Identification via Argument-Context Extraction
abstract
Interactive argument pair identification is an emerging research task for argument mining, aiming to identify whether two arguments are interactively related.It is pointed out that the context of the argument is essential to improve identification performance.However, current context-based methods achieve limited improvements since the entire context typically contains much irrelevant information.In this paper, we propose a simple contrastive learning framework to solve this problem by extracting valuable information from the context.This framework can construct hard argumentcontext samples and obtain a robust and uniform representation by introducing contrastive learning.We also propose an argument-context extraction module to enhance information extraction by discarding irrelevant blocks.The experimental results show that our method achieves the state-of-the-art performance on the benchmark dataset.Further analysis demonstrates the effectiveness of our proposed modules and visually displays more compact semantic representations.The code is available at GitHub 1 .
Lida Shi, Fausto Giunchiglia, Rui Song 0008, Daqian Shi, Xiaolei Diao, Hao Xu 0012
EMNLP1
2022 CharFormer: A Glyph Fusion based Attentive Framework for High-precision Character Image Denoising
abstract
Degraded images commonly exist in the general sources of character images, leading to unsatisfactory character recognition results. Existing methods have dedicated efforts to restoring degraded character images. However, the denoising results obtained by these methods do not appear to improve character recognition performance. This is mainly because current methods only focus on pixel-level information and ignore critical features of a character, such as its glyph, resulting in character-glyph damage during the denoising process. In this paper, we introduce a novel generic framework based on glyph fusion and attention mechanisms, i.e., CharFormer, for precisely recovering character images without changing their inherent glyphs. Unlike existing frameworks, CharFormer introduces a parallel target task for capturing additional information and injecting it into the image denoising backbone, which will maintain the consistency of character glyphs during character image denoising. Moreover, we utilize attention-based networks for global-local feature interaction, which will help to deal with blind denoising and enhance denoising performance. We compare CharFormer with state-of-the-art methods on multiple datasets. The experimental results show the superiority of CharFormer quantitatively and qualitatively.
Daqian Shi, Xiaolei Diao, Lida Shi, Hao Tang 0005, Yang Chi, Hao Xu 0012
ACM Multimedia3