Dong Shu

dblp:361/2245 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 32% Language models and text generation · 32% Information extraction and text analysis · 16%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding
1.012026
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › document analysis
financial text analysis
1.012026
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction · ACL (1) 2026
Natural language and speech › Language models and text generation › in-context learning
demonstration selection
0.912025
Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract) · AAAI 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract) · AAAI 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.912025
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312025
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

large vision-language model · 1.0large language model · 1.0gradient-based attribution · 0.9
YearPublicationVenuePosition
2026 FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
abstract
Predicting corporate earnings surprises is a profitable yet challenging task, as accurate forecasts can inform significant investment decisions.However, progress in this domain has been constrained by a reliance on expensive, proprietary, and text-only data, limiting the development of advanced models.To address this gap, we introduce FinCall-Surprise (Financial Conference Call for Earning Surprise Prediction), the first large-scale, open-source, and multi-modal dataset for earnings surprise prediction.Comprising 2,688 unique corporate conference calls from 2019 to 2021, our dataset features word-to-word conference call textual transcripts, full audio recordings, and corresponding presentation slides.We establish a comprehensive benchmark by evaluating 26 state-of-the-art unimodal and multimodal LLMs.Our findings reveal that (1) while many models achieve high accuracy, this performance is often an illusion caused by significant class imbalance in the realworld data.(2) Some specialized financial models demonstrate unexpected weaknesses in instruction-following and language generation.(3) Although incorporating audio and visual modalities provides some performance gains, current models still struggle to leverage these signals effectively.These results highlight critical limitations in the financial reasoning capabilities of existing LLMs and establish a challenging new baseline for future research.The FinCall-Surprise dataset is available at https://github.com/Tizzzzy/ FinCall-Surprise.
Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du
ACL (1)1
2026 FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
abstract
Large vision-language models (LVLMs) have made significant progress in chart understanding.However, financial charts, characterized by complex temporal structures and domainspecific terminology, remain notably underexplored.We introduce FinChart-Bench, the first benchmark specifically focused on realworld financial charts.FinChart-Bench comprises 1,200 financial chart images collected from 2015 to 2024, each annotated with True/-False (TF), Multiple Choice (MC), and Question Answering (QA) questions, totaling 7,016 questions.We conducted a comprehensive evaluation of 26 state-of-the-art LVLMs on FinChart-Bench.Our evaluation reveals critical insights: (1) the performance gap between open-source and closed-source models is narrowing, (2) performance degradation occurs in upgraded models within families, (3) many models struggle with instruction following, (4) both advanced models show significant limitations in spatial reasoning abilities, and (5) current LVLMs are not reliable enough to serve as automated evaluators.These findings highlight important limitations in current LVLM capabilities for financial chart understanding.The FinChart-Bench dataset is available at https: //github.com/Tizzzzy/FinChart-Bench.
Dong Shu, Haoyang Yuan, Yanguang Liu, Huopu Zhang, Mengnan Du
ACL (1)1
2025 Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract)
abstract
Demonstration selection algorithms play a crucial role in optimizing Large Language Models' (LLMs) in-context learning performance. Despite numerous proposed algorithms, their comparative effectiveness remains understudied. We present a comprehensive evaluation of six state-of-the-art demonstration selection algorithms across five datasets, examining both their effectiveness and computational efficiency. Our findings reveal significant trade-offs: while some demonstration selection algorithms achieve superior accuracy, they incur substantial computational costs. We also discover that increasing demonstration examples doesn't consistently improve performance, and some sophisticated algorithms struggle to outperform random selection in certain scenarios. These insights provide valuable benchmarks for future algorithm development and practical implementation. Our code is available at https://github.com/Tizzzzy/Demonstration_Selection_Overview.
Dong Shu, Mengnan Du
AAAI1
2025 Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
abstract
Sparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs).However, conventional approaches to analyzing SAEs typically rely solely on input-side activations, without considering the causal influence between each latent feature and the model's output.This work is built on two key hypotheses: (1) activated latents do not contribute equally to the construction of the model's output, and (2) only latents with high causal influence are effective for model steering.To validate these hypotheses, we propose Gradient Sparse Autoencoder (GradSAE), a simple yet effective method that identifies the most influential latents by incorporating output-side gradient information.Our code is available at https: //github.com/Tizzzzy/sae_gradient.
Dong Shu, Xuansheng Wu, Haiyan Zhao 0003, Mengnan Du, Ninghao Liu 0001
EMNLP1
2024 Knowledge Graph Large Language Model (KG-LLM) for Link Prediction
Dong Shu, Mingyu Jin, Chong Zhang 0006, Mengnan Du, Yongfeng Zhang 0003
ACML1
2024 LawLLM: Law Large Language Model for the US Legal System
abstract
In the rapidly evolving field of legal analytics, finding relevant cases and accurately predicting judicial outcomes are challenging because of the complexity of legal language, which often includes specialized terminology, complex syntax, and historical context. Moreover, the subtle distinctions between similar and precedent cases require a deep understanding of legal knowledge. Researchers often conflate these concepts, making it difficult to develop specialized techniques to effectively address these nuanced tasks. In this paper, we introduce the Law Large Language Model (LawLLM), a multi-task model specifically designed for the US legal domain to address these challenges. LawLLM excels at Similar Case Retrieval (SCR), Precedent Case Recommendation (PCR), and Legal Judgment Prediction (LJP). By clearly distinguishing between precedent and similar cases, we provide essential clarity, guiding future research in developing specialized strategies for these tasks. We propose customized data preprocessing techniques for each task that transform raw legal data into a trainable format. Furthermore, we also use techniques such as in-context learning (ICL) and advanced information retrieval methods in LawLLM. The evaluation results demonstrate that LawLLM consistently outperforms existing baselines in both zero-shot and few-shot scenarios, offering unparalleled multi-task capabilities and filling critical gaps in the legal domain. Code and data are available at https://github.com/Tizzzzy/Law_LLM.
Dong Shu, Xukun Liu, David Demeter, Mengnan Du, Yongfeng Zhang 0003
CIKM1
2024 Target-driven Attack for Large Language Models
abstract
Current large language models (LLM) provide a strong foundation for large-scale user-oriented natural language tasks. Many users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges like the language model not giving the correct answer. Although there is currently a large amount of research on black-box attacks, most of these black-box attacks use random and heuristic strategies. It is unclear how these strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we propose our target-driven black-box attack method to maximize the KL divergence between the conditional probabilities of the clean text and the attack text to redefine the attack’s goal. We transform the distance maximization problem into two convex optimization problems based on the attack goal to solve the attack text and estimate the covariance. Furthermore, the projected gradient descent algorithm solves the vector corresponding to the attack text. Our target-driven black-box attack approach includes two attack strategies: token manipulation and misinformation attack. Experimental results on multiple Large Language Models and datasets demonstrate the effectiveness of our attack method.
Chong Zhang 0006, Mingyu Jin, Dong Shu, Taowen Wang, Dongfang Liu, Xiao-Bo Jin
ECAI3
2024 ARIF: An Adaptive Attention-Based Cross-Modal Representation Integration Framework
Zihong Luo, Yifei Bi, Zile Huang, Dong Shu, Jiheng Hou, Hongchen Wang, Kaiyu Liang
ICANN (6)5