VLDB 2026 Research / reviewers in the wild / expert
Dong Shu
dblp:361/2245
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 32% Language models and text generation · 32% Information extraction and text analysis · 16% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding |
1.0 | 1 | 2026 | FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › document analysis
financial text analysis |
1.0 | 1 | 2026 | FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction · ACL (1) 2026 |
Natural language and speech › Language models and text generation › in-context learning
demonstration selection |
0.9 | 1 | 2025 | Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract) · AAAI 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract) · AAAI 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.9 | 1 | 2025 | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
large vision-language model · 1.0large language model · 1.0gradient-based attribution · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise PredictionabstractPredicting corporate earnings surprises is a profitable yet challenging task, as accurate forecasts can inform significant investment decisions.However, progress in this domain has been constrained by a reliance on expensive, proprietary, and text-only data, limiting the development of advanced models.To address this gap, we introduce FinCall-Surprise (Financial Conference Call for Earning Surprise Prediction), the first large-scale, open-source, and multi-modal dataset for earnings surprise prediction.Comprising 2,688 unique corporate conference calls from 2019 to 2021, our dataset features word-to-word conference call textual transcripts, full audio recordings, and corresponding presentation slides.We establish a comprehensive benchmark by evaluating 26 state-of-the-art unimodal and multimodal LLMs.Our findings reveal that (1) while many models achieve high accuracy, this performance is often an illusion caused by significant class imbalance in the realworld data.(2) Some specialized financial models demonstrate unexpected weaknesses in instruction-following and language generation.(3) Although incorporating audio and visual modalities provides some performance gains, current models still struggle to leverage these signals effectively.These results highlight critical limitations in the financial reasoning capabilities of existing LLMs and establish a challenging new baseline for future research.The FinCall-Surprise dataset is available at https://github.com/Tizzzzy/ FinCall-Surprise. Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du |
ACL (1) | 1 |
| 2026 | FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language ModelsabstractLarge vision-language models (LVLMs) have made significant progress in chart understanding.However, financial charts, characterized by complex temporal structures and domainspecific terminology, remain notably underexplored.We introduce FinChart-Bench, the first benchmark specifically focused on realworld financial charts.FinChart-Bench comprises 1,200 financial chart images collected from 2015 to 2024, each annotated with True/-False (TF), Multiple Choice (MC), and Question Answering (QA) questions, totaling 7,016 questions.We conducted a comprehensive evaluation of 26 state-of-the-art LVLMs on FinChart-Bench.Our evaluation reveals critical insights: (1) the performance gap between open-source and closed-source models is narrowing, (2) performance degradation occurs in upgraded models within families, (3) many models struggle with instruction following, (4) both advanced models show significant limitations in spatial reasoning abilities, and (5) current LVLMs are not reliable enough to serve as automated evaluators.These findings highlight important limitations in current LVLM capabilities for financial chart understanding.The FinChart-Bench dataset is available at https: //github.com/Tizzzzy/FinChart-Bench. Dong Shu, Haoyang Yuan, Yanguang Liu, Huopu Zhang, Mengnan Du |
ACL (1) | 1 |
| 2025 | Comparative Analysis of Demonstration Selection Algorithms for In-Context Learning in Large Language Models (Student Abstract)abstractDemonstration selection algorithms play a crucial role in optimizing Large Language Models' (LLMs) in-context learning performance. Despite numerous proposed algorithms, their comparative effectiveness remains understudied. We present a comprehensive evaluation of six state-of-the-art demonstration selection algorithms across five datasets, examining both their effectiveness and computational efficiency. Our findings reveal significant trade-offs: while some demonstration selection algorithms achieve superior accuracy, they incur substantial computational costs. We also discover that increasing demonstration examples doesn't consistently improve performance, and some sophisticated algorithms struggle to outperform random selection in certain scenarios. These insights provide valuable benchmarks for future algorithm development and practical implementation. Our code is available at https://github.com/Tizzzzy/Demonstration_Selection_Overview. Dong Shu, Mengnan Du |
AAAI | 1 |
| 2025 | Beyond Input Activations: Identifying Influential Latents by Gradient Sparse AutoencodersabstractSparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs).However, conventional approaches to analyzing SAEs typically rely solely on input-side activations, without considering the causal influence between each latent feature and the model's output.This work is built on two key hypotheses: (1) activated latents do not contribute equally to the construction of the model's output, and (2) only latents with high causal influence are effective for model steering.To validate these hypotheses, we propose Gradient Sparse Autoencoder (GradSAE), a simple yet effective method that identifies the most influential latents by incorporating output-side gradient information.Our code is available at https: //github.com/Tizzzzy/sae_gradient. Dong Shu, Xuansheng Wu, Haiyan Zhao 0003, Mengnan Du, Ninghao Liu 0001 |
EMNLP | 1 |
| 2024 | Knowledge Graph Large Language Model (KG-LLM) for Link Prediction
Dong Shu, Mingyu Jin, Chong Zhang 0006, Mengnan Du, Yongfeng Zhang 0003 |
ACML | 1 |
| 2024 | LawLLM: Law Large Language Model for the US Legal SystemabstractIn the rapidly evolving field of legal analytics, finding relevant cases and accurately predicting judicial outcomes are challenging because of the complexity of legal language, which often includes specialized terminology, complex syntax, and historical context. Moreover, the subtle distinctions between similar and precedent cases require a deep understanding of legal knowledge. Researchers often conflate these concepts, making it difficult to develop specialized techniques to effectively address these nuanced tasks. In this paper, we introduce the Law Large Language Model (LawLLM), a multi-task model specifically designed for the US legal domain to address these challenges. LawLLM excels at Similar Case Retrieval (SCR), Precedent Case Recommendation (PCR), and Legal Judgment Prediction (LJP). By clearly distinguishing between precedent and similar cases, we provide essential clarity, guiding future research in developing specialized strategies for these tasks. We propose customized data preprocessing techniques for each task that transform raw legal data into a trainable format. Furthermore, we also use techniques such as in-context learning (ICL) and advanced information retrieval methods in LawLLM. The evaluation results demonstrate that LawLLM consistently outperforms existing baselines in both zero-shot and few-shot scenarios, offering unparalleled multi-task capabilities and filling critical gaps in the legal domain. Code and data are available at https://github.com/Tizzzzy/Law_LLM. Dong Shu, Xukun Liu, David Demeter, Mengnan Du, Yongfeng Zhang 0003 |
CIKM | 1 |
| 2024 | Target-driven Attack for Large Language ModelsabstractCurrent large language models (LLM) provide a strong foundation for large-scale user-oriented natural language tasks. Many users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges like the language model not giving the correct answer. Although there is currently a large amount of research on black-box attacks, most of these black-box attacks use random and heuristic strategies. It is unclear how these strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we propose our target-driven black-box attack method to maximize the KL divergence between the conditional probabilities of the clean text and the attack text to redefine the attack’s goal. We transform the distance maximization problem into two convex optimization problems based on the attack goal to solve the attack text and estimate the covariance. Furthermore, the projected gradient descent algorithm solves the vector corresponding to the attack text. Our target-driven black-box attack approach includes two attack strategies: token manipulation and misinformation attack. Experimental results on multiple Large Language Models and datasets demonstrate the effectiveness of our attack method. Chong Zhang 0006, Mingyu Jin, Dong Shu, Taowen Wang, Dongfang Liu, Xiao-Bo Jin |
ECAI | 3 |
| 2024 | ARIF: An Adaptive Attention-Based Cross-Modal Representation Integration Framework
Zihong Luo, Yifei Bi, Zile Huang, Dong Shu, Jiheng Hou, Hongchen Wang, Kaiyu Liang |
ICANN (6) | 5 |