VLDB 2026 Research / reviewers in the wild / expert
Wenshuo Dong
dblp:399/8847
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit discovery |
0.9 | 1 | 2025 | EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
integrated gradients · 0.9gradpath · 0.9edge attribution patching · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationabstractUnderstanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit existing gradient-based circuit identification methods and find that their performance is either affected by the zero-gradient problem or saturation effects, where edge attribution scores become insensitive to input changes, resulting in noisy and unreliable attribution evaluations for circuit components. To address the saturation effect, we propose Edge Attribution Patching with GradPath (EAP-GP), EAP-GP introduces an integration path, starting from the input and adaptively following the direction of the difference between the gradients of corrupted and clean inputs to avoid the saturated region. This approach enhances attribution reliability and improves the faithfulness of circuit identification. We evaluate EAP-GP on 6 datasets using GPT-2 Small, GPT-2 Medium, and GPT-2 XL. Experimental results demonstrate that EAP-GP outperforms existing methods in circuit faithfulness, achieving improvements up to 17.7\%. Comparisons with manually annotated ground-truth circuits demonstrate that EAP-GP achieves precision and recall comparable to or better than previous approaches, highlighting its effectiveness in identifying accurate circuits. Wenshuo Dong, Zhuoran Zhang 0003, Shu Yang 0010, Lijie Hu, Ninghao Liu 0001, Pan Zhou 0001, Di Wang 0015 |
NeurIPS | 2 |