EDBT 2026 Demo / reviewers in the wild / expert
Da Xiao 0001
dblp:87/3515-1
· DBLP profile ↗
8ranked-venue papers
5as first author
5since 2021 · last 2025
0009-0008-2732-5776ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Deep learning architectures and training · 37% Language models and text generation · 29% Trustworthy machine learning · 24% | |
| Software engineering, system software, and programming languages
2 papers |
Software testing · 70% Program synthesis and code generation · 30% | |
| Network and information security
1 paper |
Web and mobile security · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
1.6 | 2 | 2025 | MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections · ICML 2025 Improving Transformers with Dynamically Composable Multi-Head Attention · ICML 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Benchmarking and Understanding Compositional Relational Reasoning of LLMs · AAAI 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Benchmarking and Understanding Compositional Relational Reasoning of LLMs · AAAI 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | Benchmarking and Understanding Compositional Relational Reasoning of LLMs · AAAI 2025 |
Natural language and speech › Language models and text generation
fine-tuned language models |
0.8 | 1 | 2024 | Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application Firewalls · IEEE Trans. Dependable Secur. Comput. 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention |
0.8 | 1 | 2024 | Improving Transformers with Dynamically Composable Multi-Head Attention · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.8 | 1 | 2024 | Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application Firewalls · IEEE Trans. Dependable Secur. Comput. 2024 |
Web and mobile security
web application firewall |
0.8 | 1 | 2024 | Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application Firewalls · IEEE Trans. Dependable Secur. Comput. 2024 |
Software testing
fuzzing |
0.8 | 1 | 2024 | Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application Firewalls · IEEE Trans. Dependable Secur. Comput. 2024 |
Natural language and speech › Language models and text generation
language modeling |
0.3 | 1 | 2025 | MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections · ICML 2025 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.3 | 1 | 2025 | MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
reward modeling · 2.3reinforcement learning · 2.3KL-divergence penalty · 2.3multiway dense connections · 0.9intervention experiment · 0.9dynamic connection weights · 0.9attribution patching · 0.9attention head composition · 0.8neural programmer-interpreter · 0.7combinator abstraction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmarking and Understanding Compositional Relational Reasoning of LLMsabstractCompositional relational reasoning (CRR) is a hallmark of human intelligence, but we lack a clear understanding of whether and how existing transformer large language models (LLMs) can solve CRR tasks. To enable systematic exploration of the CRR capability of LLMs, we first propose a new synthetic benchmark called Generalized Associative Recall (GAR) by integrating and generalizing the essence of several tasks in mechanistic interpretability (MI) study in a unified framework. Evaluation shows that GAR is challenging enough for existing LLMs, revealing their fundamental deficiency in CRR. Meanwhile, it is easy enough for systematic MI study. Then, to understand how LLMs solve GAR tasks, we use attribution patching to discover the core circuits reused by Vicuna-33B across different tasks, and a set of vital attention heads. Intervention experiments show that the correct functioning of these heads significantly impacts task performance. Especially, we identify two classes of heads whose activations represent the abstract notion of true and false in GAR tasks respectively. They play fundamental roles in CRR across various models and tasks. Ruikang Ni, Da Xiao 0001, Qingye Meng, Shihui Zheng, Hongliang Liang |
AAAI | 2 |
| 2025 | MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense ConnectionsabstractWe propose MUltiway Dynamic Dense (MUDD) connections, a simple yet effective method to address the limitations of residual connections and enhance cross-layer information flow in Transformers. Unlike existing dense connection approaches with static and shared connection weights, MUDD generates connection weights dynamically depending on hidden states at each sequence position and for each decoupled input stream (the query, key, value or residual) of a Transformer block. MUDD connections can be seamlessly integrated into any Transformer architecture to create MUDDFormer. Extensive experiments show that MUDDFormer significantly outperforms Transformers across various model architectures and scales in language modeling, achieving performance of Transformers trained with ~1.8x--2.4x compute. Notably, MUDDPythia-2.8B matches Pythia-6.9B in pretraining ppl and downstream tasks and even rivals Pythia-12B in five-shot settings, while adding only 0.23% parameters and 0.4% computation. Code in JAX and PyTorch and pre-trained models are available at https://github.com/Caiyun-AI/MUDDFormer. Da Xiao 0001, Qingye Meng, Shengping Li, Xingyuan Yuan |
ICML | 1 |
| 2024 | Improving Transformers with Dynamically Composable Multi-Head AttentionabstractMulti-Head Attention (MHA) is a key component of Transformer. In MHA, attention heads work independently, causing problems such as low-rank bottleneck of attention score matrices and head redundancy. We propose Dynamically Composable Multi-Head Attention (DCMHA), a parameter and computation efficient attention architecture that tackles the shortcomings of MHA and increases the expressive power of the model by dynamically composing attention heads. At the core of DCMHA is a Compose function that transforms the attention score and weight matrices in an input-dependent way. DCMHA can be used as a drop-in replacement of MHA in any transformer architecture to obtain the corresponding DCFormer. DCFormer significantly outperforms Transformer on different architectures and model scales in language modeling, matching the performance of models with 1.7x-2.0x compute. For example, DCPythia-6.9B outperforms open source Pythia-12B on both pretraining perplexity and downstream task evaluation. Da Xiao 0001, Qingye Meng, Shengping Li, Xingyuan Yuan |
ICML | 1 |
| 2024 | Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application FirewallsabstractWeb Application Firewalls (WAFs) are widely deployed to protect key web applications against multiple security threats, so it is important to test WAFs regularly to prevent attackers from bypassing them easily. Machine-learning-based black-box WAF testing is gaining more attention, though existing learning-based approaches have strict requirements on the source and scale of payload data and suffer from the local optimum problem, limiting their effectiveness and practical application. We propose GPTFuzzer, apracticalandeffectivegeneration-based approach to test WAFs by generating attack payloads token-by-token. Specifically, we fine-tune a Generative Pre-trained Transformer language model with reinforcement learning to make GPTFuzzer have the least restrictions on payload data and thus more applicable in practice, and we use reward modeling and KL-divergence penalty to improve the effectiveness of our approach and mitigate the local optimum issue. We implement GPTFuzzer and evaluate it on two well-known open-source WAFs against three kinds of common attacks. Experimental results show that GPTFuzzer significantly outperforms state-of-the-art approaches,i.e.ML-Driven and RAT, finding up to 7.8× (3.2× on average) more bypassing payloads within 1,250,000 requests, or finding out all bypassing payloads using up to 8.1× (3.3× on average) fewer requests. Hongliang Liang, Da Xiao 0001, Yanjie Zhou, Aibo Wang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Path context augmented statement and network for learning programs
Da Xiao 0001, Dengji Hang, Lu Ai, Shengping Li, Hongliang Liang |
Empir. Softw. Eng. | 1 |
| 2018 | Improving the Universality and Learnability of Neural Programmer-Interpreters with Combinator Abstraction
Da Xiao 0001, Jo-Yu Liao, Xingyuan Yuan |
ICLR (Poster) | 1 |
| 2017 | An end-to-end model for Android malware detectionabstractMalware detection has been a difficult problem for a very long time. Since the wide use of smart devices in recent years, the number of malwares is increasing rapidly. Most existing methods for malware detection rely too much on manual interventions (e.g. pre-defined features and patterns), which can be easily deceived. In this paper, we propose a novel end-to-end deep learning model to detect Android malwares. Our model takes the raw system call sequence, which is generated during the application's runtime, as input and decides whether the sequence is malicious without any manual intervention. We evaluate the model on 14231 Android applications and obtain a detection accuracy of 93.16%, which is 2.81% higher than the contrast experiment in which we implement the method proposed by other researchers. Hongliang Liang, Da Xiao 0001 |
ISI | 3 |
| 2012 | Multiple-File Remote Data Checking for cloud storage
Da Xiao 0001, Wenbin Yao, Chunhua Wu, Yixian Yang |
Comput. Secur. | 1 |