Nai Ding

dblp:128/4756 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-3428-2723ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 69% Question answering and dialogue systems · 13% Trustworthy machine learning · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › pre-trained language model
decoder-only language model
1.012026
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs · ACL (1) 2026
Machine learning › Trustworthy machine learning
language model interpretability
0.912025
Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain · NeurIPS 2025
Natural language and speech › Language models and text generation › text representation
syntactic representation
0.912025
Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › lexical semantics
adjective semantics
0.712023
Adjective Scale Probe: Can Language Models Encode Formal Semantics Information? · AAAI 2023
Natural language and speech › Language models and text generation
language model analysis
0.712023
Adjective Scale Probe: Can Language Models Encode Formal Semantics Information? · AAAI 2023
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference
0.712023
Adjective Scale Probe: Can Language Models Encode Formal Semantics Information? · AAAI 2023
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.612022
On Tracking Dialogue State by Inheriting Slot Values in Mentioned Slot Pools · IJCAI 2022
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.612022
On Tracking Dialogue State by Inheriting Slot Values in Mentioned Slot Pools · IJCAI 2022
Compilers and program optimization › parsing
constituency parsing
0.312026
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs · ACL (1) 2026
Compilers and program optimization
parsing
0.312026
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

token update mask · 2.0staged training · 2.0gated tree cross-attention · 2.0representational similarity analysis · 0.9jensen-shannon divergence · 0.9intracranial recordings · 0.9frequency-domain analysis · 0.9fine-tuning · 0.7diagnostic dataset · 0.7memory network · 0.6
YearPublicationVenuePosition
2026 Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
abstract
Decoder-only large language models achieve strong broad performance but are brittle to minor grammatical perturbations, undermining reliability for downstream reasoning.However, directly injecting explicit syntactic structure into an existing checkpoint can interfere with its pretrained competence.We introduce a checkpoint-compatible gated tree crossattention (GTCA) branch that reads precomputed constituency chunk memory while leaving backbone architecture unchanged.Our design uses a token update mask and staged training to control the scope and timing of structural updates.Across benchmarks and Transformer backbones, GTCA strengthens syntactic robustness beyond continued training baselines without compromising Multiple-Choice QA performance or commonsense reasoning, providing a practical checkpoint-compatible route to more syntax-robust decoder-only LLMs.Our code is available at https://github.com/ Pineandgrass/GatedTreeCrossAttention.
Shaonan Wang, Nai Ding
ACL (1)3
2025 Information Integration in Large Language Models is Gated by Linguistic Structural Markers
abstract
Language comprehension relies on integrating information across both local words and broader context.We propose a method to quantify the information integration window of large language models (LLMs) and examine how sentence and clause boundaries constrain this window.Specifically, LLMs are required to predict a target word based on either a local window (local prediction) or the full context (global prediction), and we use Jensen-Shannon (JS) divergence to measure the information loss from relying solely on the local window, termed the local-prediction deficit.Results show that integration windows of both humans and LLMs are strongly modulated by sentence boundaries, and predictions primarily rely on words within the same sentence or clause: The localprediction deficit follows a power-law decay as the window length increases and drops sharply at the sentence boundary.This boundary effect is primarily attributed to linguistic structural markers, e.g., punctuation, rather than implicit syntactic or semantic cues.Together, these results indicate that LLMs rely on explicit structural cues to guide their information integration strategy.
Nai Ding
EMNLP2
2025 Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain
abstract
Large Language Models (LLMs) demonstrate human-level or even superior language abilities, effectively modeling syntactic structures, yet the specific computational units responsible remain unclear. A key question is whether LLM behavioral capabilities stem from mechanisms akin to those in the human brain. To address these questions, we introduce the Hierarchical Frequency Tagging Probe (HFTP), a tool that utilizes frequency-domain analysis to identify neuron-wise components of LLMs (e.g., individual Multilayer Perceptron (MLP) neurons) and cortical regions (via intracranial recordings) encoding syntactic structures. Our results show that models such as GPT-2, Gemma, Gemma 2, Llama 2, Llama 3.1, and GLM-4 process syntax in analogous layers, while the human brain relies on distinct cortical regions for different syntactic levels. Representational similarity analysis reveals a stronger alignment between LLM representations and the left hemisphere of the brain (dominant in language processing). Notably, upgraded models exhibit divergent trends: Gemma 2 shows greater brain similarity than Gemma, while Llama 3.1 shows less alignment with the brain compared to Llama 2. These findings offer new insights into the interpretability of LLM behavioral improvements, raising questions about whether these advancements are driven by human-like or non-human-like mechanisms, and establish HFTP as a valuable tool bridging computational linguistics and cognitive neuroscience. This project is available at https://github.com/LilTiger/HFTP.
Jingmin An, Yilong Song, Ruolin Yang 0006, Nai Ding, Lingxi Lu, Chu Zhuang
NeurIPS4
2024 EMDSQA: A Neural Speech Quality Assessment Model With Speaker Embedding
abstract
We present a neural speech quality assessment model with speaker embedding. This model, i.e., EMDSQA, can precisely predict the Mean Opinion Score (MOS) of speech quality during online communications. Intrusive speech quality assessment methods such as perceptual objective listening quality analysis (POLQA) are not practical for online communications because every piece of degraded speech requires a corresponding clean reference. Non-intrusive methods can assess the quality of online speech, but have not reached the accuracy and robustness required for real-world applications. EMDSQA extracts the speaker embedding using an independent pipeline and feeds it as a prior feature to a self-attention-based MOS prediction model. Since EMDSQA does not need the corresponding clean reference, it is practical for real-world communication applications. An open-source test corpus, featuring real-world data, was also developed. Experimental results show that EMDSQA achieves a 0.92 Pearson correlation coefficient with the MOS measured from humans, surpassing other state-of-the-art intrusive or non-intrusive methods.
Yiya Hao, Feifei Xiong, Nai Ding, Jinwei Feng
IEEE Signal Process. Lett.4
2023 Adjective Scale Probe: Can Language Models Encode Formal Semantics Information?
abstract
It is an open question what semantic representations transformer-based language models can encode and whether they have access to more abstract aspects of semantic meaning. Here, we propose a diagnostic dataset to investigate how well language models understand the degree semantics of adjectives. In the dataset, referred as the Adjective Scale Probe (ASP), we semi-automatically generate 8 tests of Natural Language Inference (NLI) questions to test 8 key capabilities of adjective interpretation. We apply the ASP dataset to evaluate the performance of 3 language models, i.e., BERT, DeBERTa, and T0. It is found that language models perform below the majority baseline for most tests of the ASP, even when the models have been fine-tuned to achieve high performance on the large-scale MNLI dataset. But after we fine-tune the pre-trained models on a subset of the ASP, DeBERTa can achieve high performance on the untrained adjectives and untrained tests, suggesting that DeBERTa may have captured degree semantic information of adjectives through pre-training but it needs specific training data to learn how to apply such information to the current tasks. In sum, the ASP provides an easy-to-use method to test fine-grained formal semantic properties of adjectives, and reveals language models' abilities to access formal semantic information.
Wei Liu 0184, Ming Xiang, Nai Ding
AAAI3
2022 On Tracking Dialogue State by Inheriting Slot Values in Mentioned Slot Pools
abstract
Dialogue state tracking (DST) is a component of the task oriented dialogue system. It is responsible for extracting and managing slots, where each slot represents a part of the information to accomplish a task, and slot value is updated recurrently in each dialogue turn. However, many DST models cannot update slot values appropriately. These models may repeatedly inherit wrong slot values extracted in previous turns, resulting in the fail of the entire DST task. They cannot update indirectly mentioned slots well, either. This study designed a model with a mentioned slot pool (MSP) to tackle the update problem. The MSP is a slot specific memory that records all mentioned slot values that may be inherited, and our model updates slot values according to the MSP and the dialogue context. Our model rejects inheriting the previous slot value when it predicates the value is wrong. Then, it extracts the slot value from the current dialogue context. As the contextual information accumulates, the new value is more likely to be correct. It also can track the indirectly mentioned slot by picking a value from the MSP. Experimental results showed our model reached state of the art DST performance on MultiWOZ datasets.
Zhoujian Sun, Zhengxing Huang, Nai Ding
IJCAI3
2013 Validation of acoustic models of auditory neural prostheses
abstract
Acoustic models have been used in numerous studies over the past thirty years to simulate the percepts elicited by auditory neural prostheses. In these acoustic models, incoming signals are processed the same way as in a cochlear implant speech processor. The percepts that would be caused by electrical stimulation in a real cochlear implant are simulated by modulating the amplitude of either noise bands or sinusoids. Despite their practical usefulness these acoustic models have never been convincingly validated. This study presents a tool to conduct such validation using subjects who have a cochlear implant in one ear and have near perfect hearing in the other ear, allowing for the first time a direct perceptual comparison of the output of acoustic models to the stimulation provided by a cochlear implant.
Mario A. Svirsky, Nai Ding, Elad Sagi, Chin-Tuan Tan, Matthew Fitzgerald, E. Katelyn Glassman, Keena Seward, Arlene C. Neuman
ICASSP2