Jingcheng Niu

dblp:245/8596 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-2358-9466ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 43% Efficient and distributed learning · 36% Language models and text generation · 15%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.532025
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity · EMNLP 2025
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs · ACL (1) 2025
What does the Knowledge Neuron Thesis Have to do with Knowledge? · ICLR 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
1.722025
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity · EMNLP 2025
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs · ACL (1) 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit discovery
0.912025
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity · EMNLP 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning · EMNLP 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression
pruning
0.912025
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression › sparsity
sparse parameterization
0.912025
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning · EMNLP 2025
Natural language and speech › Language models and text generation
knowledge editing
0.812024
What does the Knowledge Neuron Thesis Have to do with Knowledge? · ICLR 2024
Natural language and speech › Information extraction and text analysis
slot filling
0.412019
Rationally Reappraising ATIS-based Dialogue Systems · ACL (1) 2019
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.412019
Rationally Reappraising ATIS-based Dialogue Systems · ACL (1) 2019
Natural language and speech › Language models and text generation › large language model › knowledge in language models
factual recall
0.212024
What does the Knowledge Neuron Thesis Have to do with Knowledge? · ICLR 2024

Methods — techniques the papers use, named apart from their topics

low-rank approximation · 0.9gradient-based pruning · 0.9differentiable masking · 0.9counterfactual prompting · 0.9LoRA · 0.9rule-based grammar · 0.8neural slot filling · 0.8model editing · 0.8
YearPublicationVenuePosition
2025 Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
abstract
We observe a novel phenomenon, contextual entrainment, across a wide range of language models (LMs) and prompt settings, providing a new mechanistic perspective on how LMs become distracted by “irrelevant” contextual information in the input prompt. Specifically, LMs assign significantly higher logits (or probabilities) to any tokens that have previously appeared in the context prompt, even for random tokens. This suggests that contextual entrainment is a mechanistic phenomenon, occurring independently of the relevance or semantic relation of the tokens to the question or the rest of the sentence. We find statistically significant evidence that the magnitude of contextual entrainment is influenced by semantic factors. Counterfactual prompts have a greater effect compared to factual ones, suggesting that while contextual entrainment is a mechanistic phenomenon, it is modulated by semantic factors.We hypothesise that there is a circuit of attention heads — the entrainment heads — that corresponds to the contextual entrainment phenomenon. Using a novel entrainment head discovery method based on differentiable masking, we identify these heads across various settings. When we “turn off” these heads, i.e., set their outputs to zero, the effect of contextual entrainment is significantly attenuated, causing the model to generate output that capitulates to what it would produce if no distracting context were provided. Our discovery of contextual entrainment, along with our investigation into LM distraction via the entrainment heads, marks a key step towards the mechanistic analysis and mitigation of the distraction problem.
Jingcheng Niu, Xingdi Yuan, Hamidreza Saghir, Amir H. Abdi
ACL (1)1
2025 Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
abstract
In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs).Sheaves extend the concept of functional circuits, a unit widely explored in interpretability research, by considering not only subsets of edges in an LM's computation graph but also the model's weight parameters.Our framework identifies sheaves through a gradient-based pruning algorithm that operates on both of these in such a way that reduces the original LM to a sparse skeleton that preserves certain core capabilities.Experimental results demonstrate that, across a range of linguistic and reasoning tasks, DiscoGP extracts sheaves that preserve 93-100% of a model's performance on the identified task while comprising only 1-7% of the original weights and connections.Furthermore, our analysis reveals that, compared to previously identified LM circuits, the sheaves discovered by DiscoGP exhibit superior modularity and functional fidelity.Extending our method to the neuron level also unveils novel insights into the inner workings of LLMs. 1 * Equal contribution. 1 The code and results of DiscoGP are available online: https://github.com/frankniujc/disco_gp.
Jingcheng Niu, Zining Zhu 0001, Gerald Penn
EMNLP2
2025 Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning
abstract
In this work, we propose FoRA-UA, a novel method that, using only 1-5% of the standard LoRA's parameters, achieves state-ofthe-art performance across a wide range of tasks.Specifically, we explore scenarios with extremely limited parameter budgets and derive two key insights: (1) fix-sized sparse frequency representations approximate small matrices more accurately; and (2) with a fixed number of trainable parameters, introducing a smaller intermediate representation to approximate larger matrices results in lower construction error.These findings form the foundation of our FoRA-UA method.By inserting a small intermediate parameter set, we achieve greater model compression without sacrificing performance.We evaluate FoRA-UA across diverse tasks, including natural language understanding (NLU), natural language generation (NLG), instruction tuning, and image classification, demonstrating strong generalisation and robustness under extreme compression. 1
Jinman Zhao, Jiaru Li, Jingcheng Niu, Yulan Hu, Erxue Min, Gerald Penn
EMNLP4
2024 What does the Knowledge Neuron Thesis Have to do with Knowledge?
abstract
We reassess the Knowledge Neuron (KN) Thesis: an interpretation of the mechanism underlying the ability of large language models to recall facts from a training corpus. This nascent thesis proposes that facts are recalled from the training corpus through the MLP weights in a manner resembling key-value memory, implying in effect that "knowledge" is stored in the network. Furthermore, by modifying the MLP modules, one can control the language model's generation of factual information. The plausibility of the KN thesis has been demonstrated by the success of KN-inspired model editing methods (Dai et al., 2022; Meng et al., 2022). We find that this thesis is, at best, an oversimplification. Not only have we found that we can edit the expression of certain linguistic phenomena using the same model editing methods but, through a more comprehensive evaluation, we have found that the KN thesis does not adequately explain the process of factual expression. While it is possible to argue that the MLP weights store complex patterns that are interpretable both syntactically and semantically, these patterns do not constitute "knowledge." To gain a more comprehensive understanding of the knowledge representation process, we must look beyond the MLP weights and explore recent models' complex layer structures and attention mechanisms.
Jingcheng Niu, Zining Zhu 0001, Gerald Penn
ICLR1
2022 Does BERT Rediscover a Classical NLP Pipeline?
abstract
Does BERT store surface knowledge in its bottom layers, syntactic knowledge in its middle layers, and semantic knowledge in its upper layers? In re-examining Jawahar et al. (2019) and Tenney et al.’s (2019a) probes into the structure of BERT, we have found that the pipeline-like separation that they asserted lacks conclusive empirical support. BERT’s structure is, however, linguistically founded, although perhaps in a way that is more nuanced than can be explained by layers alone. We introduce a novel probe, called GridLoc, through which we can also take into account token positions, training rounds, and random seeds. Using GridLoc, we are able to detect other, stronger regularities that suggest that pseudo-cognitive appeals to layer depth may not be the preferable mode of explanation for BERT’s inner workings.
Jingcheng Niu, Gerald Penn
COLING1
2020 Temporal Histories of Epidemic Events (THEE): A Case Study in Temporal Annotation for Public Health
abstract
We present a new temporal annotation standard, THEE-TimeML, and a corpus TheeBank enabling precise temporal information extraction (TIE) for event-based surveillance (EBS) systems in the public health domain. Current EBS must estimate the occurrence time of each event based on coarse document metadata such as document publication time. Because of the complicated language and narration style of news articles, estimated case outbreak times are often inaccurate or even erroneous. Thus, it is necessary to create annotation standards and corpora to facilitate the development of TIE systems in the public health domain to address this problem. We will discuss the adaptations that have proved necessary for this domain as we present THEE-TimeML and TheeBank. Finally, we document the corpus annotation process, and demonstrate the immediate benefit to public health applications brought by the annotations.
Jingcheng Niu, Victoria Ng, Gerald Penn, Erin E. Rees
LREC1
2019 Rationally Reappraising ATIS-based Dialogue Systems
abstract
The Air Travel Information Service (ATIS) corpus has been the most common benchmark for evaluating Spoken Language Understanding (SLU) tasks for more than three decades since it was released.Recent state-of-the-art neural models have obtained F1-scores near 98% on the task of slot filling.We developed a rule-based grammar for the ATIS domain that achieves a 95.82% F1-score on our evaluation set.In the process, we furthermore discovered numerous shortcomings in the ATIS corpus annotation, which we have fixed.This paper presents a detailed account of these shortcomings, our proposed repairs, our rulebased grammar and the neural slot-filling architectures associated with ATIS.We also rationally reappraise the motivations for choosing a neural architecture in view of this account.Fixing the annotation errors results in a relative error reduction of between 19.4 and 52% across all architectures.We nevertheless argue that neural models must play a different role in ATIS dialogues because of the latter's lack of variety.
Jingcheng Niu, Gerald Penn
ACL (1)1