VLDB 2026 Research / reviewers in the wild / expert
Zining Zhu 0001
dblp:188/5709
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-9285-9378ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts ModelsabstractMixture of Experts models are widely assumed to achieve domain specialization through sparse routing. In this work, we question this assumption by introducing COMMITTEEAUDIT, a post hoc framework that analyzes routing behavior at the level of expert groups rather than individual experts. Across three representative models and the MMLU benchmark, we uncover a domain invariant Standing Committee. This is a compact coalition of routed experts that consistently captures the majority of routing mass across domains, layers, and routing budgets, even when architectures already include shared experts. Qualitative analysis further shows that Standing Committees anchor reasoning structure and syntax, while peripheral experts handle domain-specific knowledge. These findings reveal a strong structural bias toward centralized computation, suggesting that specialization in Mixture of Experts models is far less pervasive than commonly believed. Crucially, this inherent bias indicates that current training objectives, such as load-balancing losses that enforce uniform expert utilization, may be working against the model’s natural optimization path, thereby limiting training efficiency and performance. Yan Wang 0015, Nanhan Shen, Jinyan Su, Jimin Huang, Zining Zhu 0001 |
ACL (1) | 6 |
| 2025 | INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based AgentabstractHaohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu, K.p. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu 0001, K. P. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie |
ACL (1) | 8 |
| 2025 | Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceabstractAutoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt.In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit than others.Specifically, for any target next-token distribution over the vocabulary, we attempt to find a prompt that induces the LM to output a distribution as close as possible to the target, using either soft (Li and Liang, 2021) or hard (Wallace et al., 2019) gradient-based prompt tuning.We find that (1) in general, distributions with very low or very high entropy are easier to approximate than those with moderate entropy; (2) among distributions with the same entropy, those containing "outlier tokens" are easier to approximate; (3) target distributions generated by LMs-even LMs with different tokenizers-are easier to approximate than randomly chosen targets.These results offer insights into the expressiveness of LMs and the challenges of using them as probability distribution proposers. Haojin Wang, Zining Zhu 0001, Freda Shi |
EMNLP | 2 |
| 2025 | Sheaf Discovery with Joint Computation Graph Pruning and Flexible GranularityabstractIn this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs).Sheaves extend the concept of functional circuits, a unit widely explored in interpretability research, by considering not only subsets of edges in an LM's computation graph but also the model's weight parameters.Our framework identifies sheaves through a gradient-based pruning algorithm that operates on both of these in such a way that reduces the original LM to a sparse skeleton that preserves certain core capabilities.Experimental results demonstrate that, across a range of linguistic and reasoning tasks, DiscoGP extracts sheaves that preserve 93-100% of a model's performance on the identified task while comprising only 1-7% of the original weights and connections.Furthermore, our analysis reveals that, compared to previously identified LM circuits, the sheaves discovered by DiscoGP exhibit superior modularity and functional fidelity.Extending our method to the neuron level also unveils novel insights into the inner workings of LLMs. 1 * Equal contribution. 1 The code and results of DiscoGP are available online: https://github.com/frankniujc/disco_gp. Jingcheng Niu, Zining Zhu 0001, Gerald Penn |
EMNLP | 3 |
| 2025 | ACCORD: Closing the Commonsense Measurability GapabstractFrançois Roewer-Després, Jinyue Feng, Zining Zhu, Frank Rudzicz. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. François Roewer-Després, Jinyue Feng, Zining Zhu 0001, Frank Rudzicz |
NAACL (Long Papers) | 3 |
| 2024 | What does the Knowledge Neuron Thesis Have to do with Knowledge?abstractWe reassess the Knowledge Neuron (KN) Thesis: an interpretation of the mechanism underlying the ability of large language models to recall facts from a training corpus. This nascent thesis proposes that facts are recalled from the training corpus through the MLP weights in a manner resembling key-value memory, implying in effect that "knowledge" is stored in the network. Furthermore, by modifying the MLP modules, one can control the language model's generation of factual information. The plausibility of the KN thesis has been demonstrated by the success of KN-inspired model editing methods (Dai et al., 2022; Meng et al., 2022).
We find that this thesis is, at best, an oversimplification. Not only have we found that we can edit the expression of certain linguistic phenomena using the same model editing methods but, through a more comprehensive evaluation, we have found that the KN thesis does not adequately explain the process of factual expression. While it is possible to argue that the MLP weights store complex patterns that are interpretable both syntactically and semantically, these patterns do not constitute "knowledge." To gain a more comprehensive understanding of the knowledge representation process, we must look beyond the MLP weights and explore recent models' complex layer structures and attention mechanisms. Jingcheng Niu, Zining Zhu 0001, Gerald Penn |
ICLR | 3 |
| 2023 | A State-Vector Framework for Dataset EffectsabstractThe impressive success of recent deep neural network (DNN)-based systems is significantly influenced by the high-quality datasets used in training.However, the effects of the datasets, especially how they interact with each other, remain underexplored.We propose a statevector framework to enable rigorous studies in this direction.This framework uses idealized probing test results as the bases of a vector space.This framework allows us to quantify the effects of both standalone and interacting datasets.We show that the significant effects of some commonly-used language understanding datasets are characteristic and are concentrated on a few linguistic dimensions.Additionally, we observe some "spill-over" effects: the datasets could impact the models along dimensions that may seem unrelated to the intended tasks.Our state-vector framework paves the way for a systematic understanding of the dataset effects, a crucial component in responsible and robust model development. Esmat Sahak, Zining Zhu 0001, Frank Rudzicz |
EMNLP | 2 |
| 2022 | Neural reality of argument structure constructionsabstractIn lexicalist linguistic theories, argument structure is assumed to be predictable from the meaning of verbs.As a result, the verb is the primary determinant of the meaning of a clause.In contrast, construction grammarians propose that argument structure is encoded in constructions (or form-meaning pairs) that are distinct from verbs.Decades of psycholinguistic research have produced substantial empirical evidence in favor of the construction view.Here we adapt several psycholinguistic studies to probe for the existence of argument structure constructions (ASCs) in Transformerbased language models (LMs).First, using a sentence sorting experiment, we find that sentences sharing the same construction are closer in embedding space than sentences sharing the same verb.Furthermore, LMs increasingly prefer grouping by construction with more input data, mirroring the behaviour of non-native language learners.Second, in a "Jabberwocky" priming-based experiment, we find that LMs associate ASCs with meaning, even in semantically nonsensical sentences.Our work offers the first evidence for ASCs in LMs and highlights the potential to devise novel probing methods grounded in psycholinguistic research. Transitive DitransitiveCaused-motion Resultative Throw Anita threw the hammer.Chris threw Linda the pencil.Pat threw the keys onto the roof.Lyn threw the box apart. Zining Zhu 0001, Guillaume Thomas, Frank Rudzicz, Yang Xu 0023 |
ACL (1) | 2 |
| 2022 | Predicting Fine-Tuning Performance with ProbingabstractLarge NLP models have recently shown impressive performance in language understanding tasks, typically evaluated by their finetuned performance.Alternatively, probing has received increasing attention as being a lightweight method for interpreting the intrinsic mechanisms of large NLP models.In probing, post-hoc classifiers are trained on "out-ofdomain" datasets that diagnose specific abilities.While probing the language models has led to insightful findings, they appear disjointed from the development of models.This paper explores the utility of probing deep NLP models to extract a proxy signal widely used in model development -the fine-tuning performance.We find that it is possible to use the accuracies of only three probing tests to predict the fine-tuning performance with errors 40% -80% smaller than baselines.We further discuss possible avenues where probing can empower the development of deep NLP models. Zining Zhu 0001, Soroosh Shahtalebi, Frank Rudzicz |
EMNLP | 1 |
| 2021 | How is BERT surprised? Layerwise detection of linguistic anomaliesabstractBai Li, Zining Zhu, Guillaume Thomas, Yang Xu, Frank Rudzicz. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zining Zhu 0001, Guillaume Thomas, Yang Xu 0023, Frank Rudzicz |
ACL/IJCNLP (1) | 2 |
| 2020 | An information theoretic view on selecting linguistic probesabstractThere is increasing interest in assessing the linguistic knowledge encoded in neural representations.A popular approach is to attach a diagnostic classifier -or "probe" -to perform supervised classification from internal representations.However, how to select a good probe is in debate.Hewitt and Liang (2019) showed that a high performance on diagnostic classification itself is insufficient, because it can be attributed to either "the representation being rich in knowledge", or "the probe learning the task", which Pimentel et al. (2020) challenged.We show this dichotomy is valid informationtheoretically.In addition, we find that the methods to construct and select good probes proposed by the two papers, control task (Hewitt and Liang, 2019) and control function (Pimentel et al., 2020), are equivalent -the errors of their approaches are identical (modulo irrelevant terms).Empirically, these two selection criteria lead to results that highly agree with each other. Zining Zhu 0001, Frank Rudzicz |
EMNLP (1) | 1 |
| 2017 | Deep neural networks for improved, impromptu trajectory tracking of quadrotorsabstractTrajectory tracking control for quadrotors is important for applications ranging from surveying and inspection, to film making. However, designing and tuning classical controllers, such as proportional-integral-derivative (PID) controllers, to achieve high tracking precision can be time-consuming and difficult, due to hidden dynamics and other non-idealities. The Deep Neural Network (DNN), with its superior capability of approximating abstract, nonlinear functions, proposes a novel approach for enhancing trajectory tracking control. This paper presents a DNN-based algorithm as an add-on module that improves the tracking performance of a classical feedback controller. Given a desired trajectory, the DNNs provide a tailored reference input to the controller based on their gained experience. The input aims to achieve a unity map between the desired and the output trajectory. The motivation for this work is an interactive “fly-as-you-draw” application, in which a user draws a trajectory on a mobile device, and a quadrotor instantly flies that trajectory with the DNN-enhanced control system. Experimental results demonstrate that the proposed approach improves the tracking precision for user-drawn trajectories after the DNNs are trained on selected periodic trajectories, suggesting the method's potential in real-world applications. Tracking errors are reduced by around 40-50% for both training and testing trajectories from users, highlighting the DNNs' capability of generalizing knowledge. Qiyang Li, Jingxing Qian, Zining Zhu 0001, Xuchan Bao, Mohamed K. Helwa, Angela P. Schoellig |
ICRA | 3 |