VLDB 2026 Research / reviewers in the wild / expert
Zhaofeng Wu
dblp:168/7994
· DBLP profile ↗
17ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Implicit Representations of Grammaticality in Language ModelsabstractGrammaticality and likelihood are distinct notions in human language.Pretrained language models (LMs), which are probabilistic models of language fitted to maximize corpus likelihood, generate grammatically well-formed text and discriminate well between grammatical and ungrammatical sentences in tightly controlled minimal pairs.However, their string probabilities do not sharply discriminate between grammatical and ungrammatical sentences overall.But do LMs implicitly acquire a grammaticality distinction distinct from string probability?We explore this question through studying internal representations of LMs, by training a linear probe on a dataset of grammatical and (synthetic) ungrammatical sentences obtained by applying perturbations to a naturalistic text corpus.We find that this simple grammaticality probe generalizes to human-curated grammaticality judgment benchmarks and outperforms LM probability-based grammaticality judgments.When applied to semantic plausibility benchmarks, in which both members of a minimal pair are grammatical and differ in only plausibility, the probe however performs worse than string probability.The English-trained probe also exhibits nontrivial cross-lingual generalization, outperforming string probabilities on grammaticality benchmarks in numerous other languages.Additionally, probe scores correlate only weakly with string probabilities.These results collectively suggest that LMs acquire to some extent an implicit grammaticality distinction within their hidden layers. 1 Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger Levy |
ACL (1) | 3 |
| 2025 | reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsabstractReward models have become a staple in modern NLP, serving as not only a scalable text evaluator, but also an indispensable component in many alignment recipes and inference-time algorithms.However, while recent reward models increase performance on standard benchmarks, this may partly be due to overfitting effects, which would confound an understanding of their true capability.In this work, we scrutinize the robustness of reward models and the extent of such overfitting.We build re-WordBench, which systematically transforms reward model inputs in meaning-or rankingpreserving ways.We show that state-of-theart reward models suffer from substantial performance degradation even with minor input transformations, sometimes dropping to significantly below-random accuracy, suggesting brittleness.To improve reward model robustness, we propose to explicitly train them to assign similar scores to paraphrases, and find that this approach also improves robustness to other distinct kinds of transformations.For example, our robust reward model reduces such degradation by roughly half for the Chat Hard subset in RewardBench.Furthermore, when used in alignment, our robust reward models demonstrate better utility and lead to higher-quality outputs, winning in up to 59% of instances against a standardly trained RM. Zhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Asli Celikyilmaz, Marjan Ghazvininejad |
EMNLP | 1 |
| 2025 | The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and ModalitiesabstractModern language models can process inputs across diverse languages and modalities. We hypothesize that models acquire this capability through learning a _shared representation space_ across heterogeneous data types (e.g., different languages and modalities), which places semantically similar inputs near one another, even if they are from different modalities/languages. We term this the _semantic hub hypothesis_, following the hub-and-spoke model from neuroscience (Patterson et al., 2007) which posits that semantic knowledge in the human brain is organized through a transmodal semantic "hub" which integrates information from various modality-specific ``spokes'' regions. We first show that model representations for semantically equivalent inputs in different languages are similar in the intermediate layers, and that this space can be interpreted using the model's dominant pretraining language via the logit lens. This tendency extends to other data types, including arithmetic expressions, code, and visual/audio inputs. Interventions in the shared representation space in one data type also predictably affect model outputs in other data types, suggesting that this shared representations space is not simply a vestigial byproduct of large-scale training on broad data, but something that is actively utilized by the model during input processing. Zhaofeng Wu, Xinyan Yu 0001, Dani Yogatama, Jiasen Lu |
ICLR | 1 |
| 2025 | SelfCite: Self-Supervised Alignment for Context Attribution in Large Language ModelsabstractWe introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly and labor-intensive annotations, SelfCite leverages a reward signal provided by the LLM itself through context ablation: If a citation is necessary, removing the cited text from the context should prevent the same response; if sufficient, retaining the cited text alone should preserve the same response. This reward can guide the inference-time best-of-N sampling strategy to improve citation quality significantly, as well as be used in preference optimization to directly fine-tune the models for generating better citations. The effectiveness of SelfCite is demonstrated by increasing citation F1 up to 5.3 points on the LongBench-Cite benchmark across five long-form question answering tasks. The source code is available at https://github.com/facebookresearch/SelfCite. Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Shen 0001, Zhaofeng Wu, Hu Xu 0001, Xi Victoria Lin, James R. Glass, Shang-Wen Li 0001, Scott Yih |
ICML | 4 |
| 2024 | Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentabstractAligning language models (LMs) based on human-annotated preference data is a crucial step in obtaining practical and performant LMbased systems.However, multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages.In this work, we evaluate a simple approach for zero-shot crosslingual alignment, where a reward model is trained on preference data in one source language and directly applied to other target languages.On summarization and open-ended dialog generation, we show that this method is consistently successful under comprehensive evaluation settings, including human evaluation: cross-lingually aligned models are preferred by humans over unaligned models on up to >70% of evaluation instances.We moreover find that a different-language reward model sometimes yields better aligned models than a same-language reward model.We also identify best practices when there is no languagespecific data for even supervised finetuning, another component in alignment.en de en es en ru en tr en vi de en es en ru en tr en vi en 0 20 40 60 ROUGE-L (a) Summarization, unaligned SFT model Target-Language SFT Data Zhaofeng Wu, Ananth Balashankar, Jacob Eisenstein, Ahmad Beirami |
EMNLP | 1 |
| 2024 | Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual TasksabstractZhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen 0003, Bailin Wang, Najoung Kim, Jacob Andreas |
NAACL-HLT | 1 |
| 2023 | We're Afraid Language Models Aren't Modeling AmbiguityabstractAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah Smith, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, Yejin Choi 0001 |
EMNLP | 2 |
| 2023 | Transparency Helps Reveal When Language Models Learn MeaningabstractAbstract Many current NLP systems are built from language models trained to optimize unsupervised objectives on large amounts of raw text. Under what conditions might such a procedure acquire meaning? Our systematic experiments with synthetic data reveal that, with languages where all expressions have context-independent denotations (i.e., languages with strong transparency), both autoregressive and masked language models successfully learn to emulate semantic relations between expressions. However, when denotations are changed to be context-dependent with the language otherwise unmodified, this ability degrades. Turning to natural language, our experiments with a specific phenomenon—referential opacity—add to the growing body of evidence that current language models do not represent natural language semantics well. We show this failure relates to the context-dependent nature of natural language form-meaning mappings. Zhaofeng Wu, William Merrill, Hao Peng 0009, Iz Beltagy, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | ABC: Attention with Bounded-memory ControlabstractHao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, Noah Smith. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Peng 0009, Jungo Kasai, Nikolaos Pappas 0002, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz 0001, Noah A. Smith |
ACL (1) | 5 |
| 2022 | Continued Pretraining for Better Zero- and Few-Shot PromptabilityabstractRecently introduced language model prompting methods can achieve high accuracy in zeroand few-shot settings while requiring few to no learned task-specific parameters.Nevertheless, these methods still often trail behind full model finetuning.In this work, we investigate if a dedicated continued pretraining stage could improve "promptability", i.e., zero-shot performance with natural language prompts or few-shot performance with prompt tuning.We reveal settings where existing continued pretraining methods lack promptability.We also identify current methodological gaps, which we fill with thorough large-scale experiments.We demonstrate that a simple recipe, continued pretraining that incorporates a trainable prompt during multi-task learning, leads to improved promptability in both zero-and fewshot settings compared to existing methods, up to 31% relative.On the other hand, we find that continued pretraining using MAML-style metalearning, a method that directly optimizes fewshot promptability, yields subpar performance.We validate our findings with two prompt tuning methods, and, based on our results, we provide concrete recommendations to optimize promptability for different use cases. Zhaofeng Wu, Robert L. Logan IV, Pete Walsh 0001, Akshita Bhagia, Dirk Groeneveld, Sameer Singh 0001, Iz Beltagy |
EMNLP | 1 |
| 2021 | Dynamic Sparsity Neural Networks for Automatic Speech RecognitionabstractIn automatic speech recognition (ASR), model pruning is a widely adopted technique that reduces model size and latency to deploy neural network models on edge devices with resource constraints. However, multiple models with different sparsity levels usually need to be separately trained and deployed to heterogeneous target hardware with different resource specifications and for applications that have various latency requirements. In this paper, we present Dynamic Sparsity Neural Networks (DSNN) that, once trained, can instantly switch to any predefined sparsity configuration at run-time. We demonstrate the effectiveness and flexibility of DSNN using experiments on internal production datasets with Google Voice Search data, and show that the performance of a DSNN model is on par with that of individually trained single sparsity networks. Our trained DSNN model, therefore, can greatly ease the training process and simplify deployment in diverse scenarios with resource constraints. Zhaofeng Wu, Ding Zhao, Qiao Liang 0001, Anmol Gulati, Ruoming Pang |
ICASSP | 1 |
| 2021 | Single Sensor to Estimate DOA With Programmable MetasurfaceabstractWith the small sampling length, a novel Direction-of-Arrival (DOA) estimation method based on a single programmable metasurface sensor is proposed in this article. Serving as a physical random sampling receiver, a dynamic metasurface generates series of random radiation patterns to sense the incident signals, which is subsequently processed by the compressive sensing (CS) orthogonal matching pursuit (OMP) algorithm to recover the DOA information. A mathematical model of DOA estimation is first built for theoretical analyses. Furthermore, numerical simulations are investigated with extensive performance analyses. In comparison with the metasuface-based beam-scanning method, metasurface-based CS for DOA estimation has superior performance with a smaller sampling length. Finally, a proof-of-concept experiment is made to verify the feasibility of the programmable metasurface for DOA estimation. Compared to other traditional estimation techniques that require multiple channels, the programmable metasurface can achieve estimation with only a single channel and small sampling length. Mingtuan Lin, Ming Xu 0019, Zhaofeng Wu, Jibin Liu, Bowen Deng 0008, Dongfang Guan, Song Zha |
IEEE Internet Things J. | 5 |
| 2021 | Infusing Finetuning with Semantic DependenciesabstractAbstract For natural language processing systems, two kinds of evidence support the use of text representations from neural language models “pretrained” on large unannotated corpora: performance on application-inspired benchmarks (Peters et al., 2018, inter alia), and the emergence of syntactic abstractions in those representations (Tenney et al., 2019, inter alia). On the other hand, the lack of grounded supervision calls into question how well these representations can ever capture meaning (Bender and Koller, 2020). We apply novel probes to recent language models— specifically focusing on predicate-argument structure as operationalized by semantic dependencies (Ivanova et al., 2012)—and find that, unlike syntax, semantics is not brought to the surface by today’s pretrained models. We then use convolutional graph encoders to explicitly incorporate semantic parses into task-specific finetuning, yielding benefits to natural language understanding (NLU) tasks in the GLUE benchmark. This approach demonstrates the potential for general-purpose (rather than task-specific) linguistic supervision, above and beyond conventional pretraining and finetuning. Several diagnostics help to localize the benefits of our approach.1 Zhaofeng Wu, Hao Peng 0009, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 1 |
| 2017 | Fast counting the cardinality of flows for big traffic over sliding windows
Jingsong Shan, Yinjin Fu, Guiqiang Ni, Jianxin Luo, Zhaofeng Wu |
Frontiers Comput. Sci. | 5 |
| 2016 | CVS: Fast cardinality estimation for large-scale data streams over sliding windows
Jingsong Shan, Jianxin Luo, Guiqiang Ni, Zhaofeng Wu |
Neurocomputing | 4 |
| 2015 | A simple real-time handover management in the mobile satellite communication networksabstractLow earth orbit (LEO) satellite networks are capable of providing global or regional mobile services for a large number of users. Since the user's service duration may be greater than the coverage time of a LEO satellite, the user may be handed over to another visible satellite to prevent interruption of the ongoing communication. On the other hand, a mobile user may be covered by more than one satellite at the instant of connection handover. When the user is about to be handed over to another satellite, the serving satellite minimizing the number of handovers would in general be the one that provides the largest service time which is not necessarily equal to the coverage time of the very satellite. In this paper, we propose a new handover algorithm which exploits both the Global Positioning System (GPS) infrastructure and satellite diversity to provide a simple and real-time handover management in LEO satellite networks. The proposed algorithm not only minimizes the expected number of satellite handover, but is also efficient and easy to be implemented in hand-held devices, thus facilitating the mobile users' access to the satellite networks. Numerical simulations performed for two typical mobile satellite networks, viz. Iridium and Globalstar, corroborate the advantages gained by the proposed algorithm. Zhaofeng Wu, Guyu Hu, Younes Seyedi, Fenglin Jin |
APNOMS | 1 |
| 2015 | Tunneling-based Multi-path Routing Mechanism in Packet-Switched Non-Geostationary Satellite Networks
Guyu Hu, Zhaofeng Wu, Fenglin Jin, Bowei Yang, Yinjin Fu |
ICA3PP (4) | 2 |