Wenlong Zhao 0001

dblp:03/4555-1 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-6153-1396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
abstract
The rise of large language models (LLMs) has created a significant disparity: industrial research labs with their computational resources, expert teams, and advanced infrastructures, can effectively fine-tune LLMs, while individual developers and small organizations face barriers due to limited resources to effectively explore the experiment space. In this paper, we aim to bridge this gap by presenting a comprehensive study on supervised fine-tuning of LLMs using instruction-tuning datasets spanning diverse knowledge domains and skills. We focus on small-sized LLMs (3B to 7B parameters) for their cost-efficiency and accessibility. We explore various training configurations and strategies across four open-source pre-trained models. We provide detailed documentation of these configurations, revealing findings that challenge several common training practices, including hyperparameter recommendations from TULU and phased training recommended by Orca. The code used for the experiments can be found here: https://github.com/instructlab/training. Key insights from our work include: (i) larger batch sizes paired with lower learning rates lead to improved model performance on benchmarks such as MMLU, MTBench, and Open LLM Leaderboard; (ii) early-stage training dynamics, such as lower gradient norms and higher loss values, are strong indicators of better final model performance, allowing for early termination of sub-optimal runs and significant computational savings; (iii) through a thorough exploration of hyperparameters like warmup steps and learning rate schedules, we provide guidance for practitioners and find that certain simplifications do not compromise performance; and (iv) we observe no significant difference in performance between phased (sequentially training on data divided into phases) and stacked (training on the entire dataset at once) strategies, but stacked training is simpler and more sample efficient. With these findings holding robustly across datasets as well as model families and sizes, we hope this study serves as a guide for practitioners fine-tuning small LLMs and promotes a more inclusive research environment for LLM development.
Aldo Pareja, Nikhil Shivakumar Nayak, Hao Wang 0014, KrishnaTeja Killamsetty, Shivchander Sudalairaj, Wenlong Zhao 0001, Seungwook Han, Abhishek Bhandwaldar, Guangxuan Xu, Kai Xu 0016, Ligong Han, Luke Inglis, Akash Srivastava
ICLR6
2025 OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
abstract
Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured. Yet the task is inherently challenging, partly due to difficulties in reliably measuring whether unlearning has truly occurred. Moreover, fragmentation in current methodologies and inconsistent evaluation metrics hinder comparative analysis and reproducibility. To unify and accelerate research efforts, we introduce OpenUnlearning, a standardized and extensible framework designed explicitly for benchmarking both LLM unlearning methods and metrics. OpenUnlearning integrates 13 state-of-the-art unlearning algorithms and 16 diverse evaluations across 3 leading benchmarks (TOFU, MUSE, and WMDP) and also enables analyses of forgetting behaviors across 450+ publicly released checkpoints. Leveraging OpenUnlearning, we propose a novel meta-evaluation benchmark focused specifically on assessing the faithfulness and robustness of evaluation metrics themselves. We also benchmark diverse unlearning methods and provide a comparative analysis against an extensive evaluation suite. Overall, we establish a clear, community-driven pathway toward rigorous development in LLM unlearning research.
Vineeth Dorna, Anmol Mekala, Wenlong Zhao 0001, Andrew McCallum, J. Zico Kolter, Zachary C. Lipton, Pratyush Maini
NeurIPS3
2025 Active Measurement: Efficient Estimation at Scale
abstract
AI has the potential to transform scientific discovery by analyzing vast datasets with little human effort. However, current workflows often do not provide the accuracy or statistical guarantees that are needed. We introduce \emph{active measurement}, a human-in-the-loop AI framework for scientific measurement. An AI model is used to predict measurements for individual units, which are then sampled for human labeling using importance sampling. With each new set of human labels, the AI model is improved and an unbiased Monte Carlo estimate of the total measurement is refined. Active measurement can provide precise estimates even with an imperfect AI model, and requires little human effort when the AI model is very accurate. We derive novel estimators, weighting schemes, and confidence intervals, and show that active measurement reduces estimation error compared to alternatives in several measurement tasks.
Max Hamilton, Jinlin Lai, Wenlong Zhao 0001, Subhransu Maji, Daniel Sheldon
NeurIPS3
2024 Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
abstract
Jiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenlong Zhao 0001, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum
ACL (1)2
2024 WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models
abstract
The awareness of multi-cultural human values is critical to the ability of language models (LMs) to generate safe and personalized responses. However, this awareness of LMs has been insufficiently studied, since the computer science community lacks access to the large-scale real-world data about multi-cultural values. In this paper, we present WorldValuesBench, a globally diverse, large-scale benchmark dataset for the multi-cultural value prediction task, which requires a model to generate a rating response to a value question based on demographic contexts. Our dataset is derived from an influential social science project, World Values Survey (WVS), that has collected answers to hundreds of value questions (e.g., social, economic, ethical) from 94,728 participants worldwide. We have constructed more than 20 million examples of the type "(demographic attributes, value question) → answer” from the WVS responses. We perform a case study using our dataset and show that the task is challenging for strong open and closed-source models. On merely 11.1%, 25.0%, 72.2%, and 75.0% of the questions, Alpaca-7B, Vicuna-7B-v1.5, Mixtral-8x7B-Instruct-v0.1, and GPT-3.5 Turbo can respectively achieve <0.2 Wasserstein 1-distance from the human normalized answer distributions. WorldValuesBench opens up new research avenues in studying limitations and opportunities in multi-cultural value awareness of LMs.
Wenlong Zhao 0001, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, Yuling Gu
LREC/COLING1
2024 Comparing Neighbors Together Makes it Easy: Jointly Comparing Multiple Candidates for Efficient and Effective Retrieval
abstract
A common retrieve-and-rerank paradigm involves retrieving relevant candidates from a broad set using a fast bi-encoder (BE), followed by applying expensive but accurate crossencoders (CE) to a limited candidate set.However, relying on this small subset is often susceptible to error propagation from the biencoders, which limits the overall performance.To address these issues, we propose the Comparing Multiple Candidates (CMC) framework.CMC compares a query and multiple embeddings of similar candidates (i.e., neighbors) through shallow self-attention layers, delivering rich representations contextualized to each other.Furthermore, CMC is scalable enough to handle multiple comparisons simultaneously.For example, comparing 10K candidates with CMC takes a similar amount of time as comparing 16 candidates with CE.Experimental results on the ZeSHEL dataset demonstrate that CMC, when plugged in between bi-encoders and cross-encoders as a seamless intermediate reranker (BE-CMC-CE), can effectively improve recall@k (+6.7%-p, +3.5%-p for R@16, R@64) compared to using only bi-encoders (BE-CE), with negligible slowdown (<7%).Additionally, to verify CMC's effectiveness as the final-stage reranker in improving top-1 accuracy, we conduct experiments on downstream tasks such as entity, passage, and dialogue ranking.The results indicate that CMC is not only faster (11x) but also often more effective than cross-encoders with improved prediction accuracy in Wikipedia entity linking (+0.7%-p) and DSTC7 dialogue ranking (+3.3%-p).
Jonghyun Song, Cheyon Jin, Wenlong Zhao 0001, Andrew McCallum, Jay-Yoon Lee
EMNLP3
2024 Learning Representations for Hierarchies with Minimal Support
abstract
When training node embedding models to represent large directed graphs (digraphs), it is impossible to observe all entries of the adjacency matrix during training. As a consequence most methods employ sampling. For very large digraphs, however, this means many (most) entries may be unobserved during training. In general, observing every entry would be necessary to uniquely identify a graph, however if we know the graph has a certain property some entries can be omitted - for example, only half the entries would be required for a symmetric graph. In this work, we develop a novel framework to identify a subset of entries required to uniquely distinguish a graph among all transitively-closed DAGs. We give an explicit algorithm to compute the provably minimal set of entries, and demonstrate empirically that one can train node embedding models with greater efficiency and performance, provided the energy function has an appropriate inductive bias. We achieve robust performance on synthetic hierarchies and a larger real-world taxonomy, observing improved convergence rates in a resource-constrained setting while reducing the set of training examples by as much as 99%.
Benjamin Rozonoyer, Michael Boratko, Dhruvesh Patel, Wenlong Zhao 0001, Shib Sankar Dasgupta, Andrew McCallum
NeurIPS4
2023 Editing Common Sense in Transformers
abstract
Editing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023).However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer.Commonsense knowledge with multiple correct answers, e.g., an apple can be green or red but not transparent, has not been studied but is as essential for enhancing transformers' reliability and usefulness.In this paper, we investigate whether commonsense judgments are causally associated with localized, editable parameters in Transformers, and we provide an affirmative answer.We find that directly applying the MEMIT editing algorithm results in sub-par performance, and propose to improve it for the commonsense domain by varying edit tokens and improving the layer selection strategy, i.e., MEMIT CSK .GPT-2 Large and XL models edited using MEMIT CSK outperform best-fine-tuned baselines by 10.97% and 10.73% F1 scores on PEP3k and 20Q datasets.In addition, we propose a novel evaluation dataset, PROBE SET, that contains unaffected and affected neighborhoods, affected paraphrases, and affected reasoning challenges.MEMIT CSK performs well across the metrics while fine-tuning baselines show significant trade-offs between unaffected and affected metrics.These results suggest a compelling future direction for incorporating feedback about common sense into Transformers through direct model editing. 1 * Co-first and last authors.Lorraine's work done at AI2. 1 Code and datasets for all experiments are available at https://github.com/anshitag/memit_csk
Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao 0001, Xiang Li 0069, Sarah Wiegreffe, Niket Tandon
EMNLP4
2022 Structured Energy Network As a Loss
abstract
Belanger & McCallum (2016) and Gygli et al. (2017) have shown that an energy network can capture arbitrary dependencies amongst the output variables in structured prediction; however, their reliance on gradient-based inference (GBI) makes the inference slow and unstable. In this work, we propose Structured Energy As Loss (SEAL) to take advantage of the expressivity of energy networks without incurring the high inference cost. This is a novel learning framework that uses an energy network as a trainable loss function (loss-net) to train a separate neural network (task-net), which is then used to perform the inference through a forward pass. We establish SEAL as a general framework wherein various learning strategies like margin-based, regression, and noise-contrastive, could be employed to learn the parameters of loss-net. Through extensive evaluation on multi-label classification, semantic role labeling, and image segmentation, we demonstrate that SEAL provides various useful design choices, is faster at inference than GBI, and leads to significant performance gains over the baselines.
Jay-Yoon Lee, Dhruvesh Patel, Purujit Goyal, Wenlong Zhao 0001, Zhiyang Xu, Andrew McCallum
NeurIPS4
2021 IGA: An Intent-Guided Authoring Assistant
abstract
Simeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain, Vlad Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Simeng Sun, Wenlong Zhao 0001, Varun Manjunatha, Rajiv Jain, Vlad I. Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer
EMNLP (1)2