Jason Phang

dblp:227/3174 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
8since 2021 · last 2023
0000-0003-3522-1869ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 70% Information extraction and text analysis · 8% Representation and self-supervised learning · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
long document summarization
1.222023
Investigating Efficiently Extending Transformers for Long Input Summarization · EMNLP 2023
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Language models and text generation
alignment
0.712023
Pretraining Language Models with Human Preferences · ICML 2023
Natural language and speech › Language models and text generation › large language model
large language model adaptation
0.712023
HyperTuning: Toward Adapting Large Language Models without Back-propagation · ICML 2023
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.712023
Investigating Efficiently Extending Transformers for Long Input Summarization · EMNLP 2023
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.712023
HyperTuning: Toward Adapting Large Language Models without Back-propagation · ICML 2023
Machine learning › Representation and self-supervised learning
pre-training
0.712023
Pretraining Language Models with Human Preferences · ICML 2023
Natural language and speech › Language models and text generation › text summarization › controllable summarization
query-focused summarization
0.612022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Language models and text generation
text summarization
0.612022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.522020
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work? · ACL 2020
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Performance modeling and evaluation
benchmarking
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Performance modeling and evaluation › statistical analysis
item response theory
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
linguistic generalization
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › pre-trained language model › knowledge probing
linguistic knowledge probing
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Machine learning › Deep learning architectures and training
transformer
0.212023
Investigating Efficiently Extending Transformers for Long Input Summarization · EMNLP 2023
Information retrieval
evaluation
0.212022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Information retrieval › text summarization
summarization evaluation
0.212022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

survey · 1.3human annotation · 1.1automatic evaluation metrics · 1.1soft prefix · 0.7reinforcement learning from human feedback · 0.7long-sequence pretraining · 0.7hypernetwork · 0.7conditional training · 0.7block-local attention · 0.7LoRA · 0.7item response theory · 0.5
YearPublicationVenuePosition
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)10
2023 Investigating Efficiently Extending Transformers for Long Input Summarization
abstract
While large pretrained Transformer models have proven highly capable at tackling natural language tasks, handling long sequence inputs still poses a significant challenge.One such task is long input summarization, where inputs are longer than the maximum input context of most models.Through an extensive set of experiments, we investigate what model architectural changes and pretraining paradigms most efficiently adapt a pretrained Transformer for long input summarization.We find that a staggered, block-local Transformer with global encoder tokens strikes a good balance of performance and efficiency, and that an additional pretraining phase on long sequences meaningfully improves downstream summarization performance.Based on our findings, we introduce PEGASUS-X, an extension of the PE-GASUS model with additional long input pretraining to handle inputs of up to 16K tokens, which achieves strong performance on long input summarization tasks comparable with much larger models.
Jason Phang, Peter J. Liu
EMNLP1
2023 Pretraining Language Models with Human Preferences
abstract
Language models (LMs) are pretrained to imitate text from large and diverse datasets that contain content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, among others. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training.
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, Ethan Perez
ICML6
2023 HyperTuning: Toward Adapting Large Language Models without Back-propagation
abstract
Fine-tuning large language models for different tasks can be costly and inefficient, and even methods that reduce the number of tuned parameters still require full gradient-based optimization. We propose HyperTuning, a novel approach to model adaptation that uses a hypermodel to generate task-specific parameters for a fixed downstream model. We demonstrate a simple setup for hypertuning with HyperT5, a T5-based hypermodel that produces soft prefixes or LoRA parameters for a frozen T5 model from few-shot examples. We train HyperT5 in two stages: first, hyperpretraining with a modified conditional language modeling objective that trains a hypermodel to generate parameters; second, multi-task fine-tuning (MTF) on a large number of diverse language tasks. We evaluate HyperT5 on P3, MetaICL and Super-NaturalInstructions datasets, and show that it can effectively generate parameters for unseen tasks. Moreover, we show that using hypermodel-generated parameters as initializations for further parameter-efficient fine-tuning improves performance. HyperTuning can thus be a flexible and efficient way to leverage large language models for diverse downstream applications.
Jason Phang, Weizhu Chen
ICML1
2022 SQuALITY: Building a Long-Document Summarization Dataset the Hard Way
abstract
Summarization datasets are often assembled either by scraping naturally occurring publicdomain summaries-which are nearly always in difcult-to-work-with technical domainsor by using approximate heuristics to extract them from everyday text-which frequently yields unfaithful summaries.In this work, we turn to a slower but more straightforward approach to developing summarization benchmark data: We hire highly-qualied contractors to read stories and write original summaries from scratch.To amortize reading time, we collect ve summaries per document, with the rst giving an overview and the subsequent four addressing specic questions.We use this protocol to collect SQuAL-ITY, a dataset of question-focused summaries built on the same public-domain short stories as the multiple-choice dataset QuALITY (Pang et al., 2021b).Experiments with stateof-the-art summarization systems show that our dataset is challenging and that existing automatic evaluation metrics are weak indicators of quality.
Richard Yuanzhe Pang, Angelica Chen, Jason Phang, Samuel R. Bowman
EMNLP4
2022 QuALITY: Question Answering with Long Input Texts, Yes!
abstract
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman
NAACL-HLT5
2021 Comparing Test Sets with Item Response Theory
abstract
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Clara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman
ACL/IJCNLP (1)6
2021 An interpretable classifier for high-resolution breast cancer screening images utilizing weakly supervised localization
abstract
Medical images differ from natural images in significantly higher resolutions and smaller regions of interest. Because of these differences, neural network architectures that work well for natural images might not be applicable to medical image analysis. In this work, we propose a novel neural network model to address these unique properties of medical images. This model first uses a low-capacity, yet memory-efficient, network on the whole image to identify the most informative regions. It then applies another higher-capacity network to collect details from chosen regions. Finally, it employs a fusion module that aggregates global and local information to make a prediction. While existing methods often require lesion segmentation during training, our model is trained with only image-level labels and can generate pixel-level saliency maps indicating possible malignant findings. We apply the model to screening mammography interpretation: predicting the presence or absence of benign and malignant lesions. On the NYU Breast Cancer Screening Dataset, our model outperforms (AUC = 0.93) ResNet-34 and Faster R-CNN in classifying breasts with malignant findings. On the CBIS-DDSM dataset, our model achieves performance (AUC = 0.858) on par with state-of-the-art approaches. Compared to ResNet-34, our model is 4.1x faster for inference while using 78.4% less GPU memory. Furthermore, we demonstrate, in a reader study, that our model surpasses radiologist-level AUC by a margin of 0.11.
Yiqiu Shen, Nan Wu 0008, Jason Phang, Jungkyu Park 0001, Kangning Liu, Sudarshini Tyagi, Laura Heacock, Sungheon Gene Kim, Linda Moy, Kyunghyun Cho, Krzysztof J. Geras
Medical Image Anal.3
2020 Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?
abstract
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman
ACL2
2020 Deep Neural Networks Improve Radiologists' Performance in Breast Cancer Screening
abstract
We present a deep convolutional neural network for breast cancer screening exam classification, trained, and evaluated on over 200000 exams (over 1000000 images). Our network achieves an AUC of 0.895 in predicting the presence of cancer in the breast, when tested on the screening population. We attribute the high accuracy to a few technical advances. 1) Our network's novel two-stage architecture and training procedure, which allows us to use a high-capacity patch-level network to learn from pixel-level labels alongside a network learning from macroscopic breast-level labels. 2) A custom ResNet-based network used as a building block of our model, whose balance of depth and width is optimized for high-resolution medical images. 3) Pretraining the network on screening BI-RADS classification, a related task with more noisy labels. 4) Combining multiple input views in an optimal way among a number of possible choices. To validate our model, we conducted a reader study with 14 readers, each reading 720 screening mammogram exams, and show that our model is as accurate as experienced radiologists when presented with the same data. We also show that a hybrid model, averaging the probability of malignancy predicted by a radiologist with a prediction of our neural network, is more accurate than either of the two separately. To further understand our results, we conduct a thorough analysis of our network's performance on different subpopulations of the screening population, the model's design, training procedure, errors, and properties of its internal representations. Our best models are publicly available at https://github.com/nyukat/breast_cancer_classifier.
Nan Wu 0008, Jason Phang, Jungkyu Park 0001, Yiqiu Shen, Zhe Huang 0005, Masha Zorin, Stanislaw Jastrzebski, Thibault Févry, Joe Katsnelson, Eric Kim, Stacey Wolfson, Ujas Parikh, Sushma Gaddam, Leng Leng Young Lin, Kara Ho, Joshua D. Weinstein, Beatriu Reig, Yiming Gao 0003, Hildegard Toth, Kristine Pysarenko, Alana Lewin, Jiyon Lee, Krystal Airola, Eralda Mema, Stephanie Chung, Esther Hwang, Naziya Samreen, Sungheon Gene Kim, Laura Heacock, Linda Moy, Kyunghyun Cho, Krzysztof J. Geras
IEEE Trans. Medical Imaging2
2019 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
abstract
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman
EMNLP/IJCNLP (1)12
2018 Unsupervised Sentence Compression using Denoising Auto-Encoders
abstract
In sentence compression, the task of shortening sentences while retaining the original meaning, models tend to be trained on large corpora containing pairs of verbose and compressed sentences.To remove the need for paired corpora, we emulate a summarization task and add noise to extend sentences and train a denoising auto-encoder to recover the original, constructing an end-to-end training regime without the need for any examples of compressed sentences.We conduct a human evaluation of our model on a standard text summarization dataset and show that it performs comparably to a supervised baseline based on grammatical correctness and retention of meaning.Despite being exposed to no target data, our unsupervised models learn to generate imperfect but reasonably readable sentence summaries.Although we underperform supervised models based on ROUGE scores, our models are competitive with a supervised baseline based on human evaluation for grammatical correctness and retention of meaning.
Thibault Févry, Jason Phang
CoNLL2