Jingfei Du

dblp:137/3917 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2023 SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
abstract
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk
ACL (1)4
2023 Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality
abstract
Contrastively trained vision-language models have achieved remarkable progress in vision and language representation learning.However, recent research has highlighted severe limitations of these models in their ability to perform compositional reasoning over objects, attributes, and relations.Scene graphs have emerged as an effective way to understand images compositionally.These are graphstructured semantic representations of images that contain objects, their attributes, and relations with other objects in a scene.In this work, we consider the scene graph parsed from text as a proxy for the image scene graph and propose a graph decomposition and augmentation framework along with a coarse-to-fine contrastive learning objective between images and text that aligns sentences of various complexities to the same image.We also introduce novel negative mining techniques in the scene graph space for improving attribute binding and relation understanding.Through extensive experiments, we demonstrate the effectiveness of our approach that significantly improves attribute binding, relation understanding, systematic generalization, and productivity on multiple recently proposed benchmarks (For example, improvements up to 18% for systematic generalization, 16.5% for relation understanding over a strong baseline), while achieving similar or better performance than CLIP on various general multimodal tasks.
Harman Singh, Pengchuan Zhang, Qifan Wang 0001, Wenhan Xiong, Jingfei Du, Yu Chen 0022
EMNLP6
2022 Efficient Large Scale Language Modeling with Mixtures of Experts
abstract
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov
EMNLP8
2022 Few-shot Learning with Multilingual Generative Language Models
abstract
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003
EMNLP10
2022 Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained Models
abstract
Pre-trained masked language models successfully perform few-shot learning by formulating downstream tasks as text infilling.However, as a strong alternative in full-shot settings, discriminative pre-trained models like ELECTRA do not fit into the paradigm.In this work, we adapt prompt-based few-shot learning to ELECTRA and show that it outperforms masked language models in a wide range of tasks.ELECTRA is pre-trained to distinguish if a token is generated or original.We naturally extend that to prompt-based few-shot learning by training to score the originality of the target options without introducing new parameters.Our method can be easily adapted to tasks involving multi-token predictions without extra computation overhead.Analysis shows that ELECTRA learns distributions that align better with downstream tasks. 1
Mengzhou Xia, Mikel Artetxe, Jingfei Du, Danqi Chen 0001, Veselin Stoyanov
EMNLP3
2022 Improving In-Context Few-Shot Learning via Self-Supervised Training
abstract
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srini Iyer, Veselin Stoyanov, Zornitsa Kozareva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva
NAACL-HLT2
2021 Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, Veselin Stoyanov
ICLR2
2021 Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz
ICLR4
2021 Self-training Improves Pre-training for Natural Language Understanding
abstract
Jingfei Du, Edouard Grave, Beliz Gunel, Vishrav Chaudhary, Onur Celebi, Michael Auli, Veselin Stoyanov, Alexis Conneau. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jingfei Du, Edouard Grave, Beliz Gunel, Vishrav Chaudhary, Onur Celebi, Michael Auli, Veselin Stoyanov, Alexis Conneau
NAACL-HLT1
2020 Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
Wenhan Xiong, Jingfei Du, William Yang Wang, Veselin Stoyanov
ICLR2
2014 Box office prediction based on microblog
Jingfei Du, Hua Xu 0003, Xiaoqiu Huang 0002
Expert Syst. Appl.1
2013 Multi-Objective Optimization for Overlapping Community Detection
Jingfei Du, Jianyang Lai, Chuan Shi 0001
ADMA (2)1