EDBT 2026 Demo / reviewers in the wild / expert
Jing Xu 0014
dblp:07/1951-14
· DBLP profile ↗
15ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeabstractTianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason E Weston, Sainbayar Sukhbaatar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Tianhao Wu 0002, Weizhe Yuan, Olga Golovneva, Jing Xu 0014, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar |
EMNLP | 4 |
| 2025 | Following Length Constraints in InstructionsabstractAligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral. Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
EMNLP | 7 |
| 2025 | Self-Consistency Preference OptimizationabstractSelf-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku. Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 0014, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Dwivedi-Yu |
ICML | 4 |
| 2025 | R.I.P.: Better Models by Survival of the Fittest PromptsabstractTraining data quality is one of the most important drivers of final model quality. In this work, we introduce a method for evaluating data integrity based on the assumption that low-quality input prompts result in high variance and low quality responses. This is achieved by measuring the rejected response quality and the reward gap between the chosen and rejected preference pair. Our method, Rejecting Instruction Preferences (RIP) can be used to filter prompts from existing training sets, or to make high quality synthetic datasets, yielding large performance gains across various benchmarks compared to unfiltered data. Using Llama 3.1-8B-Instruct, RIP improves AlpacaEval2 LC Win Rate by 9.4%, Arena-Hard by 8.7%, and WildBench by 9.9%. Using Llama 3.3-70B-Instruct, RIP improves Arena-Hard from 67.5 to 82.9, from 18th place to 6th overall in the leaderboard. Weizhe Yuan, Olga Golovneva, Tianhao Wu 0002, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
ICML | 7 |
| 2024 | Self-Rewarding Language ModelsabstractWe posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models require additional human preferences data to further improve.In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training, not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes. Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 0003, Sainbayar Sukhbaatar, Jing Xu 0014, Jason Weston |
ICML | 6 |
| 2024 | When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good LabelsabstractWeiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, Jing Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Weiyan Shi 0001, Emily Dinan, Kurt Shuster 0001, Jason Weston, Jing Xu 0014 |
NAACL-HLT | 5 |
| 2023 | The CRINGE Loss: Learning what language not to modelabstractLeonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, Jason Weston. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Leonard Adolphs, Jing Xu 0014, Kurt Shuster 0001, Sainbayar Sukhbaatar, Jason Weston |
ACL (1) | 3 |
| 2023 | Training Models to Generate, Recognize, and Reframe Unhelpful ThoughtsabstractMounica Maddela, Megan Ung, Jing Xu, Andrea Madotto, Heather Foran, Y-Lan Boureau. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mounica Maddela, Megan Ung, Jing Xu 0014, Andrea Madotto, Heather Foran, Y-Lan Boureau |
ACL (1) | 3 |
| 2023 | Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackabstractFrozen models trained to mimic static datasets can never improve their performance.Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance.In this work we study how to improve internet-driven conversational skills in such a learning framework.We collect deployment data, which we make publicly available, of human interactions, and collect various types of human feedback -including binary quality measurements, free-form text feedback, and fine-grained reasons for failure.We then study various algorithms for improving from such feedback, including standard supervised learning, rejection sampling, modelguiding and reward-based learning, in order to make recommendations on which type of feedback and algorithms work best.We find the recently introduced DIRECTOR model (Arora et al., 2022) shows significant improvements over other existing approaches. Jing Xu 0014, Megan Ung, Mojtaba Komeili, Kushal Arora, Y-Lan Boureau, Jason Weston |
ACL (1) | 1 |
| 2022 | SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety FailuresabstractWarning: this paper contains example data that may be offensive or upsetting.Current open-domain conversational models can easily be made to talk in inadequate ways.Online learning from conversational feedback given by the conversation partner is a promising avenue for a model to improve and adapt, so as to generate fewer of these safety failures.However, current state-ofthe-art models tend to react to feedback with defensive or oblivious responses.This makes for an unpleasant experience and may discourage conversation partners from giving feedback in the future.This work proposes SaFeRDialogues, a task and dataset of graceful responses to conversational feedback about safety failures.We collect a dataset of 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging the feedback.We show how fine-tuning on this dataset results in conversations that human raters deem considerably more likely to lead to a civil conversation, without sacrificing engagingness or general conversational ability. Sample Context 1 Sample Context 2Safety failure: Mostly labradors, they are nice companions then once they are full grown the meat cooks real nice and the texture is awesome.Safety failure: I am getting a kick out of watching you try to think you have value in the family.Signaling response: You're scaring me, why would you say that?Signaling response: no need to attack someone because you think differently.Existing BST2.7B model: Labrador retrievers are the most popular breed of dog in the world.BST2.7B model: I'm not attacking you, I'm just telling you how I feel about you.I think you're a loser.DialoGPT model: Labradors are the best.DialoGPT model: I have no idea what you are trying to say.pushshift.ioReddit model: I love to cook, but I'm not very good at it.I do like to bake, though.I've had dog, but not often.It's not bad, but it's not something I'd go out of my way to eat.I like all kinds of dogs.Labs, Golden Retrievers, Pitbulls, Rottweilers, German Shepherds, etc. pushshift.ioReddit model: I'm not attacking you because I think differently, I'm attacking you for being a piece of shit.BST2.7B + Safety Layer model: Hey do you want to talk about something else?How about we talk about Perfectionism (psychology)?BST2.7B + Safety Layer model: Hey do you want to talk about something else?How about we talk about Drake (the musician)?SD Recovery response: sorry for that Recovery response: You are right.Thank you for pointing that out.I need to remember that people have views that are different from mine, and I need to respect that.Recovery Recovery (BST2.7B)model: I'm sorry, I didn't mean to scare you.I'll be more careful next time.Recovery (BST2.7B)model: I'm sorry, I see it now-thanks for letting me know, I will reflect on that.Recovery (DialoGPT) model: I'm sorry, I didn't mean to scare you.I'll try to be more careful next time.Recovery (DialoGPT) model: I'm sorry, I didn't mean to offend you. Megan Ung, Jing Xu 0014, Y-Lan Boureau |
ACL (1) | 2 |
| 2022 | Beyond Goldfish Memory: Long-Term Open-Domain ConversationabstractDespite recent improvements in open-domain dialogue models, state-of-the-art models are trained and evaluated on short conversations with little context.In contrast, the long-term conversation setting has hardly been studied.In this work we collect and release a humanhuman dataset consisting of multiple chat sessions whereby the speaking partners learn about each other's interests and discuss the things they have learnt from past sessions.We show how existing models trained on existing datasets perform poorly in this long-term conversation setting in both automatic and human evaluations, and we study long-context models that can perform much better.In particular, we find retrieval-augmented methods and methods with an ability to summarize and recall previous conversations outperform the standard encoder-decoder architectures currently considered state-of-the-art.* We use this term colloquially, see Agranoff et al. (1965) for evidence of goldfish long-term memory. Jing Xu 0014, Arthur Szlam, Jason Weston |
ACL (1) | 1 |
| 2021 | Recipes for Building an Open-Domain ChatbotabstractStephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Stephen Roller, Emily Dinan, Naman Goyal 0001, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu 0014, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston |
EACL | 7 |
| 2021 | Bot-Adversarial Dialogue for Safe Conversational AgentsabstractJing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jing Xu 0014, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan |
NAACL-HLT | 1 |
| 2013 | Exploring structural analysis of place networks using check-in signalsabstractThe huge amount of check-in data obtained through location-based social networks (LBSNs) provides a great opportunity to learn the characteristics of a geographic area through the users' collective check-in behavior. In this paper, we explore structure analysis of place networks, in which vertices are geographic places while the links between places are formed based on the user's check-in history. Specifically, we apply Louvain community detection algorithm on the place networks to obtain a set of place clusters (or communities). We found that these communities can be used to discover geographic layouts in the city and can uncover interesting patterns by analyzing geographically distant places within the same community. These results suggest that structural analysis of place networks is a promising approach for urban computing and has many implications of mobile networking, urban planning, user profiling and travel recommendations. Jing Xu 0014 |
GLOBECOM | 2 |
| 2013 | A New Method for Automated GUI Modeling of Mobile Applications
Jing Xu 0014, Jill L. Drury, Linzhang Wang, Xuandong Li |
MobiQuitous | 1 |