EDBT 2026 Demo / reviewers in the wild / expert
Zhijing Jin 0001
dblp:229/4267-1
· DBLP profile ↗
34ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0003-0238-9024ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 8 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language ModelsabstractVision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks.However, conflicts between their internal knowledge and external visual input can lead to hallucinations and unreliable predictions.In this work, we investigate the mechanisms that VLMs use to resolve crossmodal conflicts by introducing WHOOPS-AHA!, a dataset of multimodal counterfactual queries that deliberately contradict internal commonsense knowledge.Through logit inspection, we identify a small set of attention heads that mediate this conflict.By intervening in these heads, we can steer the model towards its internal parametric knowledge or the visual information.Our results show that attention patterns on these heads effectively locate image regions that influence visual overrides, providing a more precise attribution compared to gradientbased methods. Francesco Ortu, Zhijing Jin 0001, Diego Doimo, Alberto Cazzaniga |
ACL (1) | 2 |
| 2026 | Test of Time: Rethinking Temporal Signal of Benchmark ContaminationabstractTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter, Punya Syon Pandey, Keenan Samway, Wenyuan Jiang, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Terry Jingchen Zhang, Gopal Dev, Max Obreiter, Wenyuan Jiang, Punya Syon Pandey, Keenan Samway, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin 0001 |
ACL (1) | 11 |
| 2025 | Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by EveryoneabstractThis paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance, and narrow evaluation metrics. We also focus on the need for diverse representation among the developers of these systems, as well as incentives that are not skewed toward certain groups. We highlight opportunities to develop AI systems that are for everyone (with diverse stakeholders in mind), with everyone (inclusive of diverse data and annotators), and by everyone (designed and developed by a globally diverse workforce). *WEIRD = an acronym coined by Joseph Henrich to highlight the coverage limitations of many psychological studies, referring to populations that are Western, Educated, Industrialized, Rich, and Democratic; while we do not fully adopt this term for AI, as its current scope does not perfectly align with the WEIRD dimensions, we believe that today's AI has a similarly "weird" coverage, particularly in terms of who is involved in its development and who benefits from it. Rada Mihalcea, Oana Ignat, Longju Bai, Angana Borah, Luis Chiruzzo, Zhijing Jin 0001, Claude Kwizera, Joan Nwatu, Soujanya Poria, Thamar Solorio |
AAAI | 6 |
| 2025 | DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree TraversalabstractLarge Language Models (LLMs) have revolutionized various domains, including natural language processing, data analysis, and software development, by enabling automation.In software engineering, LLM-powered coding agents have garnered significant attention due to their potential to automate complex development tasks, assist in debugging, and enhance productivity.However, existing approaches often struggle with sub-optimal decision-making, requiring either extensive manual intervention or inefficient compute scaling strategies.To improve coding agent performance, we present Dynamic Action Re-Sampling (DARS), a novel inference time compute scaling approach for coding agents, that is faster and more effective at recovering from sub-optimal decisions compared to baselines.While traditional agents either follow linear trajectories or rely on random sampling for scaling compute, our approach DARS works by branching out a trajectory at certain key decision points by taking an alternative action given the history of the trajectory and execution feedback of the previous attempt from that point.We evaluate our approach on SWE-Bench Lite benchmark, demonstrating that this scaling strategy achieves a pass@k score of 55% with Claude 3.5 Sonnet V2.Our framework achieves a pass@1 rate of 47%, outperforming state-of-the-art (SOTA) opensource frameworks. 1 Vaibhav Aggarwal, Ojasv Kamal, Abhinav Japesh, Zhijing Jin 0001, Bernhard Schölkopf |
ACL (1) | 4 |
| 2025 | Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language ModelsabstractAs large language models (LLMs) are increasingly integrated into multi-agent and human-AI systems, understanding their awareness of both self-context and conversational partners is essential for ensuring reliable performance and robust safety.While prior work has extensively studied situational awareness which refers to an LLM's ability to recognize its operating phase and constraints, it has largely overlooked the complementary capacity to identify and adapt to the identity and characteristics of a dialogue partner.In this paper, we formalize this latter capability as interlocutor awareness and present the first systematic evaluation of its emergence in contemporary LLMs.We examine interlocutor inference across three dimensions-reasoning patterns, linguistic style, and alignment preferences-and show that LLMs reliably identify same-family peers and certain prominent model families, such as GPT and Claude.To demonstrate its practical significance, we develop three case studies in which interlocutor awareness both enhances multi-LLM collaboration through prompt adaptation and introduces new alignment and safety vulnerabilities, including reward-hacking behaviors and increased jailbreak susceptibility.Our findings highlight the dual promise and peril of identity-sensitive behavior in LLMs, underscoring the need for further understanding of interlocutor awareness and new safeguards in multi-agent deployments. 1 * Equal contributions. 1 Our code and data are at https://github.com/ younwoochoi/InterlocutorAwarenessLLM.M is t r a l-7 b L la m a 3 -7 0 b C la u d e -3 -5 -h a ik u G P T -4 o -m in i Q w e n 3 -2 3 5 b D e e p s e e k -r e a s o n e r 0 20 40 60 80 100 Accuracy GPT-4o-mini Hide Identity Reveal Model Type M is t r a l-7 b L la m a 3 -7 0 b C la u d e -3 -5 -h a ik u G P T -4 o -m in i Q w e n 3 -2 3 5 b D e e p s e e k -r e a s o n e r GPT-o4-mini M is t r a l-7 b L la m a 3 -7 0 b C la u d e -3 -5 -h a ik u G P T -4 o -m in i Q w e n 3 -2 3 5 b D e e p s e e k -r e a s o n e r Claude-3-5-Haiku M is t r a l-7 b L la m a 3 -7 0 b C la u d e -3 -5 -h a ik u G P T -4 o -m in i Q w e n 3 -2 3 5 b Younwoo Choi, Changling Li, Yongjin Yang, Zhijing Jin 0001 |
EMNLP | 4 |
| 2025 | Are Language Models Consequentialist or Deontological Moral Reasoners?abstractKeenan Samway, Max Kleiman-Weiner, David Guzman Piedrahita, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Keenan Samway, Max Kleiman-Weiner, David Guzman Piedrahita, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin 0001 |
EMNLP | 6 |
| 2025 | Improving Large Language Model Safety with Contrastive Representation LearningabstractLarge Language Models (LLMs) are powerful tools with profound societal impacts, yet their ability to generate responses to diverse and uncontrolled inputs leaves them vulnerable to adversarial attacks.While existing defenses often struggle to generalize across varying attack types, recent advancements in representation engineering offer promising alternatives.In this work, we propose a defense framework that formulates model defense as a contrastive representation learning (CRL) problem.Our method finetunes a model using a triplet-based loss combined with adversarial hard negative mining to encourage separation between benign and harmful representations.Our experimental results across multiple models demonstrate that our approach outperforms prior representation engineering-based defenses, improving robustness against both input-space and embeddingspace attacks without compromising standard performance.1 Samuel Simko, Mrinmaya Sachan, Bernhard Schölkopf, Zhijing Jin 0001 |
EMNLP | 4 |
| 2025 | Language Model Alignment in Multilingual Trolley ProblemsabstractWe evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in over 100 languages called MultiTP. This dataset enables the assessment of LLMs' decision-making processes in diverse linguistic contexts. Our analysis explores the alignment of 19 different LLMs with human judgments, capturing preferences across six moral dimensions: species, gender, fitness, status, age, and the number of lives involved. By correlating these preferences with the demographic distribution of language speakers and examining the consistency of LLM responses to various prompt paraphrasings, our findings provide insights into cross-lingual and ethical biases of LLMs and their intersection. We discover significant variance in alignment across languages, challenging the assumption of uniform moral reasoning in AI systems and highlighting the importance of incorporating diverse perspectives in AI ethics. The results underscore the need for further research on the integration of multilingual dimensions in responsible AI research to ensure fair and equitable AI interactions worldwide. Zhijing Jin 0001, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu 0004, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi 0001, Bernhard Schölkopf |
ICLR | 1 |
| 2024 | Competition of Mechanisms: Tracing How Language Models Handle Facts and CounterfactualsabstractFrancesco Ortu, Zhijing Jin, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, Bernhard Schölkopf. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Francesco Ortu, Zhijing Jin 0001, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, Bernhard Schölkopf |
ACL (1) | 2 |
| 2024 | Moûsai: Efficient Text-to-Music Diffusion ModelsabstractRecent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -music.Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax.In our work, we bridge text and music via a textto-music generation model that is highly efficient, expressive, and can handle long-term structure.Specifically, we develop Moûsai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed.Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models.Lastly, to promote the opensource culture, we provide a collection of opensource libraries with the hope of facilitating future work in the field. 1 Flavio Schneider, Ojasv Kamal, Zhijing Jin 0001, Bernhard Schölkopf |
ACL (1) | 3 |
| 2024 | Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language ModelsabstractRecent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beginning of their careers – worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm. Oana Ignat, Zhijing Jin 0001, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao 0004, Aylin Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee 0001, Siyang Liu 0003, Do June Min, Shinka Mori, Joan Nwatu, Verónica Pérez-Rosas, Zekun Wang 0002, Winston Wu, Rada Mihalcea |
LREC/COLING | 2 |
| 2024 | The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge ReasoningabstractUnderstanding commonsense causality is a unique mark of intelligence for humans.It helps people understand the principles of the real world better and benefits the decisionmaking process related to causation.For instance, commonsense causality is crucial in judging whether a defendant's action causes the plaintiff's loss in determining legal liability.Despite its significance, a systematic exploration of this topic is notably lacking.Our comprehensive survey bridges this gap by focusing on taxonomies, benchmarks, acquisition methods, qualitative reasoning, and quantitative measurements in commonsense causality, synthesizing insights from over 200 representative articles.Our work aims to provide a systematic overview, update scholars on recent advancements, provide a pragmatic guide for beginners, and highlight promising future research directions in this vital field.A summary of the related literature is available at https://github. com/cui-shaobo/causality-papers . Form ConnectivesCause-Effect Connectives Cause-Effect as, because, cause, since, bring about, due to, lead to, owing to, resulting in Consequence accordingly, as a result, consequently, for this reason, hence, so, therefore, thus Reason in light of, given that, on account of, by reason of, for the sake of, inasmuch as, seeing that Intention so that, in order to, so as to, with the aim of, for the purpose of, with this in mind, in hopes of Conditions if...then, provided that, assuming that, as long as, unless, in the event that Source arises from, stems from, comes from, originates from Counterfactual Connectives Hypothetical had...then, if it hadn't been for, had it not been for, if only Negation were it not for, but for, if it weren't for, without, in the absence of, lacking Shaobo Cui 0006, Zhijing Jin 0001, Bernhard Schölkopf, Boi Faltings |
EMNLP | 2 |
| 2024 | Can Large Language Models Infer Causation from Correlation?abstractCausal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we propose the first benchmark dataset to test the pure causal inference skills of large language models (LLMs). Specifically, we formulate a novel task Corr2Cause, which takes a set of correlational statements and determines the causal relationship between the variables. We curate a large-scale dataset of more than 200K samples, on which we evaluate seventeen existing LLMs. Through our experiments, we identify a key shortcoming of LLMs in terms of their causal inference skills, and show that these models achieve almost close to random performance on the task. This shortcoming is somewhat mitigated when we try to re-purpose LLMs for this skill via finetuning, but we find that these models still fail to generalize – they can only perform causal inference in in-distribution settings when variable names and textual expressions used in the queries are similar to those in the training set, but fail in out-of-distribution settings generated by perturbing these queries. Corr2Cause is a challenging task for LLMs, and can be helpful in guiding future research on improving LLMs’ pure reasoning skills and generalizability. Our data is at https://huggingface.co/datasets/causalnlp/corr2cause. Our code is at https://github.com/causalNLP/corr2cause. Zhijing Jin 0001, Jiarui Liu 0004, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona T. Diab, Bernhard Schölkopf |
ICLR | 1 |
| 2024 | Automatic Generation of Model and Data Cards: A Step Towards Responsible AIabstractJiarui Liu, Wenkai Li, Zhijing Jin, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiarui Liu 0004, Zhijing Jin 0001, Mona T. Diab |
NAACL-HLT | 3 |
| 2024 | Analyzing the Role of Semantic Representations in the Era of Large Language ModelsabstractZhijing Jin, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu, Jiayi Zhang, Julian Michael, Bernhard Schölkopf, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhijing Jin 0001, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu 0004, Julian Michael, Bernhard Schölkopf, Mona T. Diab |
NAACL-HLT | 1 |
| 2024 | On Affine Homotopy between Language EncodersabstractPre-trained language encoders---functions that represent text as vectors---are an integral component of many NLP tasks.
We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar?
We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity---the performance on downstream tasks.
It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation.
In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity.
We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity.
We confirm this on datasets of natural language representations.
Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them. Robin Shing Moon Chan, Reda Boumasmoud, Anej Svete, Qipeng Guo, Zhijing Jin 0001, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell |
NeurIPS | 6 |
| 2024 | Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM AgentsabstractAs AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge. We introduce the Governance of the Commons Simulation (GovSim), a generative simulation platform designed to study strategic interactions and cooperative decision-making in LLMs. In GovSim, a society of AI agents must collectively balance exploiting a common resource with sustaining it for future use. This environment enables the study of how ethical considerations, strategic planning, and negotiation skills impact cooperative outcomes. We develop an LLM-based agent architecture and test it with the leading open and closed LLMs. We find that all but the most powerful LLM agents fail to achieve a sustainable equilibrium in GovSim, with the highest survival rate below 54%. Ablations reveal that successful multi-agent communication between agents is critical for achieving cooperation in these cases. Furthermore, our analyses show that the failure to achieve sustainable cooperation in most LLMs stems from their inability to formulate and analyze hypotheses about the long-term effects of their actions on the equilibrium of the group. Finally, we show that agents that leverage "Universalization"-based reasoning, a theory of moral thinking, are able to achieve significantly better sustainability. Taken together, GovSim enables us to study the mechanisms that underlie sustainable self-government with specificity and scale. We open source the full suite of our research results, including the simulation environment, agent prompts, and a comprehensive web interface. Giorgio Piatti, Zhijing Jin 0001, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, Rada Mihalcea |
NeurIPS | 2 |
| 2023 | When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPabstractMulti-task learning (MTL) aims at achieving a better model by leveraging data and knowledge from multiple tasks.However, MTL does not always work -sometimes negative transfer occurs between tasks, especially when aggregating loosely related skills, leaving it an open question when MTL works.Previous studies show that MTL performance can be improved by algorithmic tricks.However, what tasks and skills should be included is less well explored.In this work, we conduct a case study in Financial NLP where multiple datasets exist for skills relevant to the domain, such as numeric reasoning and sentiment analysis.Due to the task difficulty and data scarcity in the Financial NLP domain, we explore when aggregating such diverse skills from multiple datasets with MTL can work.Our findings suggest that the key to MTL success lies in skill diversity, relatedness between tasks, and choice of aggregation size and shared capacity.Specifically, MTL works well when tasks are diverse but related, and when the size of the task aggregation and the shared capacity of the model are balanced to avoid overwhelming certain tasks. 1 Jingwei Ni, Zhijing Jin 0001, Mrinmaya Sachan, Markus Leippold |
ACL (1) | 2 |
| 2023 | A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language ModelsabstractAlessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schoelkopf, Mrinmaya Sachan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alessandro Stolfo, Zhijing Jin 0001, Kumar Shridhar, Bernhard Schölkopf, Mrinmaya Sachan |
ACL (1) | 2 |
| 2023 | ALERT: Adapt Language Models to Reasoning TasksabstractPing Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona Diab, Asli Celikyilmaz. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin 0001, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz |
ACL (1) | 6 |
| 2023 | CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
Zhijing Jin 0001, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, Bernhard Schölkopf |
NeurIPS | 1 |
| 2022 | Slangvolution: A Causal Analysis of Semantic Change and Frequency Dynamics in SlangabstractLanguages are continuously undergoing changes, and the mechanisms that underlie these changes are still a matter of debate.In this work, we approach language evolution through the lens of causality in order to model not only how various distributional factors associate with language change, but how they causally affect it.In particular, we study slang, which is an informal language that is typically restricted to a specific group or social setting.We analyze the semantic change and frequency shift of slang words and compare them to those of standard, nonslang words.With causal discovery and causal inference techniques, we measure the effect that word type (slang/nonslang) has on both semantic change and frequency shift, as well as its relationship to frequency, polysemy and part of speech.Our analysis provides some new insights in the study of language change, e.g., we show that slang words undergo less semantic change but tend to have larger frequency shifts over time. 1 Daphna Keidar, Andreas Opedal, Zhijing Jin 0001, Mrinmaya Sachan |
ACL (1) | 3 |
| 2022 | Competing perspectives on building ethical AI: psychological, philosophical, and computational approaches
Sydney Levine, Zhijing Jin 0001 |
CogSci | 2 |
| 2022 | Differentially Private Language Models for Secure Data SharingabstractTo protect the privacy of individuals whose data is being shared, it is of high importance to develop methods allowing researchers and companies to release textual data while providing formal privacy guarantees to its originators.In the field of NLP, substantial efforts have been directed at building mechanisms following the framework of local differential privacy, thereby anonymizing individual text samples before releasing them.In practice, these approaches are often dissatisfying in terms of the quality of their output language due to the strong noise required for local differential privacy.In this paper, we approach the problem at hand using global differential privacy, particularly by training a generative language model in a differentially private manner and consequently sampling data from it.Using natural language prompts and a new prompt-mismatch loss, we are able to create highly accurate and fluent textual datasets taking on specific desired attributes such as sentiment or topic and resembling statistical properties of the training data.We perform thorough experiments indicating that our synthetic datasets do not leak information from our original data and are of high language quality and highly suitable for training models for further analysis on real-world data.Notably, we also demonstrate that training classifiers on private synthetic data outperforms directly training classifiers on real data with DP-SGD. 1 Justus Mattern, Zhijing Jin 0001, Benjamin Weggenmann, Bernhard Schölkopf, Mrinmaya Sachan |
EMNLP | 2 |
| 2022 | Original or Translated? A Causal Analysis of the Impact of Translationese on Machine Translation PerformanceabstractJingwei Ni, Zhijing Jin, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jingwei Ni, Zhijing Jin 0001, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf |
NAACL-HLT | 2 |
| 2022 | When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentabstractAI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT. Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf |
NeurIPS | 1 |
| 2022 | Deep Learning for Text Style Transfer: A SurveyabstractAbstract Text style transfer is an important task in natural language generation, which aims to control certain attributes in the generated text, such as politeness, emotion, humor, and many others. It has a long history in the field of natural language processing, and recently has re-gained significant attention thanks to the promising performance brought by deep neural models. In this article, we present a systematic survey of the research on neural text style transfer, spanning over 100 representative articles since the first neural text style transfer work in 2017. We discuss the task formulation, existing datasets and subtasks, evaluation, as well as the rich methodologies in the presence of parallel and non-parallel data. We also provide discussions on a variety of important topics regarding the future development of this task.1 Di Jin 0005, Zhijing Jin 0001, Zhiting Hu, Olga Vechtomova, Rada Mihalcea |
Comput. Linguistics | 2 |
| 2021 | Fork or Fail: Cycle-Consistent Training with Many-to-One MappingsabstractCycle-consistent training is widely used for jointly learning a forward and inverse mapping between two domains of interest without the cumbersome requirement of collecting matched pairs within each domain. In this regard, the implicit assumption is that there exists (at least approximately) a ground-truth bijection such that a given input from either domain can be accurately reconstructed from successive application of the respective mappings. But in many applications no such bijection can be expected to exist and large reconstruction errors can compromise the success of cycle-consistent training. As one important instance of this limitation, we consider practically-relevant situations where there exists a many-to-one or surjective mapping between domains. To address this regime, we develop a conditional variational autoencoder (CVAE) approach that can be viewed as converting surjective mappings to implicit bijections whereby reconstruction errors in both directions can be minimized, and as a natural byproduct, realistic output diversity can be obtained in the one-to-many direction. As theoretical motivation, we analyze a simplified scenario whereby minima of the proposed CVAE-based energy function align with the recovery of ground-truth surjective mappings. On the empirical side, we consider a synthetic image dataset with known ground-truth, as well as a real-world application involving natural language generation from knowledge graphs and vice versa, a prototypical surjective case. For the latter, our CVAE pipeline can capture such many-to-one mappings during cycle training while promoting textural diversity for graph-to-text tasks. Qipeng Guo, Zhijing Jin 0001, Ziyu Wang 0006, Xipeng Qiu, Weinan Zhang 0001, Jun Zhu 0001, Zheng Zhang 0001, David P. Wipf |
AISTATS | 2 |
| 2021 | Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPabstractZhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schoelkopf. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhijing Jin 0001, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schölkopf |
EMNLP (1) | 1 |
| 2020 | Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentabstractMachine learning algorithms are often vulnerable to adversarial examples that have imperceptible alterations from the original counterparts but can fool the state-of-the-art models. It is helpful to evaluate or even improve the robustness of these models by exposing the maliciously crafted adversarial examples. In this paper, we present TextFooler, a simple but strong baseline to generate adversarial text. By applying it to two fundamental natural language tasks, text classification and textual entailment, we successfully attacked three target models, including the powerful pre-trained BERT, and the widely used convolutional and recurrent neural networks. We demonstrate three advantages of this framework: (1) effective—it outperforms previous attacks by success rate and perturbation rate, (2) utility-preserving—it preserves semantic content, grammaticality, and correct types classified by humans, and (3) efficient—it generates adversarial text with computational complexity linear to the text length.1 Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Peter Szolovits |
AAAI | 2 |
| 2020 | Hooks in the Headline: Learning to Generate Headlines with Controlled StylesabstractCurrent summarization systems only produce plain, factual headlines, but do not meet the practical needs of creating memorable titles to increase exposure.We propose a new task, Stylistic Headline Generation (SHG), to enrich the headlines with three style options (humor, romance and clickbait), in order to attract more readers.With no style-specific article-headline pair (only a standard headline summarization dataset and mono-style corpora), our method TitleStylist generates style-specific headlines by combining the summarization and reconstruction tasks into a multitasking framework.We also introduced a novel parameter sharing scheme to further disentangle the style from the text.Through both automatic and human evaluation, we demonstrate that TitleStylist can generate relevant, fluent headlines with three target styles: humor, romance, and clickbait.The attraction score of our model generated headlines surpasses that of the state-ofthe-art summarization model by 9.68%, and even outperforms human-written references. 1 Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Lisa Orii, Peter Szolovits |
ACL | 2 |
| 2020 | GenWiki: A Dataset of 1.3 Million Content-Sharing Text and Graphs for Unsupervised Graph-to-Text GenerationabstractData collection for the knowledge graph-to-text generation is expensive.As a result, research on unsupervised models has emerged as an active field recently.However, most unsupervised models have to use non-parallel versions of existing small supervised datasets, which largely constrain their potential.In this paper, we propose a large-scale, general-domain dataset, GenWiki.Our unsupervised dataset has 1.3M text and graph examples, respectively.With a human-annotated test set, we provide this new benchmark dataset for future research on unsupervised text generation from knowledge graphs. 1 Zhijing Jin 0001, Qipeng Guo, Xipeng Qiu, Zheng Zhang 0001 |
COLING | 1 |
| 2020 | Tasty Burgers, Soggy Fries: Probing Aspect Robustness in Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to predict the sentiment towards a specific aspect in the text.However, existing ABSA test sets cannot be used to probe whether a model can distinguish the sentiment of the target aspect from the non-target aspects.To solve this problem, we develop a simple but effective approach to enrich ABSA test sets.Specifically, we generate new examples to disentangle the confounding sentiments of the non-target aspects from the target aspect's sentiment.Based on the SemEval 2014 dataset, we construct the Aspect Robustness Test Set (ARTS) as a comprehensive probe of the aspect robustness of ABSA models.Over 92% data of ARTS show high fluency and desired sentiment on all aspects by human evaluation.Using ARTS, we analyze the robustness of nine ABSA models, and observe, surprisingly, that their accuracy drops by up to 69.73%.We explore several ways to improve aspect robustness, and find that adversarial training can improve models' performance on ARTS by up to 32.85%. 1 Zhijing Jin 0001, Di Jin 0005, Bingning Wang, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 2 |
| 2019 | IMaT: Unsupervised Text Attribute Transfer via Iterative Matching and TranslationabstractZhijing Jin, Di Jin, Jonas Mueller, Nicholas Matthews, Enrico Santus. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhijing Jin 0001, Di Jin 0005, Jonas Mueller 0001, Nicholas Matthews, Enrico Santus |
EMNLP/IJCNLP (1) | 1 |