Jane Dwivedi-Yu

dblp:215/3352 · also Jane A. Dwivedi-Yu, Jane A. Yu, Jane Yu 0001 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
23since 2021 · last 2025
0000-0003-2410-9396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Efficient Tool Use with Chain-of-Abstraction Reasoning
abstract
To achieve faithful reasoning that aligns with human expectations, large language models (LLMs) need to ground their reasoning to real-world knowledge (e.g., web facts, math and physical rules). Tools help LLMs access this external knowledge, but there remains challenges for fine-tuning LLM agents (e.g., Toolformer) to invoke tools in multi-step reasoning problems, where inter-connected tool calls require holistic and efficient tool usage planning. In this work, we propose a new method for LLMs to better leverage tools in multi-step reasoning. Our method, Chain-of-Abstraction (CoA), trains LLMs to first decode reasoning chains with abstract placeholders, and then call domain tools to reify each reasoning chain by filling in specific knowledge. This planning with abstract chains enables LLMs to learn more general reasoning strategies, which are robust to shifts of domain knowledge (e.g., math results) relevant to different reasoning questions. It also allows LLMs to perform decoding and calling of external tools in parallel, which avoids the inference delay caused by waiting for tool responses. In mathematical reasoning and Wiki QA domains, we show that our method consistently outperforms previous chain-of-thought and tool-augmented baselines on both in-distribution and out-of-distribution test sets, with an average ~6% absolute QA accuracy improvement. LLM agents trained with our method also show more efficient tool use, with inference speed being on average ~1.4x faster than baseline tool-augmented LLMs.
Silin Gao, Jane Dwivedi-Yu, Xiaoqing Ellen Tan, Ramakanth Pasunuru, Olga Golovneva, Koustuv Sinha, Asli Celikyilmaz, Antoine Bosselut
COLING2
2025 Extrapolating to Unknown Opinions Using LLMs
abstract
From ice cream flavors to climate change, people exhibit a wide array of opinions on various topics, and understanding the rationale for these opinions can promote healthy discussion and consensus among them. As such, it can be valuable for a large language model (LLM), particularly as an AI assistant, to be able to empathize with or even explain these various standpoints. In this work, we hypothesize that different topic stances often manifest correlations that can be used to extrapolate to topics with unknown opinions. We explore various prompting and fine-tuning methods to improve an LLM’s ability to (a) extrapolate from opinions on known topics to unknown ones and (b) support their extrapolation with reasoning. Our findings suggest that LLMs possess inherent knowledge from training data about these opinion correlations, and with minimal data, the similarities between human opinions and model-extrapolated opinions can be improved by more than 50%. Furthermore, LLM can generate the reasoning process behind their extrapolation of opinions.
Kexun Zhang, Jane Dwivedi-Yu, Zhaojiang Lin, Yuning Mao, William Yang Wang, Lei Li 0005, Yi-Chia Wang
COLING2
2025 Culture Cartography: Mapping the Landscape of Cultural Knowledge
abstract
To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training.How do we find knowledge that is (1) salient to ingroup users, but (2) unknown to LLMs?The most common solutions are single-initiative: either researchers define challenging questions that users passively answer (traditional annotation), or users actively produce data that researchers structure as benchmarks (knowledge extraction).The process would benefit from mixed-initiative collaboration, where users guide the process to meaningfully reflect their cultures, and LLMs steer the process to meet the researcher's goals.We propose CULTURE CARTOGRAPHY as a methodology that operationalizes this mixed-initiative vision.Here, an LLM initializes annotation with questions for which it has low-confidence answers, making explicit both its prior knowledge and the gaps therein.This allows a human respondent to fill these gaps and steer the model towards salient topics through direct edits.We implement CULTURE CARTOGRAPHY as a tool called CULTURE EXPLORER.Compared to a baseline where humans answer LLMproposed questions, we find that CULTURE EX-PLORER more effectively produces knowledge that strong models like DeepSeek R1, Llama-4 and GPT-4o are missing, even with web search.Fine-tuning on this data boosts the accuracy of Llama models by up to 19.2% on related culture benchmarks.
Caleb Ziems, William Barr Held, Jane Dwivedi-Yu, Amir Goldberg, David Grusky, Diyi Yang
EMNLP3
2025 Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoning
abstract
Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited datasets with simple patterns that can potentially lead to problematic blind spots in evaluation and an overestimation of model capabilities. We introduce ExploreToM, the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation. Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs. Our evaluation reveals that state-of-the-art LLMs, such as Llama-3.1-70B and GPT-4o, show accuracies as low as 0% and 9% on ExploreToM-generated data, highlighting the need for more robust theory of mind evaluation. As our generations are a conceptual superset of prior work, fine-tuning on our data yields a 27-point accuracy improvement on the classic ToMi benchmark (Le et al., 2019). ExploreToM also enables uncovering underlying skills and factors missing for models to show theory of mind, such as unreliable state tracking or data imbalances, which may contribute to models' poor performance on benchmarks.
Melanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi 0001, Asli Celikyilmaz
ICLR2
2025 Self-Consistency Preference Optimization
abstract
Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku.
Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 0014, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Dwivedi-Yu
ICML9
2025 Collaborative Reasoner: Self-Improving Social Agents with Synthetic Conversations
abstract
With increasingly powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where numerous LLM agents work seamlessly with other AI agents and humans to solve complex problems and enhance daily life. To achieve these goals, LLM agents must develop collaborative skills such as effective persuasion, assertion and disagreement, which are often overlooked in the prevalent single-turn training and evaluation of LLMs. In this work, we present Collaborative Reasoner (Coral), a framework to evaluate and improve the collaborative reasoning abilities of language models. In particular, tasks and metrics in Coral necessitate agents to disagree with incorrect solutions, convince their partners of a correct solution, and ultimately agree as a team to commit to a final solution, all through a natural multi-turn conversation. Through comprehensive evaluation on six collaborative reasoning tasks covering domains of coding, math, scientific QA and social reasoning, we show that current models cannot effectively collaborate due to undesirable social behaviors, collapsing even on problems that they can solve singlehandedly. To improve the collaborative reasoning capabilities of LLMs, we propose a self-play method to generate synthetic multi-turn preference data and further train the language models to be better collaborators. Experiments with Llama-3.1, Ministral and Qwen-2.5 models show that our proposed self-improvement approach consistently outperforms finetuned chain-of-thought performance of the same base model, yielding gains up to 16.7% absolute. Human evaluations show that the models exhibit more effective disagreement and produce more natural conversations after training on our synthetic interaction data.
Ansong Ni, Ruta Desai, Xinjie Lei, Jiemin Zhang, Jane Dwivedi-Yu, Ramya Raghavendra, Gargi Ghosh, Shang-Wen Li 0001, Asli Celikyilmaz
NeurIPS7
2025 NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
abstract
Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding.
Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003
NeurIPS2
2024 GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
abstract
State-of-the-art language models can exhibit reasoning refinement capabilities on math, science or coding tasks. However, recent work demonstrates that even the best models struggle to identify *when and where to refine* without access to external feedback. In this paper, we propose Stepwise ORMs (**SORMs**) which are trained, only on synthetic data, to approximate the expected future reward of the optimal policy or $V^{\star}$ as a form of Process-based reward modeling. Our experiments show that SORMs can more accurately detect incorrect reasoning steps compared to ORMs, thus enabling them to give precise step-level feedback to refinement models. We then train *global* refinement models, which take only the question and a draft solution as input and predict a corrected solution, and *local* refinement models which also take as input a critique indicating the location of the first reasoning error. We generate training data for both models synthetically by reusing data used to train the SORM. We find combining global and local refinements, using the ORM as a reranker, significantly outperforms either one individually, as well as a best of three sample baseline. With this strategy we can improve the accuracy of a LLaMA-2 13B model (already fine-tuned with RL) on GSM8K from 53% to 65% when greedily sampled.
Alexander Havrilla, Sharath Chandra, Christoforos Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, Roberta Raileanu
ICML4
2024 Consequences of Conflicts in Online Conversations
abstract
Interpersonal conflicts occur frequently in both offline and online groups, with conditions for conflict especially ripe online. This research attempts to understand the consequences of online group conflict and reporting it to group administrators, both for the protagonists in the conflict and observers. If group conflict is aversive, then group members should reduce their group participation after observing conflict. Theories of imitation and behavioral mimicry suggest that even onlookers will exhibit more conflict and negative language after observing conflict conversations in their group. In contrast, theories of deterrence suggest that both the instigator of the conflict and onlookers will reduce their conflict and onlookers might even increase their engagement if conflicts are reported to group administrators. The current study uses de-identified and aggregated data from Facebook group conversations and Mahalanobis distance matching to test these ideas. Results are consistent with the hypothesis that conflict in group conversations reduces engagement within the group and increases the amount of conflict and the negativity of language users express in the group. However, inconsistent with deterrence theories, conflict and language negativity increase and group engagement decreases when conflict is reported to group administrators.
Kristen M. Altenburger, Robert E. Kraut, Shirley Anugrah Hayati, Jane Dwivedi-Yu, Kaiyan Peng, Yi-Chia Wang
ICWSM4
2023 Using Comments for Predicting the Affective Response to Social Media Posts
abstract
What people see on social media influences their affective state. Predictions of the affective reaction of an audience to a post could help posters creating content and viewers searching for it. This paper examines the value of both real comments and artificially generated ones in predicting the affective responses of an audience. We built an affect prediction model based on Facebook anonymized public posts to predict affective responses (anger, amusement, and sadness affect) as indicated by three Facebook reaction clicks (Angry, Haha, and Sad). Using the content of the original post can predict reactions well (.71 to.87 F1-scores). Adding the text of real post comments improves F1-score by up to 11%. Surprisingly, generated comments improve predictions as much as real comments. These artificial comments were produced using a pre-trained sequence-to-sequence, BART natural language generation model given a post as input. Using artificial comments means that one can predict affect reactions early in the history of a discussion, before anyone has actually commented on a post.
Yi-Chia Wang, Jane Dwivedi-Yu, Robert E. Kraut, Alon Y. Halevy
ACII2
2023 NormBank: A Knowledge Bank of Situational Social Norms
abstract
We present NORMBANK, a knowledge bank of 155k situational norms.This resource is designed to ground flexible normative reasoning for interactive, assistive, and collaborative AI systems.Unlike prior commonsense resources, NORMBANK grounds each inference within a multivalent sociocultural frame, which includes the setting (e.g., restaurant), the agents' contingent roles (waiter, customer), their attributes (age, gender), and other physical, social, and cultural constraints (e.g., the temperature or the country of operation).In total, NORMBANK contains 63k unique constraints from a taxonomy that we introduce and iteratively refine here.Constraints then apply in different combinations to frame social norms.Under these manipulations, norms are non-monotonic -one can cancel an inference by updating its frame even slightly.Still, we find evidence that neural models can help reliably extend the scope and coverage of NORMBANK.We further demonstrate the utility of this resource with a series of transfer experiments.For data and code, see
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
ACL (1)2
2023 ROBBIE: Robust Bias Evaluation of Large Generative Language Models
abstract
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Smith. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
David Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Michael Smith
EMNLP7
2023 Active Retrieval Augmented Generation
abstract
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu 0033, Jane Dwivedi-Yu, Yiming Yang 0002, Jamie Callan, Graham Neubig
EMNLP6
2023 PEER: A Collaborative Language Model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick S. H. Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, Sebastian Riedel 0001
ICLR2
2023 Toolformer: Language Models Can Teach Themselves to Use Tools
abstract
Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with basic functionality, such as arithmetic or factual lookup, where much simpler and smaller specialized models excel. In this paper, we show that LMs can teach themselves to *use external tools* via simple APIs and achieve the best of both worlds. We introduce *Toolformer*, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API. We incorporate a range of tools, including a calculator, a Q&A system, a search engine, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities.
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, Thomas Scialom
NeurIPS2
2023 Atlas: Few-shot Learning with Retrieval Augmented Language Models
abstract
Large language models have shown impressive few-shot results on a wide range of tasks. However, when knowledge is key for such results, as is the case for tasks such as question answering and fact checking, massive parameter counts to store knowledge seem to be needed. Retrieval-augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings. In this work we present Atlas, a carefully designed and pre-trained retrieval-augmented language model able to learn knowledge intensive tasks with very few training examples. We perform evaluations on a wide range of tasks, including MMLU, KILT and Natural Questions, and study the impact of the content of the document index, showing that it can easily be updated. Notably, Atlas reaches over 42% accuracy on Natural Questions using only 64 examples, outperforming a 540B parameter model by 3% despite having 50x fewer parameters.
Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel 0001, Edouard Grave
J. Mach. Learn. Res.7
2023 A fast machine-learning-guided primer design pipeline for selective whole genome amplification
abstract
Addressing many of the major outstanding questions in the fields of microbial evolution and pathogenesis will require analyses of populations of microbial genomes. Although population genomic studies provide the analytical resolution to investigate evolutionary and mechanistic processes at fine spatial and temporal scales-precisely the scales at which these processes occur-microbial population genomic research is currently hindered by the practicalities of obtaining sufficient quantities of the relatively pure microbial genomic DNA necessary for next-generation sequencing. Here we present swga2.0, an optimized and parallelized pipeline to design selective whole genome amplification (SWGA) primer sets. Unlike previous methods, swga2.0 incorporates active and machine learning methods to evaluate the amplification efficacy of individual primers and primer sets. Additionally, swga2.0 optimizes primer set search and evaluation strategies, including parallelization at each stage of the pipeline, to dramatically decrease program runtime. Here we describe the swga2.0 pipeline, including the empirical data used to identify primer and primer set characteristics, that improve amplification performance. Additionally, we evaluate the novel swga2.0 pipeline by designing primer sets that successfully amplify Prevotella melaninogenica, an important component of the lung microbiome in cystic fibrosis patients, from samples dominated by human DNA.
Jane Dwivedi-Yu, Zachary J. Oppler, Matthew W. Mitchell, Yun S. Song, Dustin Brisson
PLoS Comput. Biol.1
2022 The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
abstract
Content Warning: some examples in this paper may be offensive or upsetting.Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user's trust in the moral integrity of the system.Moral deviations are difficult to mitigate because moral judgments are not universal, and there may be multiple competing judgments that apply to a situation simultaneously.In this work, we introduce a new resource, not to authoritatively resolve moral ambiguities, but instead to facilitate systematic understanding of the intuitions, values and moral judgments reflected in the utterances of dialogue systems.The MORAL INTEGRITY CORPUS, MIC , is such a resource, which captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs).Each RoT reflects a particular moral conviction that can explain why a chatbot's reply may appear acceptable or problematic.We further organize RoTs with a set of 9 moral and social attributes and benchmark performance for attribute classification.Most importantly, we show that current neural language models can automatically generate new RoTs that reasonably describe previously unseen interactions, but they still struggle with certain scenarios.Our findings suggest that MIC will be a useful resource for understanding and language models' implicit moral assumptions and flexibly benchmarking the integrity of conversational agents.
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
ACL (1)2
2022 That's so cute!: The CARE Dataset for Affective Response Detection
abstract
Social media plays an increasing role in our communication with friends and family, and in our consumption of entertainment and information.Hence, to design effective ranking functions for posts on social media, it would be useful to predict the affective responses of a post (e.g., whether it is likely to elicit feelings of entertainment, inspiration, or anger).Similar to work on emotion detection (which focuses on the affect of the publisher of the post), the traditional approach to recognizing affective response would involve an expensive investment in human annotation of training data.We create and publicly release CARE db , a dataset of 230k social media post annotations according to seven affective responses using the Common Affective Response Expression (CARE) method.The CARE method is a means of leveraging the signal that is present in comments that are posted in response to a post, providing high-precision evidence about the affective response to the post without human annotation.Unlike human annotation, the annotation process we describe here can be iterated upon to expand the coverage of the method, particularly for new affective responses.We present experiments that demonstrate that the CARE annotations compare favorably with crowdsourced annotations.Finally, we use CARE db to train competitive BERT-based models for predicting affective response as well as emotion detection, demonstrating the utility of the dataset for related tasks.
Jane Dwivedi-Yu, Alon Y. Halevy
CoNLL1
2022 Affective Signals in a Social Media Recommender System
abstract
People come to social media to satisfy a variety of needs, such as being informed, entertained and inspired, or connected to their friends and community. Hence, to design a ranking function that gives useful and personalized post recommendations, it would be helpful to be able to predict the affective response a user may have to a post (e.g., entertained, informed, angered). This paper describes the challenges and solutions we developed to apply Affective Computing to social media recommendation systems.
Jane Dwivedi-Yu, Yi-Chia Wang, Lijing Qin, Cristian Canton, Alon Y. Halevy
KDD1
2022 Quantifying Adaptability in Pre-trained Language Models with 500 Tasks
abstract
Belinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, Jacob Andreas. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Belinda Z. Li, Jane Dwivedi-Yu, Madian Khabsa, Luke Zettlemoyer, Alon Y. Halevy, Jacob Andreas
NAACL-HLT2
2022 Understanding Conflicts in Online Conversations
abstract
With the rise of social media, users from across the world are able to connect and converse with each other online. While these connections have facilitated a growth in knowledge, online discussions can also end in acrimonious conflict. Previous computational studies have focused on creating online conflict detection models from inferred labels, primarily examine disagreement but not acrimony, and do not examine the conflict’s emergence. Social science studies have investigated offline conflict, which can differ from its online form, and rarely examines its emergence. The current research aims to understand how online conflicts arise in online personal conversations. Our ground truth is a Facebook tool that allows group members to report conflict to administrators. We contrast discussions ending with a conflict report with paired non-conflict discussions from the same post. We study both user characteristics (e.g., historical user-to-user interactions) and conversation dynamics (e.g., changes in emotional intensity over the course of the conversation). We use logistic regression to identify the features that predict conflict. User characteristics such as the commenter’s gender and previous involvement in negative online activity are strong indicators of conflict. Conversational dynamics, such as an increase in person-oriented discussion, are also important signals of conflict. These results help us understand how conflicts emerge and suggest better detection models and ways to alert group administrators and members early on to mediate the conversation.
Sharon Levy, Robert E. Kraut, Jane Dwivedi-Yu, Kristen M. Altenburger, Yi-Chia Wang
WWW3
2021 Detecting Inspiring Content on Social Media
abstract
Inspiration moves a person to see new possibilities and transforms the way they perceive their own potential. Inspiration has received little attention in psychology, and has not been researched before in the NLP community. To the best of our knowledge, this work is the first to study inspiration through machine learning methods. We aim to automatically detect inspiring content from social media data. To this end, we analyze social media posts to tease out what makes a post inspiring and what topics are inspiring. We release a dataset of 5,800 inspiring and 5,800 non-inspiring English-language public post unique ids collected from a dump of Reddit public posts made available by a third party and use linguistic heuristics to automatically detect which social media English-language posts are inspiring.
Oana Ignat, Y-Lan Boureau, Jane Dwivedi-Yu, Alon Y. Halevy
ACII3
2016 Estimating Copy Number and Allelic Variation at the Immunoglobulin Heavy Chain Locus Using Short Reads
abstract
The study of genomic regions that contain gene copies and structural variation is a major challenge in modern genomics. Unlike variation involving single nucleotide changes, data on the variation of copy number is difficult to collect and few tools exist for analyzing the variation between individuals. The immunoglobulin heavy variable (IGHV) locus, which plays an integral role in the adaptive immune response, is an example of a complex genomic region that varies in gene copy number. Lack of standard methods to genotype this region prevents it from being included in association studies and is holding back the growing field of antibody repertoire analysis. Here we develop a method that takes short reads from high-throughput sequencing and outputs a genetic profile of the IGHV locus with the read coverage depth and a putative nucleotide sequence for each operationally defined gene cluster. Our operationally defined gene clusters aim to address a major challenge in studying the IGHV locus: the high sequence similarity between gene segments in different genomic locations. Tests on simulated data demonstrate that our approach can accurately determine the presence or absence of a gene cluster from reads as short as 70 bp. More detailed resolution on the copy number of gene clusters can be obtained from read coverage depth using longer reads (e.g., ≥ 100 bp). Detail at the nucleotide resolution of single copy genes (genes present in one copy per haplotype) can be determined with 250 bp reads. For IGHV genes with more than one copy, accurate nucleotide-resolution reconstruction is currently beyond the means of our approach. When applied to a family of European ancestry, our pipeline outputs genotypes that are consistent with the family pedigree, confirms existing multigene variants and suggests new copy number variants. This study paves the way for analyzing population-level patterns of variation in IGHV gene clusters in larger diverse datasets and for quantitatively handling regions of copy number variation in other structurally varying and complex loci.
Shishi Luo, Jane Dwivedi-Yu, Yun S. Song
PLoS Comput. Biol.2