Nitish Gupta

dblp:45/10343 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Language models and text generation · 24% Machine translation · 17% Question answering and dialogue systems · 14%
Network and information security
1 paper
Systems and software security · 33% Hardware security and side channels · 33% Cryptographic primitives and cryptanalysis · 33%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 84% Query processing and optimization · 16%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
neural module network
1.432021
Paired Examples as Indirect Supervision in Latent Decision Models · EMNLP (1) 2021
Neural Module Networks for Reasoning over Text · ICLR 2020
Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM · ICLR 2025
Natural language and speech › Machine translation
low-resource machine translation
0.912025
Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM · ICLR 2025
Natural language and speech › Machine translation › computer-assisted translation
post-editing
0.912025
Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM · ICLR 2025
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
compositional question answering
0.822021
Paired Examples as Indirect Supervision in Latent Decision Models · EMNLP (1) 2021
Neural Compositional Denotational Semantics for Question Answering · EMNLP 2018
Machine learning › Deep learning architectures and training › attention mechanism
cross-attention
0.812024
LLM Augmented LLMs: Expanding Capabilities through Composition · ICLR 2024
Natural language and speech › Language models and text generation › large language model
large language model augmentation
0.812024
LLM Augmented LLMs: Expanding Capabilities through Composition · ICLR 2024
Machine learning › Efficient and distributed learning
model composition
0.812024
LLM Augmented LLMs: Expanding Capabilities through Composition · ICLR 2024
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.812024
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages · ACL (1) 2024
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.812024
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages · ACL (1) 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.712023
QA Is the New KR: Question-Answer Pairs as Knowledge Bases · AAAI 2023
Natural language and speech › Question answering and dialogue systems
multi-hop reasoning
0.712023
QA Is the New KR: Question-Answer Pairs as Knowledge Bases · AAAI 2023
Data mining › text mining › text classification
sensitivity classification
0.712023
Microsoft Purview: A System for Central Governance of Data · Proc. VLDB Endow. 2023
Natural language and speech › Information extraction and text analysis
entity linking
0.622018
Joint Multilingual Supervision for Cross-lingual Entity Linking · EMNLP 2018
Entity Linking via Joint Encoding of Types, Descriptions, and Context · EMNLP 2017
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
0.412020
Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation faithfulness
0.412020
Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Robust Named Entity Recognition with Truecasing Pretraining · AAAI 2020
Natural language and speech › Language models and text generation › text representation
syntactic representation
0.412020
Overestimation of Syntactic Representation in Neural Language Models · ACL 2020
Natural language and speech › Language models and text generation › natural language reasoning
textual reasoning
0.412020
Neural Module Networks for Reasoning over Text · ICLR 2020
Systems and software security › database security
database encryption
0.412020
Azure SQL Database Always Encrypted · SIGMOD Conference 2020
Hardware security and side channels
trusted execution environments
0.412020
Azure SQL Database Always Encrypted · SIGMOD Conference 2020
Natural language and speech › Information extraction and text analysis › entity linking
cross-lingual entity linking
0.312018
Joint Multilingual Supervision for Cross-lingual Entity Linking · EMNLP 2018
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.312018
Neural Compositional Denotational Semantics for Question Answering · EMNLP 2018
Bioinformatics and computational biology › sequence analysis
read mapping
0.212016
RapMap: a rapid, sensitive and accurate tool for mapping RNA-seq reads to transcriptomes · Bioinform. 2016
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis
0.212016
RapMap: a rapid, sensitive and accurate tool for mapping RNA-seq reads to transcriptomes · Bioinform. 2016
Bioinformatics and computational biology
transcriptomics
0.212016
RapMap: a rapid, sensitive and accurate tool for mapping RNA-seq reads to transcriptomes · Bioinform. 2016
Bioinformatics and computational biology › transcriptomics
transcript quantification
0.212016
RapMap: a rapid, sensitive and accurate tool for mapping RNA-seq reads to transcriptomes · Bioinform. 2016
Natural language and speech › Language models and text generation › text summarization › multilingual summarization
cross-lingual summarization
0.212024
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems
question generation
0.112021
Paired Examples as Indirect Supervision in Latent Decision Models · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

large language model · 1.6policy authoring · 1.3automated data scanning · 1.3neural module networks · 0.9knowledge distillation · 0.9instruction tuning · 0.9trusted execution environment · 0.9column granularity encryption · 0.9fine-tuning · 0.8cross-attention · 0.8benchmarking · 0.8question generation · 0.7entity linking · 0.7consistency objective · 0.5parse chart · 0.3gradient descent · 0.3end-to-end differentiable model · 0.3quasi-mapping · 0.2
YearPublicationVenuePosition
2025 Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM
abstract
Multilingual large language models (LLMs) are great translators, but this is largely limited to high-resource languages. For many LLMs, translating in and out of low-resource languages remains a challenging task. To maximize data efficiency in this low-resource setting, we introduce Mufu, which includes a selection of automatically generated multilingual candidates and an instruction to correct inaccurate translations in the prompt. Mufu prompts turn a translation task into a postediting one, and seek to harness the LLM’s reasoning capability with auxiliary translation candidates, from which the model is required to assess the input quality, align the semantics cross-lingually, copy from relevant inputs and override instances that are incorrect. Our experiments on En-XX translations over the Flores-200 dataset show LLMs finetuned against Mufu-style prompts are robust to poor quality auxiliary translation candidates, achieving performance superior to NLLB 1.3B distilled model in 64% of low- and very-low-resource language pairs. We then distill these models to reduce inference cost, while maintaining on average 3.1 chrF improvement over finetune-only baseline in low-resource translations.
Zheng Wei Lim, Nitish Gupta, Honglin Yu, Trevor Cohn
ICLR2
2024 IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
abstract
As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world.India is a linguistically diverse country of 1.4 Billion people.To facilitate research on multilingual LLM evaluation, we release INDICGENBENCHthe largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts and 4 language families.INDICGEN-BENCH is composed of diverse generation tasks like cross-lingual summarization, machine translation, and cross-lingual question answering.INDICGENBENCH extends existing benchmarks to many Indic languages through human curation providing multi-way parallel evaluation data for many under-represented Indic languages for the first time.We evaluate a wide range of proprietary and open-source LLMs including GPT-3.5, GPT-4, PaLM-2, mT5, Gemma, BLOOM and LLaMA on IN-DICGENBENCH in a variety of settings.The largest PaLM-2 models performs the best on most tasks, however, there is a significant performance gap in all languages compared to English showing that further research is needed for the development of more inclusive multilingual language models.INDICGENBENCH is available at www.github.com/google-research- datasets/indic-gen-bench 2 INDICGENBENCH INDICGENBENCH is a high-quality, humancurated benchmark to evaluate text generation capabilities of multilingual models on Indic languages.Our benchmark consists of 5 user-facing tasks (viz., summarization, machine translation, and question answering) across 29 Indic languages spanning 13 writing scripts and 4 language families.For certain tasks, INDICGENBENCH provides the first-ever evaluation dataset for up to 18 Indic languages.Table 1 provides summary of INDICGENBENCH and examples of instances across tasks present in it.Languages in INDICGENBENCH are divided into (relatively) Higher, Medium, and Low resource categories based on the availability of web text resources (see appendix §A for details).
Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, Partha Talukdar
ACL (1)2
2024 LLM Augmented LLMs: Expanding Capabilities through Composition
abstract
Foundational models with billions of parameters which have been trained on large corpus of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to augment them or impart new skills. On the other hand, due to their adaptation abilities,several new instances of these models are being trained towards new domains and tasks. In this work, we study the problem of efficient and practical composition of existing foundation models with more specific models to enable newer capabilities. To this end, we propose CALM—Composition to Augment Language Models—which introduces cross-attention between models to compose their representations and enable new capabilities. Salient features of CALM are: (i) Scales up LLMs on new tasks by ‘re-using’ existing LLMs along with a few additional parameters and data, (ii) Existing model weights are kept intact, and hence preserves existing capabilities, and (iii) Applies to diverse domains and settings. We illustrate that augmenting PaLM2-S with a smaller model trained on low-resource languages results in an absolute improvement of up to 13% on tasks like translation into English and arithmetic reasoning for low-resource languages. Similarly,when PaLM2-S is augmented with a code-specific model, we see a relative improvement of 40% over the base model for code generation and explanation tasks—on-par with fully fine-tuned counterparts.
Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Sriram Ganapathy, Abhishek Bapna, Partha Talukdar
ICLR4
2023 QA Is the New KR: Question-Answer Pairs as Knowledge Bases
abstract
We propose a new knowledge representation (KR) based on knowledge bases (KBs) derived from text, based on question generation and entity linking. We argue that the proposed type of KB has many of the key advantages of a traditional symbolic KB: in particular, it consists of small modular components, which can be combined compositionally to answer complex queries, including relational queries and queries involving ``multi-hop'' inferences. However, unlike a traditional KB, this information store is well-aligned with common user information needs. We present one such KB, called a QEDB, and give qualitative evidence that the atomic components are high-quality and meaningful, and that atomic components can be combined in ways similar to the triples in a symbolic KB. We also show experimentally that questions reflective of typical user questions are more easily answered with a QEDB than a symbolic KB.
William W. Cohen, Wenhu Chen, Michiel de Jong, Nitish Gupta, Alessandro Presta, Patrick Verga, John Wieting
AAAI4
2023 Bootstrapping Multilingual Semantic Parsers using Large Language Models
abstract
Abhijeet Awasthi, Nitish Gupta, Bidisha Samanta, Shachi Dave, Sunita Sarawagi, Partha Talukdar. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Abhijeet Awasthi, Nitish Gupta, Bidisha Samanta, Shachi Dave, Sunita Sarawagi, Partha Talukdar
EACL2
2023 Event Linking: Grounding Event Mentions to Wikipedia
abstract
Comprehending an article requires understanding its constituent events.However, the context where an event is mentioned often lacks the details of this event.A question arises: how can the reader obtain more knowledge about this particular event in addition to what is provided by the local context in the article?This work defines Event Linking, a new natural language understanding task at the event level.Event linking tries to link an event mention appearing in an article to the most appropriate Wikipedia page.This page is expected to provide rich knowledge about what the event mention refers to.To standardize the research in this new direction, we contribute in fourfold.First, this is the first work in the community that formally defines the Event Linking task.Second, we collect a dataset for this new task.Specifically, we automatically gather the training set from Wikipedia, and then create two evaluation sets: one from the Wikipedia domain, reporting the in-domain performance, and a second from the real-world news domain, to evaluate out-of-domain performance.Third, we retrain and evaluate two state-of-theart (SOTA) entity linking models, showing the challenges of event linking, and we propose an event-specific linking system, EVELINK, to set a competitive result for the new task.Fourth, we conduct a detailed and insightful analysis to help understand the task and the limitations of the current model.Overall, as our analysis shows, Event Linking is a challenging and essential task requiring more effort from the community.1
Xiaodong Yu 0003, Wenpeng Yin 0001, Nitish Gupta, Dan Roth 0001
EACL3
2023 Microsoft Purview: A System for Central Governance of Data
abstract
Modern data estates are spread across data located on premises, on the edge and in one or more public clouds, spread across various sources like multiple relational databases, file and storage systems, and no-SQL systems, both operational and analytic; this phenomenon is referred to as data sprawl. Data administrators who wish to enforce compliance across the entire organization have to inventory their data, identify what parts of it are sensitive, and govern the sensitive data appropriately --- across the entirety of their sprawling data estate. Today, governance of data is completely siloed; each of the data subsystems has its own (and varied) governance features. Policies applied to sensitive data are applied piece-meal by iterating over all the data sources in a custom language specific to each source. This makes data governance cumbersome, error-prone (because a given policy must be manually enforced across different subsystems, inconsistencies can easily arise), and expensive. This paper presents Microsoft Purview , a service for unified governance of the entire data estate of an organization from a single central pane of glass. The Purview service consists of three parts: (1) a Data Map or metadata catalog that is populated by automated scanning of data sources in the organization, (2) a system to store and manage sensitivity classification of data, and (3) a policy system that enables data security officers to author and implement policies that span the entire organization, e.g., a policy that says, "Non-full-time employees should be denied access to data classified as PII (Personally Identifiable Information.") Purview transforms data governance across a complex data estate by offering the ability to govern centrally and automating data discovery, classification and policy enforcement. While other commercial catalog systems also build a global catalog, Purview is unique in its support for policies. It is also distinguished by covering both structured and unstructured data, thanks to its deep integration with Office 365 and its governance framework; indeed, "Microsoft Purview" represents a new unified offering that combines Office 365 governance and what was formerly a service for governing structured data called "Azure Purview". By integrating with Office 365's Rights Management Service, Purview offers central governance over structured data stored in databases and stores, reports in systems such as Power BI, as well as document data stored in Office 365. The Purview vision is to make the metadata in the Data Map increasingly richer through further automation and curation support and to use this 360 degree view of the data estate to support a wide range of governance policies, ranging from access control to lifecycle management (e.g., retention, deletion, restricting data movement). This paper covers the design and implementation challenges in building the Purview service for Attribute-Based Access Control (ABAC) policies, focusing specifically on a detailed description of its integration with Azure SQL Database. We illustrate the power of unifying Office 365 governance with structured data governance through Purview policies that enforce consistent access control even as data flows between Office 365 and structured data engines like Azure SQL Database. We also describe the results of our empirical evaluation of the performance overheads imposed by Purview.
Shafi Ahmad, Dillidorai Arumugam, Srdan Bozovic, Elnata Degefa, Sailesh Duvvuri, Steven Gott, Nitish Gupta, Joachim Hammer, Nivedita Kaluskar, Raghav Kaushik, Rakesh Khanduja, Prasad Mujumdar, Gaurav Malhotra, Pankaj Naik, Nikolas Ogg, Krishna Kumar Parthasarthy, Raghu Ramakrishnan 0001, Vlad Rodriguez, Rahul Sharma 0011, Jakub Szymaszek, Andreas Wolter
Proc. VLDB Endow.7
2021 Paired Examples as Indirect Supervision in Latent Decision Models
abstract
Compositional, structured models are appealing because they explicitly decompose problems and provide interpretable intermediate outputs that give confidence that the model is not simply latching onto data artifacts.Learning these models is challenging, however, because end-task supervision only provides a weak indirect signal on what values the latent decisions should take.This often results in the model failing to learn to perform the intermediate tasks correctly.In this work, we introduce a way to leverage paired examples that provide stronger cues for learning latent decisions.When two related training examples share internal substructure, we add an additional training objective to encourage consistency between their latent decisions.Such an objective does not require external supervision for the values of the latent output, or even the end task, yet provides an additional training signal to that provided by individual training examples themselves.We apply our method to improve compositional question answering using neural module networks on the DROP dataset.We explore three ways to acquire paired questions in DROP: (a) discovering naturally occurring paired examples within the dataset, (b) constructing paired examples using templates, and (c) generating paired examples using a question generation model.We empirically demonstrate that our proposed approach improves both in-and outof-distribution generalization and leads to correct latent decision predictions.
Nitish Gupta, Sameer Singh 0001, Matt Gardner 0001, Dan Roth 0001
EMNLP (1)1
2020 Robust Named Entity Recognition with Truecasing Pretraining
abstract
Although modern named entity recognition (NER) systems show impressive performance on standard datasets, they perform poorly when presented with noisy data. In particular, capitalization is a strong signal for entities in many languages, and even state of the art models overfit to this feature, with drastically lower performance on uncapitalized text. In this work, we address the problem of robustness of NER systems in data with noisy or uncertain casing, using a pretraining objective that predicts casing in text, or a truecaser, leveraging unlabeled data. The pretrained truecaser is combined with a standard BiLSTM-CRF model for NER by appending output distributions to character embeddings. In experiments over several datasets of varying domain and casing quality, we show that our new model improves performance in uncased text, even adding value to uncased BERT embeddings. Our method achieves a new state of the art on the WNUT17 shared task dataset.
Stephen Mayhew 0001, Nitish Gupta, Dan Roth 0001
AAAI2
2020 Overestimation of Syntactic Representation in Neural Language Models
abstract
With the advent of powerful neural language models over the last few years, research attention has increasingly focused on what aspects of language they represent that make them so successful.Several testing methodologies have been developed to probe models' syntactic representations.One popular method for determining a model's ability to induce syntactic structure trains a model on strings generated according to a template then tests the model's ability to distinguish such strings from superficially similar ones with different syntax.We illustrate a fundamental problem with this approach by reproducing positive results from a recent paper with two non-syntactic baseline language models: an n-gram model and an LSTM model trained on scrambled inputs.
Jordan Kodner, Nitish Gupta
ACL2
2020 Obtaining Faithful Interpretations from Compositional Neural Networks
abstract
Neural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy.
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001
ACL3
2020 Neural Module Networks for Reasoning over Text
Nitish Gupta, Dan Roth 0001, Sameer Singh 0001, Matt Gardner 0001
ICLR1
2020 Azure SQL Database Always Encrypted
abstract
This paper presents Always Encrypted, a recently released feature of Microsoft SQL Server that uses column granularity encryption to provide cryptographic data protection guarantees. Always Encrypted can be used to outsource database administration while keeping the data confidential from an administrator, including cloud operators. The first version of Always Encrypted was released in Azure SQL Database and as part of SQL Server 2016, and supported equality operations over deterministically encrypted columns. The second version, released as part of SQL Server 2019, uses an enclave running within a trusted execution environment to provide richer functionality that includes comparison and string pattern matching for an IND-CPA-secure (randomized) encryption scheme. We present the security, functionality, and design of Always Encrypted, and provide a performance evaluation using the TPC-C benchmark.
Panagiotis Antonopoulos, Arvind Arasu, Kunal D. Singh, Kenneth Eguro, Nitish Gupta, Rajat Jain, Raghav Kaushik, Hanuma Kodavalla, Donald Kossmann, Nikolas Ogg, Ravishankar Ramamurthy, Jakub Szymaszek, Jeffrey Trimmer, Kapil Vaswani, Ramarathnam Venkatesan, Mike Zwilling
SIGMOD Conference5
2018 Neural Compositional Denotational Semantics for Question Answering
abstract
Answering compositional questions requiring multi-step reasoning is challenging.We introduce an end-to-end differentiable model for interpreting questions about a knowledge graph (KG), which is inspired by formal approaches to semantics.Each span of text is represented by a denotation in a KG and a vector that captures ungrounded aspects of meaning.Learned composition modules recursively combine constituent spans, culminating in a grounding for the complete sentence which answers the question.For example, to interpret "not green", the model represents "green" as a set of KG entities and "not" as a trainable ungrounded vector-and then uses this vector to parameterize a composition function that performs a complement operation.For each sentence, we build a parse chart subsuming all possible parses, allowing the model to jointly learn both the composition operators and output structure by gradient descent from endtask supervision.The model learns a variety of challenging semantic operators, such as quantifiers, disjunctions and composed relations, and infers latent syntactic structure.It also generalizes well to longer questions than seen in its training data, in contrast to RNN, its treebased variants, and semantic parsing baselines.
Nitish Gupta, Mike Lewis
EMNLP1
2018 Joint Multilingual Supervision for Cross-lingual Entity Linking
abstract
Cross-lingual Entity Linking (XEL) aims to ground entity mentions written in any language to an English Knowledge Base (KB), such as Wikipedia.XEL for most languages is challenging, owing to limited availability of resources as supervision.We address this challenge by developing the first XEL approach that combines supervision from multiple languages jointly.This enables our approach to: (a) augment the limited supervision in the target language with additional supervision from a high-resource language (like English), and (b) train a single entity linking model for multiple languages, improving upon individually trained models for each language.Extensive evaluation on three benchmark datasets across 8 languages shows that our approach significantly improves over the current state-of-theart.We also provide analyses in two limited resource settings: (a) zero-shot setting, when no supervision in the target language is available, and in (b) low-resource setting, when some supervision in the target language is available.Our analysis provides insights into the limitations of zero-shot XEL approaches in realistic scenarios, and shows the value of joint supervision in low-resource settings.1
Shyam Upadhyay, Nitish Gupta, Dan Roth 0001
EMNLP2
2017 Entity Linking via Joint Encoding of Types, Descriptions, and Context
abstract
For accurate entity linking, we need to capture various information aspects of an entity, such as its description in a KB, contexts in which it is mentioned, and structured knowledge.Additionally, a linking system should work on texts from different domains without requiring domain-specific training data or hand-engineered features.In this work we present a neural, modular entity linking system that learns a unified dense representation for each entity using multiple sources of information, such as its description, contexts around its mentions, and its fine-grained types.We show that the resulting entity linking system is effective at combining these sources, and performs competitively, sometimes out-performing current state-of-theart systems across datasets, without requiring any domain-specific training data or hand-engineered features.We also show that our model can effectively "embed" entities that are new to the KB, and is able to link its mentions accurately.
Nitish Gupta, Sameer Singh 0001, Dan Roth 0001
EMNLP1
2016 Revisiting the Evaluation for Cross Document Event Coreference
abstract
Cross document event coreference (CDEC) is an important task that aims at aggregating event-related information across multiple documents. We revisit the evaluation for CDEC, and discover that past works have adopted different, often inconsistent, evaluation settings, which either overlook certain mistakes in coreference decisions, or make assumptions that simplify the coreference task considerably. We suggest a new evaluation methodology which overcomes these limitations, and allows for an accurate assessment of CDEC systems. Our new evaluation setting better reflects the corpus-wide information aggregation ability of CDEC systems by separating event-coreference decisions made across documents from those made within a document. In addition, we suggest a better baseline for the task and semi-automatically identify several inconsistent annotations in the evaluation dataset.
Shyam Upadhyay, Nitish Gupta, Christos Christodoulopoulos 0001, Dan Roth 0001
COLING2
2016 RapMap: a rapid, sensitive and accurate tool for mapping RNA-seq reads to transcriptomes
abstract
MOTIVATION: The alignment of sequencing reads to a transcriptome is a common and important step in many RNA-seq analysis tasks. When aligning RNA-seq reads directly to a transcriptome (as is common in the de novo setting or when a trusted reference annotation is available), care must be taken to report the potentially large number of multi-mapping locations per read. This can pose a substantial computational burden for existing aligners, and can considerably slow downstream analysis. RESULTS: We introduce a novel concept, quasi-mapping, and an efficient algorithm implementing this approach for mapping sequencing reads to a transcriptome. By attempting only to report the potential loci of origin of a sequencing read, and not the base-to-base alignment by which it derives from the reference, RapMap-our tool implementing quasi-mapping-is capable of mapping sequencing reads to a target transcriptome substantially faster than existing alignment tools. The algorithm we use to implement quasi-mapping uses several efficient data structures and takes advantage of the special structure of shared sequence prevalent in transcriptomes to rapidly provide highly-accurate mapping information. We demonstrate how quasi-mapping can be successfully applied to the problems of transcript-level quantification from RNA-seq reads and the clustering of contigs from de novo assembled transcriptomes into biologically meaningful groups. AVAILABILITY AND IMPLEMENTATION: RapMap is implemented in C ++11 and is available as open-source software, under GPL v3, at https://github.com/COMBINE-lab/RapMap CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Avi Srivastava, Hirak Sarkar, Nitish Gupta, Rob Patro
Bioinform.3
2015 DNA: An SDN framework for Distributed Network Analytics (Demo Paper)
abstract
Analytics of network telemetry data helps address many important operational problems. Traditional Big Data approaches run into limitations even as they push scale boundaries for processing data further. One reason for this is the fact that in many cases, the bottleneck for analytics is not analytics processing itself but the generation and export of the data on which analytics depends. The amount of data that can be reasonably collected from the network runs into inherent limitations due to bandwidth and processing constraints in the network itself. In addition, management tasks related to determining and configuring which data to generate lead to significant deployment challenges. In order to address these issues, we propose a novel distributed solution to network analytics that we have implemented as a proof of concept, called DNA (Distributed Network Analytics). In DNA, analytics processing is performed at the source of the data by specialized agents embedded within network devices, which also dynamically set up and reconfigure telemetry data sources as required by an analytics task. An SDN controller application orchestrates network analytics tasks across the network to allow users to interact with the network as a whole instead of individual devices one at a time. Our demonstration of DNA includes a GUI front end used to specify and monitor analytics tasks as well as visualize analytics results.
Alexander Clemm, Mouli Chandramouli, Nitish Gupta, Robert Lerche, Ashwin Pankaj, Manjunath Patil, Ganesan Rajam, V. Anbalagan, Joe Zhang
IM3
2011 Artificial intelligence for mixed pixel resolution
abstract
Mixed pixels are usually the biggest reason for lowered success in classification accuracy. Aiming at the characteristics of remote sensing image classification, the mixed pixel problem is one of the main factors that affect the improvement of classification precision in image. How to decompose the mixed pixels precisely and effectively for multispectral/hyper spectral remote sensing images is a critical issue for the quantitative research. As Remote sensing data is widely used for the classification of types of land cover such as vegetation, water body thus Conflicts are one of the most characteristic attributes in satellite multilayer imagery. Conflict occurs in tagging class label to mixed pixels that encompass spectral response of different land cover on the ground element. In this paper we attempted to present a new approach for resolving the mixed pixels using Biogeography based optimization. The paper deals with the idea of tagging the mixed pixel to a particular class by finding the best suitable class for it using the concept of immigration and emigration.
Nitish Gupta
IGARSS1