Siddhant Garg

dblp:82/8467 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0004-0219-511XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Memory-QA: Answering Recall Questions Based on Multimodal Memories
abstract
Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiun-Zu Kuo, Jiayang Xu, Aaron Colak, Xin Luna Dong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hongda Jiang, Siddhant Garg, Rishab Arora, Shiunzu Kuo, Aaron Colak, Xin Dong 0001
EMNLP3
2024 Towards Improved Multi-Source Attribution for Long-Form Answer Generation
abstract
Nilay Patel, Shivashankar Subramanian, Siddhant Garg, Pratyay Banerjee, Amita Misra. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Nilay Patel, Shivashankar Subramanian, Siddhant Garg, Pratyay Banerjee, Amita Misra
NAACL-HLT3
2023 Learning Answer Generation using Supervision from Automatic Question Answering Evaluators
abstract
Recent studies show that sentence-level extractive QA, i.e., based on Answer Sentence Selection (AS2), is outperformed by Generationbased QA (GenQA) models, which generate answers using the top-k answer sentences ranked by AS2 models (a la retrieval-augmented generation style).In this paper, we propose a novel training paradigm for GenQA using supervision from automatic QA evaluation models (GAVA).Specifically, we propose three strategies to transfer knowledge from these QA evaluation models to a GenQA model: (i) augmenting training data with answers generated by the GenQA model and labelled by GAVA (either statically, before training, or (ii) dynamically, at every training epoch); and (iii) using the GAVA score for weighting the generator loss during the learning of the GenQA model.We evaluate our proposed methods on two academic and one industrial dataset, obtaining a significant improvement in answering accuracy over the previous state of the art.
Matteo Gabburo, Siddhant Garg, Rik Koncel-Kedziorski, Alessandro Moschitti
ACL (1)2
2023 Cross-Shape Attention for Part Segmentation of 3D Point Clouds
abstract
Abstract We present a deep learning method that propagates point‐wise feature representations across shapes within a collection for the purpose of 3D shape segmentation. We propose a cross‐shape attention mechanism to enable interactions between a shape's point‐wise features and those of other shapes. The mechanism assesses both the degree of interaction between points and also mediates feature propagation across shapes, improving the accuracy and consistency of the resulting point‐wise feature representations for shape segmentation. Our method also proposes a shape retrieval measure to select suitable shapes for cross‐shape attention operations for each test shape. Our experiments demonstrate that our approach yields state‐of‐the‐art results in the popular PartNet dataset.
Marios Loizou, Siddhant Garg, Dmitry Petrov, Melinos Averkiou, Evangelos Kalogerakis
Comput. Graph. Forum2
2022 Knowledge Transfer from Answer Ranking to Answer Generation
abstract
Recent studies show that Question Answering (QA) based on Answer Sentence Selection (AS2) can be improved by generating an improved answer from the top-k ranked answer sentences (termed GenQA).This allows for synthesizing the information from multiple candidates into a concise, natural-sounding answer.However, creating large-scale supervised training data for GenQA models is very challenging.In this paper, we propose to train a GenQA model by transferring knowledge from a trained AS2 model, to overcome the aforementioned issue.First, we use an AS2 model to produce a ranking over answer candidates for a set of questions.Then, we use the top ranked candidate as the generation target, and the next k top ranked candidates as context for training a GenQA model.We also propose to use the AS2 model prediction scores for loss weighting and score-conditioned input/output shaping, to aid the knowledge transfer.Our evaluation on three public and one large industrial datasets demonstrates the superiority of our approach over the AS2 baseline, and GenQA trained using supervised data.
Matteo Gabburo, Rik Koncel-Kedziorski, Siddhant Garg, Luca Soldaini, Alessandro Moschitti
EMNLP3
2022 Pre-training Transformer Models with Sentence-Level Objectives for Answer Sentence Selection
abstract
An important task for designing QA systems is answer sentence selection (AS2): selecting the sentence containing (or constituting) the answer to a question from a set of retrieved relevant documents.In this paper, we propose three novel sentence-level transformer pre-training objectives that incorporate paragraph-level semantics within and across documents, to improve the performance of transformers for AS2, and mitigate the requirement of large labeled datasets.Specifically, the model is tasked to predict whether: (i) two sentences are extracted from the same paragraph, (ii) a given sentence is extracted from a given paragraph, and (iii) two paragraphs are extracted from the same document.Our experiments on three public and one industrial AS2 datasets demonstrate the empirical superiority of our pre-trained transformers over baseline models such as RoBERTa and ELECTRA for AS2.
Luca Di Liello, Siddhant Garg, Luca Soldaini, Alessandro Moschitti
EMNLP2
2022 Paragraph-based Transformer Pre-training for Multi-Sentence Inference
abstract
Luca Di Liello, Siddhant Garg, Luca Soldaini, Alessandro Moschitti. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Luca Di Liello, Siddhant Garg, Luca Soldaini, Alessandro Moschitti
NAACL-HLT2
2021 A Simple Approach to Image Tilt Correction with Self-Attention MobileNet for Smartphones
Siddhant Garg, Debi Prasanna Mohanty, Siva Prasad Thota, Sukumar Moharana
BMVC1
2021 Towards Robustness to Label Noise in Text Classification via Noise Modeling
abstract
Large datasets in NLP suffer from noisy labels, due to erroneous automatic and human annotation procedures. We study the problem of text classification with label noise, and aim to capture this noise through an auxiliary noise model over the classifier. We first assign a probability score to each training sample of having a noisy label, through a beta mixture model fitted on the losses at an early epoch of training. Then, we use this score to selectively guide the learning of the noise model and classifier. Our empirical evaluation on two text classification tasks shows that our approach can improve over the baseline accuracy, and prevent over-fitting to the noise.
Siddhant Garg, Goutham Ramakrishnan, Varun Thumbe
CIKM1
2021 Will this Question be Answered? Question Filtering via Answer Model Distillation for Efficient Question Answering
abstract
In this paper we propose a novel approach towards improving the efficiency of Question Answering (QA) systems by filtering out questions that will not be answered by them.This is based on an interesting new finding: the answer confidence scores of state-of-the-art QA systems can be approximated well by models solely using the input question text.This enables preemptive filtering of questions that are not answered by the system due to their answer confidence scores being lower than the system threshold.Specifically, we learn Transformer-based question models by distilling Transformer-based answering models.Our experiments on three popular QA datasets and one industrial QA benchmark demonstrate the ability of our question models to approximate the Precision/Recall curves of the target QA system well.These question models, when used as filters, can effectively trade off lower computation cost of QA systems for lower Recall, e.g., reducing computation by ∼60%, while only losing ∼3-4% of Recall.
Siddhant Garg, Alessandro Moschitti
EMNLP (1)1
2020 TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection
abstract
We propose TandA, an effective technique for fine-tuning pre-trained Transformer models for natural language tasks. Specifically, we first transfer a pre-trained model into a model for a general task by fine-tuning it with a large and high-quality dataset. We then perform a second fine-tuning step to adapt the transferred model to the target domain. We demonstrate the benefits of our approach for answer sentence selection, which is a well-known inference task in Question Answering. We built a large scale dataset to enable the transfer step, exploiting the Natural Questions dataset. Our approach establishes the state of the art on two well-known benchmarks, WikiQA and TREC-QA, achieving the impressive MAP scores of 92% and 94.3%, respectively, which largely outperform the the highest scores of 83.4% and 87.5% of previous work. We empirically show that TandA generates more stable and robust models reducing the effort required for selecting optimal hyper-parameters. Additionally, we show that the transfer step of TandA makes the adaptation step more robust to noise. This enables a more effective use of noisy datasets for fine-tuning. Finally, we also confirm the positive impact of TandA in an industrial setting, using domain specific datasets subject to different types of noise.
Siddhant Garg, Thuy Vu, Alessandro Moschitti
AAAI1
2020 Can Adversarial Weight Perturbations Inject Neural Backdoors
abstract
Adversarial machine learning has exposed several security hazards of neural models. Thus far, the concept of an "adversarial perturbation" has exclusively been used with reference to the input space referring to a small, imperceptible change which can cause a ML model to err. In this work we extend the idea of "adversarial perturbations" to the space of model weights, specifically to inject backdoors in trained DNNs, which exposes a security risk of publicly available trained models. Here, injecting a backdoor refers to obtaining a desired outcome from the model when a trigger pattern is added to the input, while retaining the original predictions on a non-triggered input. From the perspective of an adversary, we characterize these adversarial perturbations to be constrained within an ℓ∞ norm around the original model weights. We introduce adversarial perturbations in model weights using a composite loss on the predictions of the original model and the desired trigger through projected gradient descent. Our results show that backdoors can be successfully injected with a very small average relative change in model weight values for several CV and NLP applications.
Siddhant Garg, Adarsh Kumar 0001, Vibhor Goel, Yingyu Liang
CIKM1
2020 BAE: BERT-based Adversarial Examples for Text Classification
abstract
Modern text classification models are susceptible to adversarial examples, perturbed versions of the original text indiscernible by humans which get misclassified by the model.Recent works in NLP use rule-based synonym replacement strategies to generate adversarial examples.These strategies can lead to outof-context and unnaturally complex token replacements, which are easily identifiable by humans.We present BAE, a black box attack for generating adversarial examples using contextual perturbations from a BERT masked language model.BAE replaces and inserts tokens in the original text by masking a portion of the text and leveraging the BERT-MLM to generate alternatives for the masked tokens.Through automatic and human evaluations, we show that BAE performs a stronger attack, in addition to generating adversarial examples with improved grammaticality and semantic coherence as compared to prior work.
Siddhant Garg, Goutham Ramakrishnan
EMNLP (1)1
2020 Functional Regularization for Representation Learning: A Unified Theoretical Perspective
abstract
Unsupervised and self-supervised learning approaches have become a crucial tool to learn representations for downstream prediction tasks. While these approaches are widely used in practice and achieve impressive empirical gains, their theoretical understanding largely lags behind. Towards bridging this gap, we present a unifying perspective where several such approaches can be viewed as imposing a regularization on the representation via a learnable function using unlabeled data. We propose a discriminative theoretical framework for analyzing the sample complexity of these approaches, which generalizes the framework of (Balcan and Blum, 2010) to allow learnable regularization functions. Our sample complexity bounds show that, with carefully chosen hypothesis classes to exploit the structure in the data, these learnable regularization functions can prune the hypothesis space, and help reduce the amount of labeled data needed. We then provide two concrete examples of functional regularization, one using auto-encoders and the other using masked self-supervision, and apply our framework to quantify the reduction in the sample complexity bound of labeled data. We also provide complementary empirical results to support our analysis.
Siddhant Garg, Yingyu Liang
NeurIPS1
2018 Surprisingly Easy Hard-Attention for Sequence to Sequence Learning
abstract
In this paper we show that a simple beam approximation of the joint distribution between attention and output is an easy, accurate, and efficient attention mechanism for sequence to sequence learning.The method combines the advantage of sharp focus in hard attention and the implementation ease of soft attention.On five translation and two morphological inflection tasks we show effortless and consistent gains in BLEU compared to existing attention mechanisms.
Shiv Shankar, Siddhant Garg, Sunita Sarawagi
EMNLP2