Prateek Sircar

dblp:326/4106 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-9152-9238ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CSMAD: Hallucination Detection via Multi-Agent Debate with NLI-Verified Contradictory Statements
abstract
Large Language Models (LLMs) are prone to hallucinations, producing fluent but factually incorrect statements. Recent multi-agent debate methods improve hallucination detection by jointly improving reasoning and decision-making. However, existing approaches either collaborate which amplifies shared overconfidence, or adopt adversarial preset stances, that can inject incorrect information complicating decision making. To address this, we propose Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim. CSMAD asks independent agents to evaluate the claim and the contradictory claim, which encourages different lines of reasoning without assigning preset stances. When the outcome is non-discriminative; both the contradictory statements are either accepted or rejected; the agents exchange rationales and update their judgments after considering opposing evidence. A final judge then decides the truth of the original claim, using both arguments as context. To make contradictory statement generation reliable, we add a Natural Language Inference (NLI) based verifier that checks whether the generated statement actually contradicts the original claim; if it does not, the system falls back to an explicit negation-based contradiction. Across public benchmarks for question answering and scientific claim verification, as well as a proprietary e-commerce claims dataset, we show that CSMAD consistently outperforms the strongest baseline for both large (Claude-3.5 Sonnet) and medium-sized (Qwen3-8B) language models, improving F1 by +2.3 and +4.1 points, respectively, while reducing LLM token cost by 28%.
Swapnil Gupta, Akshay Verma, Khushi Gupta, Prateek Sircar
SIGIR4
2026 COMET: Compatibility-Oriented Multi-modal Embedding Transformer for Visual Recommendations
abstract
Recommending visually compatible products in fashion and interior design is a significant challenge, as compatibility rules are nuanced, context-dependent, and reliant on fine-grained details that traditional models fail to capture. Existing methods often struggle with heterogeneous compatibility rules (e.g., sofa-table vs. sofa-curtain) and an over-reliance on global visual features, missing critical textual cues like style or material. To address these limitations, we introduce COMET (Compatibility-Oriented Multi-modal Embedding Transformer), a scalable, vision-language framework for visual recommendations. COMET replaces rigid, category-specific subspaces with an attribute-conditioned cross-attention mechanism, reframing compatibility as a conditional retrieval problem. By formulating a textual compatibility prompt that encodes relational context and structured attributes, COMET dynamically conditions which visual features are attended to, producing a joint, context-aware representation. This multi-modal approach allows the model to leverage descriptive text (e.g., ''mid-century modern'' or ''oak finish'') to understand visually ambiguous stylistic nuances. The model is trained using a triplet loss with hard negative sample strategy to effectively distinguish between compatible and incompatible item pairs. COMET was evaluated on established benchmarks for both fashion (Polyvore) and furniture (Bonn Furniture), in addition to our in-house datasets.
Dween Rabius Sanny, Prateek Sircar
SIGIR2
2026 MEDAL: multi-modal MEta-space Distillation and ALignment for Visual Compatibility Learning
abstract
Visual compatibility recommendation systems aim to surface compatible items (e.g. pants, shoes) that harmonise with a user-selected product (e.g., shirt). Existing methods struggle in three key aspects: they rely on global CNN representations that overlook fine-grained local cues critical for visual pairing; they force all categories into a single latent space, ignoring the fact that compatibility rules differ across product-type pairs; and they demand costly, expert-annotated outfit labels. We introduce MEDAL(Meta-space Distillation and Alignment ), a self-supervised framework that addresses all three challenges simultaneously. MEDAL (i) employs a local–global augmentation curriculum inside a teacher–student ViT to emphasise patch-level texture and pattern similarities while suppressing confounding global shape cues; (ii) partitions the joint feature manifold into learnable, pair-specific meta-spaces so that, for example, {shirt,pants} and {pants,shoes} relationships are modelled with distinct projection masks; and (iii) replaces manual labels with distantly supervised KD, harvesting pseudo-compatible sets via object detection on web images, thus scaling to millions of real-world examples. We further fuse perceptually uniform LUV colour histograms to capture global colour harmony often missed by pure vision transformers. Extensive experiments on Polyvore disjoint/non-disjoint and a 2M-image in-house dataset show state-of-the-art gains of up to +3.72/+2.7FITB and +9.58R@10 over the strongest baseline, whilst cutting annotation cost to zero. Qualitative studies confirm that MEDAL retrieves stylistically coherent outfits and correctly penalises mismatched colour palettes.
Dween Rabius Sanny, Vinay Kumar Verma, Prateek Sircar
WACV3
2023 Multi-task Student Teacher Based Unsupervised Domain Adaptation for Address Parsing
Rishav Sahay, Anoop Saladi, Prateek Sircar
PAKDD (4)3
2022 Solar: Science of Entity Loss Attribution
abstract
The ability to accurately pinpoint the location of an event (e.g. loss, fault or bug) is of fundamental requirement in many systems. While we have state-of-the-art models to predict likelihood of an outcome, being able to pinpoint to the entity responsible for the outcome is also important. For example, in an e-commerce setup, a lost package detection system needs to infer the reason or location (delivery station, sort center, trucks) in case of a missing item, a network management system would like to diagnose nodes that are faulty based on end-end packet flow traces or a compiler needs to point out the exact location of a code that is erroneous. In this paper, we present an Attention based neural architecture for entity localization to accurately pinpoint the location of package loss in delivery network and bugs in erroneous programs. Our model performs well in scenarios where there is no annotation/ground truth for entities for localization. It can also adapt itself if annotations/ground truth is available for even a subset of entities by leveraging semi-supervision. The core of our model is a ladder-style architecture that helps us achieve state-of-the-art performance in both entity localization and detection. Further, to show the generality of our approach, we demonstrate its performance on a bug localization task for software programs. On a publicly available data-set, our solution outperforms the state-of-the-art technique by a significant margin.
Anshuman Mourya, Prateek Sircar, Anirban Majumder
KDD2