Nirav Pravinbhai Bhatt

dblp:415/1763 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
1 paper
Reinforcement learning · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
1.012026
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories · AAAI 2026
Information retrieval › document retrieval › domain-specific retrieval
biomedical information retrieval
1.012026
BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives · AAAI 2026
Information retrieval › retrieval models › neural retrieval
dense retrieval
1.012026
BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives · AAAI 2026
Information retrieval › document retrieval
domain-specific retrieval
1.012026
BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives · AAAI 2026
Information retrieval
hard negative mining
1.012026
BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives · AAAI 2026
Machine learning › Reinforcement learning
offline reinforcement learning
0.312026
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories · AAAI 2026

Methods — techniques the papers use, named apart from their topics

multiple instance learning · 1.0imitation learning · 1.0dense retriever fine-tuning · 1.0citation-aware negative sampling · 1.0
YearPublicationVenuePosition
2026 SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
abstract
In this work, we study the problem of offline safe imitation learning (IL). In many real-world settings, online interactions can be risky, and accurately specifying the reward and the safety cost information at each timestep can be difficult. However, it is often feasible to collect trajectories reflecting undesirable or risky behavior, implicitly conveying the behavior the agent should avoid. We refer to these trajectories as non-preferred trajectories. Unlike standard IL, which aims to mimic demonstrations, our agent must also learn to avoid risky behavior using non-preferred trajectories. In this paper, we propose a novel approach, SafeMIL, to learn a parameterized cost that predicts if the state-action pair is risky via Multiple Instance Learning. The learned cost is then used to avoid non-preferred behaviors, resulting in a policy that prioritizes safety. We empirically demonstrate that our approach can learn a safer policy that satisfies cost constraints without degrading the reward performance, thereby outperforming several baselines.
Returaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman Ravindran
AAAI2
2026 BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives
abstract
Hard negatives are essential for training effective retrieval models. Hard-negative mining typically relies on ranking documents using cross-encoders or static embedding models based on similarity metrics such as cosine distance. Hard negative mining becomes challenging for biomedical and scientific domains due to the difficulty in distinguishing between source and hard negative documents. However, referenced documents naturally share contextual relevance with the source document but are not duplicates, making them well-suited as hard negatives. In this work, we propose BiCA: Biomedical Dense Retrieval with Citation-Aware Hard Negatives, an approach for hard-negative mining by utilizing citation links in 20,000 PubMed articles for improving a domain-specific small dense retriever. We fine-tune the GTE_small and GTE_Base models using these citation-informed negatives and observe consistent improvements in zero-shot dense retrieval using nDCG@10 for both in-domain and out-of-domain tasks on BEIR and outperform baselines on long-tailed topics in LoTTE using Success@5. Our findings highlight the potential of leveraging document link structure to generate highly informative negatives, enabling state-of-the-art performance with minimal fine-tuning and demonstrating a path towards highly data-efficient domain adaptation.
Aarush Sinha, Pavan Kumar S, Roshan Balaji, Nirav Pravinbhai Bhatt
AAAI4