Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sajal Choudhary

dblp:205/3107 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Representation and self-supervised learning · 58% Language models and text generation · 33% Information extraction and text analysis · 9%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 87% Knowledge graphs · 13%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.812024
CRAG - Comprehensive RAG Benchmark · NeurIPS 2024
Information retrieval › question answering
question answering benchmark
0.812024
CRAG - Comprehensive RAG Benchmark · NeurIPS 2024
Information retrieval
retrieval evaluation
0.812024
CRAG - Comprehensive RAG Benchmark · NeurIPS 2024
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
sentence representation learning
0.712023
Ranking-Enhanced Unsupervised Sentence Representation Learning · ACL (1) 2023
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
unsupervised sentence embeddings
0.712023
Ranking-Enhanced Unsupervised Sentence Representation Learning · ACL (1) 2023
Knowledge graphs
knowledge graph exploration
0.212024
CRAG - Comprehensive RAG Benchmark · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 1.5ranking-based contrastive learning · 0.7
YearPublicationVenuePosition
2024 CRAG - Comprehensive RAG Benchmark
abstract
Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution to alleviate Large Language Model (LLM)’s deficiency in lack of knowledge. Existing RAG datasets, however, do not adequately represent the diverse and dynamic nature of real-world Question Answering (QA) tasks. To bridge this gap, we introduce the Comprehensive RAG Benchmark (CRAG), a factual question answering benchmark of 4,409 question-answer pairs and mock APIs to simulate web and Knowledge Graph (KG) search. CRAG is designed to encapsulate a diverse array of questions across five domains and eight question categories, reflecting varied entity popularity from popular to long-tail, and temporal dynamisms ranging from years to seconds. Our evaluation on this benchmark highlights the gap to fully trustworthy QA. Whereas most advanced LLMs achieve $\le 34\%$ accuracy on CRAG, adding RAG in a straightforward manner improves the accuracy only to 44%. State-of-the-art industry RAG solutions only answer 63% questions without any hallucination. CRAG also reveals much lower accuracy in answering questions regarding facts with higher dynamism, lower popularity, or higher complexity, suggesting future research directions. The CRAG benchmark laid the groundwork for a KDD Cup 2024 challenge, attracted thousands of participants and submissions. We commit to maintaining CRAG to serve research communities in advancing RAG solutions and general QA solutions. CRAG is available at https://github.com/facebookresearch/CRAG/.
Kai Sun 0006, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, Lingkun Kong, Brian Moran, Eting Yuan, Hanwen Zha, Nan Tang 0001, Lei Chen 0002, Nicolas Scheffer, Rakesh Wanga, Scott Yih, Xin Dong 0001
NeurIPS7
2023 Ranking-Enhanced Unsupervised Sentence Representation Learning
abstract
Yeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary, Jiwei Li, Xiang Li, Puyang Xu, Sunghyun Park, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yeon Seonwoo, Guoyin Wang 0002, Changmin Seo, Sajal Choudhary, Jiwei Li 0001, Puyang Xu, Alice Oh
ACL (1)4