VLDB 2026 Research / reviewers in the wild / expert
Chad DeLuca
dblp:240/9197
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2023
0009-0009-7690-7454ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Understanding Customer Requirements - An Enterprise Knowledge Graph Approach
Basel Shbita, Anna Lisa Gentile, Pengyuan Li 0001, Chad DeLuca |
ESWC | 4 |
| 2023 | Long-Form Information Retrieval for Enterprise MatchmakingabstractUnderstanding customer requirements is a key success factor for both business-to-consumer (B2C) and business-to-business (B2B) enterprises. In a B2C context, most requirements are directly related to products and therefore expressed in keyword-based queries. In comparison, B2B requirements contain more information about customer needs and as such the queries are often in a longer form. Such long-form queries pose significant challenges to the information retrieval task in B2B context. In this work, we address the long-form information retrieval challenges by proposing a combination of (i) traditional retrieval methods, to leverage the lexical match from the query, and (ii) state-of-the-art sentence transformers, to capture the rich context in the long queries. We compare our method against traditional TF-IDF and BM25 models on an internal dataset of 12,368 pairs of long-form requirements and products sold. The evaluation shows promising results and provides directions for future work. Pengyuan Li 0001, Anna Lisa Gentile, Chad DeLuca, Daniel Tan 0002, Sandeep Gopisetty |
SIGIR | 4 |
| 2021 | SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus ExpansionabstractRecent advances in text representation have shown that training on large amounts of text is crucial for natural language understanding. However, models trained without predefined notions of topical interest typically require careful fine-tuning when transferred to specialized domains. When a sufficient amount of within-domain text may not be available, expanding a seed corpus of relevant documents from large-scale web data poses several challenges. First, corpus expansion requires scoring and ranking each document in the collection, an operation that can quickly become computationally expensive as the web corpora size grows. Relying on dense vector spaces and pairwise similarity adds to the computational expense. Secondly, as the domain concept becomes more nuanced, capturing the long tail of domain-specific rare terms becomes non-trivial, especially under limited seed corpora scenarios. Muntasir Wahed, Daniel Gruhl, Alfredo Alba, Anna Lisa Gentile, Petar Ristoski, Chad DeLuca, Steve Welch, Ismini Lourentzou |
CIKM | 6 |
| 2020 | Understanding Data Centers from Logs: Leveraging External Knowledge for Distant Supervision
Chad DeLuca, Anna Lisa Gentile, Petar Ristoski, Steve Welch |
ISWC (2) | 1 |