VLDB 2026 Research / reviewers in the wild / expert
Sandip Debnath
dblp:82/58
· DBLP profile ↗
8ranked-venue papers
6as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-authorComputer networks · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 30% Web and social media mining · 30% Data mining · 30% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining
information extraction |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Information retrieval › web search
web information retrieval |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Web and social media mining › web page analysis
web page segmentation |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Algorithmic game theory and mechanism design
prediction markets |
0.0 | 1 | 2003 | Information incorporation in online in-Game sports betting markets · EC 2003 |
Distributed and cloud data management
web caching |
0.0 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Methods — techniques the papers use, named apart from their topics
feature extraction · 0.1classification · 0.1empirical analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Analysis of an optical packet switch with partially shared buffer and wavelength conversionabstractAn analytical approach based on a reduced Markov chain (RMC) model of an optical packet switch with partially shared output buffer has been presented and a combined scheme for the contention resolution in optical packet switches has been proposed. The proposed scheme takes the advantage of output buffering, shared buffering and wavelength conversion to enhance the switch performance. The performance evaluation of the switch architecture based on RMC modelling has been validated through extensive simulation. Moreover, for a self-similar input traffic, the switch performance has been shown to achieve further improvement if a traffic shaping mechanism is incorporated at the input end of the switch architecture. Sandip Debnath, Ranjan Gangopadhyay |
IET Commun. | 1 |
| 2005 | Identifying Content Blocks from Web Documents
Sandip Debnath, Prasenjit Mitra 0001, C. Lee Giles |
ISMIS | 1 |
| 2005 | A learning based model for headline extraction of news articles to find explanatory sentences for eventsabstractMetadata information plays a crucial role in augmenting document organising efficiency and archivability. News metadata includes DateLine, ByLine, HeadLine and many others. We found that HeadLine information is useful for guessing the theme of the news article. Particularly for financial news articles, we found that HeadLine can thus be specially helpful to locate explanatory sentences for any major events such as significant changes in stock prices. In this paper we explore a support vector based learning approach to automatically extract the HeadLine metadata. We find that the classification accuracy of finding the HeadLines improves if DateLines are identified first. We then used the extracted HeadLines to initiate a pattern matching of keywords to find the sentences responsible for story theme. Using this theme and a simple language model it is possible to locate any explanatory sentences for any significant price change. Sandip Debnath, C. Lee Giles |
K-CAP | 1 |
| 2005 | Learning term-relationships for ontology construction: creating business ontologies for event explanationabstractTerm-Relationship plays major role in several areas of research including document relevance, domain ontology construction or metadata extraction. As part of our work in finding explanation for market events (such as sudden or significant stock price change for a company), we use business ontologies to facilitate the relevance ranking of documents. We use the ontologies even further to rank individual sentences in a news article so that irrelevant events (in the form of sentences) in a relevant article will not get undue importance. This two-step ranking helps us in extracting important sentences which can be labelled as "responsible" or "explanatory" sentences for a significant stock price change. In this paper we show the performance evaluation of our relevance model using ontologies. We also show a few examples of sentences which can be thought of as providing explanation for some recent price changes. Sandip Debnath, Arun Upneja, C. Lee Giles |
K-CAP | 1 |
| 2005 | Automatic Identification of Informative Sections of Web PagesabstractWeb pages - especially dynamically generated ones - contain several items that cannot be classified as the "primary content," e.g., navigation sidebars, advertisements, copyright notices, etc. Most clients and end-users search for the primary content, and largely do not seek the noninformative content. A tool that assists an end-user or application to search and process information from Web pages automatically, must separate the "primary content sections" from the other content sections. We call these sections as "Web page blocks" or just "blocks." First, a tool must segment the Web pages into Web page blocks and, second, the tool must separate the primary content blocks from the noninformative content blocks. In this paper, we formally define Web page blocks and devise a new algorithm to partition an HTML page into constituent Web page blocks. We then propose four new algorithms, ContentExtractor, FeatureExtractor, K-FeatureExtractor, and L-Extractor. These algorithms identify primary content blocks by 1) looking for blocks that do not occur a large number of times across Web pages, by 2) looking for blocks with desired features, and by 3) using classifiers, trained with block-features, respectively. While operating on several thousand Web pages obtained from various Web sites, our algorithms outperform several existing algorithms with respect to runtime and/or accuracy. Furthermore, we show that a Web cache system that applies our algorithms to remove noninformative content blocks and to identify similar blocks across Web pages can achieve significant storage savings. Sandip Debnath, Prasenjit Mitra 0001, Nirmal Pal, C. Lee Giles |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Information incorporation in online in-Game sports betting marketsabstractWe analyze data from $52$ online in-game sports betting markets (where betting is allowed continuously throughout a game), including 34 markets based on soccer (European football) games from the 2002 World Cup, and 18 basketball games from the 2002 USA National Basketball Association (NBA) championship. We show that prices on average approach the correct outcome over time, and the price dynamics in the markets are closely coupled with game events, agreeing with efficient market assumptions. We also examine qualitative distinctions between the two types of games. Sandip Debnath, David M. Pennock, C. Lee Giles, Steve Lawrence |
EC | 1 |
| 2002 | Modelling Information Incorporation in Markets, with Application to Detecting and Explaining Events
David M. Pennock, Sandip Debnath, Eric J. Glover, C. Lee Giles |
UAI | 2 |
| 2000 | Combining Multiple Perspectives
Bikramjit Banerjee, Sandip Debnath, Sandip Sen |
ICML | 2 |