VLDB 2026 Research / reviewers in the wild / expert
Nirmal Pal
dblp:01/3138
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 60% Web and social media mining · 17% Data mining · 17% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining
information extraction |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Information retrieval › web search
web information retrieval |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Web and social media mining › web page analysis
web page segmentation |
0.1 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Information retrieval › search engines
domain-specific search engine |
0.0 | 1 | 2003 | eBizSearch: a niche search engine for e-business · SIGIR 2003 |
Information retrieval › document processing
metadata generation |
0.0 | 1 | 2003 | eBizSearch: a niche search engine for e-business · SIGIR 2003 |
Information retrieval
search engines |
0.0 | 1 | 2003 | eBizSearch: a niche search engine for e-business · SIGIR 2003 |
Distributed and cloud data management
web caching |
0.0 | 1 | 2005 | Automatic Identification of Informative Sections of Web Pages · IEEE Trans. Knowl. Data Eng. 2005 |
Methods — techniques the papers use, named apart from their topics
feature extraction · 0.1classification · 0.1machine learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Automatic Identification of Informative Sections of Web PagesabstractWeb pages - especially dynamically generated ones - contain several items that cannot be classified as the "primary content," e.g., navigation sidebars, advertisements, copyright notices, etc. Most clients and end-users search for the primary content, and largely do not seek the noninformative content. A tool that assists an end-user or application to search and process information from Web pages automatically, must separate the "primary content sections" from the other content sections. We call these sections as "Web page blocks" or just "blocks." First, a tool must segment the Web pages into Web page blocks and, second, the tool must separate the primary content blocks from the noninformative content blocks. In this paper, we formally define Web page blocks and devise a new algorithm to partition an HTML page into constituent Web page blocks. We then propose four new algorithms, ContentExtractor, FeatureExtractor, K-FeatureExtractor, and L-Extractor. These algorithms identify primary content blocks by 1) looking for blocks that do not occur a large number of times across Web pages, by 2) looking for blocks with desired features, and by 3) using classifiers, trained with block-features, respectively. While operating on several thousand Web pages obtained from various Web sites, our algorithms outperform several existing algorithms with respect to runtime and/or accuracy. Furthermore, we show that a Web cache system that applies our algorithms to remove noninformative content blocks and to identify similar blocks across Web pages can achieve significant storage savings. Sandip Debnath, Prasenjit Mitra 0001, Nirmal Pal, C. Lee Giles |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2003 | eBizSearch: a niche search engine for e-businessabstractNiche Search Engines offer an efficient alternative to traditional search engines when the results returned by general-purpose search engines do not provide a sufficient degree of relevance. By taking advantage of their domain of concentration they achieve higher relevance and offer enhanced features. We discuss a new niche search engine, eBizSearch, based on the technology of CiteSeer and dedicated to e-business and e-business documents. We present the integration of CiteSeer in the framework of eBizSearch and the process necessary to tune the whole system towards the specific area of e-business. We also discuss how using machine learning algorithms we generate metadata to make eBizSearch Open Archives compliant. eBizSearch is a publicly available service and can be reached at [3]. C. Lee Giles, Yves Petinot, Pradeep B. Teregowda, Hui Han 0001, Steve Lawrence, Arvind Rangaswamy, Nirmal Pal |
SIGIR | 7 |