Radek Burget

dblp:54/3223 · DBLP profile ↗
← Back
8ranked-venue papers in the field
5as first author
5since 2021 · last 2026
0000-0001-5233-0456ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6 (3 first)Database Systems & Data Management · 1 (1 first)Other / Interdisciplinary · 1 (1 first)
YearPublicationVenuePosition
2026 Visual-Aware Representation of Web Pages for Machine Learning Applications
Radek Burget, Radek Hranicky
ICWE1
2026 SmartScrape: A Neuro-Symbolic Web Information Extraction
Ganjali Imanov, Radek Burget
ICWE2
2026 GNN-Based Token Reduction for LLM Semantic Element Detection in E-Commerce Product Pages
Hamza Salem, Radek Burget
ICWE2
2023 Scraping Data from Web Pages Using SPARQL Queries
Radek Burget
ICWE1
2023 Creating Searchable Web Page Snapshots Using Semantic Technologies
Radek Burget, Hamza Salem
ICWE1
2017 Box clustering segmentation: A new method for vision-based web page preprocessing
Jan Zeleny, Radek Burget, Jaroslav Zendulka
Inf. Process. Manag.2
2009 Web Page Element Classification Based on Visual Features
abstract
When applying the traditional data mining methods to World Wide Web documents, the typical problem is that a normal Web page contains a variety of information of different kinds in addition to its main content. This additional information such as navigation, advertisement or copyright notices negatively influences the results of the data mining methods as for example the content classification. In this paper, we present a method of interesting area detection in a Web page. This method is inspired by an assumed human reader approach to this task. First, basic visual blocks are detected in the page and subsequently, the purpose of these blocks is guessed based on their visual appearance. We describe a page segmentation method used for the visual block detection, we propose a way of the block classification based on the visual features and finally, we provide an experimental evaluation of the method on real-world data.
Radek Burget, Ivana Burgetová
ACIIDS1
2007 Layout Based Information Extraction from HTML Documents
abstract
We propose a method of information extraction from HTML documents based on modelling the visual information in the document. A page segmentation algorithm is used for detecting the document layout and subsequently, the extraction process is based on the analysis of mutual positions of the detected blocks and their visual features. This approach is more robust that the traditional DOM-based methods and it opens new possibilities for the extraction task specification.
Radek Burget
ICDAR1