EDBT 2026 Demo / reviewers in the wild / expert
Luciano Barbosa
dblp:80/468
· DBLP profile ↗
34ranked-venue papers
15as first author
8since 2021 · last 2026
0000-0002-6858-4773ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 21 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Legal Information Retrieval Through Metadata-Driven Multi-Hop RAG Architecture
Eniedson Fabiano Pereira da Silva Júnior, Cláudio de Souza Baptista, André Luiz Firmino Alves, Luciano Barbosa, Clecio B. M. Araujo, Fábio Lucas Meira de Souza Barbosa |
COMPSAC | 4 |
| 2026 | SciCheck: Reasoning Distillation for Biomedical Claim Verification
Gabriel Pereira, Luciano Barbosa |
SIGIR | 2 |
| 2024 | Improving dense retrieval models with LLM augmented data for dataset search
Levy de Souza Silva, Luciano Barbosa |
Knowl. Based Syst. | 2 |
| 2023 | Matching news articles and wikipedia tables for news augmentation
Levy de Souza Silva, Luciano Barbosa |
Knowl. Inf. Syst. | 2 |
| 2022 | A distantly supervised approach for enriching product graphs with user opinions
Johny Moreira, Tiago de Melo, Luciano Barbosa, Altigran S. da Silva |
J. Intell. Inf. Syst. | 3 |
| 2022 | Correction to: Linking place records using multi-view encoders
Vinícius M. R. Cousseau, Luciano Barbosa |
Neural Comput. Appl. | 2 |
| 2021 | Attention-Based Spatial Interpolation for House Price PredictionabstractEstimating the market price of a house is important for many businesses such as real estate and mortgage lending companies. The price of a house depends not only on its structural features (e.g. area and number of bedrooms) but also on the spatial context where it is located. In this work we estimate the price of a house based solely on its structural features and the characteristics and price of its neighbors. For that, we propose a hybrid attention mechanism that weights neighbors based on their similarity to the house in terms of structural features and geographic location. For the structural features, we apply an euclidean-based attention and, for the geographic location, we propose an attention layer based on a radial basis function kernel. Those attention mechanisms are then used by a neural network regressor to predict the price of a house and to generate a vector representation of the house based on its implicit context: the house embedding, which can be used as a feature set by any regressor to perform house price prediction. We have performed an extensive experimental evaluation on real-world datasets that shows that: (1) regressors using house embedding obtained the best results on all 4 datasets, outperforming baseline models; (2) the learned house embedding improves the performance of the evaluated regressors in almost all scenarios in comparison to raw features; and (3) simple regressor models such as Linear Regression using house embedding achieved comparable results to more competitive algorithms (e.g. Random Forest and Xgboost). Darniton Viana, Luciano Barbosa |
SIGSPATIAL/GIS | 2 |
| 2021 | Linking place records using multi-view encoders
Vinícius M. R. Cousseau, Luciano Barbosa |
Neural Comput. Appl. | 2 |
| 2020 | Distantly-Supervised Neural Relation Extraction with Side Information using BERTabstractRelation extraction (RE) consists in categorizing the relationship between entities in a sentence. A recent paradigm to develop relation extractors is Distant Supervision (DS), which allows the automatic creation of new datasets by taking an alignment between a text corpus and a Knowledge Base (KB). KBs can sometimes also provide additional information to the RE task. One of the methods that adopt this strategy is the RESIDE model, which proposes a distantly-supervised neural relation extraction using side information from KBs. Considering that this method outperformed state-of-the-art baselines, in this paper, we propose a related approach to RESIDE also using additional side information, but simplifying the sentence encoding with BERT embeddings. Through experiments, we show the effectiveness of the proposed method in Google Distant Supervision and Riedel datasets concerning the BGWA and RESIDE baseline methods. Although Area Under the Curve is decreased because of unbalanced datasets, P@N results have shown that the use of BERT as sentence encoding allows superior performance to baseline methods. Johny Moreira, Chaina Santos Oliveira, David Macedo, Cleber Zanchettin, Luciano Barbosa |
IJCNN | 5 |
| 2018 | Big Data Linkage for Product Specification PagesabstractAn increasing number of product pages are available from thousands of web sources, each page associated with a product, containing its attributes and one or more product identifiers. The sources provide overlapping information about the products, using diverse schemas, making web-scale integration extremely challenging. In this paper, we take advantage of the opportunity that sources publish product identifiers to perform big data linkage across sources at the beginning of the data integration pipeline, before schema alignment. To realize this opportunity, several challenges need to be addressed: identifiers need to be discovered on product pages, made difficult by the diversity of identifiers; the main product identifier on the page needs to be identified, made difficult by the many related products presented on the page; and identifiers across pages need to beresolved, made difficult by the ambiguity between identifiers across product categories. We present our RaF (Redundancy as Friend) solution to the problem of big data linkage for product specification pages, which takes advantage of the redundancy of identifiers at a global level, and the homogeneity of structure and semantics at the local source level, to effectively and efficiently link millions of pages of head and tail products across thousands of head and tail sources. We perform a thorough empirical evaluation of our RaF approach using the publicly available Dexter dataset consisting of 1.9M product pages from 7.1k sources of 3.5k websites, and demonstrate its effectiveness in practice. Disheng Qiu, Luciano Barbosa, Valter Crescenzi, Paolo Merialdo, Divesh Srivastava |
SIGMOD Conference | 2 |
| 2017 | Harvesting Forum Pages from Seed Sites
Luciano Barbosa |
ICWE | 1 |
| 2016 | Finding seeds to bootstrap focused crawlers
Karane Vieira, Luciano Barbosa, Altigran S. da Silva, Juliana Freire, Edleno Silva de Moura |
World Wide Web | 2 |
| 2015 | Detecting Semantically Equivalent Questions in Online User ForumsabstractTwo questions asking the same thing could be too different in terms of vocabulary and syntactic structure, which makes identifying their semantic equivalence challenging. This study aims to detect semantically equivalent questions in online user forums. We perform an extensive number of experiments using data from two different Stack Exchange forums. We compare standard machine learning methods such as Support Vector Machines (SVM) with a convolutional neural network (CNN). The proposed CNN generates distributed vector representations for pairs of questions and scores them using a similarity metric. We evaluate in-domain word embeddings versus the ones trained with Wikipedia, estimate the impact of the training set size, and evaluate some aspects of domain adaptation. Our experimental results show that the convolutional neural network with in-domain word embeddings achieves high performance even with limited training data. Dasha Bogdanova, Cícero Nogueira dos Santos, Luciano Barbosa, Bianca Zadrozny |
CoNLL | 3 |
| 2015 | USapiens: A System for Urban Trajectory Data AnalyticsabstractIn the past few years a growing number of cities have started monitoring the position of public transportation vehicles using GPS devices. In this paper, we focus on a particularly important urban dataset: GPS bus data. Buses are valuable sensors and information associated with buses can provide unprecedented insight into many different aspects of city's life, from human behavior to mobility patterns. But analyzing these large urban datasets presents many challenges. Urban datasets are complex, containing location and temporal components in addition that they are commonly released in their raw format. Furthermore, urban datasets may have noisy and missing data, locations gathered in a low sampling rate and not mapped to the underlying road network, among other issues which makes it difficult for citizens, administrators and developers to get insights. In this paper, we present a system, called USapiens, for analyzing large urban trajectory data. We first describe the architecture of the proposed system for pre-processing and analyzing urban trajectory data. We then detail five use cases we build using very large GPS dataset obtained from buses operating in the city of Rio de Janeiro to get insights into various aspects of public transportation in the city. Marcos R. Vieira, Luciano Barbosa, Matthias Kormaksson, Bianca Zadrozny |
MDM (1) | 2 |
| 2015 | Extracting Records and Posts from Forum Pages with Limited Supervision
Luciano Barbosa, Guilherme Ferreira |
WISE (2) | 1 |
| 2015 | DEXTER: Large-Scale Discovery and Extraction of Product Specifications on the WebabstractThe web is a rich resource of structured data. There has been an increasing interest in using web structured data for many applications such as data integration, web search and question answering. In this paper, we present Dexter, a system to find product sites on the web, and detect and extract product specifications from them. Since product specifications exist in multiple product sites, our focused crawler relies on search queries and backlinks to discover product sites. To perform the detection, and handle the high diversity of specifications in terms of content, size and format, our system uses supervised learning to classify HTML fragments (e.g., tables and lists) present in web pages as specifications or not. To perform large-scale extraction of the attribute-value pairs from the HTML fragments identified by the specification detector, D exter adopts two lightweight strategies: a domain-independent and unsupervised wrapper method, which relies on the observation that these HTML fragments have very similar structure; and a combination of this strategy with a previous approach, which infers extraction patterns by annotations generated by automatic but noisy annotators. The results show that our crawler strategy to locate product specification pages is effective: (1) it discovered 1:46A M product specification pages from 3; 005 sites and 9 different categories; (2) the specification detector obtains high values of F-measure (close to 0:9) over a heterogeneous set of product specifications; and (3) our efficient wrapper methods for attribute-value extraction get very high values of precision (0.92) and recall (0.95) and obtain better results than a state-of-the-art, supervised rule-based wrapper. Disheng Qiu, Luciano Barbosa, Xin Dong 0001, Yanyan Shen, Divesh Srivastava |
Proc. VLDB Endow. | 2 |
| 2014 | Bus Travel Time Predictions Using Additive ModelsabstractMany factors can affect the predictability of public bus services such as traffic, weather, day of week, and hour of day. However, the exact nature of such relationships between travel times and predictor variables is, in most situations, not known. In this paper we develop a framework that allows for flexible modeling of bus travel times through the use of Additive Models. The proposed class of models provides a principled statistical framework that is highly flexible in terms of model building. The experimental results demonstrate uniformly superior performance of our best model as compared to previous prediction methods when applied to a very large GPS data set obtained from buses operating in the city of Rio de Janeiro. Matthias Kormaksson, Luciano Barbosa, Marcos R. Vieira, Bianca Zadrozny |
ICDM | 2 |
| 2014 | Semantic Traffic Diagnosis with STAR-CITY: Architecture and Lessons Learned from Deployment in Dublin, Bologna, Miami and Rio
Freddy Lécué, Robert Tucker, Simone Tallevi-Diotallevi, Yiannis Gkoufas, Giuseppe Liguori, Mauro Borioni, Alexandre Rademaker, Luciano Barbosa |
ISWC (2) | 9 |
| 2013 | Mining Enterprise Websites for Association Thesaurus Construction
Luciano Barbosa |
WebDB | 1 |
| 2012 | Harvesting Parallel Text in Multiple Languages with Limited Supervision
Luciano Barbosa, Vivek Kumar Rangarajan Sridhar, Mahsa Yarmohammadi, Srinivas Bangalore |
COLING | 1 |
| 2011 | Focusing on novelty: a crawling strategy to build diverse language modelsabstractWord prediction performed by language models has an important role in many tasks as e.g. word sense disambiguation, speech recognition, hand-writing recognition, query spelling and query segmentation. Recent research has exploited the textual content of the Web to create language models. In this paper, we propose a new focused crawling strategy to collect Web pages that focuses on novelty in order to create diverse language models. In each crawling cycle, the crawler tries to ll the gaps present in the current language model built from previous cycles, by avoiding visiting pages whose vocabulary is already well represented in the model. It relies on an information theoretic measure to identify these gaps and then learns link patterns to pages in these regions in order to guide its visitation policy. To handle constantly evolving domains, a key feature of our crawler approach is its ability to adjust its focus as the crawl progresses. We evaluate our approach in two different scenarios in which our solution can be useful. First, we demonstrate that our approach produces more effective language models than the ones created by a baseline crawler in the context of a speech recognition task of broadcast news. In fact, in some cases, our crawler was able to obtain similar results to the baseline by crawling only 12.5% of the pages collected by the latter. Secondly, since in the news domain avoiding well-represented content might lead to novelty, i.e. up-to-date pages, we show that our diversity-based crawler can also be helpful to guide the crawler for the most recent content in the news. The results show that our approach was able to obtain on average 50% more up-to-date pages than the baseline crawler. Luciano Barbosa, Srinivas Bangalore |
CIKM | 1 |
| 2011 | Crawling Back and Forth: Using Back and Out Links to Locate Bilingual Sites
Luciano Barbosa, Srinivas Bangalore, Vivek Kumar Rangarajan Sridhar |
IJCNLP | 1 |
| 2011 | SpeechForms: From Web to Speech and BackabstractThis paper describes SpeechForms, a system that uses novel techniques to automatically identify form element semantics and form element content, and to semi-automatically generate language models that allow users to fill out each web form element by voice. Preliminary experimental results show that simple per-element language models are faster and may be more accurate than statistical n-gram language models trained on large amounts of web text data. Index Terms: language modeling, form understanding, information retrieval Luciano Barbosa, Diamantino Caseiro, Giuseppe Di Fabbrizio |
INTERSPEECH | 1 |
| 2011 | A Scalable Approach to Building a Parallel Corpus from the WebabstractParallel text acquisition from the Web is an attractive way for augmenting statistical models (e.g., machine translation, crosslingual document retrieval) with domain representative data. The basis for obtaining such data is a collection of pairs of bilingual Web sites or pages. In this work, we propose a crawling strategy that locates bilingual Web sites by constraining the visitation policy of the crawler to the graph neighborhood of bilingual sites on the Web. Subsequently, we use a novel recursive mining technique that recursively extracts text and links from the collection of bilingual Web sites obtained from the crawling. Our method does not suffer from the computationally prohibitive combinatorial matching typically used in previous work that uses document retrieval techniques to match a collection of bilingual webpages. We demonstrate the efficacy of our approach in the context of machine translation in the tourism and hospitality domain. The parallel text obtained using our novel crawling strategy results in a relative improvement of 21% in BLEU score (English-to-Spanish) over an out-of-domain seed translation model trained on the European parliamentary proceedings. Vivek Kumar Rangarajan Sridhar, Luciano Barbosa, Srinivas Bangalore |
INTERSPEECH | 2 |
| 2010 | Creating and exploring web form repositoriesabstractWe present DeepPeep (http://www.deeppeep.org), a new system for discovering, organizing and analyzing Web forms. DeepPeep allows users to explore the entry points to hidden-Web sites whose contents are out of reach for traditional search engines. Besides demonstrating important features of DeepPeep and describing the infrastructure we used to build the system, we will show how this infrastructure can be used to create form collections and form search engines for different domains. We also present the analysis component of DeepPeep which allows users to explore and visualize information in form repositories, helping them not only to better search and understand forms in different domains, but also to refine the form gathering process. Luciano Barbosa, Hoa Nguyen, Thanh Hoang Nguyen, Ramesh Pinnamaneni, Juliana Freire |
SIGMOD Conference | 1 |
| 2010 | Using Latent-Structure to Detect Objects on the WebabstractAn important requirement for emerging applications which aim to locate and integrate content distributed over the Web is to identify pages that are relevant for a given domain or task. In this paper, we address the problem of identifying pages that contain objects with a latent structure, i.e., the structure is implicitly represented in the page. We propose an algorithm which, given a set of instances of an object type, derives rules by automatically extracting statistically significant patterns present inside the objects. These rules can then be used to detect the presence of these objects in new, unseen pages. Our approach has several advantages when compared against learning-based text classifiers. Because it relies only on positive examples, constructing accurate object detectors is simpler than constructing learning classifiers, which require both positive and negative examples. Also, besides providing a classification decision for the presence of an object, the derived detectors are able to pinpoint the location of the object inside a Web page. This enables our algorithm to extract additional object fragments and apply online learning to automatically update the rules as new documents become available. An experimental evaluation, using a representative set of domains, indicates that our approach is effective. It is able to learn structural patterns and derive detectors that outperform state-of-art text classifiers and the online learning component leads to substantial improvements over the initial detectors. Luciano Barbosa, Juliana Freire |
WebDB | 1 |
| 2009 | For a few dollars less: Identifying review pages sans human labels
Luciano Barbosa, Ravi Kumar 0001, Bo Pang 0001, Andrew Tomkins |
HLT-NAACL | 1 |
| 2008 | Siphon++: a hidden-webcrawler for keyword-based interfacesabstractThe hidden Web consists of data that is generally hidden behind form interfaces, and as such, it is out of reach for traditional search engines. With the goal of leveraging the high-quality information in this largely unexplored portion of the Web, in this paper, we propose a new strategy for automatically retrieving data hidden behind keyword-based form interfaces. Unlike previous approaches to this problem, our strategy adapts the query generation and selection by detecting features of the index. We describe an extensive experimental evaluation which shows that: our strategy is able to derive appropriate queries to obtain high coverage while, at the same time, avoiding the retrieval of redundant data; and it obtains higher coverage and is more efficient approaches that use a fixed strategy for query generation. Karane Vieira, Luciano Barbosa, Juliana Freire, Altigran S. da Silva |
CIKM | 2 |
| 2008 | An Exploratory Study of Information Retrieval Techniques in Domain AnalysisabstractDomain analysis involves not only looking at standard requirements documents (e.g., use case specifications) but also at customer information packs, market analyses, etc. Looking across all these documents and deriving, in a practical and scalable way, a feature model that is comprised of coherent abstractions is a fundamental and non-trivial challenge. We conduct an exploratory study to investigate the suitability of Information Retrieval (IR) techniques for scalable identification of commonalities and variabilities in requirement specifications for software product lines. Accordingly, based on observations derived from industrial experience and on state-of-the-art research and practice, we also propose an initial framework, leveraging IR to systematically abstract requirements from existing specifications of a given domain into a feature model. We evaluate this framework, present a roadmap for its further extension, and formulate hypotheses to guide future work in exploring IR techniques for domain analysis. Vander Alves, Christa Schwanninger, Luciano Barbosa, Awais Rashid, Peter Sawyer, Paul Rayson, Christoph Pohl, Andreas Rummler |
SPLC | 3 |
| 2007 | Organizing Hidden-Web Databases by Clustering Visible Web DocumentsabstractIn this paper we address the problem of organizing hidden-Web databases. Given a heterogeneous set of Web forms that serve as entry points to hidden-Web databases, our goal is to cluster the forms according to the database domains to which they belong. We propose a new clustering approach that models Web forms as a set of hyperlinked objects and considers visible information in the form context - both within and in the neighborhood of forms - as the basis for similarity comparison. Since the clustering is performed over features that can be automatically extracted, the process is scalable. In addition, because it uses a rich set of metadata, our approach is able to handle a wide range of forms, including content-rich forms that contain multiple attributes, as well as simple keyword-based search interfaces. An experimental evaluation over real Web data shows that our strategy generates high-quality clusters - measured both in terms of entropy and F-measure. This indicates that our approach provides an effective and general solution to the problem of organizing hidden-Web databases. Luciano Barbosa, Juliana Freire, Altigran S. da Silva |
ICDE | 1 |
| 2007 | Combining classifiers to identify online databasesabstractWe address the problem of identifying the domain of onlinedatabases. More precisely, given a set F of Web forms automaticallygathered by a focused crawler and an online databasedomain D, our goal is to select from F only the formsthat are entry points to databases in D. Having a set ofWebforms that serve as entry points to similar online databasesis a requirement for many applications and techniques thataim to extract and integrate hidden-Web information, suchas meta-searchers, online database directories, hidden-Webcrawlers, and form-schema matching and merging.We propose a new strategy that automatically and accuratelyclassifies online databases based on features that canbe easily extracted from Web forms. By judiciously partitioningthe space of form features, this strategy allows theuse of simpler classifiers that can be constructed using learningtechniques that are better suited for the features of eachpartition. Experiments using real Web data in a representativeset of domains show that the use of different classifiersleads to high accuracy, precision and recall. This indicatesthat our modular classifier composition provides an effectiveand scalable solution for classifying online databases. Luciano Barbosa, Juliana Freire |
WWW | 1 |
| 2007 | An adaptive crawler for locating hiddenwebentry pointsabstractIn this paper we describe new adaptive crawling strategies to efficiently locate the entry points to hidden-Web sources. The fact that hidden-Web sources are very sparsely distributedmakes the problem of locating them especially challenging. We deal with this problem by using the contents ofpages to focus the crawl on a topic; by prioritizing promisinglinks within the topic; and by also following links that may not lead to immediate benefit. We propose a new frameworkwhereby crawlers automatically learn patterns of promisinglinks and adapt their focus as the crawl progresses, thus greatly reducing the amount of required manual setup andtuning. Our experiments over real Web pages in a representativeset of domains indicate that online learning leadsto significant gains in harvest rates' the adaptive crawlers retrieve up to three times as many forms as crawlers thatuse a fixed focus strategy. Luciano Barbosa, Juliana Freire |
WWW | 1 |
| 2006 | Automatically constructing collections of online database directories
Luciano Barbosa, Juliana Freire |
CIKM | 1 |
| 2005 | Searching for Hidden-Web Databases
Luciano Barbosa, Juliana Freire |
WebDB | 1 |