EDBT 2026 Demo / reviewers in the wild / expert
Altigran S. da Silva
dblp:s/ASdaSilva · also Altigran Soares da Silva
· DBLP profile ↗
85ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0002-8992-495XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 70 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 21Applied, interdisciplinary, general and emerging computing · 8Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Tool for Explainable Pension Fund Recommendations using Large Language ModelsabstractIn this demo, we present a prototype tool designed to help financial advisors recommend private pension funds to investors based on their preferences, offering personalized investment suggestions. The tool leverages Large Language Models (LLMs), which enhance explainability by providing clear and understandable rationales for recommendations and effectively handles both sequential and cold-start scenarios. We outline the design, implementation, and results of a user-based evaluation using real-world data. The evaluation shows a high recommendation acceptance rate among financial advisors, highlighting the tool’s potential to improve decision-making in financial advisory services. Eduardo Alves da Silva, Leandro Balby Marinho, Edleno Silva de Moura, Altigran S. da Silva |
RecSys | 4 |
| 2024 | A Study on Unsupervised Question and Answer Generation for Legal Information Retrieval and Precedents UnderstandingabstractTraditional retrieval systems are hardly adequate for Legal Research, mainly because only returning the documents related to a given query is usually insufficient. Legal documents are extensive, and we posit that generating questions about them and detecting the answers provided by these documents help the Legal Research journey. This paper presents a pipeline that relates Legal Questions with documents answering them. We align features generated by Large Language Models with traditional clustering methods to find convergent and divergent answers to the same legal matter. We performed a case study with 50 legal documents on the Brazilian judiciary system. Our pipeline found convergent and divergent answers to 23 major legal questions regarding the case law for daily fines in Civil Procedural Law. The pipeline manual evaluation shows it managed to group diverse similar answers to the same question with an average precision of 0.85. It also managed to detect two divergent legal matters with an average F1 Score of 0.94. Johny Moreira, Altigran S. da Silva, Edleno Silva de Moura, Leandro Bezerra Marinho |
SIGIR | 2 |
| 2024 | SEREIA: document store exploration through keywords
Ariel Afonso, Paulo Martins 0005, Altigran S. da Silva |
Knowl. Inf. Syst. | 3 |
| 2023 | PyLatheDB - A Library for Relational Keyword Search with Support to Schema ReferencesabstractRelational Keyword Search (R-KwS) systems enable naive/informal users to explore and retrieve information from relational databases without knowing schema details or query languages. These systems take the keywords from the input query, locate the elements of the target database that correspond to these keywords, and look for ways to "connect" these elements using the information on key/foreign key pairs. Although several such systems have been proposed, most of them only support queries whose keywords refer to the contents of the target database and only a few support queries in which keywords may also refer to elements of the database schema. We showcase PyLatheDB, a Python library for Relational Keyword Search with Support to Schema References. PyLatheDB is based on Lathe, an R-KwS framework that generalizes the well-known concepts of Query Matches (QMs) and Candidate Joining Networks (CJNs) to handle keywords referring to schema elements and introduces new algorithms to generate them. Lathe also introduced a novel approach to automatically select the CJNs that are more likely to represent the user intent when issuing a keyword query. This approach includes two major innovations: a ranking algorithm for selecting better QMs, yielding the generation of fewer but better CJNs, and an eager evaluation strategy for pruning useless CJNs. We demonstrate through a Jupyter1notebook the functioning of PyLatheDB for two representative application scenarios, showing each step of the keyword query processing. The users can interact with the notebook by running keyword queries and experimenting with configuration parameters to see how they affect the results. The notebook, a video, and the code of PyLatheDB are available at https://github.com/pr3martins/PyLatheDB. Paulo Martins 0005, Ariel Afonso, Altigran S. da Silva |
ICDE | 3 |
| 2022 | Applying burst-tries for error-tolerant prefix search
Berg Ferreira, Edleno Silva de Moura, Altigran S. da Silva |
Inf. Retr. J. | 3 |
| 2022 | A distantly supervised approach for enriching product graphs with user opinions
Johny Moreira, Tiago de Melo, Luciano Barbosa, Altigran S. da Silva |
J. Intell. Inf. Syst. | 4 |
| 2022 | A distantly supervised approach for recognizing product mentions in user-generated content
Henry S. Vieira, Altigran S. da Silva, Pável Calado, Edleno Silva de Moura |
J. Intell. Inf. Syst. | 2 |
| 2022 | Efficient Match-Based Candidate Network Generation for Keyword Queries Over Relational DatabasesabstractSeveral systems proposed for processing keyword queries over relational databases rely on the generation and evaluation of Candidate Networks (CNs), i.e., networks of joined database relations that when processed as SQL queries, provide a relevant answer to the input keyword query. Although the evaluation of CNs has been extensively addressed in the literature, the problem of generating CNs efficiently and effectively has received much less attention. This challenging problem consists of automatically locating relations in the database that may contain relevant pieces of information, given a handful of keywords, and determining suitable ways of joining these relations to satisfy the implicit information needs expressed by a user while formulating his/her query. In this paper, we propose a novel approach for generating CNs, wherein the possible matches for the query in the database are efficiently enumerated at first. Thesequery matchesare then used to guide the CN generation process, avoiding the exhaustive search procedure used by the current state-of-art approaches. We show that our approach allows the generation of a compact set of CNs that leads to superior quality answers, and demands less resources in terms of processing time and memory. These claims are supported by a comprehensive set of experiments that we carried out using several query sets and datasets used in previous related works and whose results we report and analyze here. Péricles Silva de Oliveira, Altigran S. da Silva, Edleno Silva de Moura, Rosiane de Freitas |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Vallum-Med: Protecting Medical Data in Cloud EnvironmentsabstractDespite the many advantages of cloud computing, keeping information in such an environment increases the risk of cyber attacks, as well as the possibility of unauthorized access by cloud provider employees. Another critical concern is privacy protection, since depending on data access control, confidential information may be exposed even through authorized access. To solve these issues we have previously proposed Vallum, a platform that leverages Intel SGX protection to ensure the security, confidentiality, and integrity of data at rest and during processing. It also provides tools for privacy protection, following policies set by the data owner. In this demo we present Vallum-Med, an application of Vallum for the protection of medical patient personal data, including imaging results of their cardiac examinations. We will demonstrate that this system fully supports cloud protection of such sensitive data as well as the definition of privacy policies and ensuring that all results of queries are compliant to these policies. All processing, data storage and network traffic are protected using SCONE, a docker container-based technology for seamlessly incorporating SGX protection for applications, which provides a fully encrypted memory environment. Ronny Peterson, Altigran S. da Silva, Christof Fetzer, André Martin, Ignacio Blanquer |
CIKM | 2 |
| 2020 | A Focused Crawler for Web Feature Service and Web Map Service Discovering
Víctor Macêdo Alexandrino, Giovanni Comarela, Altigran S. da Silva, Jugurta Lisboa Filho |
W2GIS | 3 |
| 2020 | LESSQL: Dealing with Database Schema Changes in Continuous DeploymentabstractThe adoption of Continuous Deployment (CD) aims at allowing software systems to quickly evolve to accommodate new features. However, structural changes to the database schema are frequent and may incur in systems' services downtime. This encompasses the proper maintenance of both schema and source code, including rewrites of all outdated queries that use the same database. Previous solutions try to mitigate the burdening task of manually rewriting outdated queries. Unfortunately, a software team must still interact with some tools to properly fix the affected queries. Moreover, the team still has to locate and modify all the impacted code, which are often error-prone tasks. Thus, a project may not experience CD benefits when changes impact various code regions. In this paper we present an alternative approach, called LESSQL, whose goal is to improve queries' stability in the presence of structural schema changes over time. LESSQL supports queries that are less dependent on the database schema since they do not include the FROM clause. An underlying framework intercepts each LESSQL query and generates a corresponding SQL query for the current schema. It also locates the query attributes in the current schema and generates proper expressions to join the required tables. LESSQL supports unsupervised, supervised and hybrid configurations to process mappings of attributes to a newer schema version. We conducted experiments in the context of a popular open-source project, which experienced many diverse structural schema changes. Experiments outcomes indicate that our approach is effective in significantly reducing the modifications required for applying schema changes, allowing to better reap the benefits of CD. While supervised and hybrid configurations achieved a success rate higher than 95% with a minor query generation overhead, the unsupervised configuration was also successful for certain types of structural schema changes. These results show that LESSQL effectively favours CD and keeps queries running after database schema changes without services interruption. Ariel Afonso, Altigran S. da Silva, Tayana Conte, Paulo Martins 0005, João M. B. Cavalcanti, Alessandro F. Garcia 0001 |
SANER | 2 |
| 2020 | Federated and secure cloud services for building medical image classifiers on an intercontinental infrastructure
Ignacio Blanquer, Francisco Vilar Brasileiro, Andrey Brito, Amanda Calatrava, Christof Fetzer, Flavio Figueiredo, Ronny Petterson Guimarães, Leandro Bezerra Marinho, Wagner Meira Jr., Altigran S. da Silva, Angel Alberich-Bayarri, Eduardo Camacho-Ramos, Ana Jimenez-Pastor, Antonio Luiz L. Ribeiro, Bruno Ramos Nascimento |
Future Gener. Comput. Syst. | 11 |
| 2019 | Vallum: Privacy, Confidentiality and Access Controlfor Sensitive Data in Cloud EnvironmentsabstractManaging sensitive data in shared environments such as public clouds is an enduring challenge. While several approaches exist to protect data at rest such as end-to-end encryption, there exist only a few solutions such as homomorphic encryption that offer secure data processing. Unfortunately, these solutions cannot be used in practice as they incur non-negligible run-time overheads and security risks. Moreover, as the majority of data management systems were designed to operate in private cloud environments, which are under the control of the data owner, they often lack appropriate mechanisms for access control as well as privacy assurance. In this paper we propose Vallum, a data access and protection layer that closes these gaps while enabling users to operate data management systems in shared environments such securely as public clouds. Vallum utilizes Intel SGX and remote attestation to ensure confidentiality and integrity of the data being stored and processed. Furthermore, it provides access protection and privacy assurance through a extensible architecture. Our performance evaluation indicates that the overhead introduced by Vallum makes it viable to be deployed in cloud infrastructures. Ronny Peterson, Altigran S. da Silva, Gabriel Fernandez 0001, André Martin, Christof Fetzer, Andrey Brito |
CloudCom | 3 |
| 2019 | Contender: Leveraging User Opinions for Purchase Decision-Making
Tiago de Melo, Altigran S. da Silva, Edleno Silva de Moura, Pável Calado |
ECIR (2) | 2 |
| 2019 | OpinionLink: Leveraging user opinions for product catalog enrichment
Tiago de Melo, Altigran S. da Silva, Edleno Silva de Moura, Pável Calado |
Inf. Process. Manag. | 2 |
| 2018 | Match-Based Candidate Network Generation for Keyword Queries over Relational DatabasesabstractSeveral systems for processing keyword queries over relational databases rely on the generation and evaluation of Candidate Networks (CNs), i.e., networks of joined relations that when processed as SQL queries, provide a relevant answer to the input keyword query. Although the evaluation of CNs has been extensively addressed in the literature, the problem of generating CNs has received much less attention. We propose a novel approach for generating CNs, wherein the possible matches for the query in the database are efficiently enumerated at first. These query matches are then used to guide the CN generation process, avoiding the exhaustive search procedure used by the current state-of-art approaches. We experimentally show that our approach allows the generation of a compact set of CNs that results in superior quality answers and demands less resources in terms of processing time and memory. Pericles de Oliveira, Altigran S. da Silva, Edleno Silva de Moura, Rosiane de Freitas |
ICDE | 2 |
| 2017 | Waves: a fast multi-tier top-k query processing algorithm
Caio Moura Daoud, Edleno Silva de Moura, David Fernandes de Oliveira, Altigran S. da Silva, Cristian Rossi, André Luiz da Costa Carvalho |
Inf. Retr. J. | 4 |
| 2017 | Color and texture applied to a signature-based bag of visual words method for image retrieval
Joyce Miranda dos Santos, Edleno Silva de Moura, Altigran S. da Silva, Ricardo da Silva Torres |
Multim. Tools Appl. | 3 |
| 2016 | Towards the Effective Linking of Social Media Contents to Products in E-Commerce CatalogsabstractOnline social media has become an essential part of our life. This media is often characterized by its diverse content, which is produced by ordinary users. The potential to easily express ideas and opinions has made social media a source of valuable information on a variety of topics. In particular, information containing comments about consumer products has become prevalent. Here, we are interested in linking products mentioned in unstructured user-generated content, namely open discussion forums, to their respective entities in consumer product catalogs. Among the issues associated with this task, ambiguity is a particularly hard problem, as users typically refer to the same product using many different forms and different products may share the same form. We argue that this problem can be effectively solved using a set of evidences that can be easily extracted from social media content and product descriptions. To achieve this, we show which features should be used, how they can be extracted, and then how to combine them through machine learning techniques. Experiments in three different product categories and two different datasets demonstrate that all the sources of evidence here proposed are important, while contextual information is fundamental to achieve higher levels of precision. In fact, our method, although straightforward, was able to achieve an average improvement of 0.17 in precision and 0.13 in F1, when compared to the current state-of-the-art solution. Henry S. Vieira, Altigran S. da Silva, Pável Calado, Marco Cristo, Edleno Silva de Moura |
CIKM | 2 |
| 2016 | LCA-based algorithms for efficiently processing multiple keyword queries over XML streams
Evandrino G. Barros, Alberto H. F. Laender, Mirella M. Moro, Altigran S. da Silva |
Data Knowl. Eng. | 4 |
| 2016 | Fast top-k preserving query processing using two-tier indexes
Caio Moura Daoud, Edleno Silva de Moura, André Luiz da Costa Carvalho, Altigran S. da Silva, David Fernandes de Oliveira, Cristian Rossi |
Inf. Process. Manag. | 4 |
| 2016 | Finding seeds to bootstrap focused crawlers
Karane Vieira, Luciano Barbosa, Altigran S. da Silva, Juliana Freire, Edleno Silva de Moura |
World Wide Web | 3 |
| 2015 | A Self-training CRF Method for Recognizing Product Model Mentions in Web Forums
Henry S. Vieira, Altigran S. da Silva, Marco Cristo, Edleno Silva de Moura |
ECIR | 2 |
| 2015 | Ranking Candidate Networks of relations to improve keyword search over relational databasesabstractRelational keyword search (R-KwS) systems based on schema graphs take the keywords from the input query, find the tuples and tables where these keywords occur and look for ways to “connect” these keywords using information on referential integrity constraints, i.e., key/foreign key pairs. The result is a number of expressions, called Candidate Networks (CNs), which join relations where keywords occur in a meaningful way. These CNs are then evaluated, resulting in a number of join networks of tuples (JNTs) that are presented to the user as ranked answers to the query. As the number of CNs is potentially very high, handling them is very demanding, both in terms of time and resources, so that, for certain queries, current systems may take too long to produce answers, and for others they may even fail to return results (e.g., by exhausting memory). Moreover, the quality of the CN evaluation may be compromised when a large number of CNs is processed. Based on observations made by other researchers and in our own findings on representative workloads, we argue that, although the number of possible Candidate Networks can be very high, only very few of them produce answers relevant to the user and are indeed worth processing. Thus, R-KwS systems can greatly benefit from methods for accessing the relevance of Candidate Networks, so that only those deemed relevant might be evaluated. We propose in this paper an approach for ranking CNs, based on their probability of producing relevant answers to the user. This relevance is estimated based on the current state of the underlying database using a probabilistic Bayesian model we have developed. Experiments that we performed indicate that this model is able to assign the relevant CNs among the top-4 in the ranking produced. In these experiments we also observed that processing only a few relevant CNs has a considerable positive impact, not only on the performance of processing keyword queries, but also on the quality of the results obtained. Pericles de Oliveira, Altigran S. da Silva, Edleno Silva de Moura |
ICDE | 2 |
| 2015 | Using active learning techniques for improving database schema matching methodsabstractThe schema matching problem consists of finding semantic correspondences between elements (e.g., attributes) of two database schemas. Typically, methods to solve this problem first use pair-wise functions called matchers to generate similarity scores values between pairs of elements from the two schemas. These scores are used as estimations of the correspondence between elements. Next, matchers are combined to establish which pairs of elements must be mapped when integrating the two schemas. For this, the best-known schema matching methods rely on fixed heuristics. We consider that using fixed heuristics is not always helpful to cope with a variety of database and element mismatch cases and argue that machine learning (ML) techniques comprise suitable alternatives to properly combine matchers. In this paper we propose ALMa (Active Learning Matching), a novel method for combining matchers based on active learning, which is an interesting and effective machine learning technique that efficiently exploits the users' expertise on the matching task. We report the results of experiments we carried out comparing ALMa with COMA, a well-known schema matching method in the literature based on fixed heuristics, and with YAM, a recent schema matching method based on supervised learning. In the experiments, ALMa achieved results better or at least similar to the baselines, while demanded less user effort, confirming the suitability of using active learning for combining matchers. Diego Rodrigues, Altigran S. da Silva, Rosiane de Freitas, Eulanda M. dos Santos |
IJCNN | 2 |
| 2015 | A genetic programming framework to schedule webpage updates
Aécio S. R. Santos, Cristiano R. de Carvalho, Jussara M. Almeida, Edleno Silva de Moura, Altigran S. da Silva, Nivio Ziviani |
Inf. Retr. J. | 5 |
| 2015 | A signature-based bag of visual words method for image indexing and search
Joyce Miranda dos Santos, Edleno Silva de Moura, Altigran S. da Silva, João M. B. Cavalcanti, Ricardo da Silva Torres, Márcio L. A. Vidal |
Pattern Recognit. Lett. | 3 |
| 2015 | Removing DUST Using Multiple Alignment of SequencesabstractA large number of URLs collected by web crawlers correspond to pages with duplicate or near-duplicate contents. To crawl, store, and use such duplicated data implies a waste of resources, the building of low quality rankings, and poor user experiences. To deal with this problem, several studies have been proposed to detect and remove duplicate documents without fetching their contents. To accomplish this, the proposed methods learn normalization rules to transform all duplicate URLs into the same canonical form. A challenging aspect of this strategy is deriving a set of general and precise rules. In this work, we present DUSTER, a new approach to derive quality rules that take advantage of a multi-sequence alignment strategy. We demonstrate that a full multi-sequence alignment of URLs with duplicated content, before the generation of the rules, can lead to the deployment of very effective rules. By evaluating our method, we observed it achieved larger reductions in the number of duplicate URLs than our best baseline, with gains of 82 and 140.74 percent in two different web collections. Kaio Wagner Lima Rodrigues, Marco Cristo, Edleno Silva de Moura, Altigran S. da Silva |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2014 | MKStream: An Efficient Algorithm for Processing Multiple Keyword Queries over XML Streams
Evandrino G. Barros, Alberto H. F. Laender, Mirella M. Moro, Altigran S. da Silva |
ER | 4 |
| 2014 | Learning to expand queries using entitiesabstractA substantial fraction of web search queries contain references to entities, such as persons, organizations, and locations. Recently, methods that exploit named entities have been shown to be more effective for query expansion than traditional pseudorelevance feedback methods. In this article, we introduce a supervised learning approach that exploits named entities for query expansion using Wikipedia as a repository of high‐quality feedback documents. In contrast with existing entity‐oriented pseudorelevance feedback approaches, we tackle query expansion as a learning‐to‐rank problem. As a result, not only do we select effective expansion terms but we also weigh these terms according to their predicted effectiveness. To this end, we exploit the rich structure of Wikipedia articles to devise discriminative term features, including each candidate term's proximity to the original query terms, as well as its frequency across multiple article fields and in category and infobox descriptors. Experiments on three Text REtrieval Conference web test collections attest the effectiveness of our approach, with gains of up to 23.32% in terms of mean average precision, 19.49% in terms of precision at 10, and 7.86% in terms of normalized discounted cumulative gain compared with a state‐of‐the‐art approach for entity‐oriented query expansion. Wladmir Cardoso Brandão, Rodrygo L. T. Santos, Nivio Ziviani, Edleno Silva de Moura, Altigran S. da Silva |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2013 | Fast document-at-a-time query processing using two-tier indexesabstractIn this paper we present two new algorithms designed to reduce the overall time required to process top-k queries. These algorithms are based on the document-at-a-time approach and modify the best baseline we found in the literature, Blockmax WAND (BMW), to take advantage of a two-tiered index, in which the first tier is a small index containing only the higher impact entries of each inverted list. This small index is used to pre-process the query before accessing a larger index in the second tier, resulting in considerable speeding up the whole process. The first algorithm we propose, named BMW-CS, achieves higher performance, but may result in small changes in the top results provided in the final ranking. The second algorithm, named BMW-t, preserves the top results and, while slower than BMW-CS, it is faster than BMW. In our experiments, BMW-CS was more than 40 times faster than BMW when computing top 10 results, and, while it does not guarantee preserving the top results, it preserved all ranking results evaluated at this level. Cristian Rossi, Edleno Silva de Moura, André Luiz da Costa Carvalho, Altigran S. da Silva |
SIGIR | 4 |
| 2013 | Learning URL Normalization Rules Using Multiple Alignment of Sequences
Kaio Wagner Lima Rodrigues, Marco Cristo, Edleno Silva de Moura, Altigran S. da Silva |
SPIRE | 4 |
| 2013 | Learning to Schedule Webpage Updates Using Genetic Programming
Aécio S. R. Santos, Nivio Ziviani, Jussara M. Almeida, Cristiano R. de Carvalho, Edleno Silva de Moura, Altigran S. da Silva |
SPIRE | 6 |
| 2013 | An evolutionary approach to complex schema matching
Moisés G. de Carvalho, Alberto H. F. Laender, Marcos André Gonçalves, Altigran S. da Silva |
Inf. Syst. | 4 |
| 2012 | Named Entity Disambiguation in Streaming Data
Alexandre Davis, Adriano Veloso, Altigran S. da Silva, Alberto H. F. Laender, Wagner Meira Jr. |
ACL (1) | 3 |
| 2012 | Sorted dominant local color for searching large and heterogeneous image databases
Márcio L. A. Vidal, João M. B. Cavalcanti, Edleno Silva de Moura, Altigran S. da Silva, Ricardo da Silva Torres |
ICPR | 4 |
| 2012 | LePrEF: Learn to precompute evidence fusion for efficient query evaluationabstractState‐of‐the‐art search engine ranking methods combine several distinct sources of relevance evidence to produce a high‐quality ranking of results for each query. The fusion of information is currently done at query‐processing time, which has a direct effect on the response time of search systems. Previous research also shows that an alternative to improve search efficiency in textual databases is to precompute term impacts at indexing time. In this article, we propose a novel alternative to precompute term impacts, providing a generic framework for combining any distinct set of sources of evidence by using a machine‐learning technique. This method retains the advantages of producing high‐quality results, but avoids the costs of combining evidence at query‐processing time. Our method, called Learn to Precompute Evidence Fusion (LePrEF), uses genetic programming to compute a unified precomputed impact value for each term found in each document prior to query processing, at indexing time. Compared with previous research on precomputing term impacts, our method offers the advantage of providing a generic framework to precompute impact using any set of relevance evidence at any text collection, whereas previous research articles do not. The precomputed impact values are indexed and used later for computing document ranking at query‐processing time. By doing so, our method effectively reduces the query processing to simple additions of such impacts. We show that this approach, while leading to results comparable to state‐of‐the‐art ranking methods, also can lead to a significant decrease in computational costs during query processing. André Luiz da Costa Carvalho, Cristian Rossi, Edleno Silva de Moura, Altigran S. da Silva, David Fernandes de Oliveira |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2012 | A Genetic Programming Approach to Record DeduplicationabstractSeveral systems that rely on consistent data to offer high-quality services, such as digital libraries and e-commerce brokers, may be affected by the existence of duplicates, quasi replicas, or near-duplicate entries in their repositories. Because of that, there have been significant investments from private and government organizations for developing methods for removing replicas from its data repositories. This is due to the fact that clean and replica-free repositories not only allow the retrieval of higher quality information but also lead to more concise data and to potential savings in computational time and resources to process this data. In this paper, we propose a genetic programming approach to record deduplication that combines several different pieces of evidence extracted from the data content to find a deduplication function that is able to identify whether two entries in a repository are replicas or not. As shown by our experiments, our approach outperforms an existing state-of-the-art method found in the literature. Moreover, the suggested functions are computationally less demanding since they use fewer evidence. In addition, our genetic programming approach is capable of automatically adapting these functions to a given fixed replica identification boundary, freeing the user from the burden of having to choose and tune this parameter. Moisés G. de Carvalho, Alberto H. F. Laender, Marcos André Gonçalves, Altigran S. da Silva |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2011 | Multiple keyword-based queries over XML streamsabstractIn this paper, we propose that various keyword-based queries be processed over XML streams in a multi-query processing way. Our algorithms rely on parsing stacks designed for simultaneously matching terms from several distinct queries and use new query indexes to speed up search operations when processing a large number of queries. Besides defining a new problem and novel solutions, we perform experiments in which aspects related to performance and scalability are examined. Felipe da C. Hummel, Altigran S. da Silva, Mirella M. Moro, Alberto H. F. Laender |
CIKM | 2 |
| 2011 | Semi-supervised genetic programming for classificationabstractLearning from unlabeled data provides innumerable advantages to a wide range of applications where there is a huge amount of unlabeled data freely available. Semi-supervised learning, which builds models from a small set of labeled examples and a potential large set of unlabeled examples, is a paradigm that may effectively use those unlabeled data. Here we propose KGP, a semi-supervised transductive genetic programming algorithm for classification. Apart from being one of the first semi-supervised algorithms, it is transductive (instead of inductive), i.e., it requires only a training dataset with labeled and unlabeled examples, which should represent the complete data domain. The algorithm relies on the three main assumptions on which semi-supervised algorithms are built, and performs both global search on labeled instances and local search on unlabeled instances. Periodically, unlabeled examples are moved to the labeled set after a weighted voting process performed by a committee. Results on eight UCI datasets were compared with Self-Training and KNN, and showed KGP as a promising method for semi-supervised learning. Filipe de Lima Arcanjo, Gisele L. Pappa, Paulo Viana Bicalho, Wagner Meira Jr., Altigran S. da Silva |
GECCO | 5 |
| 2011 | A site oriented method for segmenting web pagesabstractInformation about how to segment a Web page can be used nowadays by applications such as segment aware Web search, classification and link analysis. In this research, we propose a fully automatic method for page segmentation and evaluate its application through experiments with four separate Web sites. While the method may be used in other applications, our main focus in this article is to use it as input to segment aware Web search systems. Our results indicate that the proposed method produces better segmentation results when compared to the best segmentation method we found in literature. Further, when applied as input to a segment aware Web search method, it produces results close to those produced when using a manual page segmentation method. David Fernandes de Oliveira, Edleno Silva de Moura, Altigran S. da Silva, Berthier A. Ribeiro-Neto, Edisson Braga Araújo |
SIGIR | 3 |
| 2011 | Joint unsupervised structure discovery and information extractionabstractIn this paper we present JUDIE (Joint Unsupervised Structure Discovery and Information Extraction), a new method for automatically extracting semi-structured data records in the form of continuous text (e.g., bibliographic citations, postal addresses, classified ads, etc.) and having no explicit delimiters between them. While in state-of-the-art Information Extraction methods the structure of the data records is manually supplied the by user as a training step, JUDIE is capable of detecting the structure of each individual record being extracted without any user assistance. This is accomplished by a novel Structure Discovery algorithm that, given a sequence of labels representing attributes assigned to potential values, groups these labels into individual records by looking for frequent patterns of label repetitions among the given sequence. We also show how to integrate this algorithm in the information extraction process by means of successive refinement steps that alternate information extraction and structure discovery. Through an extensively experimental evaluation with different datasets in distinct domains, we compare JUDIE with state-of-the-art information extraction methods and conclude that, even without any user intervention, it is able to achieve high quality results on the tasks of discovering the structure of the records and extracting information from them. Eli Cortez, Daniel Oliveira 0007, Altigran S. da Silva, Edleno Silva de Moura, Alberto H. F. Laender |
SIGMOD Conference | 3 |
| 2011 | A New Approach for Verifying URL Uniqueness in Web Crawlers
Wallace Favoreto Henrique, Nivio Ziviani, Marco Cristo, Edleno Silva de Moura, Altigran S. da Silva, Cristiano R. de Carvalho |
SPIRE | 5 |
| 2011 | Lightweight methods for large-scale product categorizationabstractIn this article, we present a study about classification methods for large-scale categorization of product offers on e-shopping web sites. We present a study about the performance of previously proposed approaches and deployed a probabilistic approach to model the classification problem. We also studied an alternative way of modeling information about the description of product offers and investigated the usage of price and store of product offers as features adopted in the classification process. Our experiments used two collections of over a million product offers previously categorized by human editors and taxonomies of hundreds of categories from a real e-shopping web site. In these experiments, our method achieved an improvement of up to 9% in the quality of the categorization in comparison with the best baseline we have found. Eli Cortez, Mauro Rojas Herrera, Altigran S. da Silva, Edleno Silva de Moura, Marden S. Neubert |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2010 | Active Learning Genetic programming for record deduplicationabstractThe great majority of genetic programming (GP) algorithms that deal with the classification problem follow a supervised approach, i.e., they consider that all fitness cases available to evaluate their models are labeled. However, in certain application domains, a lot of human effort is required to label training data, and methods following a semi-supervised approach might be more appropriate. This is because they significantly reduce the time required for data labeling while maintaining acceptable accuracy rates. This paper presents the Active Learning GP (AGP), a semi-supervised GP, and instantiates it for the data deduplication problem. AGP uses an active learning approach in which a committee of multi-attribute functions votes for classifying record pairs as duplicates or not. When the committee majority voting is not enough to predict the class of the data pairs, a user is called to solve the conflict. The method was applied to three datasets and compared to two other deduplication methods. Results show that AGP guarantees the quality of the deduplication while reducing the number of labeled examples needed. Junio de Freitas, Gisele L. Pappa, Altigran S. da Silva, Marcos André Gonçalves, Edleno Silva de Moura, Adriano Veloso, Alberto H. F. Laender, Moisés G. de Carvalho |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | ONDUX: on-demand unsupervised learning for information extractionabstractInformation extraction by text segmentation (IETS) applies to cases in which data values of interest are organized in implicit semi-structured records available in textual sources (e.g. postal addresses, bibliographic information, ads). It is an important practical problem that has been frequently addressed in the recent literature. In this paper we introduce ONDUX (On Demand Unsupervised Information Extraction), a new unsupervised probabilistic approach for IETS. As other unsupervised IETS approaches, ONDUX relies on information available on pre-existing data to associate segments in the input string with attributes of a given domain. Unlike other approaches, we rely on very effective matching strategies instead of explicit learning strategies. The effectiveness of this matching strategy is also exploited to disambiguate the extraction of certain attributes through a reinforcement step that explores sequencing and positioning of attribute values directly learned on-demand from test data, with no previous human-driven training, a feature unique to ONDUX. This assigns to ONDUX a high degree of flexibility and results in superior effectiveness, as demonstrated by the experimental evaluation we report with textual sources from different domains, in which ONDUX is compared with a state-of-art IETS approach. Eli Cortez, Altigran S. da Silva, Marcos André Gonçalves, Edleno Silva de Moura |
SIGMOD Conference | 2 |
| 2010 | A Self-Supervised Approach for Extraction of Attribute-Value Pairs from Wikipedia Articles
Wladmir Cardoso Brandão, Edleno Silva de Moura, Altigran S. da Silva, Nivio Ziviani |
SPIRE | 3 |
| 2010 | Exploring features for the automatic identification of user goals in web search
Mauro Rojas Herrera, Edleno Silva de Moura, Marco Cristo, Thomaz Philippe Cavalcante Silva, Altigran S. da Silva |
Inf. Process. Manag. | 5 |
| 2010 | Information Systems Special Issue on SBBD 2007
Altigran S. da Silva |
Inf. Syst. | 1 |
| 2010 | Using structural information to improve search in Web collectionsabstractAbstract In this work, we investigate the problem of using the block structure of Web pages to improve ranking results. Starting with basic intuitions provided by the concepts of term frequency (TF) and inverse document frequency (IDF), we propose nine block‐weight functions to distinguish the impact of term occurrences inside page blocks, instead of inside whole pages. These are then used to compute a modified BM25 ranking function. Using four distinct Web collections, we ran extensive experiments to compare our block‐weight ranking formulas with two other baselines: (a) a BM25 ranking applied to full pages, and (b) a BM25 ranking that takes into account best blocks. Our methods suggest that our block‐weighting ranking method is superior to all baselines across all collections we used and that average gain in precision figures from 5 to 20% are generated. Edleno Silva de Moura, David Fernandes de Oliveira, Berthier A. Ribeiro-Neto, Altigran S. da Silva, Marcos André Gonçalves |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2010 | A Probabilistic Approach for Automatically Filling Form-Based Web InterfacesabstractIn this paper we present a proposal for the implementation and evaluation of a novel method for automatically using data-rich text for filling form-based input interfaces. Our solution takes a text as input, extracts implicit data values from it and fills appropriate fields. For this task, we rely on knowledge obtained from values of previous submissions for each field, which are freely obtained from the usage of the interfaces. Our approach, called iForm , exploits features related to the content and the style of these values, which are combined through a Bayesian framework. Through extensive experimentation, we show that our approach is feasible and effective, and that it works well even when only a few previous submissions to the input interface are available. Guilherme A. Toda, Eli Cortez, Altigran S. da Silva, Edleno Silva de Moura |
Proc. VLDB Endow. | 3 |
| 2009 | Automatically filling form-based web interfaces with free text inputsabstractOn the web of today the most prevalent solution for users to interact with data-intensive applications is the use of form-based interfaces composed by several data input fields, such as text boxes, radio buttons, pull-down lists, check boxes, etc. Although these interfaces are popular and effective, in many cases, free text interfaces are preferred over form-based ones. In this paper we discuss the proposal and the implementation of a novel IR-based method for using data rich free text to interact with form-based interfaces. Our solution takes a free text as input, extracts implicitly data values from it and fills appropriate fields using them. For this task, we rely on values of previous submissions for each field, which are freely obtained from the usage of form-based interfaces Guilherme A. Toda, Eli Cortez, Filipe de Sá Mesquita, Altigran S. da Silva, Edleno Silva de Moura, Marden S. Neubert |
WWW | 4 |
| 2009 | A strategy for allowing meaningful and comparable scores in approximate matching
Carina F. Dorneles, Marcos Freitas Nunes, Carlos Alberto Heuser, Viviane Pereira Moreira, Altigran S. da Silva, Edleno Silva de Moura |
Inf. Syst. | 5 |
| 2009 | An evolutionary approach for combining different sources of evidence in search engines
Thomaz Philippe Cavalcante Silva, Edleno Silva de Moura, João M. B. Cavalcanti, Altigran S. da Silva, Moisés G. de Carvalho, Marcos André Gonçalves |
Inf. Syst. | 4 |
| 2009 | A flexible approach for extracting metadata from bibliographic citationsabstractAbstract In this article we present FLUX‐CiM, a novel method for extracting components (e.g., author names, article titles, venues, page numbers) from bibliographic citations. Our method does not rely on patterns encoding specific delimiters used in a particular citation style. This feature yields a high degree of automation and flexibility, and allows FLUX‐CiM to extract from citations in any given format. Differently from previous methods that are based on models learned from user‐driven training, our method relies on a knowledge base automatically constructed from an existing set of sample metadata records from a given field (e.g., computer science, health sciences, social sciences, etc.). These records are usually available on the Web or other public data repositories. To demonstrate the effectiveness and applicability of our proposed method, we present a series of experiments in which we apply it to extract bibliographic data from citations in articles of different fields. Results of these experiments exhibit precision and recall levels above 94% for all fields, and perfect extraction for the large majority of citations tested. In addition, in a comparison against a state‐of‐the‐art information‐extraction method, ours produced superior results without the training phase required by that method. Finally, we present a strategy for using bibliographic data resulting from the extraction process with FLUX‐CiM to automatically update and expand the knowledge base of a given domain. We show that this strategy can be used to achieve good extraction results even if only a very small initial sample of bibliographic records is available for building the knowledge base. Eli Cortez, Altigran S. da Silva, Marcos André Gonçalves, Filipe de Sá Mesquita, Edleno Silva de Moura |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | A Genre-Aware Approach to Focused Crawling
Guilherme Tavares de Assis, Alberto H. F. Laender, Marcos André Gonçalves, Altigran S. da Silva |
World Wide Web | 4 |
| 2009 | On Finding Templates on Web Collections
Karane Vieira, André Luiz da Costa Carvalho, Klessius Berlt, Edleno Silva de Moura, Altigran S. da Silva, Juliana Freire |
World Wide Web | 5 |
| 2008 | Siphon++: a hidden-webcrawler for keyword-based interfacesabstractThe hidden Web consists of data that is generally hidden behind form interfaces, and as such, it is out of reach for traditional search engines. With the goal of leveraging the high-quality information in this largely unexplored portion of the Web, in this paper, we propose a new strategy for automatically retrieving data hidden behind keyword-based form interfaces. Unlike previous approaches to this problem, our strategy adapts the query generation and selection by detecting features of the index. We describe an extensive experimental evaluation which shows that: our strategy is able to derive appropriate queries to obtain high coverage while, at the same time, avoiding the retrieval of redundant data; and it obtains higher coverage and is more efficient approaches that use a fixed strategy for query generation. Karane Vieira, Luciano Barbosa, Juliana Freire, Altigran S. da Silva |
CIKM | 4 |
| 2008 | Locality-Based pruning methods for web searchabstractThis article discusses a novel approach developed for static index pruning that takes into account the locality of occurrences of words in the text. We use this new approach to propose and experiment on simple and effective pruning methods that allow a fast construction of the pruned index. The methods proposed here are especially useful for pruning in environments where the document database changes continuously, such as large-scale web search engines. Extensive experiments are presented showing that the proposed methods can achieve high compression rates while maintaining the quality of results for the most common query types present in modern search engines, namely, conjunctive and phrase queries. In the experiments, our locality-based pruning approach allowed reducing search engine indices to 30% of their original size, with almost no reduction in precision at the top answers. Furthermore, we conclude that even an extremely simple locality-based pruning method can be competitive when compared to complex methods that do not rely on locality information. Edleno Silva de Moura, Célia Francisca dos Santos, Bruno Dos Santos de Araujo, Altigran S. da Silva, Pável Calado, Mario A. Nascimento |
ACM Trans. Inf. Syst. | 4 |
| 2007 | A strategy for allowing meaningful and comparable scores in approximate matchingabstractThe goal of approximate data matching is to assess whether two distinct data instances represent the same real world object. This is usually achieved through the use of a similarity function, which returns a score that defines how similar two data instances are. If this score surpasses a given threshold, both data instances are considered as representing the same real world object. The score values returned by a similarity function depend on the algorithm that implements the function and have no meaning to the user (apart from the fact that a higher similarity value means that two data instances are more similar). In this paper, we propose that instead of defining the threshold in terms of the scores returned by a similarity function, the user specifies the precision that is expected from the matching process. Precision is a well known quality measure and has a clear interpretation from the user's point of view. Our approach relies on mapping between similarity scores and precision values based on a training data set. Experimental results show the training may be executed against a representative data set, and reused for other databases from the same domain. Carina F. Dorneles, Carlos Alberto Heuser, Viviane Pereira Moreira, Altigran S. da Silva, Edleno Silva de Moura |
CIKM | 4 |
| 2007 | Computing block importance for searching on web sitesabstractIn this paper we consider the problem of using the block structure of a Web page to improve ranking results when searching for information on Web sites. Given the block structure of the Web pages as input, we propose a method for computing the importance of each block (in the form of block weights) in a Web collection. As we show through experiments, the deployment of our method may allow a significant improvement in the quality of search results. We ran experiments to compare the quality of search results when using our method to the quality obtained when using no structure information. When compared to a ranking method that considered pages as monolithic units, our block-based ranking method led to improvements in the quality of search results in experiments with two sites with heterogeneous structures. Further, our method does not increase the cost of processing queries when compared to the systems using no structural information. David Fernandes de Oliveira, Edleno Silva de Moura, Berthier A. Ribeiro-Neto, Altigran S. da Silva, Marcos André Gonçalves |
CIKM | 4 |
| 2007 | Organizing Hidden-Web Databases by Clustering Visible Web DocumentsabstractIn this paper we address the problem of organizing hidden-Web databases. Given a heterogeneous set of Web forms that serve as entry points to hidden-Web databases, our goal is to cluster the forms according to the database domains to which they belong. We propose a new clustering approach that models Web forms as a set of hyperlinked objects and considers visible information in the form context - both within and in the neighborhood of forms - as the basis for similarity comparison. Since the clustering is performed over features that can be automatically extracted, the process is scalable. In addition, because it uses a rich set of metadata, our approach is able to handle a wide range of forms, including content-rich forms that contain multiple attributes, as well as simple keyword-based search interfaces. An experimental evaluation over real Web data shows that our strategy generates high-quality clusters - measured both in terms of entropy and F-measure. This indicates that our approach provides an effective and general solution to the problem of organizing hidden-Web databases. Luciano Barbosa, Juliana Freire, Altigran S. da Silva |
ICDE | 3 |
| 2007 | A Scalable Parallel Deduplication AlgorithmabstractThe identification of replicas in a database is fundamental to improve the quality of the information. Deduplication is the task of identifying replicas in a database that refer to the same real world entity. This process is not always trivial, because data may be corrupted during their gathering, storing or even manipulation. Problems such as misspelled names, data truncation, data input in a wrong format, lack of conventions (like how to abbreviate a name), missing data or even fraud may lead to the insertion of replicas in a database. The deduplication process may be very hard, if not impossible, to be performed manually, since actual databases may have hundreds of millions of records. In this paper, we present our parallel deduplication algorithm, called FER- APARDA. By using probabilistic record linkage, we were able to successfully detect replicas in synthetic datasets with more than 1 million records in about 7 minutes using a 20- computer cluster, achieving an almost linear speedup. We believe that our results do not have similar in the literature when it comes to the size of the data set and the processing time. Walter Santos, Thiago Teixeira, Carla Machado, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Altigran S. da Silva |
SBAC-PAD | 7 |
| 2007 | Exploiting Genre in Focused Crawling
Guilherme Tavares de Assis, Alberto H. F. Laender, Marcos André Gonçalves, Altigran S. da Silva |
SPIRE | 4 |
| 2007 | A cost-effective method for detecting web site replicas on search engine databases
André Luiz da Costa Carvalho, Edleno Silva de Moura, Altigran S. da Silva, Klessius Berlt, Allan J. S. Bezerra |
Data Knowl. Eng. | 3 |
| 2007 | LABRADOR: Efficiently publishing relational databases on the web by using keyword-based query interfaces
Filipe de Sá Mesquita, Altigran S. da Silva, Edleno Silva de Moura, Pável Calado, Alberto H. F. Laender |
Inf. Process. Manag. | 2 |
| 2006 | A fast and robust method for web page template detection and removalabstractThe widespread use of templates on the Web is considered harmful for two main reasons. Not only do they compromise the relevance judgment of many web IR and web mining methods such as clustering and classification, but they also negatively impact the performance and resource usage of tools that process web pages. In this paper we present a new method that efficiently and accurately removes templates found in collections of web pages. Our method works in two steps. First, the costly process of template detection is performed over a small set of sample pages. Then, the derived template is removed from the remaining pages in the collection. This leads to substantial performance gains when compared to previous approaches that combine template detection and removal. We show, through an experimental evaluation, that our approach is effective for identifying terms occurring in templates - obtaining F-measure values around 0.9, and that it also boosts the accuracy of web page clustering and classification methods. Karane Vieira, Altigran S. da Silva, Nick Pinto, Edleno Silva de Moura, João M. B. Cavalcanti, Juliana Freire |
CIKM | 2 |
| 2006 | Structure-driven crawler generation by exampleabstractMany Web IR and Digital Library applications require a crawling process to collect pages with the ultimate goal of taking advantage of useful information available on Web sites. For some of these applications the criteria to determine when a page is to be present in a collection are related to the page content. However, there are situations in which the inner structure of the pages provides a better criteria to guide the crawling process than their content. In this paper, we present a structure-driven approach for generating Web crawlers that requires a minimum effort from users. The idea is to take as input a sample page and an entry point to a Web site and generate a structure-driven crawler based on navigation patterns, sequences of patterns for the links a crawler has to follow to reach the pages structurally similar to the sample page. In the experiments we have carried out, structure-driven crawlers generated by our new approach were able to collect all pages that match the samples given, including those pages added after their generation. Márcio L. A. Vidal, Altigran S. da Silva, Edleno Silva de Moura, João M. B. Cavalcanti |
SIGIR | 2 |
| 2006 | GoGetIt!: a tool for generating structure-driven web crawlersabstractWe present GoGetIt!, a tool for generating structure-driven crawlers that requires a minimum effort from the users. The tool takes as input a sample page and an entry point to a Web site and generates a structure-driven crawler based on navigation patterns, sequences of patterns for the links a crawler has to follow to reach the pages structurally similar to the sample page. In the experiments we have performed, structure-driven crawlers generated by GoGetIt! were able to collect all pages that match the samples given, including those pages added after their generation. Márcio L. A. Vidal, Altigran S. da Silva, Edleno Silva de Moura, João M. B. Cavalcanti |
WWW | 2 |
| 2005 | Improving Web search efficiency via a locality based static pruning methodabstractThe unarguably fast, and continuous, growth of the volume of indexed (and indexable) documents on the Web poses a great challenge for search engines. This is true regarding not only search effectiveness but also time and space efficiency. In this paper we present an index pruning technique targeted for search engines that addresses the latter issue without disconsidering the former. To this effect, we adopt a new pruning strategy capable of greatly reducing the size of search engine indices. Experiments using a real search engine show that our technique can reduce the indices' storage costs by up to 60% over traditional lossless compression methods, while keeping the loss in retrieval precision to a minimum. When compared to the indices size with no compression at all, the compression rate is higher than 88%, i.e., less than one eighth of the original size. More importantly, our results indicate that, due to the reduction in storage overhead, query processing time can be reduced to nearly 65% of the original time, with no loss in average precision. The new method yields significative improvements when compared against the best known static pruning method for search engine indices. In addition, since our technique is orthogonal to the underlying search algorithms, it can be adopted by virtually any search engine. Edleno Silva de Moura, Célia Francisca dos Santos, Daniel R. Fernandes, Altigran S. da Silva, Pável Calado, Mario A. Nascimento |
WWW | 4 |
| 2004 | Information Retrieval Aware Web Site Modelling and Generation
Keyla Ahnizeret, David Fernandes de Oliveira, João M. B. Cavalcanti, Edleno Silva de Moura, Altigran S. da Silva |
ER | 5 |
| 2004 | Automatic web news extraction using tree edit distanceabstractThe Web poses itself as the largest data repository ever available in the history of humankind. Major efforts have been made in order to provide efficient access to relevant information within this huge repository of data. Although several techniques have been developed to the problem of Web data extraction, their use is still not spread, mostly because of the need for high human intervention and the low quality of the extraction results.In this paper, we present a domain-oriented approach to Web data extraction and discuss its application to automatically extracting news from Web sites. Our approach is based on a highly efficient tree structure analysis that produces very effective results. We have tested our approach with several important Brazilian on-line news sites and achieved very precise results, correctly extracting 87.71% of the news in a set of 4088 pages distributed among 35 different sites. Davi de Castro Reis, Paulo Braz Golgher, Altigran S. da Silva, Alberto H. F. Laender |
WWW | 3 |
| 2004 | Automatic generation of agents for collecting hidden Web pages for data extraction
Juliano Palmieri Lage, Altigran S. da Silva, Paulo Braz Golgher, Alberto H. F. Laender |
Data Knowl. Eng. | 2 |
| 2004 | A Bayesian network approach to searching Web databases through keyword-based queries
Pável Calado, Altigran S. da Silva, Alberto H. F. Laender, Berthier A. Ribeiro-Neto, Rodrigo C. Vieira |
Inf. Process. Manag. | 2 |
| 2002 | Using Nested Tables for Representing and Querying Semistructured Web Data
Irna M. R. Evangelista Filha, Altigran S. da Silva, Alberto H. F. Laender, David W. Embley |
CAiSE | 2 |
| 2002 | Web-DL: an experience in building digital libraries from the webabstractThe Web contains a huge volume of information, almost all unstructured and, therefore, difficult to manage. In Digital Libraries, however, information is explicitly organized, described, and managed. In this paper, we propose an architecture that allows the construction of digital libraries from the Web, using standard protocols and archival technologies, and incorporating powerful digital library and data extraction tools, thus benefiting from the breadth of the Web contents, but supporting services and organization available in digital libraries. The proposed architecture was applied to the Networked Digital Library of Theses and Dissertations, providing an important first step toward rapid construction of large DLs from the Web, as well as a large-scale solution for interoperability between independent digital libraries. Pável Calado, Altigran S. da Silva, Berthier A. Ribeiro-Neto, Alberto H. F. Laender, Juliano Palmieri Lage, Davi de Castro Reis, Pablo A. Roberto, Monique V. Vieira, Marcos André Gonçalves, Edward A. Fox |
CIKM | 2 |
| 2002 | Searching web databases by structuring keyword-based queriesabstractOn-line information services have become widespread in the Web nowadays. However, Web users are non-specialized and have a great variety of interests. Thus, interfaces for Web databases must be simple and uniform. In this paper we present an approach, based on Bayesian networks, for querying Web databases using keywords only. According to this approach, the user inputs a query through a simple search-box interface. From the input query, one or more plausible structured queries are derived and submitted to Web databases. The results are then retrieved and presented to the user as ranked answers. Our approach reduces the complexity of existing on-line interfaces and offers a solution to the problem of querying several distinct Web databases with a single interface. The applicability of the proposed approach was demonstrated by experimental results with 3 databases, obtained with a prototype search system that implements it. We have found that from 77% to 95% of the time, one of the top three resulting structured queries is the proper one. Further, when the user selects one of these three top queries for processing, the ranked answers present average precision figures from 60% to about 100%. Pável Calado, Altigran S. da Silva, Rodrigo C. Vieira, Alberto H. F. Laender, Berthier A. Ribeiro-Neto |
CIKM | 2 |
| 2002 | Representing and Querying Semistructured Web Data Using Nested Tables with Structural Variants
Altigran S. da Silva, Irna M. R. Evangelista Filha, Alberto H. F. Laender, David W. Embley |
ER | 1 |
| 2002 | A Framework for Generating Attribute Extractors for Web Data Sources
Davi de Castro Reis, Robson Braga Araújo, Altigran S. da Silva, Berthier A. Ribeiro-Neto |
SPIRE | 3 |
| 2002 | DEByE - Data Extraction By Example
Alberto H. F. Laender, Berthier A. Ribeiro-Neto, Altigran S. da Silva |
Data Knowl. Eng. | 3 |
| 2001 | Bootstrapping for Example-Based Data ExtractionabstractThe effortless generation of wrappers for Web data sources is a crucial task if proper access to the huge amount of semi-structured data on the Web is to be granted. In particular, the development of strategies for wrapper generation based on user-given examples is currently one of the most promising research directions in Web data extraction. In this paper we show how to use a pre-existing data repository to automatically generate examples and allow full automated example-based data extraction. To demonstrate the feasibility of our approach we provide a number of results obtained from experiments we carried out and discuss how our ideas can be used to improve extraction rates and for providing resilience and adaptiveness for example-based generated wrappers. Paulo Braz Golgher, Altigran S. da Silva, Alberto H. F. Laender, Berthier A. Ribeiro-Neto |
CIKM | 2 |
| 2001 | Storing Semistructured Data in Relational DatabasesabstractThis paper presents an approach to storing semistructured data in relational databases. We focus on semistructured data as extracted from Web pages by a tool called DEBYE (Data Extraction By Example), and organized according to its data model, the DEByE Object Model (DEByEOM). The approach presented here consists in representing the structure of the objects extracted by DEByE by a relational schema and populating the corresponding database accordingly. We also show how to retrieve such objects by automatically transforming high-level query specifications (query patterns) into SQL queries that are executed over the relational database. Experiments results carried out to evaluate our approach are also described. Karine V. Magalhães, Alberto H. F. Laender, Altigran S. da Silva |
SPIRE | 3 |
| 2000 | On the relational representation of complex specialization structures
Altigran S. da Silva, Alberto H. F. Laender, Marco A. Casanova |
Inf. Syst. | 1 |
| 1999 | Extracting Semi-Structured Data Through ExamplesabstractIn this paper, we describe an innovative approach to extracting semi-structured data from Web sources. The idea is to collect a couple of example objects from the user and to use this information to extract new objects from new pages or texts. To perform the extraction of new objects, we introduce a bottom-up extration strategy and, through experimentation, demonstrate that it works quite effectively with distinct Web sources, even if only a few examples are provided by the user. Berthier A. Ribeiro-Neto, Alberto H. F. Laender, Altigran S. da Silva |
CIKM | 3 |
| 1996 | An Approach to Maintaining Optimized Relational Representations of Entity-Relationship Schemas
Altigran S. da Silva, Alberto H. F. Laender, Marco A. Casanova |
ER | 1 |